Haystack
← Back to Jobs
Temporary/Casual
Technology
TE

Software Developer/Engineer (Mid Level experience)

TechArmyPhiladelphia, Pennsylvania🇺🇸United StatesPosted Jul 30, 2026

Why This Role Stands Out

This contract role offers an exciting opportunity to gain in-demand expertise in on-premise LLM and Vector DB implementation, with a hybrid work model providing flexibility. You'll thrive here if you have hands-on experience with open-source LLMs, Python, and vector databases, allowing you to build innovative RAG pipelines and contribute to cutting-edge AI solutions. Apply now to be at the forefront of this rapidly evolving field!

Quick Overview

Seniority
Mid Senior
Employment type
Temporary/Casual
Work mode
On Site
Location
Philadelphia, Pennsylvania, United States
Posted
7 weeks ago
DockerRustC++Hugging FaceKubernetesLLMPython

Job Description

Position Type: Contract

Location: Philadelphia | Work Mode: Hybrid, minimum 3 days in the office

Interview Schedule: 1st interview, 1-hour, in-person; 2nd interview, 1-hour, in-person

Consultant Requirements – On-Prem LLM & Vector DB Implementation

Core Experience

Hands-on experience deploying open-source LLMs such as Meta Llama 3 and Mistral / Mixtral in on-prem or private environments

Strong proficiency in Python for LLM inference, prompt engineering, and integration

Experience with CPU-based inference, model quantization, and performance tuning

Vector Databases & RAG

Practical experience with open-source vector databases such as Qdrant, Chroma, Milvus, or pgvector

Proven implementation of Retrieval-Augmented Generation (RAG) pipelines

Experience generating and managing embeddings and metadata filtering

Security & Governance

Understanding of data privacy, air-gapped deployments, and enterprise security requirements

Experience implementing access controls and audit logging

Nice to Have

Experience with LangChain or LlamaIndex

Exposure to Rust, Go, or C++ for high-performance services

Familiarity with Docker and Kubernetes for on-prem deployments

Knowledge of inference frameworks (e.g., vLLM, llama.cpp, Hugging Face Transformers)

Prior work in regulated or enterprise environments

Deliverables

Reference architecture and deployment guidance

Working prototype (LLM + vector DB + RAG)

Documentation and knowledge transfer to internal teams

Similar jobs