Haystack
← Back to Jobs
Technology

Software Developer / Engineer - Philadelphia, PA (Locals Only)

Apetan ConsultingPhiladelphia, PA🇺🇸United StatesPosted 23 Jul 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Software Developer / Engineer
Location: Philadelphia, PA 
Work Schedule: Hybrid 3 days on site, 2 remote


Position Overview

We are seeking a Software Developer / Engineer to help design and implement an on-premises Large Language Model (LLM) platform with Retrieval-Augmented Generation (RAG) capabilities. This role will focus on deploying open-source AI models, integrating vector databases, and building secure, enterprise-grade AI solutions in a private environment.

This is an excellent opportunity for a developer with hands-on experience in modern AI technologies who enjoys building scalable, high-performance systems.

Responsibilities

  • Deploy and optimize open-source large language models (LLMs) such as Meta Llama 3 and Mistral/Mixtral in on-premises or private environments.
  • Develop Python-based applications for LLM inference, prompt engineering, and model integration.
  • Optimize CPU-based model inference through quantization and performance tuning.
  • Design and implement Retrieval-Augmented Generation (RAG) (RAG) pipelines.
  • Configure and manage open-source vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Generate and manage embeddings while implementing metadata filtering strategies.
  • Support enterprise security requirements, including air-gapped deployments, access controls, data privacy, and audit logging.
  • Produce technical documentation, deployment guidance, and knowledge transfer materials for internal teams.
  • Build a working prototype integrating an LLM, vector database, and RAG architecture.

Required Qualifications

  • Professional experience deploying open-source LLMs (e.g., Meta Llama 3, Mistral/Mixtral) in on-premises or private environments.
  • Strong Python development experience.
  • Hands-on experience with LLM inference, prompt engineering, and AI application integration.
  • Experience optimizing CPU-based inference through model quantization and performance tuning.
  • Experience with vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Proven experience implementing Retrieval-Augmented Generation (RAG) solutions.
  • Understanding of enterprise security, data privacy, air-gapped environments, access controls, and audit logging.

Preferred Qualifications

  • Experience with LangChain or LlamaIndex.
  • Familiarity with Docker and Kubernetes.
  • Experience with inference frameworks such as vLLM, llama.cpp, or Hugging Face Transformers.
  • Experience with Rust, Go, or C++.
  • Previous experience working in enterprise or regulated environments.

Deliverables

  • Reference architecture and deployment guidance.
  • Working prototype integrating an LLM, vector database, and RAG solution.
  • Technical documentation and knowledge transfer to internal teams.

Skills

Docker
Rust
C++
Hugging Face
Kubernetes
LLM
Python

Similar jobs