Haystack
← Back to Jobs
Technology

Generative AI Engineer- Philadelphia, PA

QTech US IncPhiladelphia, PA🇺🇸United StatesPosted 4 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Job Title: Generative AI Engineer
Location: Philadelphia, PA - Hybrid
Looking for W2 candidates. No C2C

Job Summary:
Philadelphia Gas Works (PGW) is seeking an experienced Generative AI Engineer / On-Prem LLM & Vector Database Consultant to design, deploy, and optimize an enterprise-grade on-premises Large Language Model (LLM) and Vector Database solution in a secure environment. The ideal candidate will possess strong expertise in open-source LLM deployment, Retrieval-Augmented Generation (RAG), vector databases, semantic search, and enterprise AI infrastructure while ensuring data privacy, security, and high-performance inference in private or air-gapped environments.
Key Responsibilities:
Design and implement an on-premises LLM architecture for secure enterprise AI applications.
Deploy and optimize open-source LLMs including Meta Llama 3, Mistral, and Mixtral.
Build and implement Retrieval-Augmented Generation (RAG) pipelines.
Design, configure, and manage Vector Database solutions for semantic search.
Develop Python-based LLM inference, orchestration, prompt engineering, and integrations.
Optimize model performance using CPU inference, model quantization, and inference tuning.
Generate and manage embeddings, metadata filtering, and semantic search workflows.
Implement enterprise-grade authentication, authorization, access controls, and audit logging.
Ensure compliance with secure, private, and air-gapped deployment requirements.
Deliver deployment architecture, technical documentation, and knowledge transfer sessions.
Build a working prototype integrating LLM + Vector Database + RAG Pipeline.
Collaborate with engineering teams to deploy scalable AI infrastructure.
Required Skills:
Strong experience deploying Open-Source LLMs (Meta Llama 3, Mistral, Mixtral).
Proficient in Python, Prompt Engineering, LLM Inference, Model Orchestration, and AI Integration.
Experience with CPU-based inference, Model Quantization, and performance optimization.
Hands-on experience with Vector Databases (Qdrant, Chroma, Milvus, pgvector).
Proven expertise in building Retrieval-Augmented Generation (RAG) pipelines.
Experience with Embeddings, Metadata Filtering, and Semantic Search.
Strong knowledge of Enterprise Security, Data Privacy, and Air-Gapped Deployments.
Experience implementing Authentication, Authorization, Access Controls, and Audit Logging.
Preferred Qualifications:
Experience with LangChain and/or LlamaIndex.
Knowledge of Rust, Go, or C++.
Experience with Docker and Kubernetes for on-prem deployments.
Familiarity with inference frameworks:
o vLLM
o llama.cpp
o Hugging Face Transformers
Experience working in regulated or enterprise environments.
Experience designing enterprise AI reference architectures.
Best Regards:
Tejaswani R.
Phone: +1-
Email:

Skills

Docker
Rust
C++
Generative AI
Hugging Face
Kubernetes
LLM
Python

Similar jobs