← Back to Jobs
Technology
Generative AI Engineer- Philadelphia, PA
QTech US IncPhiladelphia, PA🇺🇸United StatesPosted 4 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
Job Title: Generative AI Engineer
Location: Philadelphia, PA - Hybrid
Looking for W2 candidates. No C2C
Location: Philadelphia, PA - Hybrid
Looking for W2 candidates. No C2C
Job Summary:
Philadelphia Gas Works (PGW) is seeking an experienced Generative AI Engineer / On-Prem LLM & Vector Database Consultant to design, deploy, and optimize an enterprise-grade on-premises Large Language Model (LLM) and Vector Database solution in a secure environment. The ideal candidate will possess strong expertise in open-source LLM deployment, Retrieval-Augmented Generation (RAG), vector databases, semantic search, and enterprise AI infrastructure while ensuring data privacy, security, and high-performance inference in private or air-gapped environments.
Key Responsibilities:
Design and implement an on-premises LLM architecture for secure enterprise AI applications.
Deploy and optimize open-source LLMs including Meta Llama 3, Mistral, and Mixtral.
Build and implement Retrieval-Augmented Generation (RAG) pipelines.
Design, configure, and manage Vector Database solutions for semantic search.
Develop Python-based LLM inference, orchestration, prompt engineering, and integrations.
Optimize model performance using CPU inference, model quantization, and inference tuning.
Generate and manage embeddings, metadata filtering, and semantic search workflows.
Implement enterprise-grade authentication, authorization, access controls, and audit logging.
Ensure compliance with secure, private, and air-gapped deployment requirements.
Deliver deployment architecture, technical documentation, and knowledge transfer sessions.
Build a working prototype integrating LLM + Vector Database + RAG Pipeline.
Collaborate with engineering teams to deploy scalable AI infrastructure.
Design and implement an on-premises LLM architecture for secure enterprise AI applications.
Deploy and optimize open-source LLMs including Meta Llama 3, Mistral, and Mixtral.
Build and implement Retrieval-Augmented Generation (RAG) pipelines.
Design, configure, and manage Vector Database solutions for semantic search.
Develop Python-based LLM inference, orchestration, prompt engineering, and integrations.
Optimize model performance using CPU inference, model quantization, and inference tuning.
Generate and manage embeddings, metadata filtering, and semantic search workflows.
Implement enterprise-grade authentication, authorization, access controls, and audit logging.
Ensure compliance with secure, private, and air-gapped deployment requirements.
Deliver deployment architecture, technical documentation, and knowledge transfer sessions.
Build a working prototype integrating LLM + Vector Database + RAG Pipeline.
Collaborate with engineering teams to deploy scalable AI infrastructure.
Required Skills:
Strong experience deploying Open-Source LLMs (Meta Llama 3, Mistral, Mixtral).
Proficient in Python, Prompt Engineering, LLM Inference, Model Orchestration, and AI Integration.
Experience with CPU-based inference, Model Quantization, and performance optimization.
Hands-on experience with Vector Databases (Qdrant, Chroma, Milvus, pgvector).
Proven expertise in building Retrieval-Augmented Generation (RAG) pipelines.
Experience with Embeddings, Metadata Filtering, and Semantic Search.
Strong knowledge of Enterprise Security, Data Privacy, and Air-Gapped Deployments.
Experience implementing Authentication, Authorization, Access Controls, and Audit Logging.
Strong experience deploying Open-Source LLMs (Meta Llama 3, Mistral, Mixtral).
Proficient in Python, Prompt Engineering, LLM Inference, Model Orchestration, and AI Integration.
Experience with CPU-based inference, Model Quantization, and performance optimization.
Hands-on experience with Vector Databases (Qdrant, Chroma, Milvus, pgvector).
Proven expertise in building Retrieval-Augmented Generation (RAG) pipelines.
Experience with Embeddings, Metadata Filtering, and Semantic Search.
Strong knowledge of Enterprise Security, Data Privacy, and Air-Gapped Deployments.
Experience implementing Authentication, Authorization, Access Controls, and Audit Logging.
Preferred Qualifications:
Experience with LangChain and/or LlamaIndex.
Knowledge of Rust, Go, or C++.
Experience with Docker and Kubernetes for on-prem deployments.
Familiarity with inference frameworks:
o vLLM
o llama.cpp
o Hugging Face Transformers
Experience working in regulated or enterprise environments.
Experience designing enterprise AI reference architectures.
Experience with LangChain and/or LlamaIndex.
Knowledge of Rust, Go, or C++.
Experience with Docker and Kubernetes for on-prem deployments.
Familiarity with inference frameworks:
o vLLM
o llama.cpp
o Hugging Face Transformers
Experience working in regulated or enterprise environments.
Experience designing enterprise AI reference architectures.
Best Regards:
Tejaswani R.
Phone: +1-
Email:
Tejaswani R.
Phone: +1-
Email:
Skills
Docker
Rust
C++
Generative AI
Hugging Face
Kubernetes
LLM
Python
Similar jobs
GEN AI Developer
Stanley David and Associates · Minneapolis, United States
Just nowSenior Java AI Engineer
QUANTUM TECHNOLOGIES LLC · Dallas, United States
4 minutes ago€82/hrAgentic AI Architect - Remote / Telecommute
Cynet Systems · Dallas, United States
28 minutes ago$50 - $55/hrAI Engineer / AI Architect
VeridianTech · San Jose, United States
44 minutes agoGen AI Engineer - Minneapolis, MN, Detroit, MI, Madison, WI, Columbus, OH, Chicago, IL.
TechniPros, LLC · Minneapolis, United States
1 hour agoAI Engineer
eTeam, Inc. · Minneapolis, United States
1 hour ago