Haystack
← Back to Jobs
Engineering
TB

LLM Engineer

The Brixton GroupDallas, TX🇺🇸United StatesPosted 15 Sept 2026

Why This Role Stands Out

You will gain invaluable experience building production-grade AI applications and agentic workflows with a focus on LLMops, RAG, and prompt engineering. This hybrid role is perfect for a mid-senior engineer with a strong software engineering background and hands-on LLM experience who thrives in collaborative environments. Apply now to contribute to cutting-edge AI solutions!

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Dallas, TX, United States
Posted
Yesterday

Job Description

Duration: 6+ Months

Location: Dallas, TX / New York, NY / Louisville, KY / Washington, DC

 

The LLM Engineer will design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities supporting an enterprise Agent Factory. This role focuses on production-grade AI applications, agentic workflows, Retrieval-Augmented Generation (RAG), prompt and context engineering, model evaluation, fine-tuning, model serving, private AI hosting, and LLMOps.

The engineer will collaborate with AI architects, engineering leads, platform teams, security, enterprise architecture, product owners, and domain teams to develop secure, reliable, reusable, and enterprise-ready AI capabilities.

 

Required Skillset:

  • 5+ years of software engineering, AI engineering, machine learning engineering, or platform engineering experience.
  • 2+ years of hands-on LLM, Generative AI, or agentic AI development experience.
  • Production experience building LLM applications using RAG, prompts, APIs, and cloud or on-premise platforms.
  • Experience with model evaluation, prompt testing, and AI quality measurement.
  • Experience integrating AI capabilities into enterprise applications, workflows, or developer platforms.

 

Responsibilities:

  • Build enterprise-grade LLM applications and intelligent agent capabilities.
  • Develop reusable prompts, context, retrieval, memory, evaluation, and agent components.
  • Design and implement enterprise RAG architectures and retrieval pipelines.
  • Optimize chunking, embeddings, indexing, ranking, reranking, and retrieval strategies.
  • Evaluate, fine-tune, deploy, and optimize LLMs and SLMs.
  • Apply LoRA, QLoRA, quantization, distillation, and model compression techniques.
  • Build private, hybrid, cloud, and on-premise model hosting solutions.
  • Develop scalable model-serving and inference architectures.
  • Optimize latency, throughput, concurrency, GPU utilization, and cost.
  • Implement model routing, load balancing, caching, and fallback strategies.
  • Build LLMOps, ModelOps, and AgentOps capabilities.
  • Develop observability for prompts, retrieval, responses, latency, cost, and failures.
  • Build automated evaluation and regression-testing frameworks.
  • Implement Responsible AI, security, governance, and data-protection controls.

 

Preferred Skillset:

  • Experience with open-source models such as Llama, Mistral, Mixtral, Phi, Gemma, Qwen, DeepSeek, Granite, or Falcon.
  • Experience with GPU infrastructure such as NVIDIA H100, H200, B200, B300, A100, L40S, GH200, or AMD MI300X.
  • Experience with private AI, hybrid AI, or air-gapped AI environments.
  • Experience with MCP, tool registries, agent runtimes, or enterprise integration patterns.
  • Experience in healthcare, financial services, insurance, or other regulated industries.
  • Experience with Responsible AI, model governance, model risk management, or AI compliance.

Similar jobs