Haystack
← Back to Jobs
Remote
Engineering
KE

LLM Engineer

KeylentUnited States🇺🇸United StatesPosted Sep 17, 2026

Why This Role Stands Out

This remote LLM Engineer role offers a fantastic opportunity to shape enterprise AI capabilities by deploying and optimizing cutting-edge open-source models on advanced GPU infrastructure. You'll thrive here if you have hands-on experience with model deployment and optimization techniques, and are eager to build reusable LLM patterns that drive innovation. Apply now to join a forward-thinking team and make a significant impact on the future of AI!

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
Yesterday

Job Description

LLM Engineer
Location: REMOTE

Project Duration: 1 year+

Top skills required:
1. Experience deploying open-source models such as Llama, Mistral, Mixtral, Phi, Gemma, Qwen, DeepSeek, Granite, Falcon, or domain-specific models.
2. Experience hosting models on GPU infrastructure such as NVIDIA H100, H200, B200, B300, A100, L40S, GH200, or AMD MI300X.
3. Design reusable LLM patterns, services, APIs, and accelerators for Agent Factory adoption.


Role & Responsibilities:

The LLM Engineer will design, build, optimize, deploy, and operate Large Language Model and Small Language Model capabilities that power the enterprise Agent Factory. This role is responsible for transforming foundation models into secure, reliable, reusable, and enterprise-ready AI capabilities across agentic workflows, AI for SDLC, knowledge retrieval, model evaluation, private AI hosting, and AgentOps.

LLM and SLM Model Engineering
•Evaluate, build, fine-tune, deploy, and optimize LLMs and SLMs for enterprise use cases.
•Support domain-specific model development using internal and approved datasets.
•Build supervised fine-tuning and model adaptation pipelines.
•Apply model optimization techniques such as LoRA, QLoRA, distillation, quantization, and model compression.
•Evaluate commercial, open-source, and internally hosted models for suitability, quality, cost, and operational fit.
•Support model selection strategies based on use case sensitivity, latency, accuracy, cost, and data residency requirements.

Private AI and On-Prem Model Hosting
•Build and support private AI capabilities for hosting SLMs and LLMs in enterprise-controlled environments.
•Deploy models on on-prem, hybrid, and private cloud infrastructure.
•Support GPU-enabled model hosting using enterprise AI infrastructure.
•Optimize model serving for latency, throughput, concurrency, resiliency, and GPU utilization.
•Build secure inference endpoints for internal agent and application consumption.
•Support air-gapped or restricted AI environments where required by security or compliance needs.
•Partner with infrastructure and platform teams to operationalize private model hosting patterns.

Model Serving and Inference Optimization
•Implement scalable model serving using modern inference frameworks.
•Build high-availability inference patterns for production workloads.
•Optimize inference performance, token throughput, response latency, and cost efficiency.
•Implement model routing, load balancing, caching, and fallback strategies.
•Support batch inference and real-time inference use cases.
•Develop reusable deployment templates for multiple model families and serving patterns.

LLMOps, ModelOps, and AgentOps
•Build operational practices for managing models and agents across the lifecycle.
•Implement observability for prompts, retrieval, model responses, latency, cost, and failures.
•Develop evaluation pipelines for regression testing and continuous quality improvement.
•Monitor model drift, response quality, hallucination indicators, and safety risks.
•Support CI/CD and release management for prompts, models, agents, and retrieval pipelines.
•Build dashboards and metrics for AI quality, reliability, adoption, and operational readiness.

AI Evaluation and Benchmarking
•Define and implement LLM evaluation frameworks.
•Measure accuracy, groundedness, relevance, hallucination rate, toxicity risk, safety compliance, task completion, and user satisfaction.
•Build automated test suites for prompts, agents, tools, and RAG pipelines.
•Benchmark models across enterprise use cases.
•Compare cloud-hosted, open-source, and on-prem models based on performance, cost, quality, and risk.
•Support go/no-go quality gates for production AI releases.


7+ years of software engineering
2+ years of hands-on experience building LLM, GenAI, or agentic AI solutions.

Similar jobs