Haystack
← Back to Jobs
Remote
Technology

LLM DevOps/Inference Engineer-Remote

Apetan ConsultingUnited States🇺🇸United StatesPosted 3 Aug 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

LLM DevOps/Inference Engineer

Location:  REMOTE

Duration 12-18mth+

 

Must Have:

·       Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CD.

·       Strong AI & LLM

·       Recent healthcare industry exp (HIPAA, hl7, etc)

·       LinkedIn Page

·       Strong communication

·       LinkedIn Page

 

 

Context:

We are looking for a DevOps/Inference Engineer for one of our clients building a healthcare-focused AI benchmark and evaluation suite. The initial target is for clinical prediction tasks including sepsis onset, days-to-death, and lab value trend forecasting, evaluated across multiple frontier and vertical-specific models. Role Summary You will be responsible for the infrastructure the benchmark harness runs on, the selfhosted model serving stack, and the reliability of the platform. This is a role for someone who can provision a GPU cluster in the morning and tune the inferencing model engine in the afternoon.

 

 

What You will Own

• Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CD.

• Stand up the self-hosted inference track for the long tail of vertical healthcare models. This involves provisioning infrastructure for models serving on GPU compute with sensible batching, quantization where appropriate, autoscaling, and a standard onboarding path so adding new models takes hours, not weeks.

• Build the provider abstraction layer alongside the AI engineers so that APIbased models (OpenAI, Anthropic, Gemini, and the growing list beyond) and self-hosted models present a uniform interface to the harness. Rate limiting, retry and backoff, quota management, request/response logging, and cost attribution per run are your responsibility.

• Make benchmark runs reproducible and cost-optimized with pinned model and container versions, captured configuration, spot and reserved capacity strategy, and idle GPU elimination.

• Build the observability story with throughput, latency, token and GPU-hour cost, failure taxonomy, and per-model dashboards for monitoring.

• Support the surge model that the platform must let a burst of AI engineers land, run experiments, and leave without breaking anything or leaving orphaned resources behind.

• Contribute to Trusted Execution Environment (TEE) architecture. Evaluate AWS Nitro Enclaves and comparable confidential computing approaches for the bring-your-own-data / bring-your-own-model scenario, including attestation- gated key release and the practical constraints of running model inference inside an enclave.

 

Required Skills

• AWS infrastructure at production scale: EKS or ECS, EC2 GPU instance families (G5/G6, P4d/P5) and their capacity realities, VPC design, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service Quotas.

• Infrastructure as code: Terraform. No console-clicked production resources.

• Model serving and inference optimization: Hands-on experience working with LLMs. Practical command of batching strategy, KV cache behavior, quantization tradeoffs, and multi-GPU sharding.

• Container orchestration and GPU scheduling: ECS/EKS with GPU workloads, node autoscaling, and image build pipelines for CUDA-dependent stacks.

• Reliability and cost engineering. SLOs, alerting, and a demonstrated track record of optimizing cloud spend without cutting capability.

 

Desirable Skills

• AWS SageMaker endpoints and Bedrock.

• Hands-on experience with Python to contribute directly to the harness and the provider adapter layer.

• Confidential computing fundamentals: Enclaves, remote attestation, sealed key release, and the security boundaries of TEEs.

• Healthcare compliance posture: HIPAA-eligible service selection, BAA scope, audit logging, and the access-control mechanisms for PHI data.

• Security hardening, including image scanning and network egress control for a closed-loop environment.

 

Nice to Have

• Prior experience hosting medical imaging or multimodal models.

• Nitro Enclaves in production, or comparable TEE work.

• Experience supporting self-service environments, clean tenancy boundaries, and fast credential provisioning.

Skills

AWS
CUDA
HIPAA
LLM
Python
Terraform

Similar jobs