Haystack
← Back to Jobs
Remote
Technology

Senior Lead AI Engineer (Foundation Model Hosting & LLM Inference)/Remote

Apetan ConsultingUnited States🇺🇸United StatesPosted 6 Aug 2026

Why This Role Stands Out

This remote Senior Lead AI Engineer role offers a unique opportunity to shape the future of foundation model hosting and LLM inference, driving innovation in a rapidly evolving field. You will thrive here if you possess strong expertise in AI infrastructure and distributed systems, eager to lead impactful projects and continuously develop your skills. Apply now to join a forward-thinking team and make your mark in cutting-edge AI technology.

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Job Description: Senior Lead AI Engineer (Foundation Model Hosting & LLM Inference)

Location-Remote

Position Overview

We are seeking a Senior Lead AI Engineer to lead the design, deployment, and optimization of infrastructure for hosting foundation models and serving large language model (LLM) inference workloads. The ideal candidate has strong expertise in AI infrastructure, distributed systems, cloud platforms, and production-scale model serving.

Key Responsibilities

  • Lead the design and implementation of scalable infrastructure for hosting foundation models and LLM inference services.
  • Build and optimize high-performance inference pipelines for latency, throughput, reliability, and cost efficiency.
  • Deploy, monitor, and manage AI workloads across cloud and on-premises environments.
  • Collaborate with AI researchers and ML engineers to productionize machine learning models.
  • Design APIs and backend services for model serving and inference.
  • Optimize GPU utilization, model parallelism, batching, and resource scheduling.
  • Implement monitoring, logging, security, and observability for AI platforms.
  • Drive architecture decisions, engineering best practices, and technical standards.
  • Mentor engineers and lead technical initiatives across cross-functional teams.
  • Stay current with advancements in AI infrastructure, model serving, and inference optimization.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Software Engineering, or a related field.
  • Strong programming skills in Python and/or C++, with experience in backend systems.
  • Experience deploying and managing LLMs or other foundation models in production.
  • Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Strong understanding of distributed systems, microservices, and container orchestration using Docker and Kubernetes.
  • Experience with GPU computing and model serving frameworks.
  • Familiarity with REST APIs, networking, and scalable backend architectures.

Preferred Skills

  • Experience with inference frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, or similar technologies.
  • Knowledge of model optimization techniques, including quantization, batching, and caching.
  • Experience with distributed inference, autoscaling, and GPU scheduling.
  • Familiarity with MLOps tools, CI/CD pipelines, and Infrastructure as Code.
  • Experience with monitoring and observability tools for production AI systems.

Skills

Docker
Microservices
AWS
MLOps
Machine Learning
Azure
C++
Google Cloud
Kubernetes
LLM
Python
REST

Similar jobs