Senior Lead AI Engineer (Foundation Model Hosting & LLM Inference)/Remote
Why This Role Stands Out
This remote Senior Lead AI Engineer role offers a unique opportunity to shape the future of foundation model hosting and LLM inference, driving innovation in a rapidly evolving field. You will thrive here if you possess strong expertise in AI infrastructure and distributed systems, eager to lead impactful projects and continuously develop your skills. Apply now to join a forward-thinking team and make your mark in cutting-edge AI technology.
Quick Overview
Job Description
Job Description: Senior Lead AI Engineer (Foundation Model Hosting & LLM Inference)
Location-Remote
Position Overview
We are seeking a Senior Lead AI Engineer to lead the design, deployment, and optimization of infrastructure for hosting foundation models and serving large language model (LLM) inference workloads. The ideal candidate has strong expertise in AI infrastructure, distributed systems, cloud platforms, and production-scale model serving.
Key Responsibilities
- Lead the design and implementation of scalable infrastructure for hosting foundation models and LLM inference services.
- Build and optimize high-performance inference pipelines for latency, throughput, reliability, and cost efficiency.
- Deploy, monitor, and manage AI workloads across cloud and on-premises environments.
- Collaborate with AI researchers and ML engineers to productionize machine learning models.
- Design APIs and backend services for model serving and inference.
- Optimize GPU utilization, model parallelism, batching, and resource scheduling.
- Implement monitoring, logging, security, and observability for AI platforms.
- Drive architecture decisions, engineering best practices, and technical standards.
- Mentor engineers and lead technical initiatives across cross-functional teams.
- Stay current with advancements in AI infrastructure, model serving, and inference optimization.
Qualifications
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Software Engineering, or a related field.
- Strong programming skills in Python and/or C++, with experience in backend systems.
- Experience deploying and managing LLMs or other foundation models in production.
- Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Strong understanding of distributed systems, microservices, and container orchestration using Docker and Kubernetes.
- Experience with GPU computing and model serving frameworks.
- Familiarity with REST APIs, networking, and scalable backend architectures.
Preferred Skills
- Experience with inference frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, or similar technologies.
- Knowledge of model optimization techniques, including quantization, batching, and caching.
- Experience with distributed inference, autoscaling, and GPU scheduling.
- Familiarity with MLOps tools, CI/CD pipelines, and Infrastructure as Code.
- Experience with monitoring and observability tools for production AI systems.
Skills
Similar jobs
AI Code Review Engineer
Gridiron IT Solutions · United States
1 minute ago$85k - $120k/yrAI Engineer
ITBMS Inc. · Plano, United States
52 minutes agoCloud/AI Developer (Remote)
DivIHN Integration Inc. · United States
53 minutes agoAI Engineer
Lorven Technologies, Inc. · Columbus, United States
53 minutes agoAI Architect
Libsys, Inc. · Houston, United States
57 minutes agoGenerative AI Engineer
Aventine software · San Francisco, United States
57 minutes ago