Why This Role Stands Out
You'll have the opportunity to lead the deployment and optimization of cutting-edge foundation models, directly impacting cloud applications and gaining invaluable experience in LLMOps and Azure infrastructure. This remote role is perfect for a mid-senior AI/ML Engineer passionate about bridging machine learning and software engineering, offering significant career growth within a reputable technology company.
Quick Overview
Job Description
Job Title: Senior AI/ML Engineer (LLMOps & Model Deployment)
Long Term
Remote
Job Overview
We are seeking a highly skilled AI/ML Engineer to lead the deployment, optimization, and scaling of open-weights foundation models within our Azure cloud ecosystem. In this role, you will bridge the gap between machine learning and core software engineering, turning raw models into highly available, low-latency APIs that power our cloud applications.
The ideal candidate has a deep understanding of LLMOps, hands-on experience optimizing model inference (including knowledge distillation), and a proven track record of architecting production-grade infrastructure on Azure.
Key Responsibilities
Model Deployment & API Development
Host & Scale Open-Weights Models: Deploy and maintain state-of-the-art open-weights models (e.g., Llama, Mistral, Phi) on Azure infrastructure.
API Engineering: Design, build, and secure high-performance, low-latency REST/gRPC APIs (using frameworks like FastAPI) to serve models to downstream cloud applications.
Inference Optimization: Implement advanced optimization techniques (e.g., vLLM, TensorRT-LLM, DeepSpeed) to maximize throughput and minimize time-to-first-token (TTFT).
LLMOps & Infrastructure
Pipeline Automation: Build and manage end-to-end LLMOps pipelines for continuous integration, deployment, and monitoring of models.
Azure Cloud Architecture: Leverage Azure AI Studio, Azure Machine Learning (Azure ML), Azure Kubernetes Service (AKS), and managed compute (GPUs like A100/H100) efficiently.
Monitoring & Observability: Implement comprehensive logging, tracing, and evaluations for LLM outputs (tracking drift, latency, costs, and hallucination rates).
Model Efficiency & Distillation
Knowledge Distillation: Train smaller, task-specific student models from larger, high-performing teacher models to reduce operational costs and latency without sacrificing accuracy.
Quantization & Fine-Tuning: Apply quantization techniques (AWQ, GPTQ, GGUF) and parameter-efficient fine-tuning (PEFT/LoRA) to adapt models to specific business domains.
Required Skills & Qualifications
Technical Requirements
Experience: 4+ years of professional experience as an ML Engineer, Data Scientist, or Backend Engineer with a heavy focus on AI deployment.
Cloud Proficiency: Strong hands-on experience with Microsoft Azure (Azure ML, AKS, Azure Container Apps, Key Vault).
AI/ML Frameworks: Deep proficiency with PyTorch, Hugging Face ecosystem (Transformers, Accelerate), and LangChain or LlamaIndex.
API & Backend: Strong Python programming skills and experience with containerization (Docker, Kubernetes) and API development (FastAPI).
Model Optimization: Demonstrated experience with model distillation, pruning, quantization, and utilizing inference engines like vLLM or TGI.
Soft Skills & Culture Fit
Problem Solver: Ability to triage infrastructure bottlenecks, memory constraints (OOM errors), and CUDA-related issues independently.
Collaborator: Comfortable working cross-functionally with backend engineers, product managers, and security teams.
Cost-Conscious Mindset: A sharp focus on balancing model accuracy with cloud spend and compute efficiency.
Preferred Qualifications
Certifications such as Azure AI Engineer Associate or Azure Solutions Architect.
Experience implementing secure guardrails (e.g., NeMo Guardrails, Azure AI Content Safety).
Contributions to open-source ML/LLMOps projects.
Similar jobs
- LT
AI/ML Engineer
Ledgent Technology
Davis, CA🇺🇸Hybrid2 months agoSQLMachine LearningAzure+3Technology - AT
Staff Machine Learning Operations Engineer - Computer Vision
NewATI
Woburn, Massachusetts🇺🇸On-site12 hours agoDockerGCPSQL+13Technology - ST
AI/ML Engineer Wilmington, DE-Onsite
NewStoneGate-Technologies LLC
Wilmington, DE🇺🇸HybridYesterdayDockerSQLAWS+15Technology - AH
AI/ML Engineer
NewAivra Health LLC
San Francisco, CA🇺🇸HybridYesterdayDockerSQLAWS+17Technology - CO
Lead Machine Learning Engineer (Python, AWS, SQL, GenAI) (Enterprise Platforms Technology)
NewCapital One
Mc Lean, Virginia🇺🇸$197.3k - $225.1k/yrHybrid3 hours agoSQLScalaAWS+9Technology - CO
Senior Lead Machine Learning Engineer
NewCapital One
Richmond, Virginia🇺🇸$229.9k - $262.4k/yrHybrid3 hours agoScalaAWSMachine Learning+8Technology