Why This Role Stands Out
You'll have the opportunity to lead the deployment and optimization of cutting-edge foundation models, directly impacting cloud applications and gaining invaluable experience in LLMOps and Azure infrastructure. This remote role is perfect for a mid-senior AI/ML Engineer passionate about bridging machine learning and software engineering, offering significant career growth within a reputable technology company.
Quick Overview
Job Description
Job Title: Senior AI/ML Engineer (LLMOps & Model Deployment)
Long Term
Remote
Job Overview
We are seeking a highly skilled AI/ML Engineer to lead the deployment, optimization, and scaling of open-weights foundation models within our Azure cloud ecosystem. In this role, you will bridge the gap between machine learning and core software engineering, turning raw models into highly available, low-latency APIs that power our cloud applications.
The ideal candidate has a deep understanding of LLMOps, hands-on experience optimizing model inference (including knowledge distillation), and a proven track record of architecting production-grade infrastructure on Azure.
Key Responsibilities
Model Deployment & API Development
Host & Scale Open-Weights Models: Deploy and maintain state-of-the-art open-weights models (e.g., Llama, Mistral, Phi) on Azure infrastructure.
API Engineering: Design, build, and secure high-performance, low-latency REST/gRPC APIs (using frameworks like FastAPI) to serve models to downstream cloud applications.
Inference Optimization: Implement advanced optimization techniques (e.g., vLLM, TensorRT-LLM, DeepSpeed) to maximize throughput and minimize time-to-first-token (TTFT).
LLMOps & Infrastructure
Pipeline Automation: Build and manage end-to-end LLMOps pipelines for continuous integration, deployment, and monitoring of models.
Azure Cloud Architecture: Leverage Azure AI Studio, Azure Machine Learning (Azure ML), Azure Kubernetes Service (AKS), and managed compute (GPUs like A100/H100) efficiently.
Monitoring & Observability: Implement comprehensive logging, tracing, and evaluations for LLM outputs (tracking drift, latency, costs, and hallucination rates).
Model Efficiency & Distillation
Knowledge Distillation: Train smaller, task-specific student models from larger, high-performing teacher models to reduce operational costs and latency without sacrificing accuracy.
Quantization & Fine-Tuning: Apply quantization techniques (AWQ, GPTQ, GGUF) and parameter-efficient fine-tuning (PEFT/LoRA) to adapt models to specific business domains.
Required Skills & Qualifications
Technical Requirements
Experience: 4+ years of professional experience as an ML Engineer, Data Scientist, or Backend Engineer with a heavy focus on AI deployment.
Cloud Proficiency: Strong hands-on experience with Microsoft Azure (Azure ML, AKS, Azure Container Apps, Key Vault).
AI/ML Frameworks: Deep proficiency with PyTorch, Hugging Face ecosystem (Transformers, Accelerate), and LangChain or LlamaIndex.
API & Backend: Strong Python programming skills and experience with containerization (Docker, Kubernetes) and API development (FastAPI).
Model Optimization: Demonstrated experience with model distillation, pruning, quantization, and utilizing inference engines like vLLM or TGI.
Soft Skills & Culture Fit
Problem Solver: Ability to triage infrastructure bottlenecks, memory constraints (OOM errors), and CUDA-related issues independently.
Collaborator: Comfortable working cross-functionally with backend engineers, product managers, and security teams.
Cost-Conscious Mindset: A sharp focus on balancing model accuracy with cloud spend and compute efficiency.
Preferred Qualifications
Certifications such as Azure AI Engineer Associate or Azure Solutions Architect.
Experience implementing secure guardrails (e.g., NeMo Guardrails, Azure AI Content Safety).
Contributions to open-source ML/LLMOps projects.
Similar jobs
- TH
Staff ML Engineer
NewThe Hіrеxрrеѕѕ
San Jose, CA🇺🇸Hybrid21 hours agoMLOpsDeep LearningLLM+1Technology - MW
ML Ops Engineer
NewMind Ware Inc
Atlanta, GA🇺🇸Hybrid21 hours agoShellAWSMachine Learning+3 - AC
AI/ML Engineer (Palantir Foundry) 100% Remote
NewApetan Consulting
United States🇺🇸Remote21 hours agoSQLSnowflakeTableau+10Technology - DE
AI/ ML Engineer
NewDevfi
Washington, DC🇺🇸On-site21 hours agoDockerSQLAWS+12Technology - TF
Data Scientist or Machine Learning Engineer (U.S. Citizen, TS/SC with Security Clearance
NewTask Force Talent
Chantilly, VA🇺🇸Hybrid21 hours agoSQLMachine LearningNLP+4Technology - TF
Data Scientist or Machine Learning Engineer (U.S. Citizen, TS/SC with Security Clearance
NewTask Force Talent
Arlington, VA🇺🇸$150k - $200k/yrOn-site21 hours agoSQLMachine LearningNLP+4Technology - RE
Principal Machine Learning Engineer
NewResourcesoft, Inc.
Raleigh, NC🇺🇸Hybrid21 hours agoMicroservicesAWSMLOps+6Technology - SI
AI/ML engineer with Security Clearance
NewSIXGEN
Columbia, MD🇺🇸Remote21 hours agoAgileLLMPythonTechnology - GS
Machine Learning Engineer
NewGlobal Soft Systems
Parsippany-Troy Hills, NJ🇺🇸On-site21 hours agoAWSMachine LearningDatabricks+4Technology - PI
Staff Machine Learning Engineer, Content Visual AI
NewAuto ApplyPinterest
San Francisco🇺🇸Remote5 hours agoMachine LearningHiveLLMTechnology - HE
Inference Optimization Engineer
NewAuto ApplyHedra
San Francisco🇺🇸Hybrid7 hours agoCUDADeep LearningC+++2Engineering - MM
Head of AI
NewMitchell Martin, Inc.
New York, NY🇺🇸$225k/yrHybrid21 hours agoOnboarding