Haystack
← Back to Jobs
Engineering
CI

LLMOps / MLOps Engineer (Need Locals only)

Cardinal Integrated Technologies IncSanta Clara, CA🇺🇸United StatesPosted 29 Jul 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Hi,

Senior LLMOps / MLOps Engineer

Location: Santa Clara, CA (Onsite)

Duration: 6 - 12 Months

Must Have Skills

Skill 1 Strong proficiency in Python and software engineering best practices

Skill 2 14+ years of experience in MLOps, LLMOps, AI/ML Platform Engineering

Skill 3 Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving

Good To have Skills

Skill 1 Exposure to AI Observability, Governance, and Responsible AI practices

Mandatory if Applicable

Domain Experience (If any) Senior LLMOps / MLOps Engineer

Summary

We are looking for a highly skilled Senior LLMOps / MLOps Engineer with strong expertise in LLM inferencing, model hosting, and serving Large Language Models (LLMs) at scale. The ideal candidate should be a hands-on engineer with proven experience deploying and optimizing open-source LLMs, building high-performance inference platforms using technologies such as vLLM, SGLang, TGI, Triton, and Ray Serve, and driving GPU utilization, latency, throughput, and cost optimization. This is a highly technical role requiring active involvement in designing, building, troubleshooting, and optimizing production AI systems. Experience in MLOps platforms and scalable AI infrastructure is essential.

Must-Have Skills

  • 5-7 years of experience in MLOps, LLMOps, AI/ML Platform Engineering.
  • Strong proficiency in Python and software engineering best practices.
  • Experience working with open-source LLMs such as Llama, Mistral, Gemma, or Qwen.
  • Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving.
  • Experience with Kubernetes, Docker, Azure ML, Databricks, and MLflow.
  • Good understanding of RAG, Vector Databases, GPU Optimization, Quantization, KV Cache, PagedAttention, and ContinuoDynamic Batching.
  • Demonstrated hands-on experience building, deploying, troubleshooting, and optimizing production-grade LLM and GenAI solutions.
  • Experience deploying, scaling, and monitoring production-grade GenAI/LLM applications.
  • Exposure to AI Observability, Governance, and Responsible AI practices.

Good-to-Have Skills

  • Hands-on experience with LLM Fine-Tuning using PEFT, SFT, CPT, LoRA, and QLoRA techniques.
  • Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT.
  • Knowledge of distributed training and multi-GPU environments.
  • Experience with Agentic AI frameworks such as LangGraph, AutoGen, or CrewAI.
  • Understanding of simulation platforms, digital twins, modeling & simulation workflows, or scientific computing.

Similar jobs