Haystack
← Back to Jobs
Other
SI

Inference Lead

Stellent IT LLCCharlotte, NC🇺🇸United StatesPosted 4 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Charlotte, NC, United States
Posted
5 days ago
MicroservicesLoad BalancingMachine LearningCapacity PlanningKubernetesRESTgRPC

Job Description

Job Title:- Inference Lead (Machine Learning Platform Engineer Lead Real-Time Inference)

Location:- Charlotte, North Carolina (Hybrid Onsite - local candidates preferred. )

Duration:- 6 months

All visa except OPT/CPT

Contract

Real-Time Services Real-Time Inference Engineering Lead

Role Summary

The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments.

Key Responsibilities

  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and cost.
  • Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines for repeatable model and service releases.
  • Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and mentor inference and platform engineers.

Required Skills

  • Online inference and real-time model-serving architecture.
  • REST/gRPC APIs and distributed microservices.
  • Kubernetes, containers, autoscaling, and traffic management.
  • Performance engineering, latency optimization, and load testing.
  • Monitoring, SLOs, capacity planning, and production operations.
  • CI/CD and progressive-deployment approaches.
  • Resilience and high-availability engineering.
  • Cloud and on-premises deployment experience.

Preferred Qualifications

  • Degree in computer science, engineering, or a related discipline.
  • Experience with enterprise model-serving platforms and inference runtimes.
  • Cloud, Kubernetes, SRE, or ML engineering certification.

Skill

Years of experience

Last used/Worked (year)

Candidate self-rating(out of 10)

Navya Gupta
Sr. IT Technical Recruiter

Email:

Gtalk:
Phone: +1

Linkedin id:
Address: 505 Knolle Court, Saint Augustine| FL 32092

Similar jobs