Haystack
← Back to Jobs
Engineering
AC

Real Time Inference Engineering Consultant

Alltech Consulting Services, Inc.Concord, CA🇺🇸United StatesPosted 26 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Concord, CA, United States
Posted
Yesterday

Job Description

Job Title: Real Time Inference Engineering Consultant
Location: : Concord CA (ONSITE 3 DAYS A WEEK)
CONTRACT

JOB DESCRIPTION :

Real-Time Inference Engineering Consultant

The Real-Time Inference Engineering Consultant is responsible for designing, building, and operationalizing scalable, low-latency model serving platforms that power real-time AI and predictive decisioning use cases. This role leads the architecture, deployment, performance optimization, resiliency, and operational governance of online inference services across cloud and on-premises environments.

* Design and implement highly available, low-latency model serving architectures for real-time inference workloads.
* Develop and standardize deployment patterns for scalable AI/ML services across Kubernetes-based platforms.
* Lead API-based inference service design, integration, and lifecycle management.
* Optimize model serving performance through latency tuning, caching strategies, autoscaling, and resource management.
* Establish monitoring, observability, SLOs/SLAs, alerting, and operational runbooks for production services.
* Drive load testing, capacity planning, resiliency engineering, and disaster recovery readiness.
* Integrate inference platforms with CI/CD pipelines to enable automated deployments and controlled releases.
* Partner with Data Science, Platform Engineering, MLOps, and Infrastructure teams to ensure reliable production model operations.
* Govern operational best practices, security, reliability, and performance standards for enterprise AI deployments.

* Online inference and model-serving architectures
* Real-time APIs and distributed systems design
* Kubernetes, container orchestration, and service mesh technologies
* Autoscaling, capacity management, and workload optimization
* Performance engineering, load testing, and latency optimization
* Monitoring, observability, logging, tracing, and SLO management
* CI/CD, DevOps, and Infrastructure-as-Code practices
* Reliability engineering, fault tolerance, and resiliency patterns
* Cloud and on-premises platform operations
* Python, Java, Go, or similar backend development experience

* Experience with MLOps platforms and enterprise AI deployment frameworks.
* Hands-on experience with real-time recommendation, fraud, risk, personalization, or predictive analytics platforms.
* Familiarity with GPU-based inference, model optimization, and multi-cloud deployments.

A successful candidate combines AI platform engineering, distributed systems expertise, and operational excellence to deliver resilient, high-performance inference platforms that enable enterprise-scale real-time AI solutions.

Similar jobs