Haystack
← Back to Jobs
Engineering
MS

AIOps Engineering Lead (W2 candidate only)

Metalight Solutions IncChicago, IL🇺🇸United StatesPosted 21 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Chicago, IL, United States
Posted
5 days ago
Stakeholder Management

Job Description

Job Title: AIOps Engineering Lead
Location: Chicago, IL (3days On-Site, 2days Remote)
Position Type: Contract (W2 Candidate Only)

Strong Experience Combining: Google Cloud Platform + Vertex AI + Databricks + Kubernetes + Python + CI/CD + LLM/RAG + Production ML + Observability + Data Engineering + Security.

Position Overview

We are seeking a highly experienced AIOps Engineering Lead to drive the deployment, operational reliability, observability, and continuous delivery of enterprise AI systems across Google Cloud Platform (Google Cloud Platform) and Databricks.

The ideal candidate will combine strong AI/ML engineering, cloud infrastructure, DevOps/MLOps, and production operations expertise. This person will be responsible for operationalizing LLM, RAG, Agentic AI, multimodal, and predictive ML workloads in enterprise production environments.

This is a hands-on technical leadership role requiring strong experience with production AI systems, Google Cloud Platform, Kubernetes, Python, CI/CD, observability, security, data engineering, and incident management.

Key Responsibilities

  • Own the SLO/SLA, production readiness, reliability, and operational posture of enterprise AI services.
  • Lead production incident management including triage, mitigation, root-cause analysis (RCA), remediation, and prevention automation.
  • Deploy and operationalize LLM, RAG, Agentic AI, multimodal, and ML systems using scalable Google Cloud Platform-native and Databricks architectures.
  • Build and standardize automated AI/ML pipelines covering versioning, testing, deployment, monitoring, rollback, and release management.
  • Establish comprehensive observability across application performance, model quality, data quality, drift, safety, latency, errors, throughput, and cost.
  • Optimize AI workloads for GPU/CPU utilization, autoscaling, performance, throughput, and cloud cost efficiency.
  • Implement enterprise security and governance controls covering IAM, secrets management, auditability, lineage, approvals, and Responsible AI.
  • Partner with AI/ML engineers, data engineers, DevOps teams, architects, product teams, and offshore engineering teams.
  • Act as the technical lead for production AI operations and drive cross-functional execution.
  • Develop reusable platform templates, reference architectures, golden paths, deployment patterns, and operational runbooks.
  • Continuously improve automation, reliability, deployment processes, monitoring, and incident prevention.

Required Technical Skills

AI / ML / GenAI

  • Strong hands-on experience operating LLM and RAG applications in production.
  • Experience with Agentic AI / AI agents and enterprise AI workloads.
  • Experience supporting production machine learning systems and pipelines.
  • Understanding of model monitoring, data drift, model quality, AI safety, and AI governance.

Google Cloud Platform

Strong hands-on experience with:

  • Google Cloud Platform (Google Cloud Platform)
  • Vertex AI
  • Cloud Run
  • Pub/Sub
  • Cloud Build
  • BigQuery

Data & ML Platforms

  • Databricks
  • Apache Airflow / Cloud Composer
  • Data engineering and schema governance
  • Batch and streaming data processing

DevOps / Platform Engineering

  • Python
  • Docker
  • Kubernetes
  • CI/CD
  • Production deployment and release automation
  • Infrastructure/platform automation
  • Monitoring and observability

Security & Governance

  • Cloud IAM
  • Secrets management
  • Auditability and compliance
  • Data/model lineage
  • Production approvals and governance
  • Responsible AI controls

Preferred Qualifications

  • Experience with Google ADK or similar agentic AI frameworks.
  • Experience building internal AI/ML platforms or reusable engineering frameworks.
  • Contributions to internal platforms, reusable templates, reference architectures, or open-source tooling.
  • Experience with enterprise-scale AI operations and production reliability.
  • Strong stakeholder management and communication skills.

Similar jobs