Haystack
← Back to Jobs
Technology
II

Senior AI/ML Ops Engineer Agent Evaluation, Observability & Production Reliability

iPeople Infosystems LLCAustin, TX🇺🇸United StatesPosted 19 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Job Title: Senior AI/ML Ops Engineer Agent Evaluation, Observability & Production Reliability
Location: Cupertino, CA or Austin, TX (onsite)
Type: Contract Position

Job Description

Must Have

  • 5+ years of experience in ML engineering, MLOps, platform engineering, or SRE, including 2+ years working hands-on with LLMs or LLM-powered applications in production.
  • Demonstrated experience building evaluation systems for ML or LLM applications: test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.
  • Strong software engineering skills in Python (and ideally TypeScript), with a track record of building reliable, well-tested internal platforms and tooling.
  • Deep familiarity with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite) and experience embedding automated quality gates into deployment pipelines.
  • Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus) and, ideally, LLM-specific observability tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave).
  • Proven ability to debug complex distributed systems under pressure, including production incident response, root-cause analysis, and blameless postmortems.
  • Excellent cross-functional communication: able to translate evaluation results into clear findings and recommendations for both engineers and non-technical stakeholders.
  • Comfort with ambiguity and a builder s mindset: this role starts with a blank page and ends with the evaluation platform the whole organization relies on.
  • Experience with agentic frameworks and orchestration patterns (e.g., multi-agent systems, tool use, RAG pipelines) and their distinct failure modes.
  • Experience with prompt management, model routing, or fine-tuning workflows and evaluating changes across model versions and providers.
  • Background in statistics or experimentation (A/B testing, significance testing, sampling strategies for human review).
  • Design and build reusable AI agent skills, plugins and maintain internal marketplace infrastructure to extend and scale Data, AIML capabilities across the organization.
  • Expertise in causal inference and measurement strategy including causal graphs, ontologies, and knowledge graphs to drive rigorous, decision grade data analysis.
  • Experience operating in regulated or high-stakes domains where agent errors carry real business or customer impact.

Skills

MLOps
Datadog
GitHub Actions
GitLab CI
Grafana
Jenkins
LLM
Phoenix
Prometheus
Python
TypeScript

Similar jobs