Haystack
← Back to Jobs
Technology
RI

AI DevOps/Observability Engineer

Rivago infotech incCharlotte, NC🇺🇸United StatesPosted Oct 2, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Charlotte, NC, United States
Posted
21 hours ago
DockerAWSMLOpsNew RelicAzureDatadogGenerative AIGoogle CloudKubernetesLLMPhoenixPython

Job Description

Role : AI DevOps/Observability Engineer

Location : Charlotte, NC (Onsite)

Persistent Systems

We are seeking a highly skilled AI DevOps/Observability Engineer to join our production operations team. In this role, you will be responsible for the reliability, performance, and operational readiness of our priority Generative AI and LLM agent releases. You will bridge the gap between AI development and production operations by implementing robust telemetry, automated evaluation pipelines, and comprehensive monitoring systems. Your work will ensure our intelligent agents are stable, accurate, efficient, and scale seamlessly in production environments.

 

Key Responsibilities

  • Observability & Telemetry: Implement and maintain deep tracing, logging, and telemetry solutions specifically tailored for LLM application architectures and multi-agent workflows.
  • Monitoring & Insights: Design, build, and maintain production dashboards that track systemic health, infrastructure metrics, and specialized AI performance indicators.
  • Operational Readiness: Establish actionable alerting systems, define Service Level Objectives (SLOs), curate runtime runbooks, and provide concrete engineering evidence for production readiness.
  • Evaluation & Testing Pipelines: Develop automated continuous evaluation suites to assess LLM agent behavior, safety, and output quality prior to and during deployment.
  • Performance Analysis: Analyze prompt efficiency, token usage, latency, and overall model performance to optimize cost, speed, and accuracy.
  • Production Operations: Support the deployment pipeline, participate in incident management, and continually improve the resilience of our live AI services.

 

Required Skills and Qualifications

Core Technical Skills

  • LLM & Agent Evaluation: Experience with framework-based evaluation tools (e.g., Ragas, DeepEval, TruLens) to measure hallucination, faithfulness, and relevancy.
  • Tracing & Telemetry: Proficiency with LLM-specific tracing tools (e.g., LangSmith, LangFuse, Phoenix, Arize) and open standards like OpenTelemetry.
  • Metrics & Dashboards: Hands-on experience building monitoring views in platforms like Datadog, PrometheGrafana, New Relic, or cloud-native suites.
  • Production Alerting & SLOs: Proven ability to define meaningful Service Level Indicators (SLIs) and SLOs, minimizing alert fatigue while maximizing system reliability.
  • Test Automation: Strong background in integrating automated test frameworks into CI/CD pipelines for continuous integration of AI features.
  • Prompt & Model Analysis: Analytical mindset to benchmark prompt variants, track regression in model behavior, and profile latency across model providers.

Programming & Operations

  • Python Mastery: Advanced Python programming skills, including experience with async execution, API integration, and AI frameworks (e.g., LangChain, LlamaIndex).
  • Production Operations: Solid understanding of cloud infrastructure (AWS/Google Cloud Platform/Azure), containerization (Docker, Kubernetes), and DevOps best practices.

 

Preferred Qualifications

  • 3+ years of experience operationalizing LLMs or generative AI applications in production.
  • Experience managing vector databases (e.g., Pinecone, Milvus, Chroma) and tracking RAG pipeline performance.
  • Background in Site Reliability Engineering (SRE) or specialized MLOps roles.

Similar jobs