Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Charlotte, NC, United States
Posted
17 hours ago
Job Description
Observability and Evaluation Engineer
Charlotte, NC
12 months
Implements telemetry, traces, dashboards, evaluation suites, s, service objectives, runbooks, and readiness evidence for priority agent releases.
LLM and agent evaluation; tracing and telemetry; metrics and dashboards; ing; SLOs; test automation; prompt and model performance analysis; Python; production operations.
We are seeking an Observability and Evaluation Engineer to build the monitoring, evaluation, and reliability infrastructure behind our priority AI agent releases. This role owns the full observability lifecycle — from telemetry and tracing to evaluation suites, alerting, and production readiness — ensuring agent systems are measurable, reliable, and accountable at scale.
Key Responsibilities
-
Design and implement telemetry pipelines and distributed tracing for LLM and agent-based systems
-
Build dashboards and metrics to monitor model/agent performance, latency, cost, and reliability
-
Develop and maintain evaluation suites for LLM and agent quality, accuracy, and regression testing
-
Define and track Service Level Objectives (SLOs) and error budgets for agent services
-
Implement alerting systems to proactively detect degradation, drift, or failures
-
Create and maintain runbooks for incident response and operational troubleshooting
-
Produce production readiness evidence and documentation for agent release approvals
-
Conduct prompt and model performance analysis to identify optimization opportunities
-
Build and maintain test automation frameworks supporting continuous evaluation
-
Collaborate with engineering teams to support production operations of agent-based systems
Required Skills & Experience
-
Strong hands-on experience with LLM and agent evaluation methodologies and frameworks
-
Proficiency in tracing and telemetry tools (e.g., OpenTelemetry, Datadog, Grafana, or similar)
-
Experience building metrics, dashboards, and alerting systems for production systems
-
Solid understanding of SLOs, error budgets, and reliability engineering practices
-
Strong Python programming skills
-
Experience with test automation for ML/LLM systems
-
Background in production operations and incident/runbook management
-
Analytical skills for prompt and model performance evaluation
Similar jobs
- JM
Sr Lead Software Engineer
NewJ.P. Morgan
Columbus, Ohio🇺🇸On-site3 minutes agoMachine LearningAgileJenkins+2Technology - GR
Staff Generalist Engineer
NewGreptile
San Francisco🇺🇸11 hours agoFigmaComplianceJavaScript+2Engineering - GR
Senior Generalist Engineer
NewGreptile
San Francisco🇺🇸11 hours agoFigmaJavaScriptLLM+1Engineering - SI
Software Engineer, Agent Architecture
NewSierra
San Francisco, California🇺🇸On-site31 minutes agoAR/VRGoogle WorkspaceLLM+1Technology - JM
Software Engineer III- Python, Databricks
NewJ.P. Morgan
Jersey City, New Jersey🇺🇸On-site31 minutes agoMongoDBOracleSQL+9Technology - ST
Staff Software Engineer, AI
NewStord
United States🇺🇸Hybrid31 minutes agoDrizzleElixirGCP+12Technology