Haystack
← Back to Jobs
Technology

AI/ML Engineer

Esvee Technologies IncSan Jose, CA🇺🇸United StatesPosted 4 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Key Responsibilities

  1. Foundations & Local Sandbox Development
  • Build robust golden datasets extracted from UAT logs and utilize frontier models to synthetically generate variations (typos, phrasing, syntax) for robust testing.
  • Construct local developer sandbox environments using the Google ADK framework with version-controlled prompts in Cider.
  • Generate test scripts and code snippets mapped to defined success metrics for continuous local unit testing.

  1. Production Pipeline Automation & System Architecture
  • Architect and deploy language-agnostic RPC endpoints to systematically invoke GTM agents within the ecosystem.
  • Build resilient production pipelines supporting parallel inference execution across 1,000+ trajectory datasets in under 15 minutes.
  • Stand up a centralized Model Context Protocol (MCP) logging server to capture raw prompts, tool trajectories, SQL queries, and token costs using strict JSON schemas.
  • Implement asynchronous message queues (Pub/Sub) for rate-limiting/backpressure, along with retry policies for network and generation failures.
  • Establish CI/CD Pull Request (PR) gatesthat block commits causing capability regressions, and enable pre-production shadow deployments using production mirror traffic.

  1. Skill Benchmarking & Trajectory Validation
  • Author eval test suites to isolate specific agent competencies (e.g., CRM writes, SQL analytics).
  • Inject sandboxed mocks to validate tool-calling logic without producing live CRM side effects or executing heavy database reads.
  • Validate multi-turn trajectories, checking chronological tool order, API loop prevention, and exact parameter payload assertions (e.g., date ranges, seller regions).
  • Stream execution logs for baseline delta analysis and capability scoring.

  1. Enterprise Analytics, Governance & Security (Phase 4)
  • Build low-latency Hydra ETL pipelines to stream structured evaluation JSON records into data warehouses and construct Plx analytics dashboards.
  • Enforce automated PII masking/redaction layers for sensitive seller and financial data, while configuring retention and purging policies (e.g., 90-day trajectory logs).

Required Qualifications & Technical Skills

  • Core Language:Advanced proficiency in Python.
  • Software & Systems Architecture:Deep expertise in enterprise software architectures, distributed computing, async task processing, and load balancing.
  • AI/LLM Telemetry & Evaluations:Proven experience in LLM performance telemetry, prompt engineering, LLM-as-a-Judge systems, and statistical inter-rater agreement (IRR/Kappa).

Skills

SQL
ETL
Load Balancing
LLM
Python

Similar jobs