Haystack
← Back to Jobs
Full time
Technology

Staff Machine Learning Engineer

Zendesk Pty Ltd (Australia)Melbourne, Victoria🇦🇺AustraliaPosted 24 Jul 2026

Why This Role Stands Out

This role offers a fantastic opportunity to shape the future of AI-powered customer service, working on cutting-edge agent architectures with direct customer impact. You'll thrive here if you're a skilled Machine Learning Engineer eager to build reliable, scalable AI systems, develop expertise in agent execution and knowledge retrieval, and contribute to a dynamic team in a hybrid work environment. Apply to join Zendesk and be at the forefront of innovation!

Quick Overview

Work Type
Hybrid
Schedule
Full Time
Level
Mid Senior

Job Description

Job Description

The Custom Agent Service manages the complete agent lifecycle from configuration to execution. Admins set up agents with instructions, knowledge articles, and actions through a UI. A customer ticket triggers execution; the agent plans, acts via real APIs, and resolves the issue. The backend is Python, running on Kubernetes on AWS, and the agent architectures range from single pass ReAct loops to our iterative multi plan executor. We are onboarding early customers so the work ships to real accounts with real tickets.

What We Need Help With Agent Execution Core

The planning loop, tool dispatch, memory integration, and error recovery that form the main execution path are core to the service. You will work on the code that decides what the agent does next, improving speed, reliability, and capability. This includes integrating new architectures into production and hardening them for real traffic.

Knowledge Retrieval

Agents retrieve and reason over customer knowledge bases at runtime. The pipeline (embedding, reranking, context assembly) must balance answer quality against latency and token cost across thousands of heterogeneous knowledge bases per deployment.

Actions and Connectors

Agents call Zendesk APIs, third party connectors (Shopify, Salesforce, etc.), custom actions from admins, and other agents via A2A. You will ensure reliable retries, timeouts, schema validation, and graceful degradation when connectors fail mid execution, and you will build self service connector integrations through the Connector SDK.

Production Instrumentation for Model Training

Every agent execution generates a trajectory (reasoning steps, tool calls, outcomes, user feedback). We are building domain specialized models, which requires clean and actionable production data. You will instrument the pipeline to capture implicit reward signals such as resolution success, escalation patterns, and user satisfaction to feed the ML training pipeline.

Security and Compliance

PII filtering, audit logging, action versioning, and governance patterns keep agents within admin configured bounds. You will work directly with Product Security on security review items as the platform scales.

What We Are Looking For

5+ years of backend engineering with strong Python skills. You have shipped production systems, not just models, and understand the difference between a locally working agent and one that runs across 100,000 accounts. You are comfortable across the full agent stack: LLM APIs, prompt engineering, tool calling, memory management, and evaluation.

Requirements:

  • Experience building agent loops from scratch and deciding when a framework is helpful versus obtrusive.
  • Design for failure: model timeouts, garbage output, connector failures, and customer latency.
  • Deliver production ready code, review PRs thoroughly, and communicate status, blockers, and risks clearly.
Tech Stack
  • Languages: Python (primary), Go (for platform services)
  • Agent Frameworks: Custom iterative architectures, ReAct, integrations with open source tooling
  • Infrastructure: Kubernetes, Spinnaker, AWS (ECS, S3, ElastiCache)
  • Data: Postgres, ElastiCache/Redis (vector + KV), Kafka
  • Evaluation: Braintrust (experiment tracking, scoring, CI/CD)
  • Protocols: MCP, REST, gRPC
Benefits

Our hybrid work model allows in office collaboration in Zendesk offices worldwide while providing remote flexibility for part of the week. We value work life balance, continuous learning, and a culture of inclusivity.

Equal Opportunity & Inclusion

Zendesk is an equal opportunity employer, and we are proud of our ongoing efforts to foster global diversity, equity, and inclusion in the workplace. Individuals seeking employment and employees at Zendesk are considered without regard to race, color, religion, national origin, age, sex, gender, gender identity, gender expression, sexual orientation, marital status, medical condition, ancestry, disability, military or veteran status, or any other characteristic protected by applicable law. Zendesk is an AA/EEO/Veterans/Disabled employer. If you are based in the United States and would like more information about your EEO rights under the law, please contact the HR team. Zendesk endeavors to make reasonable accommodations for applicants with disabilities and disabled veterans pursuant to applicable federal and state law. If you require a reasonable accommodation to submit this application, complete any pre employment testing, or otherwise participate in the employee selection process, please send an email to with your specific accommodation request.

Skills

Spinnaker
AWS
Assembly
Kafka
Kubernetes
LLM
Python
REST
React
Redis
gRPC

Similar jobs