Haystack
← Back to Jobs
Other

AI Ops Engineer II

SysazzleDenver, CO🇺🇸United StatesPosted 17 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Title: AI Ops Engineer II

Duration: 4-month contract

Location: Denver, CO, 80202(Remote)

 

Description:

While modern generative AI and LLM orchestration tools have only been mainstream for a short time, the 3 to 5 years of required experience for this Level II role reflects the total engineering maturity needed for enterprise deployment. Candidates are expected to have 1 to 2 years of direct, hands-on experience with modern LLMs, multi-agent frameworks, and AI coding assistants, built on top of a solid 2 to 3-year foundational background in backend software engineering (Python), API integration, data engineering, or traditional AIOps. This combined 3-to-5-year tenure ensures the candidate not only understands how to orchestrate AI workflows, but also possesses the established systems thinking required to securely and reliably integrate these tools into complex, enterprise-grade infrastructure platforms like ServiceNow, Dynatrace, and Google Cloud Platform.The AI Automation Engineer will be the foundational builder for client’s Enterprise Operational AI Platform, bridging the gap between advanced multi-agent AI architecture and our core IT operations. Their initial focus will be designing robust data integrations that connect our LLM framework to essential infrastructure systems like ServiceNow, Dynatrace, Zabbix, and Google Cloud Platform. Using tools like the Model Context Protocol (MCP) and AI-assisted development (Devin/Windsurf), they will execute a phased rollout strategy—starting with read-only workflows for incident triage and root cause analysis (RCA), and progressively advancing toward fully autonomous Help Desk automation and proactive self-healing capabilities. Ultimately, their success will be measured by improved AI observability, reduced mean time to resolution (MTTR), and the seamless orchestration of domain-specific AI agents that provide a unified, intelligent interface for our I&O engineers.

Position Overview

·        Client’s IT Operations team is seeking an innovative and highly technical AI Automation Engineer to help build and scale our unified operational AI interface.

·        This role is completely focused on the design, data integration, and agentic framework of our enterprise AI platform, which serves as the single, centralized entry point for I&O engineers to access operational data and insights.

·        You will partner closely with our lead AI architects to operationalize this vision. Your work will center on building the foundational data integrations, developing multi-agent reasoning workflows, and implementing AI-driven use cases for Help Desk automation, incident triage, and root cause analysis (RCA).

About the Operational AI Platform & Architecture

·        Our Enterprise Operational AI Platform is client’s unified operational AI interface. It acts as the centralized entry point for I&O engineers to access enterprise-wide operational data, insights, and automation capabilities—eliminating the need to navigate fragmented tools and dashboards.

From an architectural standpoint, the platform is designed as a multi-agent reasoning system powered by advanced LLMs. Key architectural components include:

·        Agentic Framework & MCP: It utilizes a dynamic agent registry and the Model Context Protocol (MCP) to orchestrate real-time, interactive operational tasks across different domain-specific AI agents.

·        Foundational Data Integration: The platform integrates directly with I&O systems of record and telemetry (e.g., ServiceNow CMDB, Dynatrace, Zabbix, Google Cloud Platform) to ingest infrastructure state, logs, and monitoring data.

·        Phased Execution Model: The architecture supports a scalable adoption path, beginning with query-only/read-only contextualization (e.g., incident triage, RCA) and advancing into fully autonomous, agentic actions (e.g., proactive outage prevention, maintenance suppression, and coordinated self-healing).

·        Complementary Enterprise Search: It is architected to work alongside and cross-validate data with enterprise search tools (like Glean) while providing the dynamic orchestration capabilities required for real-time operations.

Key Responsibilities

Operational AI Interface Development

·        Assist in the architecture and development of the unified AI assistant, ensuring a seamless experience for I&O engineers.

·        Develop and register AI agents and Model Context Protocol (MCP) integrations to handle real-time, interactive operational tasks.

·        Build backend workflows that support discovery, platform engineering sizing decisions, and proactive outage prevention.

Foundational Data & Integration

·        Develop robust data integrations connecting the AI platform to foundational infrastructure sources, including CMDB, Google Cloud Platform, ServiceNow, Zabbix, and Dynatrace.

·        Automate the ingestion and contextualization of infrastructure data across the entire I&O organization, leveraging AI-assisted development tools like Devin and Windsurf to accelerate engineering and accurately perform reactive and scheduled operational tasks.

·        Collaborate with observability and enterprise search (e.g., Glean) teams to cross-validate data, manage knowledge quality, and prevent LLM hallucinations.

Use Case Delivery & ROI

·        Deliver initial "low-hanging fruit" use cases, specifically focusing on Help Desk automation and Incident Management (MI) support.

·        Design AI capabilities to assist with change windows, maintenance suppression, and coordination of self-healing actions.

·        Establish AI observability metrics to measure confidence levels, token usage, and improvements to Mean Time to Resolution (MTTR).

Required Skill Sets & Qualifications

Technical Proficiencies:

·        AI & LLM Development: Deep understanding of Large Language Models (LLMs), advanced prompt engineering, and Retrieval-Augmented Generation (RAG).

·        AI Developer Tools: Hands-on experience or familiarity with advanced AI coding assistants and autonomous engineering platforms such as Devin and Windsurf.

·        Agentic Frameworks: Hands-on experience building multi-agent AI systems (e.g., LangChain, AutoGen, CrewAI) and utilizing Model Context Protocols (MCP) or tool-calling architectures.

·        Programming & APIs: Strong software engineering skills, primarily in Python, with extensive experience building and consuming RESTful APIs to connect disparate systems.

·        AIOps & Infrastructure Data: Familiarity with broader I&O data, including ITSM platforms (ServiceNow for ticketing/CMDB) and monitoring/telemetry systems (Dynatrace, Zabbix, Google Cloud Platform).

·        Data Quality & Governance: Experience handling unstructured and structured knowledge bases, improving data quality, and implementing guardrails for AI safety and accuracy.

Professional Competencies:

·        Systems Thinking: Ability to understand complex, enterprise-wide I&O environments and translate those environments into contextual schemas for AI consumption.

·        Cross-Functional Collaboration: Proven ability to work closely with principal architects, IT support teams, and security teams to safely deploy AI in a governed, enterprise environment.

·        Iterative Delivery: A practical approach to AI, focusing on delivering measurable value through phased adoption (e.g., starting with query-only workflows before moving to fully agentic actions).

Skills

Generative AI
Google Cloud
Help Desk
LLM
Python
Root Cause Analysis
ServiceNow
Triage
Zabbix

Similar jobs