Haystack
← Back to Jobs
Technology

AI DevOps Engineer

Learn Beyond Consulting LLCAtlanta, GA🇺🇸United StatesPosted 21 Jul 2026

Why This Role Stands Out

This hybrid role offers a fantastic opportunity to shape cutting-edge AI platforms, leveraging your expertise in cloud architecture and DevSecOps to drive innovation. You'll thrive here if you're a seasoned engineer passionate about building scalable, secure, and observable AI infrastructure within a collaborative environment. Apply now to make a significant impact and grow your career in the exciting field of AI operations!

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

AI DevOps Engineer

Cloud Architecture & Automation: 10+ years of experience designing and deploying cloud infrastructure using Google Cloud Platform (preferred), Azure, or Databricks with strong Terraform/Bicep expertise.

DevSecOps & Governance: Proven ability to implement secure CI/CD pipelines, cloud security controls (IAM, encryption, secrets management), and governance frameworks.

AI/ML Operations & Observability: Hands-on experience supporting production AI/ML environments with model monitoring, drift detection, logging, alerting, and observability solutions

Job Description

We are seeking a Senior DevOps Engineer IV to support the design, implementation, and optimization of enterprise AI platforms and cloud infrastructure. This role will focus on cloud architecture, Infrastructure as Code (IaC), AI/ML operations, observability, security, governance, and Agentic AI systems. The ideal candidate will have extensive experience building scalable cloud environments, implementing DevOps best practices, and enabling production-grade AI solutions in regulated enterprise environments.

Key Responsibilities
Design, deploy, and maintain enterprise cloud infrastructure supporting AI/ML workloads.
Implement Infrastructure as Code using Terraform, Bicep, or similar automation tools.
Develop and manage CI/CD pipelines with integrated security, governance, and compliance controls.
Architect scalable, highly available AI platforms across cloud environments.
Drive FinOps practices, including cloud cost optimization, resource utilization, and enterprise cost visibility.
Implement AI observability frameworks covering model performance, drift detection, reliability, business KPIs, and operational monitoring.
Design and support AI/ML lifecycle management including monitoring, retraining strategies, logging, alerting, and incident response.
Embed security, compliance, and model risk management controls into AI development and deployment processes.
Support development and operationalization of Agentic AI solutions, including orchestration, monitoring, testing, and governance.
Establish best practices for AgentOps, model governance, and AI platform reliability

Additional Skills & Qualifications

Design and implementation of multi-cloud AI infrastructure with integrated governance and policy controls.
Experience embedding security, compliance, and governance controls directly into IaC and deployment pipelines.
Strong understanding of AI FinOps, including token optimization, cost-performance tradeoffs, and enterprise cost visibility.
Experience implementing model risk management controls, including auditability, explainability, and access governance.
Knowledge of designing AI systems for regulated environments and enforcing runtime guardrails and policy controls.
Ability to develop enterprise-wide AI observability strategies, covering model performance, data drift, bias detection, reliability, and business KPIs.
Experience implementing centralized monitoring frameworks and automated response mechanisms across AI platforms.
Exposure to LLM-based applications, AI agents, prompt engineering, API integrations, and orchestration frameworks such as LangChain.
Experience designing and supporting agent-based systems at scale, including multi-agent coordination, tool orchestration, memory management, and state management.
Knowledge of AgentOps practices, including AI deployment, testing, monitoring, iteration, and governance.
Understanding of autonomous AI system failure modes and mitigation strategies.

Skills

Encryption
Azure
Databricks
Google Cloud
LLM
Terraform

Similar jobs