Haystack
← Back to Jobs
Technology

Senior AI Ops / DevOps Engineer

VDart, Inc.Atlanta, GA🇺🇸United StatesPosted 22 Jul 2026

Why This Role Stands Out

This hybrid role offers a unique opportunity to blend cutting-edge AI technologies with robust DevOps practices, providing significant career growth in a rapidly evolving field. You'll thrive here if you're a skilled engineer with expertise in cloud, Kubernetes, and CI/CD, eager to innovate and build intelligent, secure software delivery ecosystems. Apply now to shape the future of AI-driven operations.

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Job Title: Senior AI Ops / DevOps Engineer

Location: Atlanta, GA(Hybrid)

Type: Contract

Description:

  • The ideal candidate is a senior hands-on DevOps engineer with strong cloud, automation, Kubernetes, and CI/CD expertise, combined with practical experience applying AI, LLM agents, and MCP-based integrations to modern engineering workflows.
  • This role will help create a secure AI-driven delivery ecosystem that accelerates software engineering velocity while maintaining strong governance, reliability, auditability, and operational control. This role is crucial to show the efficiency.

Day to Day Job Duties:

  • The Senior AI Ops / DevOps Engineer will architect, build, and manage next-generation AI-driven CI/CD and cloud operations ecosystems. This role will go beyond traditional DevOps automation by integrating LLM agents, Model Context Protocol servers, intelligent observability, and secure AI-assisted workflows into the software delivery lifecycle.
  • Architect, build, and manage AI-enabled CI/CD pipelines that improve developer productivity, code quality, release reliability, and deployment speed.
  • Design and deploy production-grade Model Context Protocol clients and servers to securely connect enterprise LLMs with engineering tools, repositories, cloud infrastructure, and observability platforms.
  • Develop custom MCP servers using Python, TypeScript, Node.js, or JavaScript to expose logs, infrastructure metrics, deployment data, and internal tools to authorized AI agents.
  • Integrate LLM agents into developer workflows to support automated code review, vulnerability detection, test generation, release validation, and infrastructure recommendations.
  • Build and maintain robust CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, ArgoCD, Jenkins, or similar tools.
  • Implement ChatOps 2.0 capabilities that allow engineers to interact with deployment pipelines, cloud environments, logs, and operational workflows using secure conversational interfaces.
  • Create safe autonomous remediation workflows for log analysis, incident triage, root-cause analysis, and infrastructure issue resolution.
  • Build guardrails that allow AI agents to generate, inspect, and safely execute Infrastructure as Code using Terraform, OpenTofu, Terragrunt, Pulumi, Crossplane, or similar tools.
  • Manage containerized workloads using Docker and Kubernetes platforms such as AWS EKS, Azure AKS, or Google GKE.
  • Integrate AI-driven observability workflows with platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK.
  • Implement AI safety controls including role-based access control, least-privilege execution, human-in-the-loop approvals, audit logging, rollback mechanisms, and secure tool access.
  • Partner with software engineering, DevOps, SRE, security, platform, and data/AI teams to identify opportunities for intelligent automation.
  • Create reusable automation frameworks, runbooks, dashboards, documentation, and enablement materials for engineering teams.
  • Drive an “automate everything” culture by reducing manual toil and improving operational efficiency across cloud and software delivery processes.

Basic Qualifications:

  • Minimum 7+ years of experience in DevOps, Cloud Engineering, SRE, Platform Engineering, or Infrastructure Automation.
  • Minimum 4+ years of hands-on experience designing and managing CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD, or similar platforms.
  • Minimum 3+ years of experience managing scalable cloud environments in AWS, Azure, or Google Cloud Platform, with strong preference for AWS.
  • Strong hands-on experience with Kubernetes, Docker, and production container orchestration platforms such as EKS, AKS, or GKE.
  • Advanced proficiency with Infrastructure as Code tools such as Terraform, OpenTofu, Terragrunt, Pulumi, CloudFormation, or Crossplane.
  • Strong programming and scripting experience using Python, TypeScript, JavaScript, Bash, or Go.
  • Practical experience working with LLM APIs such as OpenAI, Anthropic, or similar enterprise AI platforms.
  • Experience with AI orchestration or agentic frameworks such as LangChain, CrewAI, LlamaIndex, or similar tools.
  • Strong understanding of the Model Context Protocol ecosystem and experience designing or integrating MCP clients and servers.
  • Experience integrating DevSecOps controls into CI/CD pipelines, including SAST, DAST, dependency scanning, container scanning, secrets scanning, and vulnerability management.
  • Strong knowledge of secret management and security tooling such as HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, or similar platforms.
  • Experience with observability, monitoring, logging, and alerting platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK.
  • Familiarity with security and compliance frameworks such as SOC2, ISO27001, or enterprise audit control environments.
  • Ability to troubleshoot complex pipeline, infrastructure, deployment, and production issues across cloud-native environments.
  • Preferred / Nice to Have
  • Experience building AI-assisted infrastructure provisioning workflows.
  • Experience implementing autonomous or semi-autonomous incident response and remediation capabilities.
  • Experience with MLOps, model deployment pipelines, model monitoring, MLflow, SageMaker, or equivalent platforms.
  • Experience implementing human-in-the-loop approval models for AI-generated operational actions.
  • Experience with policy-as-code tools such as Open Policy Agent, Sentinel, Checkov, or similar solutions.
  • Experience working in regulated industries such as banking, financial services, healthcare, or insurance.
  • Experience with GitOps operating models using ArgoCD, Flux, or similar tools.
  • AWS, Kubernetes, DevOps, Security, or AI/ML certifications are a plus.
  • Soft Skills & Mindset
  • Strong “automate everything” mindset with a passion for reducing repetitive manual tasks and operational toil.
  • Security-first approach with practical skepticism of autonomous AI actions and a focus on validation, boundaries, approvals, and rollback.
  • Ability to bridge traditional software engineering, DevOps, SRE, security, and data/AI teams.
  • Strong communication skills with the ability to explain complex AI-enabled DevOps concepts to both technical and leadership audiences.
  • Collaborative educator who can help upskill engineering teams on AI-assisted delivery, secure automation, and modern DevOps practices.
  • Ownership mindset with the ability to design solutions, implement them hands-on, and support them in production.

Degree:

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or equivalent work experience.

 

Skills

Docker
Node.js
AWS
ELK
MLOps
MLflow
Splunk
ArgoCD
Azure
Bash
CircleCI
CloudFormation
Datadog
GitHub Actions
GitLab CI
Google Cloud
Grafana
JavaScript
Jenkins
Kubernetes
LLM
Prometheus
Pulumi
Python
SAFe
Terraform
TypeScript
Vault

Similar jobs