← Back to Jobs
Technology
Senior AI DevOps Engineer (AI Ops / Platform Engineering)
VDart, Inc.Atlanta, GA🇺🇸United StatesPosted 7 Aug 2026
Quick Overview
Work Type
On Site
Level
Mid Senior
Job Description
Role: Senior AI DevOps Engineer (AI Ops / Platform Engineering)
Location: Atlanta (Onsite)
Type: Contract
Position Summary
- We're looking for an experienced Senior AI DevOps Engineer to help build the next generation of AI-powered software delivery and cloud operations. In this role, you'll combine modern DevOps practices with Generative AI, LLM agents, Model Context Protocol (MCP), and intelligent automation to transform how engineering teams build, deploy, and operate software.
- You'll partner with Platform Engineering, DevOps, Security, SRE, and AI teams to design secure, scalable, cloud-native solutions that accelerate software delivery while improving reliability, observability, and operational efficiency.
- This is an opportunity to work on cutting-edge AI technologies that are redefining modern software engineering.
What You'll Do:
- Design, build, and optimize AI-enabled CI/CD pipelines that improve developer productivity, deployment speed, and software quality.
- Develop and deploy Model Context Protocol (MCP) clients and servers that securely connect enterprise LLMs with engineering tools, cloud infrastructure, and operational platforms.
- Build custom MCP services using Python, TypeScript, JavaScript, or Node.js to expose infrastructure, deployment, monitoring, and operational data to authorized AI agents.
- Integrate LLM-powered workflows for automated code reviews, testing, security analysis, release validation, and infrastructure recommendations.
- Build and maintain enterprise CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, CircleCI, ArgoCD, or similar platforms.
- Implement AI-driven ChatOps capabilities that enable engineers to interact with deployment pipelines, cloud environments, and operational tools through secure conversational interfaces.
- Design intelligent remediation workflows for incident detection, root cause analysis, log analysis, and operational troubleshooting.
- Develop secure Infrastructure-as-Code automation using Terraform, OpenTofu, Pulumi, Terragrunt, CloudFormation, or similar technologies.
- Deploy and manage containerized applications using Kubernetes and Docker across AWS, Azure, or Google Cloud.
- Build AI-powered observability solutions leveraging Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, ELK, or similar platforms.
- Implement security guardrails including RBAC, least-privilege access, approval workflows, audit logging, rollback mechanisms, and secure AI tool access.
- Partner with Engineering, Platform, Security, SRE, and AI teams to identify and implement intelligent automation opportunities.
- Create reusable automation frameworks, documentation, dashboards, and engineering best practices that scale across the organization.
Required Qualifications
- 7+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering (SRE), Cloud Engineering, or Infrastructure Automation.
- 4+ years designing and supporting enterprise CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, CircleCI, ArgoCD, or similar tools.
- 3+ years of cloud engineering experience in AWS, Azure, or Google Cloud (AWS preferred).
- Strong experience deploying and managing Kubernetes and Docker in production environments (EKS, AKS, or GKE).
- Hands-on experience with Infrastructure as Code using Terraform, OpenTofu, Pulumi, Terragrunt, CloudFormation, or similar tools.
- Strong programming skills in Python, TypeScript, JavaScript, Bash, Go, or similar languages.
- Experience integrating enterprise LLM platforms such as OpenAI, Anthropic, or equivalent AI services into engineering workflows.
- Experience with AI orchestration frameworks such as LangChain, CrewAI, LlamaIndex, or similar technologies.
- Experience designing or implementing Model Context Protocol (MCP) clients and servers.
- Experience implementing DevSecOps practices including SAST, DAST, dependency scanning, container security, secrets management, and vulnerability management.
- Experience with secrets management solutions such as HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault.
- Strong experience with monitoring, logging, and observability platforms such as Datadog, Grafana, Prometheus, CloudWatch, Splunk, Dynatrace, or ELK.
- Excellent troubleshooting skills across cloud infrastructure, CI/CD pipelines, Kubernetes, and production systems.
- Strong communication skills and the ability to collaborate across engineering, security, and AI teams.
Preferred Qualifications
- Experience building AI-assisted infrastructure provisioning and deployment workflows.
- Experience implementing autonomous or AI-assisted incident response and operational remediation.
- Experience with MLOps platforms including MLflow, Amazon SageMaker, Vertex AI, Azure ML, or similar technologies.
- Experience implementing human-in-the-loop approval workflows for AI-generated operational actions.
- Knowledge of Policy-as-Code frameworks such as Open Policy Agent (OPA), Sentinel, or Checkov.
- Experience with GitOps platforms such as ArgoCD or Flux.
- Experience working within regulated industries such as financial services, healthcare, insurance, or government.
- AWS, Kubernetes, DevOps, Security, or AI/ML certifications.
What Makes You Successful
- Passion for automation and continuously improving engineering productivity.
- Security-first mindset with practical experience implementing safe, responsible AI automation.
- Ability to bridge DevOps, Platform Engineering, AI, Security, and Software Engineering disciplines.
- Strong problem-solving skills with an ownership mentality from design through production support.
- Comfortable leading technical initiatives and mentoring engineering teams on modern AI-enabled development practices.
Education
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline (or equivalent professional experience).
Skills
Docker
Node.js
AWS
ELK
MLOps
MLflow
Splunk
ArgoCD
Azure
Bash
CircleCI
CloudFormation
Datadog
Generative AI
GitHub Actions
GitLab CI
Google Cloud
Grafana
JavaScript
Jenkins
Kubernetes
LLM
Prometheus
Pulumi
Python
SAFe
Terraform
TypeScript
Vault
Similar jobs
W2 ONLY - Need SRE Cloud Foundations - REMOTE PST Hours
Digitive LLC · United States
12 minutes agoMiddleware Engineer / Platform Engineer
Cynet Systems · Charlotte, United States
15 minutes ago$53 - $58/hrPrincipal Network Engineer, onsite, Tucson, AZ with Security Clearance
RTX · Tucson, United States
16 minutes ago$107.5k - $204.5k/yrMiddleware Engineer / Platform Engineer
Talent Groups · Charlotte, United States
41 minutes agoSenior DevOps Engineer AI & Automation- 5+ yrs- Plano, Texas- Onsite
iMedhas Consulting Services · Plano, United States
42 minutes agoSr DevOps Eng
Vaco by Highspring · Charlotte, United States
42 minutes ago