← Back to Jobs
Technology
AWS DevOps Datadog
NTT DATA Americas, IncSandy Springs, GA🇺🇸United StatesPosted 27 Jul 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
Company Overview:
Req ID: 382633
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.
NTT DATA''s Client is currently seeking an AWS DevOps / SRE Engineer – Datadog & AIOps
Location: Atlanta, Georgia — Preferred Onsite/Hybrid
Employment Type: Full-Time
Experience Level: 5+ years in DevOps or Cloud Engineering with production SRE experience
Role Summary
We are seeking a hands-on AWS DevOps Engineer with strong Site Reliability Engineering capabilities and deep Datadog experience. This role will design and improve secure, scalable CI/CD pipelines; increase platform reliability through observability, automation, and SLO-driven practices; and introduce practical AIOps and generative AI capabilities that improve build quality, deployment safety, incident response, and engineering productivity.
Day-to-Day Job Duties
Design, build, and maintain resilient AWS environments using services such as EKS, EC2, S3, IAM, Lambda, RDS, CloudWatch, Route 53, ALB/NLB, and Secrets Manager.
Build, standardize, and optimize CI/CD pipelines using GitLab CI, GitHub Actions, Jenkins, or similar platforms, with automated testing, quality gates, approvals, rollback, and progressive-delivery controls.
Apply SRE practices by defining service-level indicators, service-level objectives, error budgets, availability targets, and operational-readiness criteria.
Implement and administer Datadog capabilities including infrastructure monitoring, APM, log management, Real User Monitoring, synthetics, dashboards, monitors, service maps, and incident workflows.
Create actionable observability and alerting strategies that reduce noise, improve mean time to detect and recover, and support rapid root-cause analysis.
Automate infrastructure provisioning and configuration using Terraform, CloudFormation, Ansible, or equivalent Infrastructure as Code tools.
Operate containerized workloads using Docker and Kubernetes/EKS, including autoscaling, health checks, resource optimization, and cluster reliability.
Integrate security and compliance controls into CI/CD, including secrets management, IAM least privilege, vulnerability scanning, SAST/DAST, dependency checks, artifact integrity, and audit evidence.
Use AIOps and generative AI to improve pipeline efficiency through intelligent failure analysis, configuration review, test generation, anomaly detection, change-risk scoring, and remediation recommendations.
Develop automation and operational tooling using Python, Bash, PowerShell, or similar scripting languages.
Lead production troubleshooting, incident response, post-incident reviews, problem management, and permanent corrective-action tracking.
Partner with application, platform, security, QA, and product teams to improve deployment frequency, change-failure rate*** lead time, reliability, and recovery performance.
Maintain runbooks, architecture diagrams, operational procedures, and engineering standards in Confluence, Jira, or similar tools.
Provide technical guidance and mentor engineers on cloud reliability, observability, automation, DevOps, and SRE practices.
Basic Qualifications
Minimum 5+ years of experience in DevOps, Cloud Engineering, Platform Engineering, or a related role supporting enterprise production systems.
Minimum 3+ years of hands-on experience designing, deploying, and operating solutions on AWS.
Minimum 5+ years of Strong experience building and supporting production-grade CI/CD pipelines; GitLab CI experience is preferred.
Minimum 5+ years of Demonstrated SRE experience with SLOs/SLIs, error budgets, incident response, on-call operations, reliability engineering, capacity planning, and blameless post-incident reviews.
Minimum 5+ years of Strong hands-on Datadog experience across metrics, logs, APM/tracing, dashboards, alerting, monitors, synthetics, integrations, and service-level reporting.
Minimum 5+ years of Experience with Terraform or CloudFormation and repeatable Infrastructure as Code practices.
Minimum 5+ years of Experience with Docker and Kubernetes; AWS EKS experience is strongly preferred.
Minimum 5+ years of Proficiency in at least one scripting or programming language such as Python, Bash, PowerShell, Go, or JavaScript/TypeScript.
Strong understanding of Linux, networking, IAM, secrets management, cloud security, and production troubleshooting.
Minimum 5+ years of Experience working in Agile environments and using tools such as Jira, Confluence, and Git.
Strong communication, collaboration, documentation, and problem-solving skills with a high degree of ownership.
Preferred Qualifications
Hands-on experience applying AIOps, machine learning, or generative AI to CI/CD, observability, incident management, automated testing, code review, or root-cause analysis.
Experience integrating LLM-based assistants or agents with developer platforms, repositories, ticketing systems, observability tools, or operational runbooks using secure enterprise controls.
Experience with progressive delivery and GitOps tools such as Argo CD, Flux, feature flags, canary deployments, and blue/green deployments.
Experience with OpenTelemetry, Prometheus, Grafana, Splunk, CloudWatch, or other observability platforms in addition to Datadog.
Knowledge of DORA metrics and experience improving deployment frequency, lead time for changes, change-failure rate*** and mean time to recovery.
Experience supporting regulated, financial-services, or other highly controlled enterprise environments.
AWS, Kubernetes, Terraform, Datadog, or relevant DevOps/SRE certifications.
Experience leading technical initiatives or mentoring engineering teams.
Key Success Measures
Improved CI/CD speed, stability, reuse, and developer adoption.
Reduced deployment failures, alert noise, incident recurrence, and recovery time.
Clear service-health reporting through Datadog dashboards, SLOs, and actionable alerts.
Increased automation across provisioning, testing, release controls, and operational remediation.
Safe, measurable adoption of AIOps and AI capabilities without weakening security or governance.
Travel
Expectation is 3 Days in office.
Degree
Bachelor''s degree in Computer Science, Engineering, Information Technology, or equivalent work experience.
What We Are Looking For
A hands-on engineer who balances delivery velocity with production reliability and operational discipline.
A proactive problem-solver who uses data, automation, and observability to prevent recurring issues.
A collaborative technical leader who can influence teams to adopt secure cloud, DevOps, SRE, and AI-assisted engineering practices.
Someone comfortable owning high-visibility production platforms and driving improvements from design through operations.
About NTT DATA:
NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com
NTT DATA endeavors to make accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here. If you''d like more information on your EEO rights under the law, please click here. For Pay Transparency information, please click here.
Req ID: 382633
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.
NTT DATA''s Client is currently seeking an AWS DevOps / SRE Engineer – Datadog & AIOps
Location: Atlanta, Georgia — Preferred Onsite/Hybrid
Employment Type: Full-Time
Experience Level: 5+ years in DevOps or Cloud Engineering with production SRE experience
Role Summary
We are seeking a hands-on AWS DevOps Engineer with strong Site Reliability Engineering capabilities and deep Datadog experience. This role will design and improve secure, scalable CI/CD pipelines; increase platform reliability through observability, automation, and SLO-driven practices; and introduce practical AIOps and generative AI capabilities that improve build quality, deployment safety, incident response, and engineering productivity.
Day-to-Day Job Duties
Design, build, and maintain resilient AWS environments using services such as EKS, EC2, S3, IAM, Lambda, RDS, CloudWatch, Route 53, ALB/NLB, and Secrets Manager.
Build, standardize, and optimize CI/CD pipelines using GitLab CI, GitHub Actions, Jenkins, or similar platforms, with automated testing, quality gates, approvals, rollback, and progressive-delivery controls.
Apply SRE practices by defining service-level indicators, service-level objectives, error budgets, availability targets, and operational-readiness criteria.
Implement and administer Datadog capabilities including infrastructure monitoring, APM, log management, Real User Monitoring, synthetics, dashboards, monitors, service maps, and incident workflows.
Create actionable observability and alerting strategies that reduce noise, improve mean time to detect and recover, and support rapid root-cause analysis.
Automate infrastructure provisioning and configuration using Terraform, CloudFormation, Ansible, or equivalent Infrastructure as Code tools.
Operate containerized workloads using Docker and Kubernetes/EKS, including autoscaling, health checks, resource optimization, and cluster reliability.
Integrate security and compliance controls into CI/CD, including secrets management, IAM least privilege, vulnerability scanning, SAST/DAST, dependency checks, artifact integrity, and audit evidence.
Use AIOps and generative AI to improve pipeline efficiency through intelligent failure analysis, configuration review, test generation, anomaly detection, change-risk scoring, and remediation recommendations.
Develop automation and operational tooling using Python, Bash, PowerShell, or similar scripting languages.
Lead production troubleshooting, incident response, post-incident reviews, problem management, and permanent corrective-action tracking.
Partner with application, platform, security, QA, and product teams to improve deployment frequency, change-failure rate*** lead time, reliability, and recovery performance.
Maintain runbooks, architecture diagrams, operational procedures, and engineering standards in Confluence, Jira, or similar tools.
Provide technical guidance and mentor engineers on cloud reliability, observability, automation, DevOps, and SRE practices.
Basic Qualifications
Minimum 5+ years of experience in DevOps, Cloud Engineering, Platform Engineering, or a related role supporting enterprise production systems.
Minimum 3+ years of hands-on experience designing, deploying, and operating solutions on AWS.
Minimum 5+ years of Strong experience building and supporting production-grade CI/CD pipelines; GitLab CI experience is preferred.
Minimum 5+ years of Demonstrated SRE experience with SLOs/SLIs, error budgets, incident response, on-call operations, reliability engineering, capacity planning, and blameless post-incident reviews.
Minimum 5+ years of Strong hands-on Datadog experience across metrics, logs, APM/tracing, dashboards, alerting, monitors, synthetics, integrations, and service-level reporting.
Minimum 5+ years of Experience with Terraform or CloudFormation and repeatable Infrastructure as Code practices.
Minimum 5+ years of Experience with Docker and Kubernetes; AWS EKS experience is strongly preferred.
Minimum 5+ years of Proficiency in at least one scripting or programming language such as Python, Bash, PowerShell, Go, or JavaScript/TypeScript.
Strong understanding of Linux, networking, IAM, secrets management, cloud security, and production troubleshooting.
Minimum 5+ years of Experience working in Agile environments and using tools such as Jira, Confluence, and Git.
Strong communication, collaboration, documentation, and problem-solving skills with a high degree of ownership.
Preferred Qualifications
Hands-on experience applying AIOps, machine learning, or generative AI to CI/CD, observability, incident management, automated testing, code review, or root-cause analysis.
Experience integrating LLM-based assistants or agents with developer platforms, repositories, ticketing systems, observability tools, or operational runbooks using secure enterprise controls.
Experience with progressive delivery and GitOps tools such as Argo CD, Flux, feature flags, canary deployments, and blue/green deployments.
Experience with OpenTelemetry, Prometheus, Grafana, Splunk, CloudWatch, or other observability platforms in addition to Datadog.
Knowledge of DORA metrics and experience improving deployment frequency, lead time for changes, change-failure rate*** and mean time to recovery.
Experience supporting regulated, financial-services, or other highly controlled enterprise environments.
AWS, Kubernetes, Terraform, Datadog, or relevant DevOps/SRE certifications.
Experience leading technical initiatives or mentoring engineering teams.
Key Success Measures
Improved CI/CD speed, stability, reuse, and developer adoption.
Reduced deployment failures, alert noise, incident recurrence, and recovery time.
Clear service-health reporting through Datadog dashboards, SLOs, and actionable alerts.
Increased automation across provisioning, testing, release controls, and operational remediation.
Safe, measurable adoption of AIOps and AI capabilities without weakening security or governance.
Travel
Expectation is 3 Days in office.
Degree
Bachelor''s degree in Computer Science, Engineering, Information Technology, or equivalent work experience.
What We Are Looking For
A hands-on engineer who balances delivery velocity with production reliability and operational discipline.
A proactive problem-solver who uses data, automation, and observability to prevent recurring issues.
A collaborative technical leader who can influence teams to adopt secure cloud, DevOps, SRE, and AI-assisted engineering practices.
Someone comfortable owning high-visibility production platforms and driving improvements from design through operations.
About NTT DATA:
NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com
NTT DATA endeavors to make accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here. If you''d like more information on your EEO rights under the law, please click here. For Pay Transparency information, please click here.
Skills
Docker
AWS
Machine Learning
Splunk
Agile
Ansible
Bash
CloudFormation
Confluence
Datadog
Generative AI
Git
GitHub Actions
GitLab CI
Grafana
JavaScript
Jenkins
Jira
Kubernetes
LLM
PowerShell
Prometheus
Python
SAFe
Terraform
TypeScript
Similar jobs
SRE Engineer (Interview type: L1 Video and L2 -Face to Face)
American IT Systems · O'Fallon, United States
39 minutes agoDevOps Engineer
Booz Allen Hamilton · Quantico, United States
40 minutes ago$77.6k - $176k/yrOpenShift Platform Engineer
Apex Systems · Charlotte, United States
40 minutes agoSenior DevOps Engineer - Kubernetes & Azure Cloud
EPAM Systems · New York, United States
42 minutes agoFTE :: DevOps Engineer || Berkeley Heights, NJ / Alpharetta, GA
IT First Source · Alpharetta, United States
56 minutes agoSoftware Engineering - SRE
Altitude Technology Solutions Inc · Plano, United States
56 minutes ago