Quick Overview
Job Description
DTS is looking for experienced Sr SRE for our Direct Client position based in Atlanta, GA
Job Description:
Position Summary
We are seeking an experienced Enterprise Site Reliability Engineer to advance reliability, resilience, observability, automation, and operational excellence across the organization.
This is a highly visible, hands-on engineering role that will work across SRE, application, architecture, cloud, platform, infrastructure, security, data, and other IT domains. The successful candidate will combine deep technical expertise with curiosity, thorough analysis, and strong engineering judgment to identify risks others may overlook, connect technical findings to business impact, and develop scalable enterprise solutions.
The role requires someone who can move effectively between detailed technical analysis, software development, application architecture reviews, complex incident leadership, enterprise standards, and executive-level communication. The candidate must be able to influence SRE and engineering practices across organizational boundaries without relying on direct authority.
Key Responsibilities
- Define and advance enterprise SRE standards for SLIs, SLOs, error budgets, production readiness, incident management, and operational excellence.
- Establish practical observability standards for OpenTelemetry, traces, logs, metrics, events, telemetry correlation, tagging, data quality, and service ownership.
- Conduct thorough application maturity assessments across observability, reliability, resilience, operability, automation, and incident readiness.
- Analyze application architecture and production operations to uncover dependencies, failure modes, capacity constraints, performance bottlenecks, and risks that may not be immediately visible.
- Connect technical and operational findings to customer experience, business impact, service risk, and investment priorities.
- Translate assessments and operational data into clear, prioritized, and measurable improvement roadmaps.
- Influence architecture and engineering decisions by recommending reliability patterns such as fault isolation, graceful degradation, circuit breakers, retries, rate limiting, load shedding, high availability, and disaster recovery.
- Design end-to-end observability across distributed services, APIs, business transactions, customer journeys, cloud platforms, and cross-domain dependencies.
- Develop production-grade automation, applications, APIs, and integrations that reduce toil, improve consistency, and scale reliability practices across the enterprise.
- Automate service onboarding, telemetry validation, SLO reporting, production readiness checks, incident enrichment, and remediation workflows.
- Integrate observability, cloud, CI/CD, IT service management, incident management, and configuration platforms using APIs, SDKs, webhooks, and event-driven patterns.
- Lead complex incident investigations and post-incident reviews, challenge assumptions, identify contributing factors, and drive corrective actions through completion.
- Analyze operational and observability data to detect patterns, quantify risks, identify systemic gaps, and generate actionable insights for engineers and leaders.
- Proactively identify opportunities to improve application operability, resilience, telemetry, support readiness, and operational processes before issues affect customers.
- Define observability for AI agents and AI-enabled applications, including workflows, model and tool interactions, dependencies, latency, failures, quality, token consumption, cost, and reliability.
- Use AI-assisted engineering tools such as Kiro, GitHub Copilot, and AI agents to accelerate development, incident analysis, correlation, pattern detection, and operational insights.
- Develop reusable reference architectures, engineering patterns, assessment frameworks, maturity models, scorecards, and implementation guidance.
- Work with SREs and engineering teams across the organization to resolve cross-domain reliability issues and promote consistent practices.
- Mentor engineers, facilitate difficult technical decisions, constructively challenge existing approaches, and build alignment across teams.
- Communicate complex technical risks, business impact, recommendations, and progress clearly to engineering teams and senior leadership.
Required Qualifications
- 8+ years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or software engineering for large-scale production systems.
- Deep knowledge of SRE principles, operational excellence, distributed systems, microservices, APIs, and cloud-native architecture.
- Deep hands-on experience architecting, operating, and troubleshooting large-scale AWS environments across compute, containers, serverless, networking, databases, storage, identity, and cloud observability, including EC2, EKS, Lambda, VPC, Elastic Load Balancing, Route 53, RDS/Aurora, DynamoDB, S3, IAM, and CloudWatch.
- Strong hands-on experience with Kubernetes and Red Hat OpenShift Service on AWS (ROSA) architecture, operations, performance, and troubleshooting.
- Strong experience with OpenTelemetry, distributed tracing, logs, metrics, events, and Dynatrace or a comparable enterprise observability platform.
- Experience defining and operationalizing SLIs, SLOs, error budgets, production readiness criteria, and reliability scorecards.
- Demonstrated ability to assess application architecture and operational maturity from observability, reliability, resilience, and operability perspectives.
- Strong software development skills in Python, Java, Go, JavaScript/TypeScript, or similar languages.
- Experience developing production-grade APIs, integrations, automation services, and internal engineering tools.
- Proficiency with REST APIs, SDKs, Git, automated testing, CI/CD, Infrastructure as Code, secure development, and software lifecycle practices.
- Experience leading complex incidents, technical investigations, root cause analysis, post-incident reviews, and corrective-action programs.
- Strong analytical and investigative skills, with the curiosity and technical judgment to look beyond immediate symptoms and uncover systemic problems.
- Demonstrated ability to convert technical findings and operational data into clear, actionable insights and enterprise recommendations.
- Ability to connect reliability risks and technical decisions to customer experience, business impact, and organizational priorities.
- Excellent written, verbal, technical, and executive communication skills.
- Proven ability to mentor experienced engineers, challenge assumptions constructively, facilitate decisions, and influence outcomes without direct authority.
- Ability to collaborate effectively with SREs, architects, application teams, and specialists across multiple IT domains.
Preferred Qualifications
- AWS certification, preferably AWS Certified Solutions Architect Professional, AWS Certified DevOps Engineer Professional, or a relevant specialty certification.
- Kubernetes or OpenShift certification, such as CKA, CKS, or Red Hat Certified OpenShift Administrator.
- Experience integrating enterprise platforms such as Dynatrace, AWS, ServiceNow, GitLab, Kubernetes, and ROSA.
- Experience developing self-service reliability capabilities, internal developer platforms, or enterprise engineering products.
- Experience with performance engineering, capacity planning, resilience testing, and chaos engineering.
- Experience establishing observability and reliability controls for AI agents, LLM-enabled applications, or AI-driven workflows.
- Experience working with large-scale, highly available, business-critical enterprise systems.
DTS offers excellent compensation package.
Contact:
Pankaj Kumar
Digital Technology Solutions
Similar jobs
- KT
AWS Federated Cloud Platform Engineer
NewKforce Technology Staffing
Princeton, NJ🇺🇸Remote23 hours agoDockerAWSSplunk+9Technology - SR
Sr. DevSecOps Engineer with Security Clearance
NewSeneca Resources, LLC
Reston, VA🇺🇸Hybrid23 hours agoEngineering - NT
W2: DevOps Engineer - F2F interview (local GA and CA)
NewNoblesoft Technologies Inc.
Milton, GA🇺🇸On-site23 hours agoDockerShellAzure+7Technology - PE
Site Reliability Engineer
NewPercient Inc.
Plano, TX🇺🇸On-site23 hours agoDockerDynamoDBMicroservices+26Technology - PC
DevOps/Platform Engineer (Kubernetes/ API & AI Gateway)
NewPyramid Consulting, Inc.
Minneapolis, MN🇺🇸$65 - $68/hrOn-site23 hours agoAPI GatewayAWSEnvoy+11Technology - JM
DevOps Engineer + HPC (High Performance Computing)
NewJean Martin, Inc
United States🇺🇸HybridYesterdayDockerAWSMachine Learning+11Technology