Haystack
← Back to Jobs
Technology

SRE Engineer

Georgia ITWashington, DC🇺🇸United StatesPosted 20 Jul 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Job Title:   SRE Engineer

Location:   Washington, DC

Duration:   12+ Months

Rate:          DOE

 

Job Description

Reliability Engineer (SRE) to champion system availability, performance, and automation across their enterprise cloud infrastructure. In this role, you will bridge the gap between development and operations by implementing robust CI/CD pipelines and Infrastructure-as-Code (IaC), while heavily leveraging the Dynatrace observability platform to drive deep-dive distributed tracing, build intelligent dashboards, and tune anomaly detection. As a core member of the reliability team, you will apply formal SRE principles—such as defining SLIs/SLOs and managing error budgets—to optimize capacity and resiliency, ensure strict security compliance, and participate in an on-call rotation using ITIL frameworks to minimize incident response times.

 

Responsibilities

  • Observability & Monitoring: Standardize and automate Dynatrace installations, integrate telemetry collection into CI/CD pipelines, enforce tagging/metadata standards, configure distributed tracing with context propagation, and optimize custom dashboards and anomaly alerts.
  • Deployment & Automation: Design, implement, and maintain CI/CD pipelines using GitHub Actions, AWS CodePipeline, or Jenkins; provision scalable cloud infrastructure using Terraform, CloudFormation, or AWS CDK.
  • Incident & Problem Management: Serve as a production on-call responder using ITIL frameworks and ServiceNow; lead troubleshooting efforts, conduct deep root-cause analysis (RCA), and author comprehensive knowledge base articles.
  • Reliability Engineering: Champion SRE metrics including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets; design and execute resiliency test plans and support performance testing.
  • Performance & Capacity Optimization: Drive operational cost optimization initiatives across cloud environments and configure robust auto-scaling policies and thresholds.
  • Security & Compliance: Manage service accounts, access permissions, and digital certificates; respond rapidly to security incidents and execute remediation protocols.

Qualifications & Requirements

  • Education & Experience: Bachelor’s degree in Computer Science, Engineering, or a related technical field, paired with 2 to 4 years of hands-on experience in SRE, DevOps, or infrastructure-focused roles.
  • Cloud & Containerization: Practical, hands-on experience managing multi-tenant environments within AWS and Azure, alongside a solid understanding of container technologies like Docker, Kubernetes, and Amazon ECS.
  • Automation & Scripting: Mid-level proficiency in Python (or similar scripting languages) and practical experience with configuration management tools like Ansible to build automated self-service tools.
  • Systems & Networking Architecture: Strong foundational knowledge of Linux systems engineering, core networking concepts, and navigating relational, cloud-native, and NoSQL databases.
  • Professional Competencies: Excellent written and verbal communication skills for cross-functional collaboration, a proven ability to work independently, and the flexibility to participate in an on-call rotation outside standard business hours

 

Skills

Docker
AWS
Ansible
Azure
CDK
CloudFormation
GitHub Actions
Jenkins
Kubernetes
Python
Terraform

Similar jobs