Quick Overview
Job Description
Job Title: SRE Engineer
Location: Washington, DC
Duration: 12+ Months
Rate: DOE
Job Description
Reliability Engineer (SRE) to champion system availability, performance, and automation across their enterprise cloud infrastructure. In this role, you will bridge the gap between development and operations by implementing robust CI/CD pipelines and Infrastructure-as-Code (IaC), while heavily leveraging the Dynatrace observability platform to drive deep-dive distributed tracing, build intelligent dashboards, and tune anomaly detection. As a core member of the reliability team, you will apply formal SRE principles—such as defining SLIs/SLOs and managing error budgets—to optimize capacity and resiliency, ensure strict security compliance, and participate in an on-call rotation using ITIL frameworks to minimize incident response times.
Responsibilities
- Observability & Monitoring: Standardize and automate Dynatrace installations, integrate telemetry collection into CI/CD pipelines, enforce tagging/metadata standards, configure distributed tracing with context propagation, and optimize custom dashboards and anomaly alerts.
- Deployment & Automation: Design, implement, and maintain CI/CD pipelines using GitHub Actions, AWS CodePipeline, or Jenkins; provision scalable cloud infrastructure using Terraform, CloudFormation, or AWS CDK.
- Incident & Problem Management: Serve as a production on-call responder using ITIL frameworks and ServiceNow; lead troubleshooting efforts, conduct deep root-cause analysis (RCA), and author comprehensive knowledge base articles.
- Reliability Engineering: Champion SRE metrics including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets; design and execute resiliency test plans and support performance testing.
- Performance & Capacity Optimization: Drive operational cost optimization initiatives across cloud environments and configure robust auto-scaling policies and thresholds.
- Security & Compliance: Manage service accounts, access permissions, and digital certificates; respond rapidly to security incidents and execute remediation protocols.
Qualifications & Requirements
- Education & Experience: Bachelor’s degree in Computer Science, Engineering, or a related technical field, paired with 2 to 4 years of hands-on experience in SRE, DevOps, or infrastructure-focused roles.
- Cloud & Containerization: Practical, hands-on experience managing multi-tenant environments within AWS and Azure, alongside a solid understanding of container technologies like Docker, Kubernetes, and Amazon ECS.
- Automation & Scripting: Mid-level proficiency in Python (or similar scripting languages) and practical experience with configuration management tools like Ansible to build automated self-service tools.
- Systems & Networking Architecture: Strong foundational knowledge of Linux systems engineering, core networking concepts, and navigating relational, cloud-native, and NoSQL databases.
- Professional Competencies: Excellent written and verbal communication skills for cross-functional collaboration, a proven ability to work independently, and the flexibility to participate in an on-call rotation outside standard business hours
Skills
Similar jobs
Sr Android Platform Engineer
SDV International · Sterling, United States
5 minutes agoDevOps / Infrastructure Engineer - Rockville, MD or McLean, VA
FutureTech Consultants LLC · Rockville, United States
5 minutes agoAWS Cloud Platform Engineer Lead
Cynet Systems · Reston, United States
5 minutes ago$82 - $87/hrDevOps Architect with TS/SCI with Poly
SolveIT Services Inc · Palm Bay, United States
7 minutes ago$180k/yrAWS Cloud DevOps Engineer
iPeople Infosystems LLC · Oakland, United States
53 minutes agoDevOps Engineer
Congensys Corp. · Charlotte, United States
54 minutes ago