Quick Overview
Job Description
Site Reliability Engineer
Veterans Affairs ESOM
Overview
We are seeking an experienced Site Reliability Engineer (SRE) to support enterprise Problem Management and Root Cause Analysis (RCA) activities. In this senior-level role, you will investigate complex production outages and incidents across a large, diverse environment, identifying technical root causes and supporting corrective actions. This opportunity is ideal for a technically versatile professional comfortable working across multiple infrastructure and application domains to ensure system reliability and sustainable resolution.
Education Requirements
- Bachelor's Degree in Engineering, Computer Science, Systems, Business, or a related scientific/technical discipline.
Certification Requirements
There are no certification requirements for this role.
Clearance Requirements
Clearance Level: Public Trust
Work Arrangement
Remote
Responsibilities
- Investigate complex production outages and major incidents using logs, telemetry, monitoring data, traces, incident history, change records, and other technical evidence.
- Assess unfamiliar systems to determine the information needed for effective technical investigations.
- Query, extract, correlate, and analyze data using enterprise observability platforms and scripting/query languages.
- Develop and test root cause hypotheses, challenge unsupported conclusions, and validate findings against technical evidence.
- Develop technical mitigation and corrective action recommendations focused on sustainable resolution.
- Support verification that proposed fixes address the original failure conditions through testing, simulation, or other defensible methods.
- Identify recurring patterns and systemic risks across systems and teams.
- Collaborate with technical teams, investigation coordinators, analysts, and stakeholders throughout RCA activities.
- Leverage approved AI-assisted tools to accelerate analysis while independently validating results.
Required Qualifications
- Demonstrated experience diagnosing complex, enterprise-scale production outages across multiple technology domains.
- Hands-on experience querying enterprise observability and log analysis platforms such as Splunk, Dynatrace, Elastic/ELK, Grafana, or Datadog.
- Proficiency with scripting and query languages such as SQL, Python, or PowerShell.
- Ability to correlate logs, metrics, changes, incidents, and operational data into clear technical timelines and evidence-based conclusions.
- Strong analytical judgment with the ability to independently evaluate, validate, or challenge proposed root causes.
- Excellent communication and collaboration skills across engineering, support, and technical teams.
Desired Skills
- Experience working with enterprise observability and monitoring platforms.
- Ability to work across diverse technical environments and analyze telemetry data effectively.
- Strong problem-solving skills and attention to detail.
- Proven ability to work independently and handle complex investigations.
Pay Range
Pay range: $80,000.00–$110,000.00 (annual). This pay range is based on qualifications and experience, and assumes all role requirements are met.
Why Apply
This is a unique opportunity to contribute to enterprise-wide system stability and reliability through complex technical investigations. If you thrive in challenging environments and enjoy solving complex production issues, we encourage you to apply and join a dedicated team committed to operational excellence.
Similar jobs
- SU
Systems & Platform Engineer - TS/SCI with Security Clearance
NewSunayu, LLC
Bethesda, MD🇺🇸On-siteYesterdayDockerOracleAWS+7Technology - CI
Senior Java Platform Engineer - Senior Vice President - Citi
Citi
Tampa, FL🇺🇸$141.4k - $212.2k/yrRemote4 days agoDockerMicroservicesMongoDB+14Technology - LI
DevOps Engineer
NewLIGHTFEATHER IO LLC
United States🇺🇸HybridYesterdayDockerRubyAWS+4Technology - HA
Senior AI Platform Engineer
NewHasbro
Renton🇺🇸Hybrid11 hours agoAWSCDKPython+2Technology - DE
Staff DevOps Engineer, Developer Experience (DevEx)
NewDemandbase
US - Remote🇺🇸Remote5 hours agoGCPAWSService Mesh+14Technology - NT
Systems Dev/DevOps Engineer with Security Clearance
NewNewGen Technologies
Herndon, VA🇺🇸HybridYesterdayDockerRubyShell+15Technology