Haystack
← Back to Jobs
Technology
TC

Site Reliability Engineer

Take2 ConsultingUnited States🇺🇸United StatesPosted Sep 30, 2026

Quick Overview

Salary
$80k - $110k/yr
Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
SQLELKSplunkDatadogGrafanaPowerShellPython

Job Description

Site Reliability Engineer

Veterans Affairs ESOM

Overview

We are seeking an experienced Site Reliability Engineer (SRE) to support enterprise Problem Management and Root Cause Analysis (RCA) activities. In this senior-level role, you will investigate complex production outages and incidents across a large, diverse environment, identifying technical root causes and supporting corrective actions. This opportunity is ideal for a technically versatile professional comfortable working across multiple infrastructure and application domains to ensure system reliability and sustainable resolution.

Education Requirements

  • Bachelor's Degree in Engineering, Computer Science, Systems, Business, or a related scientific/technical discipline.

Certification Requirements

There are no certification requirements for this role.

Clearance Requirements

Clearance Level: Public Trust

Work Arrangement

Remote

Responsibilities

  • Investigate complex production outages and major incidents using logs, telemetry, monitoring data, traces, incident history, change records, and other technical evidence.
  • Assess unfamiliar systems to determine the information needed for effective technical investigations.
  • Query, extract, correlate, and analyze data using enterprise observability platforms and scripting/query languages.
  • Develop and test root cause hypotheses, challenge unsupported conclusions, and validate findings against technical evidence.
  • Develop technical mitigation and corrective action recommendations focused on sustainable resolution.
  • Support verification that proposed fixes address the original failure conditions through testing, simulation, or other defensible methods.
  • Identify recurring patterns and systemic risks across systems and teams.
  • Collaborate with technical teams, investigation coordinators, analysts, and stakeholders throughout RCA activities.
  • Leverage approved AI-assisted tools to accelerate analysis while independently validating results.

Required Qualifications

  • Demonstrated experience diagnosing complex, enterprise-scale production outages across multiple technology domains.
  • Hands-on experience querying enterprise observability and log analysis platforms such as Splunk, Dynatrace, Elastic/ELK, Grafana, or Datadog.
  • Proficiency with scripting and query languages such as SQL, Python, or PowerShell.
  • Ability to correlate logs, metrics, changes, incidents, and operational data into clear technical timelines and evidence-based conclusions.
  • Strong analytical judgment with the ability to independently evaluate, validate, or challenge proposed root causes.
  • Excellent communication and collaboration skills across engineering, support, and technical teams.

Desired Skills

  • Experience working with enterprise observability and monitoring platforms.
  • Ability to work across diverse technical environments and analyze telemetry data effectively.
  • Strong problem-solving skills and attention to detail.
  • Proven ability to work independently and handle complex investigations.

Pay Range

Pay range: $80,000.00–$110,000.00 (annual). This pay range is based on qualifications and experience, and assumes all role requirements are met.

Why Apply

This is a unique opportunity to contribute to enterprise-wide system stability and reliability through complex technical investigations. If you thrive in challenging environments and enjoy solving complex production issues, we encourage you to apply and join a dedicated team committed to operational excellence.

Similar jobs