Haystack
← Back to Jobs
Technology
EX

Site Reliability Engineer with active Clearance

Excelque, IncChantilly, VA🇺🇸United StatesPosted 1 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Chantilly, VA, United States
Posted
Yesterday
DockerMicroservicesAWSELKConfluenceGitGrafanaJenkinsJiraKubernetesPrometheusPythonTerraform

Job Description

Hi,

Job Title: Site Reliability Engineer

Location: 5 days per week onsite in Chantilly, VA

Clearance Required: TS/SCI w CI Poly required

Client: Booz Allen

If you are interested in this position, please share your updated resume at:

The Opportunity

As a Lead Site Reliability Engineer (SRE) on our team, you ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and implementation of observability, automation, incident response, and operational best practices across cloud and air-gapped environments, while partnering closely with DevOps, infrastructure, and security teams to improve system resilience and reduce operational risk. The Lead SRE also drives root cause analysis, capacity planning, reliability standards, and continuous improvement initiatives to support highly available, efficient, and scalable services. This is your chance to further your skills in cloud infrastructure and technologies while continuing to grow your SRE experience. You'll build and support a reliable site for the environment in order to meet the development and maintenance requirements of systems and platforms. Work with the development and operation teams to evaluate the health, stability and reliability of systems and platforms. Design and develop technical tools to debug problems that occur in the deployment of applications, within specific platforms and systems.

Join our efforts to strengthen our security posture and safeguard national interests.

Qualifications

8+ years of experience with monitoring, logging, and observability platforms, such as Prometheus, Grafana, and ELK stack

8+ years of experience with Linux systems administration and networking fundamentals within AWS

Experience with Python scripting and automation

Experience with Infrastructure as Code using Terraform and Terragrunt

Knowledge of Kubernetes administration, troubleshooting, and operations.

TS/SCI clearance with a polygraph

Bachelor s degree and 8+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering, or 12+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering in lieu of a degree

Ability to obtain a Security+ CE, SSCP, CCNA-Security, or GSEC Certification within 6 months of start date

Nice to Have Skills

Experience with deploying and managing OpenTelemetry.

Experience with AWS CloudWatch, AWS EKS, and related AWS services

Experience managing Kubernetes environments through Rancher

Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management

Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence

Knowledge of distributed systems, microservices architectures, and containerized workloads

Knowledge of NIST 800-53 and NIST-190

Master s degree in a relevant field

Security+ CE, SSCP, CCNA-Security, or GSEC Certification

Similar jobs