Why This Role Stands Out
This role offers a fantastic opportunity to lead the design and implementation of critical production systems, leveraging your expertise in observability, automation, and cloud environments to enhance system resilience. If you're a seasoned Site Reliability Engineer with a passion for ensuring high availability and a desire to contribute to safeguarding national interests, this position provides significant growth potential and impactful work. Apply to join a dedicated team and elevate your career in a secure and vital sector.
Quick Overview
Job Description
Job Title: Site Reliability Engineer
Location: 5 days per week onsite in Chantilly, VA
Clearance Required: TS/SCI w CI Poly required Top Skills Senior/Lead-level Site Reliability Engineering / Production Reliability experience
Strong, hands-on ELK / Elastic Stack experience:
Elasticsearch
Logstash
Kibana
Hands-on Prometheus and/or Grafana
Kubernetes The Opportunity As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms.
This role leads the design and implementation of observability, automation, incident response, and operational best practices across cloud and air-gapped environments, while partnering closely with DevOps, infrastructure, and security teams to improve system resilience and reduce operational risk. The Lead SRE also drives root cause analysis, capacity planning, reliability standards, and continuous improvement initiatives to support highly available, efficient, and scalable services.
This is your chance to further your skills in cloud infrastructure and technologies while continuing to grow your SRE experience. You'll build and support a reliable site for the environment in order to meet the development and maintenance requirements of systems and platforms. Work with the development and operation teams to evaluate the health, stability and reliability of systems and platforms. Design and develop technical tools to debug problems that occur in the deployment of applications, within specific platforms and systems.
Join our efforts to strengthen our security posture and safeguard national interests. Qualifications 8+ years of experience with monitoring, logging, and observability platforms, such as Prometheus, Grafana, and ELK stack
8+ years of experience with Linux systems administration and networking fundamentals within AWS
Experience with Python scripting and automation
Experience with Infrastructure as Code using Terraform and Terragrunt
Knowledge of Kubernetes administration, troubleshooting, and operations.
TS/SCI clearance with a polygraph
Bachelor’s degree and 8+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering, or 12+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering in lieu of a degree
Ability to obtain a Security+ CE, SSCP, CCNA-Security, or GSEC Certification within 6 months of start date Nice to Have Skills Experience with deploying and managing OpenTelemetry.
Experience with AWS CloudWatch, AWS EKS, and related AWS services
Experience managing Kubernetes environments through Rancher
Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management
Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence Knowledge of distributed systems, microservices architectures, and containerized workloads
Knowledge of NIST 800-53 and NIST-190 Master’s degree in a relevant field
Security+ CE, SSCP, CCNA-Security, or GSEC Certification If interested, Please reach out to us at 202.828.3494 or email at or
Similar jobs
- AP
Network Automation Engineer
NewApton Inc
Alpharetta, GA🇺🇸On-site21 hours agoTCP/IPAnsibleGit+3Technology - AT
Sr. DevSecOps Engineer with Security Clearance
NewArena Technical Resources
Sterling, VA🇺🇸$210k - $225k/yrHybrid21 hours agoAgileJiraEngineering - DS
Sr Azure DevOps Cloud Engineer
NewDia Software Solutions
United States🇺🇸Hybrid21 hours agoDockerMicroservicesSQL+7Technology - SO
Azure Integration & DevOps Engineer
System One
Arlington, VA🇺🇸$90/hrHybrid5 weeks agoAzureC#Git+5Technology - WI
Azure DevOps Cloud Engineer (812275)
NewWiserHunt Inc.
Mechanicsville, VA🇺🇸Hybrid21 hours agoDockerMicroservicesSQL+7Technology - SM
Azure DevOps Cloud Engineer
NewSmallArc, Inc
Mechanicsville, VA🇺🇸Hybrid21 hours agoDockerMicroservicesSQL+6Technology