Haystack
← Back to Jobs
Technology

Senior Site Reliability Engineer

GovCIOUnited States🇺🇸United StatesPosted 11 Aug 2026

Why This Role Stands Out

You'll thrive as a Senior Site Reliability Engineer at GovCIO, designing and implementing resilient, scalable cloud infrastructure with a highly competitive salary of $210,000-$230,000. This hybrid role offers significant opportunities for career growth and skill development in automation and multi-cloud environments, perfect for those who enjoy bridging development and operations. You'll be instrumental in ensuring system reliability and performance, making this an exciting opportunity to make a substantial impact.

Quick Overview

Salary
$210k - $230k/yr
Work Type
Hybrid
Level
Mid Senior

Job Description

GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation, reliability, and performance optimization across multi-cloud environments. This position is located in Arlington, VA and is a hybrid remote/onsite position.

Responsibilities

Key Responsibilities:

Infrastructure & Automation
Design, deploy, and manage cloud infrastructure using Infrastructure as Code (IaC) principles
Develop and maintain Terraform modules for AWS and Azure environments
Create and manage Ansible playbooks for configuration management and application deployment
Implement CI/CD pipelines using GitHub Actions to automate build, test, and deployment processes
Implement GitOps workflows for declarative infrastructure and application delivery
Build self-service tools and platforms to enable development teams

Reliability & Performance
Establish and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
Implement comprehensive monitoring, logging, and alerting solutions
Conduct capacity planning and performance tuning
Perform root cause analysis and implement preventive measures
Design and execute chaos engineering experiments to validate system resilience

Disaster Recovery & Business Continuity
Design and implement disaster recovery strategies across multi-cloud environments
Develop and maintain backup and restore procedures
Create and test business continuity plans
Implement automated failover mechanisms
Document recovery time objectives (RTO) and recovery point objectives (RPO)

Cloud Operations
Manage and optimize AWS services (EC2, S3, RDS, Lambda, ECS, EKS, CloudWatch, etc.)
Manage and optimize Azure services (VMs, Storage, SQL Database, AKS, Monitor, etc.)
Implement cost optimization strategies and resource tagging
Ensure security best practices and compliance requirements
Manage identity and access management (IAM) policies

Collaboration & Leadership
Participate in on-call rotation and incident response
Collaborate with development teams on architecture and design decisions
Mentor team members on SRE practices and tools
Document systems, processes, and runbooks
Drive continuous improvement initiatives

Qualifications

Required Education and Experience
Bachelor's Degree with 12+ yrs experience
Clearance Level: Active Secret with the ability to obtain and hold DEA suitability

Technical Skills
Cloud Platforms: 3+ years of hands-on experience with AWS and Azure
Infrastructure as Code: Expert-level proficiency with Terraform
Configuration Management: Strong experience with Ansible
Scripting: Proficiency in Python, Bash, or PowerShell
Containerization: Experience with Docker and Kubernetes
Version Control: Strong Git and GitHub workflow knowledge
GitOps: Experience implementing GitOps practices and workflows
Monitoring Tools: Experience with Prometheus, Grafana, ELK Stack, or similar
CI/CD: Hands-on experience with GitHub Actions, Jenkins, GitLab CI, or Azure DevOps

Core Competencies
Deep understanding of Microsoft/Linux systems administration
Strong networking knowledge (TCP/IP, DNS, load balancing, VPN)
Experience with database administration (PostgreSQL, MySQL, SQL Server)
Knowledge of security best practices and compliance frameworks
Understanding of microservices architecture and distributed systems
Experience with disaster recovery planning and execution

Soft Skills
Excellent problem-solving and analytical abilities
Strong communication skills, both written and verbal
Ability to work independently and in team environments
Customer-focused mindset with emphasis on reliability
Adaptability to rapidly changing technologies and requirements

Preferred Qualifications
AWS Certified Solutions Architect or SysOps Administrator
Azure Administrator or Solutions Architect certification
Certified Kubernetes Administrator (CKA)
HashiCorp Certified: Terraform Associate
GitHub Certified or demonstrated expertise with GitHub Enterprise
Experience with service mesh technologies (Istio, Linkerd)
Knowledge of observability platforms (Datadog, New Relic, Dynatrace)
Experience with GitOps tools and practices (ArgoCD, Flux, GitHub Actions for GitOps)
Familiarity with compliance frameworks (SOC 2, HIPAA, FedRAMP)
Previous experience in a DevOps or Platform Engineering role

Posted Salary Range

USD $210,000.00 - USD $230,000.00 /Yr.

Skills

Docker
Microservices
MySQL
SQL
SQL Server
AWS
ELK
Load Balancing
New Relic
SOC 2
Service Mesh
TCP/IP
Ansible
ArgoCD
Azure
Bash
DNS
Datadog
Git
GitHub Actions
GitLab CI
Grafana
HIPAA
Istio
Jenkins
Kubernetes
PostgreSQL
PowerShell
Prometheus
Python
Terraform

Similar jobs