Haystack
← Back to Jobs
Full time
Technology
VA

Site Reliability / Operations Engineer

VantorHerndon, Virginia🇺🇸United StatesPosted 11 Aug 2026

Why This Role Stands Out

This hybrid role at Vantor offers a fantastic opportunity to shape secure, scalable cloud infrastructure for impactful digital transformation projects, leveraging your expertise in SRE and DevOps to drive innovation. You'll thrive here if you're a mid-senior engineer passionate about automation, collaboration, and continuous improvement within a low-ego, agile team environment. Apply now to contribute to cutting-edge technology and grow your career in a dynamic setting.

Quick Overview

Seniority
Mid Senior
Employment type
Full Time
Work mode
Hybrid
Location
Herndon, Virginia, United States
DockerGCPAWSELKAgileAzureBashGrafanaKubernetesPrometheusPythonTerraform

Job Description

Vantor seeks a Site Reliability / Operations Engineer to design, automate, and operate secure, scalable cloud infrastructure for modern web and data platforms in a TS/SCI environment. You'll build and manage CI/CD pipelines, infrastructure as code, containers, and Kubernetes; implement monitoring, logging, and alerting; and lead incident response to keep services reliable and performant. You'll collaborate closely with developers and clients in an agile, low ego culture, champion automation and best practices, and help evolve our SRE and DevOps standards across impactful digital transformation projects.

Responsibilities

  • Design, build, and maintain secure, scalable cloud infrastructure for client applications.
  • Implement and manage infrastructure as code, CI/CD pipelines, and automation tooling.
  • Operate and optimize containerized and Kubernetes-based platforms.
  • Set up and refine monitoring, logging, and alerting for reliability and performance.
  • Lead and participate in incident response, troubleshooting, and on-call rotations.
  • Collaborate with development teams to embed SRE and Dev
  • Ops best practices.
  • Harden systems and configurations to meet TS/SCI security and compliance requirements.
  • Continuously improve reliability, scalability, and cost efficiency of services.
  • Document architectures, runbooks, and operational procedures.
  • Contribute to evolving SRE standards, tooling, and knowledge sharing across teams.

Required Skills

  • Site Reliability Engineering (SRE)
  • Linux/Unix administration
  • Cloud platforms (AWS/Azure/GCP)
  • Infrastructure as Code (Terraform/Cloud
  • Formation)
  • CI/CD pipelines
  • Containerization (Docker)
  • Kubernetes operations
  • Monitoring and observability (Prometheus/Grafana/ELK)
  • Scripting (Python/Bash)
  • Incident response and on-call operations

Similar jobs