Haystack
← Back to Jobs
Employee
Technology

Cloud SRE Systems Engineer with Security Clearance

Integral FederalTysons, VA🇺🇸United StatesPosted 18 Aug 2026

Quick Overview

Work Type
Hybrid
Schedule
Employee
Level
Mid Senior

Job Description

Overview The Cloud Site Reliability Engineering (SRE) Systems Engineer is responsible for ensuring reliability, availability, performance, and operational excellence of VA product environments deployed in the VA Enterprise Cloud (VAEC) for the Department of Veterans Affairs (VA), Office of Information and Technology (OIT), Product Delivery Services (PDS), Benefits and Memorials (BAM), Memorial and Benefit Services (MBS) deliver secure, reliable, and effective Information Technology (IT) solutions that support the Department's mission . Responsibilities

  • Applies SRE practices, manages cloud environments, performs patching and automation, ensures compliance with required Service Level Targets (SLTs), supports deployments, handles incidents, and maintains monitoring, observability, and operational documentation
  • Maintain all environments in an operational state across VAEC, including application/system hardware 24x7x365.
  • Apply OS and application patches and perform backup and restoration activities.
  • Support software releases (often outside core hours) and validate maximum load after production deployment.
  • Manage and optimize cloud resource utilization, including capacity planning and forecasting
  • Implement SRE practices and patterns to improve reliability, performance, automation, fault tolerance, telemetry, logging, and alerting across distributed and microservices architectures.
  • Improve fault tolerance for distributed/microservices systems when components or resources become temporarily unavailable.
  • Conduct trend analysis of system performance and resource utilization to forecast future demand and identify bottlenecks.
  • Implement real ‑ time proactive monitoring dashboards enabling visibility into system health and alerting.

Qualifications Required:

  • Bachelor's Degree in computer science or IT related degree with 7-10 years' experience
  • Knowledge of modern SRE practices: automation, resilience engineering, observability, fault tolerance, SLT/SLA adherence, and incident analysis.
  • Experience producing monitoring reports, SLT assessments, postmortems, and operational dashboards
  • Experience with Jira and GitHub
  • Competency maintaining IRPs, DRPs, backup/restore plans, and executing continuity operations
  • Strong documentation and communication skills for reporting outages, SLT performance, and incident updates
  • Public Trust Preferred:
  • Experience with AWS VA Enterprise Cloud (VAEC)
  • Experience with Azure DevOps Company Overview Integral partners with federal defense, intelligence, and civilian leaders to tackle their most important challenges and deliver positive outcomes.

Since our founding in 1998, we have helped clients leverage existing and emerging technologies to transform their enterprises, empower growth, drive innovation, and build sustainable success. The forward-leaning solutions we deliver are tailored to each mission with a focus on keeping our nation safe and secure. Integral is headquartered in McLean, VA and serves clients throughout the country. We offer a comprehensive total rewards package including paid parental leave and immediate vesting in our 401(k). Give us a try and become part of a curated group of professionals at Integral Federal!

Our package also includes: · Medical, Dental && Vision Insurance · Flexible Spending Accounts · Short-Term and Long-Term Disability Insurance · Life Insurance · Paid Time Off && Holidays · Earned Bonuses && Awards · Professional Training Reimbursement · Employee Assistance Program Equal Opportunity Employer/Protected Veteran/Disability

Skills

Microservices
AWS
Azure
Jira
SAFe

Similar jobs