Haystack
← Back to Jobs
Technology
AT

DevOps Engineer

Aivanta Tech IncChicago, IL🇺🇸United StatesPosted 8 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Chicago, IL, United States
Posted
21 hours ago
AWSELKEncryptionMFASplunkAnsibleAzureBashGrafanaKubernetesPowerShellPrometheusPythonTerraformVMwareVault

Job Description

Required Education
• Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent work experience

Required Experience
• 7+ years in Backup Engineering, Infrastructure Engineering, or Site Reliability Engineering
• 5+ years designing enterprise backup solutions
• 3+ years supporting cyber recovery architectures
• Experience implementing SRE principles within enterprise infrastructure environments
• Strong understanding of distributed systems and high availability architectures
• Cohesity
• Dell PowerProtect Data Manager
• Dell Data Domain
• Dell Cyber Recovery
• Rubrik
• Commvault
• Veritas NetBackup
• Veeam
• Air-gapped vaults
• Immutable backups
• Clean Rooms
• Isolated Recovery Environments (IRE)
• Recovery orchestration
• Cyber resilience testing
• Ransomware recovery
• Recovery validation
• Microsoft Azure
• AWS
• Google Cloud Platform
• Cloud-native backup
• Cross-region recovery
• Hybrid cloud resiliency
• VMware
• Hyper-V
• Kubernetes
• OpenShift
• Linux
• Windows Server
• Active Directory
• Enterprise storage platforms
• Ansible
• Terraform
• Python
• PowerShell
• Bash
• GitHub
• GitHub Actions
• CI/CD pipelines
• Dynatrace
• Grafana
• Prometheus
• Splunk
• ELK Stack
• ServiceNow
• Zero Trust architecture
• NIST Cybersecurity Framework
• CIS Controls
• Encryption and key management
• Identity and Access Management (IAM)
• Multi-factor authentication (MFA)
• Secure recovery processes

Preferred Qualifications
• Experience in financial services or another highly regulated industry
• Experience supporting GSIB cyber resiliency programs
• Knowledge of regulatory expectations from agencies such as the Federal Reserve, OCC, or FFIEC
• Experience with chaos engineering and resilience testing
• Familiarity with SRE tooling and reliability metrics
• Experience implementing AI-assisted operations (AIOps) and predictive analytics
• Strong systems thinking and engineering mindset
• Excellent troubleshooting and root cause analysis skills
• Ability to lead cross-functional technical recovery efforts
• Strong communication and executive presentation skills
• Proven ability to influence engineering standards and drive operational excellence
• Commitment to continuous improvement through automation and reliability engineering

Job Description – Project Overview
• Seeking a highly technical Senior Site Reliability Engineer (SRE) with deep expertise in enterprise backup engineering, cyber recovery, and platform resiliency
• Responsible for engineering highly available, secure, and automated recovery capabilities that protect against operational failures, ransomware, and other cyber threats
• Combines traditional SRE principles (automation, observability, reliability engineering, and resilience) with experience designing and operating enterprise backup platforms, immutable storage, air-gapped cyber vaults, isolated recovery environments (IREs), and recovery orchestration
• Partners closely with Infrastructure, Cyber Security, Cloud Engineering, Application Development, and Disaster Recovery teams to ensure critical services remain recoverable, resilient, and continuously validated
• Engineer and maintain highly available, resilient enterprise platforms using SRE principles
• Define and measure Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery services
• Develop automation to reduce operational toil and improve reliability
• Perform root cause analysis (RCA) and implement permanent corrective actions
• Continuously improve platform reliability, scalability, performance, and recoverability
• Establish proactive monitoring, alerting, and observability for backup and cyber recovery platforms
• Participate in incident response and major incident recovery activities
• Design, implement, and administer enterprise backup and recovery solutions across on-premises, cloud, and SaaS platforms
• Engineer immutable backup architectures that support ransomware resilience
• Design backup strategies for virtual environments, physical servers, databases, Kubernetes/OpenShift, cloud-native workloads, NAS/Object Storage, and enterprise applications
• Optimize backup performance, retention, replication, encryption, and recovery objectives
• Implement policy-based backup automation and lifecycle management
• Ensure compliance with enterprise RPO and RTO requirements
• Design and implement enterprise cyber recovery solutions including air-gapped recovery vaults, clean rooms, Isolated Recovery Environments (IRE), and immutable storage architectures
• Develop secure recovery workflows following cyberattack scenarios
• Engineer automated malware scanning and recovery validation processes
• Design and test recovery orchestration for severe-but-plausible cyber events
• Support recovery point validation and promotion into production recovery environments
• Collaborate with Cyber Security teams on ransomware resilience strategies
• Develop Infrastructure as Code (IaC) and Recovery as Code automation
• Build automated recovery runbooks using Ansible, Terraform, PowerShell, Python, and GitHub Actions
• Automate recovery validation, reporting, and compliance evidence generation
• Eliminate manual recovery processes wherever possible
• Implement monitoring for backup success rates, replication health, recovery readiness, storage utilization, cyber vault health, and infrastructure dependencies
• Build dashboards for executive and operational visibility
• Integrate with enterprise observability platforms (Dynatrace, Grafana, Splunk, Prometheus)
• Plan and execute cyber recovery exercises, clean room validation, air-gap recovery testing, full isolated recovery environment exercises, Bare Metal Recovery (BMR) testing, and Disaster Recovery testing
• Validate application recoverability against defined RTO/RPO objectives
• Produce executive reporting on recovery readiness and testing outcomes

Similar jobs