Haystack
← Back to Jobs
Technology
XG

Information Technology - DevOps Engineer (Cloud Engineer)

Xcelo Group IncChicago, IL🇺🇸United StatesPosted 3 Sept 2026

Why This Role Stands Out

This hybrid DevOps Engineer role offers a fantastic opportunity to build and operate resilient cloud platforms, focusing on cutting-edge cyber recovery solutions. You'll thrive here if you have extensive experience in enterprise backup, SRE principles, and a passion for automation. Apply now to contribute to a secure and highly available infrastructure!

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Chicago, IL, United States
Posted
Yesterday
AWSELKEncryptionMFASplunkActive DirectoryAnsibleAzureBashGitHub ActionsGoogle CloudGrafanaKubernetesPowerShellPrometheusPythonTerraformVMwareVaultZero Trust

Job Description

Hi,



We are having a immediate requirement for the below mentioned role:



Job Title:Information Technology - DevOps Engineer (Cloud Engineer)

Work Location: Chicago, IL,

Work Mode: Hybrid

Work Auth:All Visa Accepted(No H1 and No Fake Profile)

Job Summary

We are seeking a highly experienced Senior Site Reliability Engineer (SRE) with strong expertise in enterprise backup engineering, cyber recovery, infrastructure resiliency, automation, and observability.

The ideal candidate will design and operate secure, highly available recovery platforms across on-premises, cloud, virtualized, and containerized environments, with a strong focus on ransomware resilience, immutable backups, cyber recovery vaults, recovery automation, and SRE best practices.

Required Education

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent professional experience.


Required Experience

  • 7+ years of experience in Backup Engineering, Infrastructure Engineering, Site Reliability Engineering, or related infrastructure roles.

  • 5+ years designing and supporting enterprise backup and recovery solutions.

  • 3+ years supporting cyber recovery, cyber resiliency, or ransomware recovery architectures.

  • Hands-on experience implementing SRE principles, automation, reliability engineering, monitoring, and operational resilience.

  • Strong knowledge of distributed systems, high availability, disaster recovery, RPO/RTO, and enterprise infrastructure architecture.


Backup & Recovery Technologies

Hands-on experience with one or more of the following:

  • Cohesity

  • Dell PowerProtect Data Manager

  • Dell Data Domain

  • Dell Cyber Recovery

  • Rubrik

  • Commvault

  • Veritas NetBackup

  • Veeam


Cyber Recovery & Resiliency

Strong experience with:

  • Air-gapped cyber recovery vaults

  • Immutable backups and immutable storage

  • Clean Rooms

  • Isolated Recovery Environments (IRE)

  • Recovery orchestration

  • Cyber resilience testing

  • Ransomware recovery

  • Recovery validation

  • Bare Metal Recovery (BMR)

  • Secure recovery workflows

  • Disaster Recovery testing

  • RPO/RTO validation


Cloud & Infrastructure

Experience supporting:

  • Microsoft Azure

  • Amazon Web Services (AWS)

  • Google Cloud Platform (Google Cloud Platform)

  • Cloud-native backup and recovery

  • Cross-region recovery

  • Hybrid-cloud resiliency

  • VMware

  • Hyper-V

  • Kubernetes

  • OpenShift

  • Linux

  • Windows Server

  • Active Directory

  • Enterprise storage platforms


Automation / Infrastructure as Code

Strong hands-on experience with:

  • Ansible

  • Terraform

  • Python

  • PowerShell

  • Bash

  • GitHub

  • GitHub Actions

  • CI/CD pipelines

  • Infrastructure as Code (IaC)

  • Recovery as Code


Observability & Monitoring

Experience with:

  • Dynatrace

  • Grafana

  • Prometheus

  • Splunk

  • ELK Stack

  • ServiceNow


Security & Compliance

Strong understanding of:

  • Zero Trust architecture

  • NIST Cybersecurity Framework

  • CIS Controls

  • Encryption and key management

  • Identity and Access Management (IAM)

  • Multi-Factor Authentication (MFA)

  • Secure recovery processes


Key Responsibilities

  • Design, engineer, and maintain highly available enterprise backup and recovery platforms using SRE principles.

  • Define and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery services.

  • Develop automation to reduce operational toil and improve platform reliability.

  • Perform Root Cause Analysis (RCA) and implement permanent corrective actions.

  • Improve platform reliability, scalability, performance, availability, and recoverability.

  • Build proactive monitoring, alerting, and observability for backup and cyber recovery platforms.

  • Participate in incident response, major incident management, and recovery operations.

  • Design and administer enterprise backup solutions across on-premises, cloud, SaaS, virtual machines, physical servers, databases, Kubernetes/OpenShift, NAS, object storage, and enterprise applications.

  • Engineer immutable backup architectures supporting ransomware resilience.

  • Optimize backup performance, retention, replication, encryption, and recovery objectives.

  • Implement policy-based backup automation and lifecycle management.

  • Ensure backup and recovery environments meet defined RPO and RTO requirements.

  • Design and implement air-gapped recovery vaults, Clean Rooms, IRE environments, and immutable storage architectures.

  • Develop secure recovery workflows for cyberattack and ransomware scenarios.

  • Automate malware scanning, recovery-point validation, and recovery readiness checks.

  • Design and test recovery orchestration for severe cyber disruption scenarios.

  • Work closely with Cyber Security teams to develop ransomware resilience strategies.

  • Develop Infrastructure as Code and Recovery as Code solutions.

  • Create automated recovery runbooks using Ansible, Terraform, Python, PowerShell, and GitHub Actions.

  • Automate recovery validation, compliance reporting, and evidence generation.

  • Implement monitoring for backup success rates, replication health, cyber vault health, recovery readiness, storage utilization, and infrastructure dependencies.

  • Build dashboards for operational teams and executive leadership.

  • Integrate backup and recovery platforms with Dynatrace, Grafana, Prometheus, Splunk, and other enterprise monitoring tools.

  • Plan and execute cyber recovery exercises, Clean Room validation, air-gap recovery testing, isolated recovery exercises, BMR testing, and Disaster Recovery testing.

  • Validate application recoverability against defined business RTO/RPO requirements.

  • Prepare executive-level reporting on cyber recovery readiness, resilience testing, risks, and remediation activities.


Preferred Qualifications

  • Experience working within financial services, banking, insurance, or another highly regulated industry.

  • Experience supporting GSIB cyber resiliency programs.

  • Understanding of regulatory expectations from organizations such as the Federal Reserve, OCC, and FFIEC.

  • Experience with chaos engineering and resilience testing.

  • Strong understanding of SRE reliability metrics and operational excellence practices.

  • Experience implementing AIOps, intelligent monitoring, or predictive analytics.

  • Excellent troubleshooting and root cause analysis skills.

  • Ability to lead cross-functional technical recovery initiatives.

  • Strong communication and executive presentation skills.

  • Proven ability to influence engineering standards and improve operational reliability.

  • Strong commitment to automation, continuous improvement, and resilience engineering.


Key Skills

SRE | Backup Engineering | Cyber Recovery | Cohesity | Dell PowerProtect | Data Domain | Dell Cyber Recovery | Rubrik | Commvault | NetBackup | Veeam | Immutable Backup | Air-Gapped Vault | Clean Room | IRE | Ransomware Recovery | RPO/RTO | AWS | Azure | Google Cloud Platform | VMware | Kubernetes | OpenShift | Ansible | Terraform | Python | PowerShell | GitHub Actions | Dynatrace | Grafana | Prometheus | Splunk | ServiceNow | Zero Trust | NIST | Disaster Recovery

Similar jobs