Haystack
← Back to Jobs
Technology

Site Reliability Engineer (SRE) Cloud Migration & Operational Excellence

Infinite Computer Solutions (ICS)United States🇺🇸United StatesPosted 10 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Site Reliability Engineer (SRE) Cloud Migration & Operational Excellence
As a Site Reliability Engineer (SRE), you will play a critical role in the migration of business-critical applications from on-premises infrastructure to AWS and Google Cloud Platform (Google Cloud Platform). You will work at the intersection of software engineering, cloud infrastructure, platform engineering, observability, automation, and operations to ensure migrated applications are secure, scalable, reliable, and supportable.
This role is ideal for a hands-on engineer who combines strong infrastructure and cloud expertise with software engineering practices. You will help establish reliability standards, automate operational processes, implement observability solutions, and drive operational excellence throughout the migration and modernization journey.
What You'll Do
Reliability Engineering & Cloud Operations
Design and implement reliability strategies supporting migration of enterprise applications to AWS and Google Cloud Platform.
Establish Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics for critical business services.
Partner with application, infrastructure, database, security, and cloud architecture teams to improve platform stability and operational readiness.
Develop operational runbooks, recovery procedures, and automated response mechanisms.
Observability & Monitoring
Design and implement enterprise monitoring and observability solutions.
Develop dashboards, alerts, health checks, synthetic monitoring, and operational reporting.
Utilize Dynatrace, Splunk, Moogsoft, CloudWatch, and Google Cloud Platform Operations Suite to improve visibility into application and infrastructure performance.
Create proactive alerting strategies that reduce operational risk and improve incident response.
Infrastructure Automation & Platform Engineering
Develop and maintain Infrastructure as Code (IaC) using Terraform.
Support enterprise landing zones, cloud provisioning, configuration management, and platform automation.
Implement CI/CD and GitOps practices supporting cloud-native deployments.
Partner with cloud platform teams to standardize reusable infrastructure patterns.
Incident Response & Resiliency
Participate in production support, incident response, and root cause analysis activities.
Lead efforts to reduce recurring incidents through automation and engineering improvements.
Develop resiliency testing, failover validation, disaster recovery, and operational readiness procedures.
Drive reliability improvements through performance testing, capacity planning, and fault analysis.
AI-Assisted Operations
Leverage Copilot, Codex, and similar AI-enabled tools to accelerate automation development, observability configuration, scripting, troubleshooting, documentation, and operational analysis.
Use AI to improve incident analysis, runbook generation, alert optimization, and reliability engineering workflows.
________________________________________
Required Qualifications
Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience.
8+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, Infrastructure Engineering, or Systems Engineering.
Demonstrated experience supporting migrations from on-premises environments to AWS and/or Google Cloud Platform.
Deep hands-on experience with Terraform and Infrastructure as Code practices.
Experience working within enterprise landing zones, cloud governance models, IAM, networking, security controls, and operational guardrails.
Experience with Splunk, Dynatrace, and Moogsoft.
Strong scripting and automation experience using PowerShell, Python, Bash, or similar technologies.
Experience with GitLab, CI/CD pipelines, deployment automation, and DevSecOps practices.
Strong understanding of cloud networking, DNS, load balancing, certificates, secrets management, and security best practices.
Experience using Jira and ServiceNow in enterprise operational environments.
________________________________________
Preferred Qualifications
Experience within fintech, payments, banking, or highly regulated industries.
AWS, Google Cloud Platform, Terraform, Kubernetes, or SRE certifications.
Experience with Kubernetes, containers, service mesh, and cloud-native architectures.
Familiarity with SQL Server, database performance monitoring, and application dependency mapping.
Experience supporting enterprise-scale production systems with stringent uptime requirements.
________________________________________
Technologies You'll Work With
AWS, Google Cloud Platform (Google Cloud Platform) ,Terraform, GitLab, Jira, ServiceNow
Dynatrace, Splunk, Moogsoft
Kubernetes, PowerShell, Python, CloudWatch
Google Cloud Platform Operations Suite
Enterprise Landing Zones
CI/CD & DevSecOps Toolchains
________________________________________
Success in the First 12 Months
Establish SRE standards and operational readiness frameworks for migrated cloud workloads.
Implement automated monitoring, alerting, and incident response capabilities.
Improve application availability, reliability, observability, and recovery times.
Deliver reusable Terraform modules and operational automation assets.
Reduce operational toil through automation and AI-assisted engineering practices.
Build strong partnerships across cloud engineering, development, infrastructure, database, security, and operations teams.
Help ensure cloud migrations are completed with minimal disruption and maximum reliability.
Ideal Candidate
A highly technical Site Reliability Engineer who combines cloud infrastructure expertise, Terraform-based automation, enterprise observability, and software engineering practices to ensure business-critical applications operate reliably in AWS and Google Cloud Platform. The ideal candidate has experience with enterprise landing zones, Splunk, Dynatrace, Moogsoft, GitLab, Jira, and ServiceNow, and actively leverages AI-assisted tools to improve operational efficiency, automation, and system reliability.

Skills

SQL
SQL Server
AWS
Load Balancing
Service Mesh
Splunk
Bash
DNS
Google Cloud
Jira
Kubernetes
PowerShell
Python
Terraform

Similar jobs