← Back to Jobs
Technology
Site Reliability Engineer (SRE) Cloud Migration & Operational Excellence
Infinite Computer Solutions (ICS)United States🇺🇸United StatesPosted 10 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
| Site Reliability Engineer (SRE) Cloud Migration & Operational Excellence As a Site Reliability Engineer (SRE), you will play a critical role in the migration of business-critical applications from on-premises infrastructure to AWS and Google Cloud Platform (Google Cloud Platform). You will work at the intersection of software engineering, cloud infrastructure, platform engineering, observability, automation, and operations to ensure migrated applications are secure, scalable, reliable, and supportable. This role is ideal for a hands-on engineer who combines strong infrastructure and cloud expertise with software engineering practices. You will help establish reliability standards, automate operational processes, implement observability solutions, and drive operational excellence throughout the migration and modernization journey. What You'll Do Reliability Engineering & Cloud Operations Design and implement reliability strategies supporting migration of enterprise applications to AWS and Google Cloud Platform. Establish Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics for critical business services. Partner with application, infrastructure, database, security, and cloud architecture teams to improve platform stability and operational readiness. Develop operational runbooks, recovery procedures, and automated response mechanisms. Observability & Monitoring Design and implement enterprise monitoring and observability solutions. Develop dashboards, alerts, health checks, synthetic monitoring, and operational reporting. Utilize Dynatrace, Splunk, Moogsoft, CloudWatch, and Google Cloud Platform Operations Suite to improve visibility into application and infrastructure performance. Create proactive alerting strategies that reduce operational risk and improve incident response. Infrastructure Automation & Platform Engineering Develop and maintain Infrastructure as Code (IaC) using Terraform. Support enterprise landing zones, cloud provisioning, configuration management, and platform automation. Implement CI/CD and GitOps practices supporting cloud-native deployments. Partner with cloud platform teams to standardize reusable infrastructure patterns. Incident Response & Resiliency Participate in production support, incident response, and root cause analysis activities. Lead efforts to reduce recurring incidents through automation and engineering improvements. Develop resiliency testing, failover validation, disaster recovery, and operational readiness procedures. Drive reliability improvements through performance testing, capacity planning, and fault analysis. AI-Assisted Operations Leverage Copilot, Codex, and similar AI-enabled tools to accelerate automation development, observability configuration, scripting, troubleshooting, documentation, and operational analysis. Use AI to improve incident analysis, runbook generation, alert optimization, and reliability engineering workflows. ________________________________________ Required Qualifications Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience. 8+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, Infrastructure Engineering, or Systems Engineering. Demonstrated experience supporting migrations from on-premises environments to AWS and/or Google Cloud Platform. Deep hands-on experience with Terraform and Infrastructure as Code practices. Experience working within enterprise landing zones, cloud governance models, IAM, networking, security controls, and operational guardrails. Experience with Splunk, Dynatrace, and Moogsoft. Strong scripting and automation experience using PowerShell, Python, Bash, or similar technologies. Experience with GitLab, CI/CD pipelines, deployment automation, and DevSecOps practices. Strong understanding of cloud networking, DNS, load balancing, certificates, secrets management, and security best practices. Experience using Jira and ServiceNow in enterprise operational environments. ________________________________________ Preferred Qualifications Experience within fintech, payments, banking, or highly regulated industries. AWS, Google Cloud Platform, Terraform, Kubernetes, or SRE certifications. Experience with Kubernetes, containers, service mesh, and cloud-native architectures. Familiarity with SQL Server, database performance monitoring, and application dependency mapping. Experience supporting enterprise-scale production systems with stringent uptime requirements. ________________________________________ Technologies You'll Work With AWS, Google Cloud Platform (Google Cloud Platform) ,Terraform, GitLab, Jira, ServiceNow Dynatrace, Splunk, Moogsoft Kubernetes, PowerShell, Python, CloudWatch Google Cloud Platform Operations Suite Enterprise Landing Zones CI/CD & DevSecOps Toolchains ________________________________________ Success in the First 12 Months Establish SRE standards and operational readiness frameworks for migrated cloud workloads. Implement automated monitoring, alerting, and incident response capabilities. Improve application availability, reliability, observability, and recovery times. Deliver reusable Terraform modules and operational automation assets. Reduce operational toil through automation and AI-assisted engineering practices. Build strong partnerships across cloud engineering, development, infrastructure, database, security, and operations teams. Help ensure cloud migrations are completed with minimal disruption and maximum reliability. Ideal Candidate A highly technical Site Reliability Engineer who combines cloud infrastructure expertise, Terraform-based automation, enterprise observability, and software engineering practices to ensure business-critical applications operate reliably in AWS and Google Cloud Platform. The ideal candidate has experience with enterprise landing zones, Splunk, Dynatrace, Moogsoft, GitLab, Jira, and ServiceNow, and actively leverages AI-assisted tools to improve operational efficiency, automation, and system reliability. |
Skills
SQL
SQL Server
AWS
Load Balancing
Service Mesh
Splunk
Bash
DNS
Google Cloud
Jira
Kubernetes
PowerShell
Python
Terraform
Similar jobs
Azure DevOps/ AWS Engineer Expert
Kodi Inc · Columbus, United States
24 minutes agoSalesforce Platform Engineer
Indus River Technologies Inc. · Dickson, United States
24 minutes agoSenior Azure DevOps Engineer
AgreeYa Solutions · Phoenix, United States
25 minutes agoAkamai Platform Engineer- Independent candidates w2 contract
Pull Skill Technologies · Fort Worth, United States
26 minutes agoSite Reliability Engineer
HPTech Inc. · Secaucus, United States
26 minutes agoMuleSoft Admin/Mulesoft Platform Engineer with DevOps
NMK Global Inc. · Wilmington, United States
57 minutes ago