Quick Overview
Job Description
Senior AWS Cloud Platform Engineer
Location: Naperville, IL
Work Schedule: Hybrid – Monday & Friday Remote; Tuesday–Thursday Onsite
Job Overview
We are seeking a highly skilled Senior AWS Cloud Platform Engineer / Senior DevOps Engineer to design, deploy, automate, and operate cloud-native platforms at scale.
The ideal candidate will have strong hands-on expertise in AWS, Kubernetes, EKS/AKS, GitOps, Argo CD, GitHub Actions, Terraform, CI/CD, Infrastructure as Code, cloud networking, and platform automation.
This role goes beyond traditional operations. We are looking for a platform-minded engineer who continuously identifies opportunities to automate processes, reduce operational effort, improve reliability, and optimize technology costs.
Key Responsibilities
Kubernetes & Container Platform Engineering
- Deploy, manage, and troubleshoot workloads across AWS EKS and Azure AKS.
- Design, maintain, and upgrade Kubernetes clusters throughout their lifecycle.
- Manage Kubernetes networking, ingress controllers, service meshes, pod communication, and traffic routing.
- Troubleshoot complex issues involving pods, nodes, networking, DNS, storage, and cluster scalability.
- Implement high-availability and disaster-recovery strategies for Kubernetes platforms.
GitOps & CI/CD
- Design and manage enterprise-scale GitOps workflows using Argo CD.
- Automate application deployments and environment promotion strategies.
- Build and maintain CI/CD pipelines using GitHub Actions.
- Define and enforce Git branching, release management, and deployment practices.
- Improve deployment reliability, rollback capabilities, and release governance.
AWS Cloud Engineering
Hands-on experience with:
- Amazon EKS
- EC2
- Route 53
- IAM
- CloudFront/CDN
- Application Load Balancer (ALB) & Network Load Balancer (NLB)
- VPC & AWS Networking
- Security Groups & NACLs
- Secrets Manager
- ElastiCache / Redis
- S3
- CloudWatch
- RDS
- Auto Scaling & High Availability architectures
Responsibilities include:
- Designing secure, scalable, and highly available AWS architectures.
- Implementing disaster recovery and business continuity solutions.
- Optimizing cloud cost, performance, security, and reliability.
- Managing multi-account AWS environments.
Infrastructure as Code
- Build and maintain reusable Terraform modules.
- Provision infrastructure using Infrastructure as Code best practices.
- Maintain environment consistency and compliance through automation.
- Troubleshoot Terraform state, drift, and large-scale deployments.
Platform Automation & Engineering Excellence
- Identify manual processes that can be automated.
- Develop solutions for deployment, infrastructure, logging, monitoring, and incident-response automation.
- Build self-service engineering capabilities.
- Develop internal developer tooling to improve engineering productivity.
- Continuously evaluate opportunities to reduce operational effort and tooling costs.
Observability & Monitoring
- Design monitoring and alerting solutions using Datadog or similar platforms.
- Implement automated alert correlation and incident-reduction mechanisms.
- Create dashboards, SLOs, and observability standards.
- Drive proactive monitoring practices across production environments.
Disaster Recovery & Resiliency
- Design highly resilient cloud architectures.
- Lead recovery planning for major production outages.
- Develop recovery strategies for:
- AWS region failures
- Kubernetes cluster failures
- Account compromises
- Infrastructure loss events
- Implement recovery solutions using IaC, backups, replication, and automation.
Required Qualifications
- 4+ years of experience in DevOps, SRE, Platform Engineering, Cloud Engineering, or a related field.
- Expert-level Kubernetes administration experience.
- Strong hands-on experience with AWS EKS and/or Azure AKS.
- Deep understanding of Argo CD and GitOps.
- Strong experience with GitHub Actions and CI/CD.
- Advanced knowledge of Git branching and release management.
- Strong AWS architecture and operational experience.
- Advanced Terraform / Infrastructure as Code skills.
- Strong Docker and containerization experience.
- Linux system administration experience.
- Strong cloud networking and security knowledge.
- Experience designing highly available and resilient cloud platforms.
Preferred Qualifications
- Azure / AKS experience.
- Service mesh technologies such as Istio or Linkerd.
- Multi-cloud platform experience.
- Datadog or comparable observability platforms.
- Programming/scripting experience with Python, Go, Bash, or Shell.
- Previous Platform Engineering experience.
- Experience with automation and internal developer platforms.
Similar jobs
- PE
HPC Software Deployment Configuration Manager, Lead with Security Clearance
Peraton
Fort Meade, MD🇺🇸$104k - $166k/yrHybrid3 weeks agoHTTPSTechnology - BI
Dev/Sec/Ops Platform Engineer
NewBigBear.ai
McLean, VA🇺🇸HybridYesterdayAWSAnsibleAzure+5Technology - ST
SRE - Nexus.corp /SonarQube Engineer
NewSPECTRAFORCE TECHNOLOGIES Inc.
United States🇺🇸RemoteYesterdayAWSSonarQubeAnsible+7Technology - AS
Senior Power Platform Engineer
Apex Systems
Charlotte, NC🇺🇸Hybrid2 weeks agoSQLJavaScriptPower BI+1Technology - AS
Systems Engineer
Apex Systems
Alexandria, VA🇺🇸$110k/yrOn-site7 weeks agoKubernetesTechnology - TT
Platform Engineer (Google Cloud Platform)
NewTrigyn Technologies, Inc.
Albany, NY🇺🇸HybridYesterdayAWSSonarQubeSplunk+6Technology