Quick Overview
Job Description
Job Summary
We are looking for a Senior Cloud Infrastructure Engineer – Data Platform to join our team in supporting and improving highly reliable AWS infrastructure and production data platforms.
The ideal candidate will be a hands-on infrastructure/platform engineer with strong experience in AWS, Terraform, Linux, networking, containers, automation, and production reliability. The role will involve infrastructure delivery, automation, troubleshooting, incident response, and production support, along with supporting data platform workloads.
Key Responsibilities
Design, implement, maintain, and troubleshoot AWS cloud infrastructure across production environments.
Manage AWS services including VPC, IAM, ECS/Fargate, Lambda, S3, and RDS.
Develop and maintain reusable Terraform modules, remote state, environment separation, and infrastructure-as-code practices.
Support infrastructure changes through CI/CD pipelines, code reviews, and safe deployment/rollback procedures.
Monitor production environments, respond to incidents, perform root cause analysis, and implement permanent fixes.
Troubleshoot Linux systems, DNS, TLS, routing, load balancing, connectivity, and resource utilization issues.
Manage containerized workloads using Docker and AWS ECS/Fargate.
Develop automation using Python and Shell scripting to reduce manual operational activities.
Implement and maintain security, access management, secrets management, backup, recovery, and disaster recovery practices.
Perform capacity planning, resource optimization, and AWS cost optimization.
Support production data platforms and troubleshoot issues involving Databricks/Redshift, Airflow, dbt, and SQL-based workloads.
Investigate data pipeline failures, data freshness issues, dependencies, retries, backfills, and recovery procedures.
Collaborate with engineering and data teams to improve platform reliability, scalability, and operational efficiency.
Required Skills
9+ years of experience in Cloud Infrastructure, SRE, Platform Engineering, DevOps, or a related field.
Strong hands-on experience with AWS in production environments.
Strong Terraform experience, including modules, remote state, and environment management.
Strong knowledge of Linux and cloud networking.
Experience with Docker and containerized environments.
Experience with Python and/or Shell scripting for automation.
Strong understanding of IAM, security, secrets management, backups, and disaster recovery.
Experience with production support, incident response, troubleshooting, and root cause analysis.
Experience with monitoring, alerting, reliability, capacity planning, and performance optimization.
Experience supporting production data workloads with SQL and data operations.
Similar jobs
- KT
Order Management Cloud Engineer
NewKRG Technologies Inc
Raleigh, NC🇺🇸Hybrid19 hours agoTechnology - LE
Workplace Network Engineer
NewLegora
New York City🇺🇸On-site3 hours agoACLSCapacity PlanningDNS+5Technology - SA
Senior Staff Network Engineer (R4843) with Security Clearance
NewShield AI Inc
Dallas, TX🇺🇸Hybrid19 hours agoAnsibleTerraformZero TrustTechnology - PR
Senior Network Engineer with Security Clearance
NewPrism, Inc.
Alexandria, VA🇺🇸Hybrid19 hours agoAnsibleZero TrustTechnology - PR
Network Engineer SME with Security Clearance
NewPrism, Inc.
Alexandria, VA🇺🇸Hybrid19 hours agoAnsiblePKIPython+1Technology - SM
Azure Cloud Engineer
NewSmallArc, Inc
United States🇺🇸Hybrid19 hours agoSSOAzurePowerShell+1Technology