Why This Role Stands Out
This hybrid Systems Architect role offers significant growth potential by allowing you to lead critical reliability initiatives and develop cutting-edge automation skills within a reputable technology company. If you thrive on complex problem-solving and collaborating with diverse engineering teams, this position is an excellent opportunity to advance your career.
Quick Overview
Job Description
Role Summary
The Senior Site Reliability Engineer (SRE) supports complex cloud-based environments with a focus on reliability, resilience, and recoverability. This role leads disaster recovery exercises, validates platform rebuild procedures, develops infrastructure automation, and supports modern deployment practices. The position collaborates with cross-functional engineering teams to enhance platform reliability, improve automation, and support large-scale cloud modernization and continuity initiatives.
Responsibilities
- Support full-lifecycle platform portability and disaster recovery (DR) exercises, including validation of platform rebuild procedures and recovery playbooks.
- Execute infrastructure-level recovery activities and validate the ability to rebuild platforms within established recovery objectives.
- Verify data completeness, integrity, and accuracy during recovery exercises and document findings and remediation recommendations.
- Identify gaps across infrastructure, deployment automation, monitoring, and data recovery processes and collaborate with engineering teams to implement corrective actions.
- Design, implement, and support Infrastructure as Code (IaC) workflows using technologies such as Terraform and AWS CloudFormation.
- Build and maintain standardized CI/CD pipelines and automated deployment processes.
- Manage and optimize Kubernetes clusters and containerized workloads, including provisioning, scaling, and reliability improvements.
- Develop and maintain observability solutions using CloudWatch, Datadog, or comparable monitoring and alerting technologies.
- Develop automation, tooling, and scripts using languages such as Python or Java to reduce manual processes and improve operational consistency.
- Collaborate with platform engineering, security, application, and data teams to support secure, reliable, and consistent platform operations.
- Participate in on-call rotations, root-cause analysis, and incident response activities to improve system resilience and operational effectiveness.
Required Qualifications
- Bachelor’s degree with 7 to 10 years of relevant experience in Site Reliability Engineering, DevOps, cloud engineering, infrastructure engineering, or a related discipline; master’s degree with 5 to 8 years of relevant experience; or equivalent professional experience in lieu of a degree.
- Advanced hands-on experience with AWS services across compute, networking, storage, identity and access management, and serverless technologies.
- Strong experience with Infrastructure as Code technologies such as Terraform and AWS CloudFormation.
- Experience developing CI/CD deployment pipelines and progressive delivery processes using GitHub Actions or comparable technologies.
- Strong knowledge of Kubernetes administration, container orchestration, and containerized deployments.
- Experience validating disaster recovery processes, performing system recovery or rebuild activities, and conducting data integrity checks.
- Experience creating and maintaining dashboards, metrics, logs, and alerts using CloudWatch, Datadog, or comparable observability platforms.
- Proficiency with programming or scripting languages such as Python, Java, C#, or Go.
- Demonstrated ability to troubleshoot complex infrastructure and distributed-system issues, including networking, cascading failures, backpressure, and consistency challenges.
- Strong analytical, problem-solving, documentation, and communication skills.
- Ability to work effectively in a fast-paced environment supporting business-critical systems.
Preferred Qualifications
- Relevant AWS DevOps, DevSecOps, Solutions Architect, SysOps Administrator, Developer, or comparable cloud certifications.
- Kubernetes certifications such as CKA or CKAD.
- Familiarity with Zero Trust security principles and cloud security best practices.
- Experience with GitLab, Jenkins, or comparable CI/CD platforms.
- Experience working within highly regulated or complex enterprise environments.
- Experience supporting large-scale cloud modernization initiatives.
- Experience supporting disaster recovery exercises, continuity planning, platform recovery, or readiness assessments.
Additional Requirements
- Ability to obtain and maintain the required Public Trust determination.
- Ability to participate in an on-call rotation as required.
- Ability to support authorized overtime or non-standard work hours based on business and project requirements.
Similar jobs
- NG
Linux Systems Architect - Level 2 (AHT) with Security Clearance
NewNorthrop Grumman
Redondo Beach, CA🇺🇸$91.8k - $137.6k/yrHybrid21 hours agoAgileTechnology - CS
Software Architect- (Senior/Expert Java + IBM WebSphere Architect + Government Taxation experience)
NewCyber Sphere LLC
Albany, NY🇺🇸Hybrid21 hours agoMicroservicesSQLSpring+3Technology - EI
Network Architect
NewEcho IT Solutions, Inc.
New York, NY🇺🇸Hybrid21 hours agoAWSAzureGoogle CloudTechnology - CE
ServiceNow Technical Architect
NewCentraprise Corp
Houston, TX🇺🇸Hybrid21 hours agoServiceNowStakeholder ManagementTechnology - ST
Network Architect SME with Security Clearance
NewSarela Technology Solutions
Fort Belvoir, VA🇺🇸$195k - $205k/yrHybrid21 hours agoPythonTerraformAnsible+3Technology - TA
Mission Architect, Rendezvous and Proximity Operations
NewTrue Anomaly
Denver🇺🇸Hybrid8 hours agoMATLABC++Greenhouse+2