Quick Overview
Salary
$68 - $72/hr
Seniority
Mid Senior
Work mode
On Site
Location
Hampton, VA, United States
Posted
21 hours ago
DockerAWSC#CloudFormationDatadogGitHub ActionsJavaJenkinsKubernetesPythonTerraformZero Trust
Job Description
Role Summary
The Senior Site Reliability Engineer supports complex cloud environments and mission-critical systems, with a focus on platform reliability, resilience, disaster recovery, infrastructure automation, and operational readiness. This role collaborates with engineering, security, application, and data teams to maintain secure, scalable, and reliable cloud-based platforms.
Responsibilities
- Support full-lifecycle platform portability and disaster recovery exercises, including validation of platform rebuild procedures and recovery playbooks.
- Execute infrastructure recovery exercises to confirm platforms can be restored within established recovery objectives.
- Verify data completeness, integrity, and accuracy during recovery exercises and document findings and remediation recommendations.
- Identify readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes and collaborate with engineering teams on corrective actions.
- Design, implement, and support Infrastructure as Code workflows using Terraform, AWS CloudFormation, and standardized CI/CD pipelines.
- Manage and optimize Kubernetes clusters and Docker-based workloads, including provisioning, scaling, and reliability improvements.
- Develop and maintain observability solutions using CloudWatch, Datadog, and similar monitoring and alerting tools.
- Develop automation, tools, and scripts using languages such as Python or Java to reduce manual processes and improve operational consistency.
- Collaborate with platform engineering, security, application, and data teams to support secure, reliable, and compliant platform operations.
- Participate in on-call rotations, root-cause analysis, and incident response activities to improve system resilience and operational performance.
Basic Qualifications
- Bachelor’s degree and 8 to 10 years of relevant experience in Site Reliability Engineering, DevOps, cloud engineering, infrastructure engineering, or a related field; master’s degree and 6 to 8 years of relevant experience; or equivalent professional experience in lieu of a degree.
- Advanced hands-on knowledge of AWS services across compute, networking, storage, identity and access management, and serverless technologies.
- Strong experience with Infrastructure as Code tools such as Terraform and CloudFormation.
- Experience building CI/CD pipelines and progressive delivery processes using GitHub Actions or similar tools.
- Strong knowledge of Kubernetes administration, container orchestration, and Docker-based deployments.
- Experience validating disaster recovery processes, performing system rebuilds, and conducting data integrity checks.
- Experience developing dashboards, metrics, logging, and alerting solutions using CloudWatch, Datadog, or similar observability tools.
- Proficiency with programming or scripting languages such as Python, Java, C#, or Go.
- Experience troubleshooting complex system failures, including cascading failures, network partitions, backpressure, and distributed-system consistency issues.
- Strong analytical, communication, and documentation skills.
- Ability to manage multiple priorities while supporting highly visible, mission-critical systems.
Preferred Qualifications
- AWS DevOps, DevSecOps, or related certification.
- Additional AWS certifications, such as Solutions Architect, SysOps Administrator, or Developer, and/or Kubernetes certifications such as CKA or CKAD.
- Familiarity with Zero Trust security principles and cloud security best practices.
- Experience with GitLab, Jenkins, or similar CI/CD platforms.
- Experience working in highly regulated environments.
- Experience supporting large-scale enterprise modernization initiatives, including legacy-to-cloud migrations.
- Experience participating in disaster recovery exercises, continuity planning, or platform readiness assessments.
Additional Requirements
- Ability to meet applicable background investigation and eligibility requirements.
- Must be authorized to work in the United States.
Location: On-site
Similar jobs
- KA
Kubernetes Platform Engineer – Voice Services Modernization
NewKaltechsoft
Philadelphia, PA🇺🇸On-site21 hours agoDockerAWSELK+11Technology - TE
AI Platform Engineer -Agentic AI
NewTekAssembly
United States🇺🇸Hybrid2 days agoDockerFastAPIMicroservices+9Technology - IS
DevOps Build and Release Engineer
NewIntelliX Software, Inc.
Lansing, MI🇺🇸Hybrid21 hours agoShellAzureBamboo+5Technology - CS
Senior DevOps Lead - Remote / Telecommute
NewCynet Systems
Richmond, VA🇺🇸Remote21 hours agoDockerAWSSOC 2+16Technology - RS
Platform Engineer
NewReal Soft, Inc / Diversity Direct
San Antonio, TX🇺🇸Hybrid21 hours agoShellAWSAgile+11Technology - SO
Platform Engineer – VDI
NewSystem One
Washington, DC🇺🇸$61/hrOn-site21 hours agoSplunkActive DirectoryAzure+1Technology