Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
MicroservicesNode.jsAWSLoad BalancingArgoCDDNSGitGitHub ActionsGoHTTPHelmJenkinsKubernetesPythonRESTRedisTerraform
Job Description
Site Reliability Engineer
Introduction
We’re seeking an experienced, highly collaborative SRE to partner with product teams and tackle our most critical infrastructure challenges. You’ll be hands-on in designing, building, and operating our cloud platform—and driving the reliability, performance, and security that empower our engineering organization.
Responsibilities
- Infrastructure as Code & CI/CD: Automate provisioning and deployments with Terraform and integrate best-practice pipelines (GitHub Actions, ArgoCD, etc.).
- Reliability Engineering: Define SLIs/SLOs, manage error budgets, and build dashboards & alerts to proactively measure and improve system health.
- Security & Compliance: Enforce least-privilege IAM policies, automate vulnerability scans, and maintain audit logging for compliance.
- Monitoring & Observability: Instrument services with metrics, logs, and distributed tracing to enable rapid troubleshooting, aid teams in alerting, custom metrics, and dashboarding.
- Incident Management: Own on-call rotations, lead real-time incident response, conduct post-mortems, and drive continuous improvements.
- Cost Optimization: Implement tagging strategies, right-size resources, and leverage concrete data to decide on optimal methods to control cloud spend at scale.
- Documentation & Mentorship: Author runbooks, standards, and best-practice guides—and coach dev teams on implementing modern DevOps, reliability, and security patterns.
Requirements
Required Skills
- 5+ years of experience running production critical systems
- Proficiency with AWS Cloud and Cloud-Native best practices
- Experience with Kubernetes (EKS, GKE) and Container Orchestration at scale
- Skilled in Terraform for infrastructure provisioning and maintenance
- Knowledge of managing and debugging databases like Redis and Postgres
- Familiarity with VPC, VPN, Load Balancing, and cloud networking components
- Proficiency with Git workflows, branching strategies, and CI/CD system integrations
- Understanding of web and network protocols and standards (HTTP, REST, TLS, DNS, etc.)
Preferred Skills
- Bachelor's degree, or equivalent in Computer Science, Engineering, or a related field
- Experience with ArgoCD, Github Actions, Jenkins, or other CI/CD pipeline solutions
- Working knowledge of Python, Golang, and Helm templating languages
- Node.js experience, including running scalable, resilient Node microservices
- Foundational security best practices for cloud infrastructure
- Awareness of Terragrunt, managing Terraform state, and optimal project structure
- Production readiness fundamentals amidst a fast-moving team
Language Requirement
Professional proficiency in English (both written and spoken) is required for this role.
Industry
Technology, Information and Internet
Similar jobs
- TC
Network Engineer
NewTEKsystems c/o Allegis Group
Cary, NC🇺🇸$45 - $60/hrHybridYesterdayZero TrustTechnology - TC
Azure Platform Engineer
NewTEKsystems c/o Allegis Group
Plano, TX🇺🇸$70/hrHybridYesterdaySnowflakeAzureBigQuery+3Technology - AG
Azure DevOps Engineer
NewAdvent Global Solutions, Inc.
Chicago, IL🇺🇸On-siteYesterdayDockerSonarQubeAzure+7Technology - SG
Platform Engineer I
NewSoftware Guidance & Assistance
Cincinnati, OH🇺🇸HybridYesterdayAgileTechnology - DE
Cloud Platform Engineer
Decisionpoint Corporation
United States🇺🇸Remote4 weeks agoAWSCloudFormationKubernetes+3Technology - C-
Site Reliability Engineer II
C-Serv
San Jose, California🇺🇸Hybrid3 days agoAWSLoad BalancingArgoCD+8Technology