Haystack
← Back to Jobs
Technology
SH

Site Reliability Engineer

ShaarproUnited States🇺🇸United StatesPosted 31 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
MicroservicesNode.jsAWSLoad BalancingArgoCDDNSGitGitHub ActionsGoHTTPHelmJenkinsKubernetesPythonRESTRedisTerraform

Job Description

Site Reliability Engineer

Introduction

We’re seeking an experienced, highly collaborative SRE to partner with product teams and tackle our most critical infrastructure challenges. You’ll be hands-on in designing, building, and operating our cloud platform—and driving the reliability, performance, and security that empower our engineering organization.

Responsibilities

  • Infrastructure as Code & CI/CD: Automate provisioning and deployments with Terraform and integrate best-practice pipelines (GitHub Actions, ArgoCD, etc.).
  • Reliability Engineering: Define SLIs/SLOs, manage error budgets, and build dashboards & alerts to proactively measure and improve system health.
  • Security & Compliance: Enforce least-privilege IAM policies, automate vulnerability scans, and maintain audit logging for compliance.
  • Monitoring & Observability: Instrument services with metrics, logs, and distributed tracing to enable rapid troubleshooting, aid teams in alerting, custom metrics, and dashboarding.
  • Incident Management: Own on-call rotations, lead real-time incident response, conduct post-mortems, and drive continuous improvements.
  • Cost Optimization: Implement tagging strategies, right-size resources, and leverage concrete data to decide on optimal methods to control cloud spend at scale.
  • Documentation & Mentorship: Author runbooks, standards, and best-practice guides—and coach dev teams on implementing modern DevOps, reliability, and security patterns.

Requirements

Required Skills

  • 5+ years of experience running production critical systems
  • Proficiency with AWS Cloud and Cloud-Native best practices
  • Experience with Kubernetes (EKS, GKE) and Container Orchestration at scale
  • Skilled in Terraform for infrastructure provisioning and maintenance
  • Knowledge of managing and debugging databases like Redis and Postgres
  • Familiarity with VPC, VPN, Load Balancing, and cloud networking components
  • Proficiency with Git workflows, branching strategies, and CI/CD system integrations
  • Understanding of web and network protocols and standards (HTTP, REST, TLS, DNS, etc.)

Preferred Skills

  • Bachelor's degree, or equivalent in Computer Science, Engineering, or a related field
  • Experience with ArgoCD, Github Actions, Jenkins, or other CI/CD pipeline solutions
  • Working knowledge of Python, Golang, and Helm templating languages
  • Node.js experience, including running scalable, resilient Node microservices
  • Foundational security best practices for cloud infrastructure
  • Awareness of Terragrunt, managing Terraform state, and optimal project structure
  • Production readiness fundamentals amidst a fast-moving team

Language Requirement

Professional proficiency in English (both written and spoken) is required for this role.

Industry

Technology, Information and Internet

Similar jobs