Haystack
← Back to Jobs
Full time
Technology

Site Reliability Engineer

Selby JenningsManhattan, NY🇺🇸United StatesPosted 29 Jul 2026

Why This Role Stands Out

This hybrid Site Reliability Engineer role offers a fantastic opportunity to shape highly available and scalable infrastructure, automate critical processes, and enhance CI/CD pipelines within a reputable company. If you thrive on tackling complex challenges in Kubernetes environments and possess strong Infrastructure as Code skills, you'll find this position a rewarding path for significant career growth and skill development. Apply today to become a key player in ensuring operational excellence!

Quick Overview

Work Type
Hybrid
Schedule
Full Time
Level
Mid Senior

Job Description



Responsibilities

  • Design, build, and maintain highly available and scalable infrastructure
  • Automate operational processes to improve efficiency and reduce manual intervention
  • Manage and support Kubernetes-based containerized environments
  • Develop and maintain Infrastructure as Code using Terraform and related tools
  • Build and enhance CI/CD pipelines to streamline deployment processes
  • Monitor system health, performance, and reliability across production environments
  • Lead incident response efforts and drive root cause analysis for production issues
  • Partner with engineering teams to improve system design, resilience, and observability
  • Implement best practices around monitoring, alerting, capacity planning, and disaster recovery
  • Continuously identify opportunities to improve reliability, performance, and operational excellence


Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience)
  • Experience in Site Reliability Engineering, Platform Engineering, Production Engineering, DevOps, or Infrastructure Engineering
  • Strong Linux systems administration experience
  • Hands-on experience with Kubernetes and containerized environments
  • Experience with Terraform or other Infrastructure as Code tools
  • Strong knowledge of cloud platforms such as AWS, Azure, or GCP
  • Proficiency in Python, Go, Bash, or similar scripting/programming languages
  • Experience building and supporting CI/CD pipelines
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK
  • Strong troubleshooting and problem-solving skills in large-scale production environments

Skills

GCP
AWS
ELK
Splunk
Azure
Bash
Datadog
Grafana
Kubernetes
Prometheus
Python
Terraform

Similar jobs