Haystack
← Back to Jobs
Manufacturing
TR

Only W2: SRE / Production Reliability Engineer, Woonsocket, RI

TrorWoonsocket, RI🇺🇸United StatesPosted 26 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Woonsocket, RI, United States
Posted
Yesterday

Job Description

Role: Senior SRE / Production Reliability Engineer
Experience: 8+ Years
Location: Woonsocket, RI
 
Job Summary
We are looking for a Senior SRE / Production Reliability Engineer to improve the reliability, performance, and availability of critical production systems.
The ideal candidate should have strong experience in SRE/DevOps, Incident Management, Observability, Kubernetes, Google Cloud Platform, and production monitoring, along with hands-on experience with time-series anomaly detection.
 
What You’ll Do
  • Own reliability and performance of critical production applications.
  • Act as an Incident Commander (IC) during P1/P2 production incidents.
  • Lead incident response, root-cause analysis, postmortems, and reliability improvements.
  • Define and manage SLIs, SLOs, and error budgets.
  • Build and improve monitoring, alerting, and observability solutions.
  • Tune and validate time-series anomaly detection models for production monitoring.
  • Develop automation and operational tools using Python, Java, and React.
  • Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
  • Work with engineering and operations teams to improve system reliability and reduce manual work.
  • Support large-scale deployments and manage production risks such as configuration drift and blast radius.
 
Must-Have Skills
  • 8+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
  • Hands-on experience as an Incident Commander for P1/P2 incidents.
  • Strong experience with time-series anomaly detection models in production observability — mandatory.
  • Strong production-level programming skills in:
    • Python
    • Java
    • React
  • Strong experience with SLI, SLO, and error budgets.
  • Hands-on observability experience with:
    • Prometheus
    • Grafana
    • OpenTelemetry
    • At least 2 log platforms such as Loki, Splunk, or Elasticsearch
  • Strong Google Cloud Platform experience.
  • Strong Kubernetes operational experience.
  • Experience with Rancher K3s.
  • Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
  • Experience with production-scale distributed systems and on-call support.
 
Preferred Skills
  • Production Readiness Reviews / service launch experience.
  • BigQuery and PostgreSQL.
  • Chaos Engineering / fault injection.
  • TIC/Technical Incident Commander certification.
  • Healthcare, pharmacy, retail, or other high-availability environments.
  • LLM/GenAI for incident management, alert summarization, or runbook recommendations.
  • Kafka.
  • Istio / Envoy.
  • Terraform / Ansible.
 
Important: The Incident Commander experience and production time-series anomaly detection should be treated as hard requirements. A candidate who only has general monitoring/observability experience but has never worked with anomaly detection models would not be a strong fit.

Similar jobs