Haystack
← Back to Jobs
Manufacturing
TR

Only W2: SRE / Production Reliability Engineer, Woonsocket, RI

TrorWoonsocket, RI🇺🇸United StatesPosted Sep 21, 2026

Why This Role Stands Out

This hybrid role offers a fantastic opportunity to elevate critical production systems at a reputable company, with significant impact and room for professional growth. You'll thrive here if you possess strong SRE/DevOps expertise and enjoy tackling complex reliability challenges. Apply today to contribute your skills and advance your career!

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Woonsocket, RI, United States
Posted
19 hours ago

Job Description

Role: Senior SRE / Production Reliability Engineer
Location: Woonsocket, RI
Duration : Long Term 
Job Summary
We are looking for a Senior SRE / Production Reliability Engineer to improve the reliability, performance, and availability of critical production systems.
The ideal candidate should have strong experience in SRE/DevOps, Incident Management, Observability, Kubernetes, Google Cloud Platform, and production monitoring, along with hands-on experience with time-series anomaly detection.
What You’ll Do
  • Own reliability and performance of critical production applications.
  • Act as an Incident Commander (IC) during P1/P2 production incidents.
  • Lead incident response, root-cause analysis, postmortems, and reliability improvements.
  • Define and manage SLIs, SLOs, and error budgets.
  • Build and improve monitoring, alerting, and observability solutions.
  • Tune and validate time-series anomaly detection models for production monitoring.
  • Develop automation and operational tools using Python, Java, and React.
  • Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
  • Work with engineering and operations teams to improve system reliability and reduce manual work.
  • Support large-scale deployments and manage production risks such as configuration drift and blast radius.
Must-Have Skills
  • 10+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
  • Hands-on experience as an Incident Commander for P1/P2 incidents.
  • Strong experience with time-series anomaly detection models in production observability — mandatory.
  • Strong production-level programming skills in:
    • Python
    • Java
    • React
  • Strong experience with SLI, SLO, and error budgets.
  • Hands-on observability experience with:
    • Prometheus
    • Grafana
    • OpenTelemetry
    • At least 2 log platforms such as Loki, Splunk, or Elasticsearch
  • Strong Google Cloud Platform experience.
  • Strong Kubernetes operational experience.
  • Experience with Rancher K3s.
  • Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
  • Experience with production-scale distributed systems and on-call support.
Preferred Skills
  • Production Readiness Reviews / service launch experience.
  • BigQuery and PostgreSQL.
  • Chaos Engineering / fault injection.
  • TIC/Technical Incident Commander certification.
  • Healthcare, pharmacy, retail, or other high-availability environments.
  • LLM/GenAI for incident management, alert summarization, or runbook recommendations.
  • Kafka.
  • Istio / Envoy.
  • Terraform / Ansible.
Important: The Incident Commander experience and production time-series anomaly detection should be treated as hard requirements. A candidate who only has general monitoring/observability experience but has never worked with anomaly detection models would not be a strong fit.

Similar jobs