Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Woonsocket, RI, United States
Posted
Yesterday
Job Description
Role: Senior SRE / Production Reliability Engineer
Experience: 8+ Years
Location: Woonsocket, RI
Experience: 8+ Years
Location: Woonsocket, RI
Job Summary
We are looking for a Senior SRE / Production Reliability Engineer to improve the reliability, performance, and availability of critical production systems.
The ideal candidate should have strong experience in SRE/DevOps, Incident Management, Observability, Kubernetes, Google Cloud Platform, and production monitoring, along with hands-on experience with time-series anomaly detection.
What You’ll Do
-
Own reliability and performance of critical production applications.
-
Act as an Incident Commander (IC) during P1/P2 production incidents.
-
Lead incident response, root-cause analysis, postmortems, and reliability improvements.
-
Define and manage SLIs, SLOs, and error budgets.
-
Build and improve monitoring, alerting, and observability solutions.
-
Tune and validate time-series anomaly detection models for production monitoring.
-
Develop automation and operational tools using Python, Java, and React.
-
Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
-
Work with engineering and operations teams to improve system reliability and reduce manual work.
-
Support large-scale deployments and manage production risks such as configuration drift and blast radius.
Must-Have Skills
-
8+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
-
Hands-on experience as an Incident Commander for P1/P2 incidents.
-
Strong experience with time-series anomaly detection models in production observability — mandatory.
-
Strong production-level programming skills in:
-
Python
-
Java
-
React
-
-
Strong experience with SLI, SLO, and error budgets.
-
Hands-on observability experience with:
-
Prometheus
-
Grafana
-
OpenTelemetry
-
At least 2 log platforms such as Loki, Splunk, or Elasticsearch
-
-
Strong Google Cloud Platform experience.
-
Strong Kubernetes operational experience.
-
Experience with Rancher K3s.
-
Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
-
Experience with production-scale distributed systems and on-call support.
Preferred Skills
-
Production Readiness Reviews / service launch experience.
-
BigQuery and PostgreSQL.
-
Chaos Engineering / fault injection.
-
TIC/Technical Incident Commander certification.
-
Healthcare, pharmacy, retail, or other high-availability environments.
-
LLM/GenAI for incident management, alert summarization, or runbook recommendations.
-
Kafka.
-
Istio / Envoy.
-
Terraform / Ansible.
Important: The Incident Commander experience and production time-series anomaly detection should be treated as hard requirements. A candidate who only has general monitoring/observability experience but has never worked with anomaly detection models would not be a strong fit.
Similar jobs
- PE
HPC Software Deployment Configuration Manager, Lead with Security Clearance
Peraton
Fort Meade, MD🇺🇸$104k - $166k/yrHybrid3 weeks agoHTTPSTechnology - BI
Dev/Sec/Ops Platform Engineer
NewBigBear.ai
McLean, VA🇺🇸HybridYesterdayAWSAnsibleAzure+5Technology - ST
SRE - Nexus.corp /SonarQube Engineer
NewSPECTRAFORCE TECHNOLOGIES Inc.
United States🇺🇸RemoteYesterdayAWSSonarQubeAnsible+7Technology - AS
Senior Power Platform Engineer
Apex Systems
Charlotte, NC🇺🇸Hybrid2 weeks agoSQLJavaScriptPower BI+1Technology - AS
Systems Engineer
Apex Systems
Alexandria, VA🇺🇸$110k/yrOn-site7 weeks agoKubernetesTechnology - TT
Platform Engineer (Google Cloud Platform)
NewTrigyn Technologies, Inc.
Albany, NY🇺🇸HybridYesterdayAWSSonarQubeSplunk+6Technology