Haystack
← Back to Jobs
Remote
Manufacturing
TR

Hiring Senior SRE / Production Reliability Engineer, 100% Remote

TrorUnited States🇺🇸United StatesPosted Oct 8, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
21 hours ago
Root Cause Analysis

Job Description

Role: Senior SRE / Production Reliability Engineer
Location: 100% Remote
Experience: 10+ Years

We need an experienced SRE who can keep production systems stable, find problems before users are affected, troubleshoot incidents, and automate repetitive checks.
The person will support Salesforce CRM (GPS) and YAVA platforms.
Looking for a strong SRE who can monitor production, troubleshoot P1/P2 incidents, perform RCA, validate deployments, test system reliability, and automate operational tasks.

What You Will Do
  • Monitor production systems and make sure they are available, fast, and reliable.
  • Build dashboards, alerts, and monitoring to detect issues early.
  • Monitor critical areas such as:
    • APIs
    • DNS
    • Load Balancers
    • WAF
    • MQ
    • Databases / DB2
    • Cloud
    • Telephony
    • Vendor systems
  • Create synthetic tests/health checks to make sure important customer and agent journeys are working.
  • Validate systems before and after deployments or major changes.
  • Check DNS, certificates, routing, connectivity, and configurations.
  • Perform load testing, failover testing, recovery testing, and resilience testing.
  • Troubleshoot P1/P2 production incidents and identify the root cause.
  • Conduct RCA/postmortems and make sure the same issue does not happen again.
  • Automate health checks, monitoring, certificate checks, DNS checks, and post-deployment validation.
  • Create and maintain runbooks and troubleshooting tools.
  • Work with application, network, security, cloud, database, telephony, and vendor teams.
Must Have
  • Strong SRE / Production Engineering / DevOps / Infrastructure experience.
  • Strong production troubleshooting experience.
  • Experience with:
    • Monitoring, logging, metrics, tracing
    • Dashboards and alerting
    • Synthetic monitoring
    • Incident management and RCA
    • Automation
  • Good knowledge of:
    • DNS
    • Load Balancers
    • TLS/SSL certificates
    • WAF
    • API Gateways
    • Cloud
    • Service-to-service/API integrations
  • Programming/scripting experience in Python, Go, PowerShell, Java, or similar.
  • Experience handling production incidents and performing root cause analysis.
  • Good communication skills and ability to work with multiple technical teams.
  • Salesforce CRM / CTI integrations
  • Contact center / telephony experience
  • Five9
  • Real-time voice or conversational AI
  • IBM MQ
  • DB2 / Mainframe
  • Imperva
  • F5
  • Azure / Google Cloud Platform
  • Performance and capacity testing
  • Chaos/resilience testing
  • Healthcare or other high-availability enterprise environments

Looking forward to qualified submissions only.

Similar jobs