Haystack
← Back to Jobs
Technology
PR

W2 - Site Reliability Engineer (SRE) Lead – IBM MQ / Kafka | Hybrid | Plano, TX / Charlotte, NC

ProhiresPlano, TX🇺🇸United StatesPosted 25 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Plano, TX, United States
Posted
Yesterday
ShellEncryptionSplunkGrafanaKafkaKubernetesPowerShellPrometheusPython

Job Description

Job Title: Site Reliability Engineer (SRE)
Job Location: Plano, TX / Charlotte, NC (3 Days onsite, 2 Days remote) Hybrid role.
Job Type: 12+ Months contract
 
 
 
Job Descriptions:
We are seeking an experienced Site Reliability Engineer (SRE) Lead - Messaging Services tdrive platform reliability, observability, and operational excellence across IBM MQ and Kafka environments.
 
This role combines:
  • Production engineering and reliability leadership for messaging platforms
  • Platform security, resilience engineering, and vulnerability remediation
  • Ownership of large-scale, distributed messaging runtimes
Key responsibilities include:
  • Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput
  • Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters
Implementing SRE best practices:
  • SLIs / SLOs focused on message delivery, latency, and availability
  • Incident management, escalation, and postmortem culture
  • Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput
  • Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery)
  • Building resilient messaging platforms capable of supporting real-time, event-driven workloads
  • Supporting global production messaging environments with on-call rotation and escalation ownership
  • Partnering with engineering, application, and security teams tensure reliability, scalability, and secure message transport
  • Strong experience in Site Reliability Engineering / Production Engineering
Hands-on expertise with:
  • IBM MQ (queue managers, clustering, channels, DLQ management)
  • Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
  • Large-scale distributed messaging systems and runtime management
Deep understanding of:
  • System reliability, scalability, and high availability design
  • Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering)
  • Incident management, root cause analysis, and problem management
Experience with:
  • Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms
  • Event and anomaly detection in high-volume systems
  • Strong scripting/automation skills:
  • Shell, Python, PowerShell
  • Experience managing Linux/Unix and Windows production environments
Knowledge of:
  • Event-driven architecture and messaging-based integration patterns
Understanding of:
  • Messaging platform security (TLS, certificates, channel auth, encryption)
  • Vulnerability remediation and risk mitigation in production systems
  • Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures
  • Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads
Familiarity with:
  • Kubernetes / containerized messaging platforms
  • Experience with:
  • Kafka ecosystem components (Schema Registry, Connect, Streams)
  • IBM MQ advanced features (Native HA, clustering)
Exposure to:
  • AI-driven operations (AIOps), anomaly detection, or automated remediation
  • Large-scale messaging modernization or migration programs
  • Messaging or middleware certifications (IBM MQ, Kafka, or equivalent)
  • Experience in regulated environments (e.g., financial services)

Similar jobs