Haystack
← Back to Jobs
Technology
TA

Site Reliability Engineer

Techgroup America Inc.Charlotte, NC🇺🇸United StatesPosted 24 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Charlotte, NC, United States
Posted
23 hours ago
ShellEncryptionSplunkGrafanaKafkaKubernetesPowerShellPrometheusPython

Job Description

Site Reliability Engineer (SRE) – Messaging Services

We are seeking an experienced SRE – Messaging Services to drive reliability, observability, security, and operational excellence across IBM MQ and Kafka/Confluent environments.

Key Responsibilities:

  • Lead reliability engineering for large-scale IBM MQ and Kafka platforms.
  • Drive patching, EOL remediation, vulnerability remediation, and platform stabilization.
  • Define and implement SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems.
  • Enhance monitoring and observability for message flows, queue depths, Kafka lag, throughput, and latency.
  • Develop proactive fault detection and automated remediation strategies.
  • Support high-availability, resilient, and scalable messaging platforms.
  • Troubleshoot production issues including message backlogs, latency spikes, and connection failures.
  • Partner with engineering, application, infrastructure, and security teams.
  • Support global production environments and participate in on-call rotations.

Required Skills:

  • Strong experience in SRE / Production Engineering.
  • Hands-on expertise with IBM MQ and Kafka/Confluent.
  • Strong knowledge of distributed messaging, high availability, scalability, and reliability patterns.
  • Experience with Dynatrace, Splunk, Prometheus, Grafana, or similar observability tools.
  • Strong scripting/automation skills using Python, Shell, or PowerShell.
  • Experience with Linux/Unix and Windows production environments.
  • Knowledge of messaging security, TLS, certificates, encryption, and vulnerability remediation.
  • Strong troubleshooting, incident management, and RCA skills.

Preferred:

  • Experience with Kubernetes/containerized messaging platforms.
  • Knowledge of Kafka Schema Registry, Connect, and Streams.
  • Experience with IBM MQ clustering or Native HA.
  • Exposure to AIOps, anomaly detection, or automated remediation.
  • Experience with messaging modernization/migration programs.
  • Financial services or other regulated-industry experience preferred.

Candidate Requirements:

  • Strong communication and leadership skills.
  • Ability to work in high-pressure production environments.

Similar jobs