Haystack
← Back to Jobs
Technology
SI

Site Reliability Engineer (SRE) || Plano, TX OR Charlotte, NC - 3 Days Onsite || 12+months || Video and F2F Interview

Stellent IT LLCPlano, TX🇺🇸United StatesPosted 26 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Plano, TX, United States
Posted
Yesterday
ShellEncryptionSplunkGrafanaKafkaKubernetesPowerShellPrometheusPython

Job Description

W2 Candidates ONLY

Title: Site Reliability Engineer (SRE)
Location: Plano, TX / Charlotte, NC (3 Days onsite, 2 Days remote) Hybrid role.
Duration: 12+ Months contract

Interview Process: 2 virtual rounds; onsite interview may also be requested
Client Type: Direct Client (MSP/VMS)

Note:

Please submit only qualified candidates who closely match the job requirements.

Genuine LinkedIn is must.

Local candidates for onsite/hybrid roles.

BOA experience first, followed by other banking/financial clients, then strong enterprise clients.

Avoid candidates with 6+ month employment gaps.

Prefer candidates with stable, long-term employment/contract history.

Speed is critical due to the MSP/VMS process.

Please share qualified candidates for the SRE Lead position with strong hands-on experience in IBM MQ and Kafka/Confluent, large-scale messaging/production engineering, SRE practices, monitoring/observability, incident management, RCA, SLI/SLO, HA/resiliency, and Shell/Python/PowerShell automation. Linux/Windows experience is required, while Kubernetes and Banking/Financial Services experience is preferred.

Job Descriptions:

We are seeking an experienced Site Reliability Engineer (SRE) Lead - Messaging Services tdrive platform reliability, observability, and operational excellence across IBM MQ and Kafka environments.

This role combines:

  • Production engineering and reliability leadership for messaging platforms
  • Platform security, resilience engineering, and vulnerability remediation
  • Ownership of large-scale, distributed messaging runtimes

Key responsibilities include:

  • Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput
  • Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters

Implementing SRE best practices:

  • SLIs / SLOs focused on message delivery, latency, and availability
  • Incident management, escalation, and postmortem culture
  • Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput
  • Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery)
  • Building resilient messaging platforms capable of supporting real-time, event-driven workloads
  • Supporting global production messaging environments with on-call rotation and escalation ownership
  • Partnering with engineering, application, and security teams tensure reliability, scalability, and secure message transport
  • Strong experience in Site Reliability Engineering / Production Engineering

Hands-on expertise with:

  • IBM MQ (queue managers, clustering, channels, DLQ management)
  • Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
  • Large-scale distributed messaging systems and runtime management

Deep understanding of:

  • System reliability, scalability, and high availability design
  • Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering)
  • Incident management, root cause analysis, and problem management

Experience with:

  • Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms
  • Event and anomaly detection in high-volume systems
  • Strong scripting/automation skills:
  • Shell, Python, PowerShell
  • Experience managing Linux/Unix and Windows production environments

Knowledge of:

  • Event-driven architecture and messaging-based integration patterns

Understanding of:

  • Messaging platform security (TLS, certificates, channel auth, encryption)
  • Vulnerability remediation and risk mitigation in production systems
  • Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures
  • Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads

Familiarity with:

  • Kubernetes / containerized messaging platforms
  • Experience with:
  • Kafka ecosystem components (Schema Registry, Connect, Streams)
  • IBM MQ advanced features (Native HA, clustering)

Exposure to:

  • AI-driven operations (AIOps), anomaly detection, or automated remediation
  • Large-scale messaging modernization or migration programs
  • Messaging or middleware certifications (IBM MQ, Kafka, or equivalent)
  • Experience in regulated environments (e.g., financial services)

Equal Opportunity Employer/Veterans/Disabled

Benefit offerings available for our associates include medical, dental, vision, life insurance, short-term disability, additional voluntary benefits, an EAP program, commuter benefits, and a 401K plan. Our benefit offerings provide employees with the flexibility to choose the type of coverage that meets their individual needs. In addition, our associates may be eligible for paid leave including Paid Sick Leave or any other paid leave required by Federal, State, or local law, as well as Holiday pay where applicable. Disclaimer: These benefit offerings do not apply to client-recruited jobs and jobs that are direct hires to a client.

The Company will consider qualified applicants with arrest and conviction records in accordance with federal, state, and local laws and/or security clearance requirements, including, as applicable:

  • The California Fair Chance Act
  • Los Angeles City Fair Chance Ordinance
  • Los Angeles County Fair Chance Ordinance for Employers
  • San Francisco Fair Chance Ordinance

Senior Talent Acquisition

E-

STELLENT IT A Nationally Recognized Minority Certified Enterprise

Similar jobs