Haystack
← Back to Jobs
Technology
JG

Site Reliability Engineer (SRE) - Mid-Level

Judge Group, Inc.Southlake, TX🇺🇸United StatesPosted 10 Sept 2026

Why This Role Stands Out

This hybrid role offers you the chance to leverage your advanced Python skills and automation-first mindset to drive significant improvements in system reliability and scalability within a reputable company. You'll thrive here if you enjoy deep operational ownership, complex problem-solving, and building lasting solutions, with competitive hourly compensation of $50-$55 USD. Don't miss this opportunity to grow your expertise in a dynamic environment.

Quick Overview

Salary
$50 - $55/hr
Seniority
Mid Senior
Work mode
On Site
Location
Southlake, TX, United States
Posted
Yesterday
AWSSplunkAnsibleAzureDatadogGoogle CloudGrafanaKubernetesPrometheusPythonTerraform

Job Description

Location: Southlake, TX Salary: $50.00 USD Hourly - $55.00 USD Hourly Description: Our client is currently seeking a Site Reliability Engineer (SRE) - Mid-Level

Location: Hybrid in Southlake, TX or Austin, TX (4 days a week onsite)

Contract: 12 Months with possibility of extension

About the job:

We are seeking a highly motivated Site Reliability Engineer to join our team on a contract basis. In this role, you will apply an automation-first mindset to tackle complex operational challenges across both on-premises and cloud-native environments. This is not a traditional build/release or pure DevOps role; it requires deep operational ownership and advanced Python development skills to build lasting solutions. You will focus on minimizing operational toil, enhancing system observability, and driving reliability initiatives to ensure our platforms scale seamlessly and meet strict availability objectives.

Responsibilities:
  • Develop and maintain robust Python-based automation solutions to reduce manual operational effort and prevent recurring issues.
  • Automate infrastructure management and integrate platforms utilizing APIs and client libraries across Linux, Windows, Kubernetes, and cloud-native environments.
  • Monitor production systems to meet reliability targets, leading incident response, troubleshooting, and post-mortem/root cause analysis activities.
  • Build and maintain comprehensive dashboards, alerts, and monitoring solutions to improve visibility into application and infrastructure health via metrics, logs, and traces.
  • Support disaster recovery efforts, failover testing, operational readiness activities, and ongoing performance analysis.
  • Assist in implementing infrastructure automation and support CI/CD deployment reliability initiatives.
  • Investigate alerts to identify noise-reduction opportunities and explore AI/ML-driven operational improvements for intelligent alerting and anomaly detection.


Minimum qualifications:
  • Bachelor's degree in Computer Science, Engineering, a related technical field, or equivalent practical experience.
  • 3-5 years of hands-on experience in Site Reliability Engineering (SRE) or Production Engineering.
  • Experience in production operations, including incident response, root cause analysis (RCA), problem remediation, and supporting large-scale production systems.
  • Strong software development experience using Python to build automation tools, frameworks, and operational solutions (beyond basic scripting).
  • Experience with Linux/Windows systems, networking fundamentals, and distributed applications.
  • Experience with monitoring, observability, and alerting platforms (e.g., Splunk, Grafana, Prometheus, Datadog, or enterprise cloud operation suites).


Preferred qualifications:
  • Experience with containerization and orchestration technologies, specifically Kubernetes, across major cloud platforms (e.g., AWS, Azure, Google Cloud Platform).
  • Experience with Infrastructure as Code (IaC) and configuration management tools (e.g., Terraform, Ansible).
  • Experience with CI/CD pipelines and deployment automation reliability.
  • Knowledge of AIOps and AI/ML-driven operational tooling (e.g., anomaly detection, intelligent alerting, log analytics).
  • Exposure to OpenTelemetry and modern observability practices.
  • Experience supporting highly available, mission-critical production systems in regulated or enterprise environments.

Medical, dental, and vision insurance are available to qualified candidates who meet eligibility requirements.
By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively "Judge") to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.
Contact:
This job and many more are available through The Judge Group. Please apply with us today!

Similar jobs