Haystack
← Back to Jobs
Technology
SC

Site Reliability Engineer (SRE

SN Cloud SolutionsPittsburgh, PA🇺🇸United StatesPosted 31 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Pittsburgh, PA, United States
Posted
19 hours ago
DockerMicroservicesShellAWSELKLoad BalancingSplunkTCP/IPAnsibleAzureBashDNSDatadogGenerative AIGitGoogle CloudGrafanaHelmKubernetesLLMPrometheusPythonRESTTerraform

Job Description

Job Description

We are looking for an experienced Site Reliability Engineer (SRE) with strong Production Support experience and hands-on expertise in AI/LLM technologies, Kubernetes, and Docker. The ideal candidate will be responsible for maintaining highly available production environments, troubleshooting critical issues, improving system reliability, and supporting AI/LLM-based applications and platforms.

Key Responsibilities

  • Provide 24x7 production support for critical applications and services, including incident response, troubleshooting, and root cause analysis.
  • Monitor system health, application performance, availability, and reliability across production environments.
  • Deploy, manage, and troubleshoot applications running on Kubernetes and Docker environments.
  • Support AI/ML and LLM-based applications, APIs, and services in production.
  • Troubleshoot application, infrastructure, container, networking, and performance-related issues.
  • Participate in incident, problem, and change management processes.
  • Perform Root Cause Analysis (RCA) and implement permanent solutions to recurring production issues.
  • Develop and maintain automation scripts using Python, Bash, or Shell to reduce manual operational activities.
  • Work with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Implement monitoring, logging, alerting, and observability solutions for production workloads.
  • Collaborate with Development, DevOps, AI/ML, and Infrastructure teams to improve system reliability and deployment processes.
  • Support CI/CD pipelines and automated deployment processes.
  • Participate in on-call rotations and handle critical production incidents within defined SLAs.
  • Continuously improve system scalability, performance, availability, and operational efficiency.

Required Skills

  • 5+ years of experience in SRE, Production Support, DevOps, or Site Reliability Engineering.
  • Strong hands-on experience with Kubernetes/K8s and Docker.
  • Experience supporting production applications and critical incidents.
  • Hands-on exposure to LLM, Generative AI, AI/ML applications, or AI platforms.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with Python, Bash, or Shell scripting.
  • Experience with at least one cloud platform: AWS, Azure, or Google Cloud Platform.
  • Knowledge of CI/CD, Git, monitoring, logging, and observability.
  • Strong understanding of application performance, system reliability, scalability, and high availability.
  • Excellent troubleshooting, analytical, and communication skills.

Preferred Skills

  • Experience supporting LLM/GenAI applications in production.
  • Knowledge of LLM APIs, model serving, inference, or AI platforms.
  • Experience with Prometheus, Grafana, ELK/EFK, Splunk, Datadog, or similar monitoring tools.
  • Experience with Terraform or Ansible.
  • Knowledge of microservices and REST APIs.
  • Experience with Helm and Kubernetes deployments.
  • Understanding of networking, DNS, TCP/IP, load balancing, and cloud infrastructure.

Similar jobs