Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Pittsburgh, PA, United States
Posted
19 hours ago
DockerMicroservicesShellAWSELKLoad BalancingSplunkTCP/IPAnsibleAzureBashDNSDatadogGenerative AIGitGoogle CloudGrafanaHelmKubernetesLLMPrometheusPythonRESTTerraform
Job Description
Job Description
We are looking for an experienced Site Reliability Engineer (SRE) with strong Production Support experience and hands-on expertise in AI/LLM technologies, Kubernetes, and Docker. The ideal candidate will be responsible for maintaining highly available production environments, troubleshooting critical issues, improving system reliability, and supporting AI/LLM-based applications and platforms.
Key Responsibilities
- Provide 24x7 production support for critical applications and services, including incident response, troubleshooting, and root cause analysis.
- Monitor system health, application performance, availability, and reliability across production environments.
- Deploy, manage, and troubleshoot applications running on Kubernetes and Docker environments.
- Support AI/ML and LLM-based applications, APIs, and services in production.
- Troubleshoot application, infrastructure, container, networking, and performance-related issues.
- Participate in incident, problem, and change management processes.
- Perform Root Cause Analysis (RCA) and implement permanent solutions to recurring production issues.
- Develop and maintain automation scripts using Python, Bash, or Shell to reduce manual operational activities.
- Work with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Implement monitoring, logging, alerting, and observability solutions for production workloads.
- Collaborate with Development, DevOps, AI/ML, and Infrastructure teams to improve system reliability and deployment processes.
- Support CI/CD pipelines and automated deployment processes.
- Participate in on-call rotations and handle critical production incidents within defined SLAs.
- Continuously improve system scalability, performance, availability, and operational efficiency.
Required Skills
- 5+ years of experience in SRE, Production Support, DevOps, or Site Reliability Engineering.
- Strong hands-on experience with Kubernetes/K8s and Docker.
- Experience supporting production applications and critical incidents.
- Hands-on exposure to LLM, Generative AI, AI/ML applications, or AI platforms.
- Strong Linux/Unix administration and troubleshooting skills.
- Experience with Python, Bash, or Shell scripting.
- Experience with at least one cloud platform: AWS, Azure, or Google Cloud Platform.
- Knowledge of CI/CD, Git, monitoring, logging, and observability.
- Strong understanding of application performance, system reliability, scalability, and high availability.
- Excellent troubleshooting, analytical, and communication skills.
Preferred Skills
- Experience supporting LLM/GenAI applications in production.
- Knowledge of LLM APIs, model serving, inference, or AI platforms.
- Experience with Prometheus, Grafana, ELK/EFK, Splunk, Datadog, or similar monitoring tools.
- Experience with Terraform or Ansible.
- Knowledge of microservices and REST APIs.
- Experience with Helm and Kubernetes deployments.
- Understanding of networking, DNS, TCP/IP, load balancing, and cloud infrastructure.
Similar jobs
- AI
Cloudera Public Cloud Platform Engineer (CDP)
NewARK Infotech Spectrum
United States🇺🇸Remote19 hours agoShellAWSEncryption+9Technology - ST
Senior DevOps Engineer / Architect|Chicago, IL (Hybrid Local Candidates Only)
NewStoneGate-Technologies LLC
Chicago, IL🇺🇸On-site19 hours agoNode.jsAWSEncryption+7Technology - AT
Lead ForgeRock Platform Engineer
NewAroha Technologies
San Antonio, TX🇺🇸Remote19 hours agoMicroservicesSpringSpring Boot+14Technology - NS
Cloud Engineer (Devops)
NewNovaLink Solutions
Raleigh, NC🇺🇸Hybrid19 hours agoAWSAzureGoogle CloudTechnology - KG
DevOps Engineer (Fulltime - New Castle, DE / Jersey City, NJ / Tampa, FL)
NewK&K Global Talent Solutions
New Castle, DE🇺🇸On-site19 hours agoHelmJenkinsKubernetes+1Technology - IT
Senior Infrastructure Engineer
NewITCAPS LLC
Houston, TX🇺🇸On-site19 hours agoAzureTechnology