← Back to Jobs
Full time
Technology
Site Reliability Engineer
Selby JenningsManhattan, NY🇺🇸United StatesPosted 29 Jul 2026
Why This Role Stands Out
This hybrid Site Reliability Engineer role offers a fantastic opportunity to shape highly available and scalable infrastructure, automate critical processes, and enhance CI/CD pipelines within a reputable company. If you thrive on tackling complex challenges in Kubernetes environments and possess strong Infrastructure as Code skills, you'll find this position a rewarding path for significant career growth and skill development. Apply today to become a key player in ensuring operational excellence!
Quick Overview
Work Type
Hybrid
Schedule
Full Time
Level
Mid Senior
Job Description
Responsibilities
- Design, build, and maintain highly available and scalable infrastructure
- Automate operational processes to improve efficiency and reduce manual intervention
- Manage and support Kubernetes-based containerized environments
- Develop and maintain Infrastructure as Code using Terraform and related tools
- Build and enhance CI/CD pipelines to streamline deployment processes
- Monitor system health, performance, and reliability across production environments
- Lead incident response efforts and drive root cause analysis for production issues
- Partner with engineering teams to improve system design, resilience, and observability
- Implement best practices around monitoring, alerting, capacity planning, and disaster recovery
- Continuously identify opportunities to improve reliability, performance, and operational excellence
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience)
- Experience in Site Reliability Engineering, Platform Engineering, Production Engineering, DevOps, or Infrastructure Engineering
- Strong Linux systems administration experience
- Hands-on experience with Kubernetes and containerized environments
- Experience with Terraform or other Infrastructure as Code tools
- Strong knowledge of cloud platforms such as AWS, Azure, or GCP
- Proficiency in Python, Go, Bash, or similar scripting/programming languages
- Experience building and supporting CI/CD pipelines
- Familiarity with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK
- Strong troubleshooting and problem-solving skills in large-scale production environments
Skills
GCP
AWS
ELK
Splunk
Azure
Bash
Datadog
Grafana
Kubernetes
Prometheus
Python
Terraform
Similar jobs
Network Engineer
Hillsdale College · Hillsdale, United States
48 minutes agoPlatform Engineer III
Apex Systems · Cincinnati, United States
2 hours ago$60 - $68/hrPython Automation Platform Engineer
GState Consulting LLC · Plano, United States
2 hours ago€60/hrLead DevOps Engineer - OH (Pref), NC, TX
Apex Systems · Columbus, United States
2 hours ago€74/hrALM Architect / Azure Devops architect
Pristine Resource · United States
2 hours agoSystems Engineer
Booz Allen Hamilton · Huntsville, United States
2 hours ago$61.9k - $141k/yr