Quick Overview
Job Description
Site Reliability Engineer (SRE) - Job Description
Job Title
Site Reliability Engineer (SRE)
Job Summary
We are seeking a highly motivated and skilled Site Reliability Engineer (SRE) to join our engineering team. The ideal candidate will be responsible for maintaining the reliability, scalability, performance, and security of mission-critical applications and infrastructure. The SRE will work closely with development, cloud, and operations teams to automate processes, improve system availability, and drive operational excellence.
Key Responsibilities
- Design, implement, and maintain highly available and scalable infrastructure.
- Monitor application and infrastructure health using observability tools.
- Automate operational tasks using scripting and Infrastructure as Code (IaC).
- Manage incident response, root cause analysis, and post-incident reviews.
- Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Optimize system performance, reliability, and resource utilization.
- Build and maintain CI/CD pipelines for seamless application deployments.
- Collaborate with software engineering teams to improve application resilience.
- Implement disaster recovery, backup, and business continuity strategies.
- Ensure security best practices across cloud and on-premises environments.
- Create and maintain technical documentation, runbooks, and operational procedures.
- Participate in on-call support and production issue management.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
- 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Administration.
- Strong experience with Linux/Unix systems administration.
- Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud Platform (Google Cloud Platform).
- Proficiency in scripting/programming languages such as Python, Bash, Go, or PowerShell.
- Experience with containerization technologies such as Docker and Kubernetes.
- Knowledge of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
- Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.
- Strong understanding of networking, DNS, load balancing, and security principles.
- Experience with monitoring tools such as Prometheus, Grafana, Datadog, New Relic, Splunk, or ELK Stack.
Preferred Qualifications
- Experience supporting large-scale distributed systems.
- Cloud certifications (AWS, Azure, or Google Cloud Platform).
- Kubernetes Administrator (CKA) or similar certification.
- Understanding of microservices architecture.
- Experience with Chaos Engineering and Reliability Engineering practices.
- Knowledge of database administration and performance tuning.
Technical Skills
Cloud & Infrastructure
- AWS, Azure, Google Cloud Platform
- VMware, Linux, Unix
Automation & IaC
- Terraform
- Ansible
- CloudFormation
- PowerShell
- Bash
- Python
Containers & Orchestration
- Docker
- Kubernetes
- OpenShift
Monitoring & Observability
- Prometheus
- Grafana
- Datadog
- Splunk
- ELK Stack
- New Relic
CI/CD
- Jenkins
- Azure DevOps
- GitHub Actions
- GitLab CI/CD
Key Competencies
- Problem-solving and troubleshooting
- Incident management
- Automation mindset
- Strong communication and collaboration skills
- Analytical thinking
- Ownership and accountability
- Continuous improvement mindset
Success Metrics
- System uptime and availability
- Mean Time to Detect (MTTD)
- Mean Time to Resolve (MTTR)
- Service reliability and performance
- Deployment success rate
- Reduction in operational toil through automation
Location: Remote/Hybrid/Onsite
Experience: 4-10+ Years
Employment Type: Full-Time
This JD can be used for hiring Mid-Level to Senior SRE professionals in enterprise cloud and production support environments.
Similar jobs
- AW
Delivery Consultant - DevOps, WWPS ProServe
Amazon Web Services, Inc.
Denver, CO🇺🇸$131.3k - $177.6k/yrHybrid4 weeks agoRubySwiftAWS+9Technology - KT
REQT: Direct Client : Senior Site Reliability Engineer @ Austin, TX – Hybrid – Only Locals
NewKSN Technologies, Inc.
Austin, TX🇺🇸Hybrid22 hours agoDockerAWSSplunk+8Technology - IN
Lead Platform Engineer (SRE)
NewInfoVision, Inc.
United States🇺🇸Hybrid22 hours agoMicroservicesELKLoad Balancing+8Technology - IF
Systems Engineer III - DevOps & Application Sustainment with Security Clearance
NewIntegral Federal
Tysons Corner, VA🇺🇸$118k - $133k/yrHybrid22 hours agoAWSAnsibleBash+7Technology - KG
DevOps SRE Engineer
NewK&K Global Talent Solutions
Austin, TX🇺🇸On-site22 hours agoDockerAWSArgoCD+12Technology - RT
Trading Application Support DevOps Engineer
Request Technology, LLC
Downers Grove, IL🇺🇸On-site3 days agoSQLTCP/IPBash+3Technology