Quick Overview
Job Description
Key Responsibilities:Lead Site Reliability Engineering initiatives focused on application availability, scalability, performance, and resiliency.Design, implement, and maintain highly reliable and scalable production environments.Develop and maintain automation for infrastructure, deployments, monitoring, and operational processes.Establish and improve SLOs, SLIs, and SLAs and drive reliability metrics across applications and services.Lead incident management, troubleshooting, root-cause analysis, and post-incident reviews.Implement and maintain comprehensive monitoring, logging, alerting, and observability solutions.Work closely with development, QA, DevOps, infrastructure, and security teams to improve system reliability.Automate repetitive operational tasks and reduce manual intervention through scripting and infrastructure automation.Support CI/CD pipelines and implement reliability and quality checks throughout the software delivery lifecycle.Analyze system performance, identify bottlenecks, and implement solutions to improve application performance and stability.Participate in production deployments, release management, capacity planning, and disaster-recovery initiatives.Mentor engineers and provide technical leadership on SRE best practices.Define and document operational procedures, runbooks, troubleshooting guides, and disaster-recovery processes.Required Skills:10+ years of experience in SRE, DevOps, Production Engineering, Infrastructure Engineering, or a related field.Strong experience with Site Reliability Engineering principles and practices.Hands-on experience with AWS, Azure, or Google Cloud Platform cloud platforms.Strong knowledge of Kubernetes and Docker.Experience with Terraform or other Infrastructure as Code tools.Strong scripting/programming experience with Python, Shell scripting, or similar languages.Hands-on experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.Strong experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, Datadog, New Relic, or ELK.Experience with production incident management, troubleshooting, and root-cause analysis.Strong understanding of Linux/Unix systems, networking, APIs, and distributed systems.Experience with Git and modern software development practices.Strong understanding of SLOs, SLIs, SLAs, error budgets, and reliability engineering.Experience working in Agile/Scrum environments.Excellent communication, problem-solving, and leadership skills.Preferred Qualifications:Experience in banking, financial services, or other highly regulated environments.Experience with microservices and cloud-native architectures.Experience implementing observability and distributed tracing using tools such as OpenTelemetry.Experience with Kafka or other distributed messaging platforms.Experience with performance engineering and capacity planning.Experience with disaster recovery, high availability, and business continuity.Experience leading SRE/DevOps teams or enterprise reliability initiatives.
Similar jobs
- PP
Senior Snowflake Platform Engineer
NewPraxis Precision Medicines, Inc.
United States - Remote🇺🇸Remote4 hours agoSQLAWSSnowflake+2Technology - PP
Senior Data Platform Engineer, Commercial
NewPraxis Precision Medicines, Inc.
United States - Remote🇺🇸Remote4 hours agoSQLETLSnowflake+5Technology - OC
Lead DevOps Engineer
NewOctus
Remote - US🇺🇸Remote6 hours agoDockerMicroservicesAWS+5Technology - ON
Site Reliability Engineering Lead
NewOneapp
United States (Remote)🇺🇸Remote3 hours agoProcurementPythonTechnology - MY
Site Reliability Engineer
NewMyFitnessPal
Remote - US🇺🇸$120k - $165k/yrRemote5 hours agoAWSPCI DSSSOC 2+8Technology - KR
AI Platform Engineer, Enablement and Governance Operations
NewKraken.com
United States🇺🇸Remote4 hours agoOAuthSSOCompliance+7Technology