Site Reliability Engineer (AWS, Terraform, Ansible, Python, GitOps, SRE, Grafana/Datadog) - Permanen
Quick Overview
Job Description
Site Reliability Engineer (AWS, Terraform, Ansible, Python, GitOps, SRE, Grafana/Datadog) - Permanent - London
An exciting opportunity has arisen for an experienced Site Reliability Engineer to join a highly regulated, enterprise-scale technology environment supporting mission-critical cloud platforms. This is an excellent opportunity for an engineer with a strong infrastructure operations background who is passionate about automation, reliability, and cloud technologies.
Working within a collaborative Platform Operations team, you'll play a key role in driving Site Reliability Engineering (SRE) best practices, improving automation, and enhancing the reliability, scalability, and observability of cloud-hosted environments.
The Role
As a Site Reliability Engineer, you will be responsible for implementing and evolving SRE methodologies across cloud infrastructure while helping to automate operational processes and reduce manual intervention. You'll work closely with infrastructure and engineering teams to improve platform reliability and operational excellence.
Key Responsibilities
- Drive the adoption and implementation of Site Reliability Engineering (SRE) principles across cloud-hosted platforms.
- Improve platform reliability through automation, observability, and continuous improvement initiatives.
- Define and implement SLAs, SLOs and SLIs to improve service performance and availability.
- Identify operational toil and automate repetitive tasks using Infrastructure as Code and automation tools.
- Develop, review and troubleshoot production automation and infrastructure code.
- Enhance GitOps capabilities using Terraform and Ansible Automation Platform.
- Support multi-region, cloud-based environments.
- Participate in an on-call rota, managing production incidents and ensuring platform stability.
- Perform root cause analysis and drive preventative improvements following incidents.
- Collaborate with infrastructure, engineering and operational teams to improve deployment processes and cloud operations.
Required Skills & Experience
- Previous experience in a Site Reliability Engineering (SRE) or Infrastructure Operations role.
- Strong operational support experience, including incident management, root cause analysis and on-call support.
- Experience implementing SRE methodologies within enterprise environments.
- Strong Scripting skills using Python, Ansible or PowerShell.
- Hands-on experience with AWS and/or GCP cloud platforms.
- Experience with Terraform, Infrastructure as Code and GitOps practices.
- Experience with observability and monitoring platforms such as Grafana, Datadog or Dynatrace.
- Strong troubleshooting and analytical skills.
- Excellent communication skills with both technical and business stakeholders.
- Experience working within regulated financial services or banking environments is highly desirable.
Desirable Experience
- Software development background.
- Experience with Ansible Automation Platform.
- Knowledge of ITIL.
- AWS or Terraform certifications.
What's on Offer
- Permanent opportunity.
- London based with 2 days per week onsite.
- Opportunity to work on highly available, enterprise-scale cloud platforms.
- Exposure to modern cloud technologies, automation and SRE best practices.
- Collaborative engineering culture focused on innovation, reliability and continuous improvement.
If you're an experienced Site Reliability Engineer looking to work on large-scale cloud infrastructure where automation, resilience and engineering excellence are at the heart of the platform, we'd love to hear from you.
Skills
Similar jobs
AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM
Scope AT Limited · United Kingdom
56 minutes agoInfrastructure Engineer - potential for hybrid
JJ Associates · UK, United Kingdom
1 hour ago£45k - £50k/yrLead Platform Engineer
Transunion · Hyde Park, United Kingdom
1 hour agoAzure Devops Engineer
VIQU IT · London, United Kingdom
1 hour agoSite Reliability Engineer (SRE)
Randstad Technologies Recruitment · UK, United Kingdom
1 hour agoPlatform Engineer
Raytheon · Gloucester, United Kingdom
3 hours ago