Quick Overview
Job Description
Senior Site Reliability Engineer (AWS focused)
Up to £75,000 + bonus + on call allowance
Milton Keynes (2 days on site a week)
VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring for a Senior Site Reliability Engineer to build stability, respond to live incidents, and assist with system upkeep. The role will also play a key part in on implementing and adopting new tooling and processes surrounding the wider transformation.
This is a genuine opportunity to own and operate how the cloud function works, and progress into a team lead position as the team grows.
Experience required for the Senior Site Reliability Engineer
- Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within a customer facing environment - eg SaaS or MSP.
- Strong hands-on experience with both AWS, and on-premise virtual machines.
- Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor).
- Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working.
- Ability to communicate across internal teams and external customers.
- Skilled in networking across both cloud (Azure) and on premise environments.
- Either Windows or Linux systems administration skills (Linux preferred).
- Previous use of AI tools to enhance efficiency.
Job Duties of the Senior Site Reliability Engineer
- Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure Servers and networks, and automate application life cycles.
- Regularly use Datadog and other observability tools for application performance monitoring.
- Implement new ways of working, helping to shape how the organisation responds and recovers to incidents.
- Take ownership of incident resolutions.
- Actively drive down key reliability metrics (MTTR, incident frequency, on-call toil) by evaluating key incidents.
- Work on an a on call rota, ensuring you are available to respond to incidents during this time.
- Identify areas for automation and help implement changes that raise the bar for reliability.
Apply now to speak with VIQU IT in confidence. Or reach out to Jack McManus via the (see below)
Do you know someone great? We'll thank you with up to £1,000 if your referral is successful (terms apply).
Similar jobs
- IA
SC/NPPV3 DevOps Engineer - Azure
NewIO Associates
London🇬🇧£500 - £550/moHybrid1 hour agoDockerAnsibleAzure+10Technology - SG
DevOps and Infrastructure Engineer
NewSanderson Government & Defence
Gloucestershire🇬🇧Hybrid1 hour agoAgileArgoCDGrafana+3Technology - HA
Principal Platform Engineer
NewHackajob Ltd
Manchester🇬🇧On-site1 hour agoTechnology - HA
Lead SRE - AWS Platform
NewHackajob Ltd
Glasgow, Lanarkshire🇬🇧Hybrid1 hour agoTechnology - HA
Product Associate - SRE Team - Chase UK
NewHackajob Ltd
London🇬🇧Hybrid1 hour agoTechnology - HA
Sr Lead AI Platform Engineer
NewHackajob Ltd
Glasgow, Lanarkshire🇬🇧Hybrid1 hour agoTechnology