Haystack
← Back to Jobs
full time
Technology
VI

Senior SRE (AWS)

VIQU ITMilton Keynes, Buckinghamshire🇬🇧United KingdomPosted 28 Sept 2026

Quick Overview

Salary
£75k/yr
Seniority
Mid Senior
Employment type
Full Time
Work mode
On Site
Location
Milton Keynes, Buckinghamshire, United Kingdom
Posted
23 hours ago
AWSAzureDatadogGrafanaKubernetesPrometheusTerraform

Job Description

Senior Site Reliability Engineer (AWS focused) Up to £75,000 + bonus + on call allowance Milton Keynes (2 days on site a week) VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring for a Senior Site Reliability Engineer to build stability, respond to live incidents, and assist with system upkeep.

The role will also play a key part in on implementing and adopting new tooling and processes surrounding the wider transformation. This is a genuine opportunity to own and operate how the cloud function works, and progress into a team lead position as the team grows.

Experience required for the Senior Site Reliability Engineer Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within a customer facing environment – e.g SaaS or MSP.<br />Strong hands-on experience with both AWS, and on-premise virtual machines.

Experience withInfrastructure as Code / Terraform, Container orchestration (Kubernetes), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor).<br />Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working.<br />Ability to communicate across internal teams and external customers.<br />Skilled in networking across both cloud (Azure) and on premise environments.<br />Either Windows or Linux systems administration skills (Linux preferred).<br />Previous use of AI tools to enhance efficiency.<br />Job Duties of the Senior Site Reliability Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles.<br />Regularly use Datadog and other observability tools for application performance monitoring.<br />Implement new ways of working, helping to shape how the organisation responds and recovers to incidents.<br />Take ownership of incident resolutions.

Actively drive down key reliability metrics (MTTR, incident frequency, on-call toil) by evaluating key incidents.<br />Work on an a on call rota, ensuring you are available to respond to incidents during this time.

Identify areas for automation and help implement changes that raise the bar for reliability.<br />Apply now to speak with VIQU IT in confidence. Or reach out to Jack McManus via the (url removed) Do you know someone great? We’ll thank you with up to £1,000 if your referral is successful (terms apply).

For more exciting roles and opportunities like this, please follow us on LinkedIn @VIQU IT Recruitment

Similar jobs