Why This Role Stands Out
This fully remote Site Reliability Engineer role offers a fantastic opportunity to contribute to global, business-critical platforms, fostering significant career growth through complex problem-solving and automation. You'll thrive here if you're an experienced SRE, DevOps, Platform, or Cloud Engineer with strong Kubernetes and Azure skills, eager to enhance system reliability and collaborate with dynamic teams. Embrace the chance to make a real impact and elevate your expertise in a globally recognized technology organization.
Quick Overview
Job Description
An exciting global technology organisation is looking for a Site Reliability Engineer (SRE) to join its growing engineering team. This is a fully remote position, offering the opportunity to work on large-scale, business-critical platforms used by customers around the world.
The role would suit an experienced Site Reliability, DevOps, Platform or Cloud Engineer with strong hands-on experience across Kubernetes and Microsoft Azure who enjoys solving complex production problems, improving reliability and automating manual processes.
You will work closely with Development and DevOps teams, helping to design, build, operate and scale highly available production environments while ensuring services remain reliable, secure and performant.
Responsibilities:
- Supporting and improving highly available, business-critical production services
- Monitoring production environments to ensure availability, scalability, performance and security
- Responding to production incidents, diagnosing complex technical issues and restoring services
- Carrying out root cause analysis and contributing to post-incident reviews
- Building automation to reduce manual and repetitive operational tasks
- Using Infrastructure as Code to improve the consistency and scalability of environments
- Working closely with Development and DevOps teams throughout the application release process
- Balancing the delivery of new functionality with platform reliability and service-level objectives
- Identifying bottlenecks and proposing improvements across infrastructure and applications
- Improving the reliability, quality and time-to-market of software solutions
- Supporting the continued development and scaling of a global product platform
- Participating in an out-of-hours on-call rota
Key Skills:
- Strong hands-on Kubernetes experience
- Strong hands-on Microsoft Azure experience
- A solid background within Site Reliability Engineering, DevOps, Platform Engineering or a similar environment
- Terraform and Infrastructure as Code experience
- Strong production troubleshooting and incident response experience
- Experience supporting highly available production environments
- A strong understanding of reliability, scalability and automation
- Experience with scripting or programming, ideally PowerShell or Python or Go
- Knowledge of database technologies such as MySQL
- A security-first approach to designing and operating infrastructure
Similar jobs
- EG
Quality Engineer
Ernest Gordon Recruitment
Telford, Shropshire🇬🇧£30k - £35k/yrHybrid6 days agoEngineering - VA
Electrical Service Engineer
VANRATH
Lisburn, Northern Ireland🇬🇧£35k - £40k/yrHybrid6 days agoEngineering - HA
Senior Principal Mechanical Engineer
Hackajob Ltd
Edinburgh City Centre, Edinburgh🇬🇧£44.2k - £56k/yrHybrid3 weeks agoCADFEAMechanical DesignEngineering - SR
IT Field Service Engineer
Sanderson Recruitment
Exeter, Devon🇬🇧£24.6k - £27.7k/yrOn-site3 months agoEngineering - DJ
Engineering Manager
NewDemob Job Ltd
Milton Keynes, Buckinghamshire🇬🇧£60k/yrHybrid9 hours agoEngineering - PS
Multi Skilled Maintenance Engineer
Pioneer Selection
Thatcham, West Berkshire🇬🇧£45k/yrOn-site3 days agoEngineering