Site Reliability & Observability Engineer – Datadog / Azure
Quick Overview
Job Description
Site Reliability & Observability Engineer / Datadog – Synthetic Monitoring, APM, RUM, Log Management, SLO's, Alerting / Azure / Azure DevOps / Cloudflare / 6-month contract / Hybrid – West Midlands / Remote / £450 – 600 per day Inside IR35.
One of our leading clients is seeking a Lead Site Reliability & Observability Engineer to build and operate a world-class monitoring, synthetic testing, and reliability platform.
Location – West Midlands / Remote – 5 days per week with 1-2 days per week onsite
Duration – 6 months +
Day rate – £450 – 600 per day Inside IR35
This role will lead the implementation of Datadog across Azure and Cloudflare, creating a comprehensive early warning system that continuously validates APIs, integrations, and customer user journeys in production.
Key Responsibilities:
·? Own and evolve the Datadog observability platform.
· Design and maintain synthetic monitoring for critical API and UI workflows.
· Build continuous production validation covering business-critical customer journeys.
· Integrate monitoring, testing, dashboards, and alerting into Azure DevOps and GitHub pipelines.
· Develop monitoring-as-code and testing-as-code practices using Terraform.
· Create actionable dashboards, SLOs, SLIs, alerts, and anomaly detection.
· Integrate Datadog with Azure, Cloudflare, and modern SaaS architectures.
· Drive reliability, performance, and root-cause analysis across production systems.
Required Experience:
· Strong hands-on Datadog expertise, including:
o Synthetic Monitoring
o APM
o RUM
o Log Management
o SLOs and Alerting
· Experience operating large-scale global SaaS platforms.
· Deep Azure experience.
· Experience integrating Cloudflare services.
·? Strong CI/CD experience with Azure DevOps and GitHub.
·? Expertise in API, integration, and browser-based testing.
· Infrastructure as Code experience using Terraform.
· Experience with distributed systems, microservices, and cloud-native architectures.
Desirable:
· Datadog certifications.
· Azure certifications.
·? Cloudflare administration experience.
· Background in Site Reliability Engineering (SRE) or Platform Engineering leadership roles.
Similar jobs
- RS
DevOps Engineer
NewRichmond Square Consulting Limited
Hereford, Herefordshire🇬🇧Hybrid1 hour agoTechnology - CG
Lead Cloud Platform Engineer
NewCircle Group
City, Leeds🇬🇧Hybrid5 hours agoScalaTechnology - TA
DevOps Engineer SFTPGo Axway Red Hat Banking
NewTelstra Associates
London🇬🇧On-site5 hours agoGitGitHub ActionsGitLab CI+4Technology - CO
Senior DBA/Platform Engineer
NewCongruity360
Belfast, Northern Ireland🇬🇧HybridYesterdaySQLETLAgile+2Technology - AM
DevOps Engineer, Veeqo
NewAmazon
Swansea, Glamorgan🇬🇧Hybrid9 hours agoTechnology - SG
DevOps Engineer (SC Cleared)
NewSanderson Government and Defence
Cheltenham, Gloucestershire🇬🇧£75k/yrHybrid9 hours agoTechnology