SRE Observability & Reliability Engineer
Why This Role Stands Out
This hybrid role offers you the chance to significantly impact enterprise-level applications and platforms by building scalable observability solutions, making it an excellent opportunity for skill development in a reputable global IT talent partner. If you thrive in fast-paced environments and possess expertise in SRE, Cloud, and observability tools, this position is perfect for you to advance your career. Apply today to join a dynamic team and contribute to operational excellence.
Quick Overview
Job Description
About Us:
CodeForce 360 is a trusted global IT talent partner helping Fortune 500 companies, system integrators, and enterprise organizations build high-performing technology teams. With over 16 years of industry expertise, we combine speed, precision, and market intelligence to deliver exceptional talent across today's most in-demand technologies.
About the Job- We are looking for an experienced SRE Observability & Reliability Engineer to join one of our enterprise client engagements. The ideal candidate should have strong expertise in SRE, Cloud, Dynatrace, KPI & Scorecard Design, Open Telemetry along with experience working in fast-paced enterprise environments.
Job Description:
We are seeking a Senior SRE Observability & Reliability Engineer to drive reliability, observability, and operational excellence across enterprise applications and platforms. This role will assess application health and operational maturity, identify gaps in monitoring, alerting, logging, tracing, and incident management processes, and implement scalable observability solutions that improve system performance, availability, and business outcomes.
The ideal candidate will have extensive experience in Site Reliability Engineering (SRE), Application Performance Monitoring (APM), observability platforms, telemetry engineering, and operational analytics. Responsibilities include designing and implementing monitoring and telemetry frameworks, optimizing alerting and incident response processes, building executive dashboards and reliability scorecards, establishing SRE governance standards, and driving continuous improvement through data-driven insights.
Key Responsibilities
- Assess application reliability, performance, telemetry coverage, and operational maturity.
- Identify gaps in observability, monitoring, logging, tracing, RCA processes, alerting, and reporting.
- Design and implement APM, distributed tracing, structured logging, and telemetry ingestion solutions.
- Define and optimize SLIs, SLOs, alerting strategies, and reliability metrics.
- Understand business process, lead RCA activities, and establish tagging, compliance, and governance standards.
- Develop executive dashboards, scorecards, heatmaps, and operational analytics using Grafana, ServiceNow Performance Analytics, and Power BI.
- Build telemetry pipelines and integrate data from Splunk, Dynatrace, AppDynamics, ServiceNow, and cloud platforms.
- Develop runbooks, monitoring standards, and SRE best practices to improve operational effectiveness.
Required Skills & Technologies
- Observability & Monitoring: Splunk, Dynatrace, Grafana, AppDynamics, Open Telemetry, ServiceNow Performance Analytics
- Programming & Automation: Python, Java, .NET, REST APIs, scripting and automation
- Cloud & Infrastructure: Kubernetes, AWS, Google Cloud Platform, Microservices, Container Platforms
- Analytics & Reporting: Power BI, ETL/Data Integration, Data Modelling, Dashboard Development, KPI & Scorecard Design
- Reliability Engineering: SRE, Incident Management, RCA, Availability Engineering, Service Health Monitoring, Operational Governance
Experience
- 8–10+ years of experience in SRE, Observability Engineering, Infrastructure Operations, or Application Support.
- Proven experience implementing enterprise observability and telemetry solutions.
- Strong expertise in executive reporting, reliability scorecards, and operational analytics.
- Experience integrating monitoring and telemetry platforms into a unified observability ecosystem.
- Strong stakeholder management and communication skills with the ability to translate technical insights into business value.
This role is ideal for a hands-on reliability engineering professional who can combine deep technical expertise with operational leadership to improve application resilience, observability maturity, and service reliability across the enterprise.
How To Apply
Job ID: JPC - 234109
Contact:
Name: Bhushan Reddy
Email:
Phone:
CodeForce 360 proudly provides equal employment opportunities to all employees and applicants and prohibits discrimination and harassment of any kind without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable federal, state, or local laws.
Skills
Similar jobs
Mid-Senior Site Reliability Engineer Kubernetes Platform
StratEdge It consulting INC · San Jose, United States
5 minutes ago$120k - $130k/yrLead Software Engineer (DevOps)
Purplejack Technologies LLC · Boston, United States
7 minutes agoDevops Engineer
Flexon Technologies Inc. · Austin, United States
8 minutes agoSRE + AI + Dynatrace - Atlanta GA [hybrid] - Contract
Exatech Inc · Atlanta, United States
27 minutes agoLead SRE observability - G10
New York Technology Partners · United States
49 minutes agoLinux System Admin with Python /SRE
New York Technology Partners · United States
49 minutes ago