Why This Role Stands Out
This hybrid SRE/Observability Engineer role at Akshaya Inc. offers a fantastic opportunity to drive significant improvements in system reliability and efficiency using cutting-edge observability platforms. You'll thrive here if you're a seasoned engineer passionate about optimizing performance and enhancing user experience, with ample scope for professional growth in a well-regarded company. Apply today to shape the future of their technological infrastructure!
Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
McLean, VA, United States
Posted
1 week ago
AWSSplunkJavaScriptJenkinsPrometheusPythonStakeholder Management
Job Description
Role: SRE/Observability Engineer
Location: McLane, VA
Job Summary
We are seeking a highly skilled Senior Observability Engineer with 9+ years of experience in designing, implementing, and optimizing enterprise observability solutions. The ideal candidate will possess deep expertise in modern observability platforms, application performance monitoring, Open Telemetry implementation and cloud technologies, with a strong focus on improving system reliability, operational efficiency, and user experience.
Key Responsibilities
Analyze the existing observability solution deployed in Elastic cloud and understand the gap;
Document ideal scenario versus existing deployment and recommend the changes required to bring the Observability solution to improve overall application monitoring
implement end-to-end observability solutions for distributed and cloud-native applications; Work with development, Infrastructure and application support team to streamline the application monitoring using Elastic Cloud
Develop comprehensive monitoring strategies covering infrastructure, applications, logs, traces, metrics, and user experience.
Migrate application monitoring from legacy monitoring platforms (Dynatrace, Splunk, Prometheus ) to modern observability platforms such as Dynatrace and Elastic.
Design and implement Elastic-based monitoring architectures, including data pipelines, storage, APM, dashboards, and advanced analytics.
Build custom extensions, automated workflows, and synthetic monitoring solutions using Python and JavaScript.
Integrate observability platforms with CI/CD pipelines (Jenkins and related tools) to automate monitoring, alerting, and incident management.
Configure OpenPipeline, Business Events, anomaly detection, and AI-driven analytics to improve operational visibility.
Optimize observability platform licensing, data ingestion, and storage costs while maintaining monitoring effectiveness.
Collaborate with Development, Infrastructure, and Operations teams to improve application reliability, performance, and operational excellence.
Conduct dashboard reviews, monitoring assessments, and observability maturity improvements across enterprise applications.
Support proactive monitoring, root cause analysis, incident response, and continuous service improvement initiatives.
Required Skills & Qualifications
9+ years of experience in Application Performance Monitoring (APM), Observability, and Monitoring Engineering.
Excellent knowledge about Deploying observably using Open Telemetry framework Strong expertise in onboarding application in Elastic Search (;/div>
Strong knowledge of:
o Distributed tracing
o Log analytics
o Infrastructure and application monitoring
o Synthetic monitoring
o Real User Monitoring (RUM)
Experience developing automation using Python and JavaScript.
Experience integrating monitoring platforms with Jenkins and CI/CD pipelines.
Hands on experience in implementing Observability for Container based application
Strong understanding of observability architecture, SRE principles, and cloud-native monitoring practices.
Hands-on experience with AWS cloud platforms.
Knowledge of anomaly detection, Open Pipeline configuration, Business Events, and AI-assisted observability.
Strong analytical, troubleshooting, and root cause analysis skills.
Experience designing scalable enterprise observability architectures.
Knowledge of license optimization and observability cost management.
Experience implementing AI-driven observability and automated incident management.
Excellent communication, stakeholder management, and cross-functional collaboration skills.
Passion for driving operational excellence through automation, proactive monitoring, and observability best practices.
Similar jobs
- IN
Member of Technical Staff, Production Site Reliability Engineer
NewInferact
San Francisco🇺🇸$200k - $400k/yrOn-site7 hours agoDockerBashKubernetes+2Technology - 2K
Lead Engineer, Dev Ops - WWE 2K
New2K
California🇺🇸$138.6k/yr10 hours agoCADEngineering - 2K
DevOps Engineer
New2K
Austin🇺🇸10 hours agoGCPAWSSplunk+10Technology - PP
Director of DevOps and Cloud, Hands-on, Onsite in Charlotte, NC
NewParallel Partners
Charlotte, NC🇺🇸$170k - $220k/yrOn-site11 hours agoEncryptionNew RelicSOC 2+13Technology - PP
Machine Learning Platform Engineer, Machine Learning (ML) and Artificial Intelligence (AI) Required, Work From Home
NewParallel Partners
San Francisco, CA🇺🇸$140k - $180k/yrRemote12 hours agoMachine LearningLLMPyTorch+1Technology - MA
Senior Site Reliability Engineer
NewMastercard
O Fallon, Missouri🇺🇸$96k - $163k/yrHybrid52 minutes agoDockerGCPMicroservices+12Technology