Why This Role Stands Out
This hybrid SRE/Observability Engineer role at Akshaya Inc. offers a fantastic opportunity to drive significant improvements in system reliability and efficiency using cutting-edge observability platforms. You'll thrive here if you're a seasoned engineer passionate about optimizing performance and enhancing user experience, with ample scope for professional growth in a well-regarded company. Apply today to shape the future of their technological infrastructure!
Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
McLean, VA, United States
Posted
3 days ago
AWSSplunkJavaScriptJenkinsPrometheusPythonStakeholder Management
Job Description
Role: SRE/Observability Engineer
Location: McLane, VA
Job Summary
We are seeking a highly skilled Senior Observability Engineer with 9+ years of experience in designing, implementing, and optimizing enterprise observability solutions. The ideal candidate will possess deep expertise in modern observability platforms, application performance monitoring, Open Telemetry implementation and cloud technologies, with a strong focus on improving system reliability, operational efficiency, and user experience.
Key Responsibilities
Analyze the existing observability solution deployed in Elastic cloud and understand the gap;
Document ideal scenario versus existing deployment and recommend the changes required to bring the Observability solution to improve overall application monitoring
implement end-to-end observability solutions for distributed and cloud-native applications; Work with development, Infrastructure and application support team to streamline the application monitoring using Elastic Cloud
Develop comprehensive monitoring strategies covering infrastructure, applications, logs, traces, metrics, and user experience.
Migrate application monitoring from legacy monitoring platforms (Dynatrace, Splunk, Prometheus ) to modern observability platforms such as Dynatrace and Elastic.
Design and implement Elastic-based monitoring architectures, including data pipelines, storage, APM, dashboards, and advanced analytics.
Build custom extensions, automated workflows, and synthetic monitoring solutions using Python and JavaScript.
Integrate observability platforms with CI/CD pipelines (Jenkins and related tools) to automate monitoring, alerting, and incident management.
Configure OpenPipeline, Business Events, anomaly detection, and AI-driven analytics to improve operational visibility.
Optimize observability platform licensing, data ingestion, and storage costs while maintaining monitoring effectiveness.
Collaborate with Development, Infrastructure, and Operations teams to improve application reliability, performance, and operational excellence.
Conduct dashboard reviews, monitoring assessments, and observability maturity improvements across enterprise applications.
Support proactive monitoring, root cause analysis, incident response, and continuous service improvement initiatives.
Required Skills & Qualifications
9+ years of experience in Application Performance Monitoring (APM), Observability, and Monitoring Engineering.
Excellent knowledge about Deploying observably using Open Telemetry framework Strong expertise in onboarding application in Elastic Search (;/div>
Strong knowledge of:
o Distributed tracing
o Log analytics
o Infrastructure and application monitoring
o Synthetic monitoring
o Real User Monitoring (RUM)
Experience developing automation using Python and JavaScript.
Experience integrating monitoring platforms with Jenkins and CI/CD pipelines.
Hands on experience in implementing Observability for Container based application
Strong understanding of observability architecture, SRE principles, and cloud-native monitoring practices.
Hands-on experience with AWS cloud platforms.
Knowledge of anomaly detection, Open Pipeline configuration, Business Events, and AI-assisted observability.
Strong analytical, troubleshooting, and root cause analysis skills.
Experience designing scalable enterprise observability architectures.
Knowledge of license optimization and observability cost management.
Experience implementing AI-driven observability and automated incident management.
Excellent communication, stakeholder management, and cross-functional collaboration skills.
Passion for driving operational excellence through automation, proactive monitoring, and observability best practices.
Similar jobs
- MI
Site Reliability Engineer II - CTJ - Poly with Security Clearance
NewMicrosoft Corporation
Reston, VA🇺🇸$102.1k - $202.2k/yrHybrid12 hours agoC#C++HTTPS+3Technology - RH
DevOps Engineer
NewRobert Half
Pittsburgh, PA🇺🇸Hybrid12 hours agoDockerAWSAzure+7Technology - BT
Cloud Platform Engineer (BT-26168) with Security Clearance
NewBastion Technologies, Inc.
Houston, TX🇺🇸HybridYesterdaySQLMachine LearningAgile+5Technology - SH
DevSecOps Engineer
ShorePoint, Inc
United States🇺🇸Hybrid6 weeks agoAgileEngineering - LM
Systems Engineer
Lockheed Martin Corporation
Mount Laurel Township, NJ🇺🇸$75k - $136.0k/yrHybrid7 weeks agoAgileC++Java+2Technology - BO
Site Reliability Engineer (Associate or Experienced)
NewBoeing
Saint Louis, Missouri🇺🇸$99.5k - $134.6k/yrOn-site2 hours agoSQLAWSSonarQube+15Technology