100% REMOTE: Site Reliability Engineer (SRE) Observability & Monitoring
Why This Role Stands Out
This 100% remote Site Reliability Engineer role offers a fantastic opportunity to build and scale critical observability frameworks, directly impacting system reliability and performance. If you are a seasoned SRE with a passion for OpenTelemetry and a knack for proactive monitoring, you will thrive in this impactful position. Apply now to contribute your expertise and advance your career in a flexible, well-compensated environment.
Quick Overview
Job Description
****** 100% REMOTE ******
Position: Site Reliability Engineer (SRE) Observability & Monitoring
Location: 100% Remote
Duration: 6-12+ Months
Rate: $63-$65/Hr
About the Role
We are seeking a Site Reliability Engineer (SRE) with strong expertise in observability, monitoring, and distributed tracing to join our SRE team. The ideal candidate will help us design, build, and scale an observability framework that provides end-to-end visibility into our systems and applications. A strong focus will be places on OpenTelemetry, as we continue to standardize our telemetry pipeline across logs, metrics, and traces.
Responsibilities
- Design, implement, and maintain observability solutions using OpenTelemetry, Prometheus, Grafana, AppDynamics, and Splunk.
- Build and manage telemetry pipelines (metrics, logs, traces) ensuring reliable data collection, transformation, and export.
- Lead initiatives to improve incident detection, response, and post-incident analysis with a strong emphasis on RCA (Root Cause Analysis).
- Define and maintain SLIs, SLOs, and error budgets to measure and improve system reliability.
- Partner with development and operations teams to instrument applications and services for better monitoring and tracing coverage.
- Develop dashboards, alerts, and visualizations to provide actionable insights into system health and performance.
- Contribute to automation and self-healing practices that improve uptime and reduce operational toil.
- Stay current with trends in observability and advocate best practices across the engineering organization.
Requirements
- 10+ years of SRE/ Devops/ Cloud/ Infrastructure engineering experience with a focus on monitoring and observability.
- Strong communication skills with the ability to articulate technical requirements and explain benefits of observability clearly to development teams.
- Experience working in an area that requires strong business knowledge and the ability to pull together business parters and development teams to perform RCA across distributed systems.
- Experience implementing an Observability Framework at an organization.
- Hands-on experience with OpenTelemetry SDKs, collectors, and exporters.
- Proficiency with observability stacks such as Prometheus, Grafana, Loki, Tempo, Elastic Stack, or Splunk Observability (Splunk/AppDynamics).
- Strong knowledge on cloud platforms (Google Cloud Platform, or Azure).
- Hand-on experience on container orchestration using Kubernetes (OCP, GKE, AKS)
- Familiarity with CI/CD pipelines like (Jenkins and Github actions), infrastructure as code (Terraform/Ansible/ARM/CloudFormation).
- Experience provisioning infrastructure and capacity planning.
- Hands-on skills in programming languages like Java, and Python.
Its a 100% Remote Opportunity
Thank You,
Augustin Ahmed
Sr Resource Manager
Parmesoft Inc,
Similar jobs
- TS
TTG-345 - Palantir Foundry Data Platform Engineer - $339,950 Tot with Security Clearance
NewTTG Solutions Inc.
Herndon, VA🇺🇸$204k - $247k/yrHybrid21 hours agoETLLESSPythonTechnology - TS
TTG-230 - DevOps Infrastructure Automation Engineer - $265,100 T with Security Clearance
NewTTG Solutions Inc.
Chantilly, VA🇺🇸$158k - $192k/yrHybrid21 hours agoDockerSOAPShell+9Technology - V1
Senior Data Platform Engineer
NewVersion 1
United States🇺🇸Hybrid6 hours agoSQLSQL ServerAWS+5Technology - OR
Senior Site Reliability Engineer
NewOracle Corporation
Nashville, TN🇺🇸$81.1k - $187k/yrHybridYesterdayOracleAnsibleBash+4Technology - OR
Principal Site Reliability Engineer
Oracle Corporation
Nashville, TN🇺🇸$84.9k - $209.5k/yrHybrid3 days agoOracleLoad BalancingAnsible+5Technology - SP
Sr. Site Reliability Engineer, Platform Infrastructure
NewSpaceX
Bastrop, TX🇺🇸Hybrid21 hours agoDockerAnsibleKubernetes+3Technology - CI
DevOps Engineer
CACI International, Inc.
Annapolis, MD🇺🇸$103.8k - $218.1k/yrHybrid4 days agoDockerMongoDBMySQL+12Technology - SY
Mid-Level Cloud Platform Engineer [$274k/yr+] with Security Clearance
NewSYSTOLIC
Annapolis Junction, MD🇺🇸$274k/yrHybrid21 hours agoDockerShellAWS+11Technology - BS
Senior DevOps Software Engineer with Security Clearance
NewBase-2 Solutions, LLC
Bethesda, MD🇺🇸$10k/yrHybridYesterdaySAFeMicroservicesLogstash+17Technology - BS
DevOps Software Engineer with Security Clearance
NewBase-2 Solutions, LLC
Bethesda, MD🇺🇸$10k/yrHybrid21 hours agoMicroservicesLogstashAgile+12Technology - BA
Senior DevOps Engineer
NewBooz Allen Hamilton
Washington, DC🇺🇸$99k - $225k/yrOn-site21 hours agoDockerShellAWS+7Technology - JM
Vice President - Senior Manager of Site Reliability Engineering (Chief Data & Analytics Office)
NewJ.P. Morgan
Jersey City, New Jersey🇺🇸On-site18 hours agoDockerAWSSplunk+8Technology