Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
AZ, United States
Posted
Yesterday
ShellSplunkGrafanaKubernetesPrometheusPythonREST
Job Description
Observability Operations Engineer
Experience: 7+ Years
Employment Type: Full Time
We are seeking an experienced Observability Operations Engineer to administer, optimize, and support enterprise-scale observability platforms. The ideal candidate will have strong hands-on experience with Dynatrace, Splunk, OpenSearch/Elasticsearch, along with expertise in monitoring, logging, tracing, alerting, dashboards, automation, and production troubleshooting.
Must Have Technical/Functional Skills- Strong Observability Administration experience with Dynatrace, Splunk, and OpenSearch/Elasticsearch.
- Hands-on experience with monitoring, logging, tracing, alerting, dashboards, and platform performance tuning.
- Strong knowledge of Linux, Kubernetes, containers, and cloud environments.
- Experience with Grafana, Prometheus, OpenTelemetry, and related observability technologies.
- Automation experience using Python, Shell scripting, and REST APIs.
- Experience supporting enterprise-scale production environments, troubleshooting incidents, and performing Root Cause Analysis (RCA).
- Knowledge of observability best practices, platform security, capacity planning, upgrades, patching, and operational governance.
- Administer, configure, and optimize enterprise Dynatrace, Splunk, and OpenSearch/Elasticsearch platforms.
- Maintain platform availability, scalability, performance, security, and reliability.
- Build and manage monitoring, logging, tracing, dashboards, alerts, and operational metrics.
- Monitor platform health and proactively identify performance and capacity issues.
- Troubleshoot complex production issues using observability tools and perform detailed Root Cause Analysis (RCA).
- Support Linux, Kubernetes, containerized, and cloud-based environments.
- Develop automation for repetitive operational activities using Python, Shell scripting, and REST APIs.
- Drive self-healing, automation, and AI-assisted operations initiatives.
- Manage platform upgrades, patching, capacity planning, backups, and operational governance.
- Collaborate closely with SRE, DevOps, Platform, Infrastructure, and Application teams.
- Establish and maintain operational standards, documentation, and best practices for observability platforms.
- Participate in incident management, problem management, and continuous improvement initiatives.
- Strong verbal and written communication skills.
- Ability to communicate clearly and assertively with technical and business stakeholders.
- Strong team player with the ability to collaborate across multiple engineering teams.
- Excellent analytical and problem-solving skills.
- Ability to work effectively in a production support environment and handle high-priority incidents.
- Strong ownership, accountability, and attention to detail.
Dynatrace | Splunk | OpenSearch/Elasticsearch | Grafana | Prometheus | OpenTelemetry | Kubernetes | Linux | Python | Shell Scripting | REST APIs | Cloud | Observability | Production Support | RCA
Similar jobs
- DE
Cloud Platform Engineer
Decisionpoint Corporation
United States🇺🇸Remote4 weeks agoAWSCloudFormationKubernetes+3Technology - C-
Site Reliability Engineer II
C-Serv
San Jose, California🇺🇸Hybrid3 days agoAWSLoad BalancingArgoCD+8Technology - SH
Site Reliability Engineer
NewShaarpro
United States🇺🇸HybridYesterdayMicroservicesNode.jsAWS+14Technology - GD
DevOps Automation Engineer
GDH
United States🇺🇸$58 - $61/hrRemote1 week agoAWSAnsibleBash+7Technology - OP
Platform Engineer-Adaptiv
NewOpTech
Cincinnati, OH🇺🇸On-siteYesterdaySQLShellAWS+4Technology - VE
Staff Data Platform Engineer - Finance
Vercel
Hybrid - San Francisco🇺🇸3 days agoGCPNext.jsAWS+15Technology