Haystack
← Back to Jobs
Technology
FT

Observability Operations Engineer

Fixity TechnologiesAZ🇺🇸United StatesPosted 31 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
AZ, United States
Posted
Yesterday
ShellSplunkGrafanaKubernetesPrometheusPythonREST

Job Description

Observability Operations Engineer

Experience: 7+ Years
Employment Type: Full Time

Job Description

We are seeking an experienced Observability Operations Engineer to administer, optimize, and support enterprise-scale observability platforms. The ideal candidate will have strong hands-on experience with Dynatrace, Splunk, OpenSearch/Elasticsearch, along with expertise in monitoring, logging, tracing, alerting, dashboards, automation, and production troubleshooting.

Must Have Technical/Functional Skills
  • Strong Observability Administration experience with Dynatrace, Splunk, and OpenSearch/Elasticsearch.
  • Hands-on experience with monitoring, logging, tracing, alerting, dashboards, and platform performance tuning.
  • Strong knowledge of Linux, Kubernetes, containers, and cloud environments.
  • Experience with Grafana, Prometheus, OpenTelemetry, and related observability technologies.
  • Automation experience using Python, Shell scripting, and REST APIs.
  • Experience supporting enterprise-scale production environments, troubleshooting incidents, and performing Root Cause Analysis (RCA).
  • Knowledge of observability best practices, platform security, capacity planning, upgrades, patching, and operational governance.
Roles & Responsibilities
  • Administer, configure, and optimize enterprise Dynatrace, Splunk, and OpenSearch/Elasticsearch platforms.
  • Maintain platform availability, scalability, performance, security, and reliability.
  • Build and manage monitoring, logging, tracing, dashboards, alerts, and operational metrics.
  • Monitor platform health and proactively identify performance and capacity issues.
  • Troubleshoot complex production issues using observability tools and perform detailed Root Cause Analysis (RCA).
  • Support Linux, Kubernetes, containerized, and cloud-based environments.
  • Develop automation for repetitive operational activities using Python, Shell scripting, and REST APIs.
  • Drive self-healing, automation, and AI-assisted operations initiatives.
  • Manage platform upgrades, patching, capacity planning, backups, and operational governance.
  • Collaborate closely with SRE, DevOps, Platform, Infrastructure, and Application teams.
  • Establish and maintain operational standards, documentation, and best practices for observability platforms.
  • Participate in incident management, problem management, and continuous improvement initiatives.
Generic Managerial / Soft Skills
  • Strong verbal and written communication skills.
  • Ability to communicate clearly and assertively with technical and business stakeholders.
  • Strong team player with the ability to collaborate across multiple engineering teams.
  • Excellent analytical and problem-solving skills.
  • Ability to work effectively in a production support environment and handle high-priority incidents.
  • Strong ownership, accountability, and attention to detail.
Primary Skills

Dynatrace | Splunk | OpenSearch/Elasticsearch | Grafana | Prometheus | OpenTelemetry | Kubernetes | Linux | Python | Shell Scripting | REST APIs | Cloud | Observability | Production Support | RCA

Similar jobs