Haystack
← Back to Jobs
Technology

Senior Observability Operations Engineer

Dminds Solutions Inc.Phoenix, AZ🇺🇸United StatesPosted 4 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Job Title: Senior Observability Operations Engineer

Location: Phoenix, AZ (Onsite)

Duration: Long term contract

Job Description:

  • We are seeking a highly skilled Senior Observability Operations Engineer to manage and enhance our enterprise observability platform.
  • The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions.
  • Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is highly desirable.
  • The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications.

Key Responsibilities

  • Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
  • Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
  • Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
  • Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM).
  • Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
  • Develop dashboards, alerts, reports, and executive operational metrics.
  • Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
  • Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
  • Perform root cause analysis for production incidents using observability platforms.
  • Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
  • Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible.
  • Participate in incident, problem, change, and release management processes.
  • Drive platform upgrades, patching, security compliance, and operational governance.
  • Improve platform reliability through automation, self-healing, and AI-assisted operations.

Required Technical Skills

Observability Platforms

  • Dynatrace Administration
  • Splunk Enterprise Administration
  • OpenSearch Administration
  • Elasticsearch Administration
  • Grafana
  • Prometheus
  • Kibana
  • Jaeger
  • OpenTelemetry
  • Kafka (preferred)

Infrastructure

  • Linux Administration
  • Kubernetes
  • Docker
  • OpenShift or Rancher
  • Networking (TCP/IP, DNS, Load Balancers, Firewalls)
  • System Administration

Cloud & DevOps

  • AWS, Azure, or Google Cloud Platform
  • CI/CD pipelines
  • Git
  • Terraform
  • Ansible
  • REST APIs

Scripting

  • Python
  • Bash/Shell
  • PowerShell (preferred)

AI & Automation Skills (Preferred)

  • Experience using Generative AI (ChatGPT, GitHub Copilot, Amazon Q, Microsoft Copilot, or similar) to improve operational efficiency.
  • Knowledge of AIOps platforms and AI-driven observability.
  • Experience with Dynatrace Davis AI for anomaly detection and root cause analysis.
  • Understanding of machine learning concepts for predictive monitoring and intelligent alerting.
  • Experience building AI-assisted operational runbooks and troubleshooting workflows.
  • Knowledge of Retrieval-Augmented Generation (RAG), vector databases, embeddings, and AI-powered knowledge search is a plus.
  • Experience integrating AI with observability platforms using APIs.
  • Familiarity with LLMs, prompt engineering, and AI-assisted automation.
  • Experience using Python with AI frameworks (LangChain, LangGraph, OpenAI APIs, or similar) is desirable.
  • Exposure to AI-driven incident summarization, log analysis, and automated ticket enrichment.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • 6 10+ years of IT infrastructure or observability operations experience.
  • 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch.
  • Strong Linux system administration experience.
  • Experience supporting enterprise-scale production environments.
  • Strong troubleshooting and analytical skills.
  • Excellent communication and stakeholder management skills.

Preferred Certifications

  • Dynatrace Associate or Professional Certification
  • Splunk Enterprise Certified Administrator
  • Elastic Certified Engineer
  • Kubernetes (CKA/CKAD)
  • AWS/Azure/Google Cloud Platform Certification
  • ITIL Foundation
  • AI/ML or Generative AI certification (preferred)

Soft Skills

  • Strong ownership and accountability
  • Excellent problem-solving and analytical thinking
  • Ability to work independently with minimal supervision
  • Strong collaboration across cross-functional teams
  • Continuous learning mindset
  • Ability to thrive in fast-paced production environments

Thanks & Regards

Saravanan

DMinds Solutions Inc.

Skills

Docker
Shell
AWS
Machine Learning
Splunk
TCP/IP
Ansible
Azure
Bash
DNS
Generative AI
Git
Google Cloud
Grafana
Kafka
Kibana
Kubernetes
Phoenix
PowerShell
Prometheus
Python
REST
Stakeholder Management
Terraform

Similar jobs