← Back to Jobs
Technology
Senior Observability Operations Engineer
Dminds Solutions Inc.Phoenix, AZ🇺🇸United StatesPosted 4 Aug 2026
Quick Overview
Work Type
On Site
Level
Mid Senior
Job Description
Job Title: Senior Observability Operations Engineer
Location: Phoenix, AZ (Onsite)
Duration: Long term contract
Job Description:
- We are seeking a highly skilled Senior Observability Operations Engineer to manage and enhance our enterprise observability platform.
- The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions.
- Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is highly desirable.
- The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications.
Key Responsibilities
- Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
- Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
- Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
- Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM).
- Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
- Develop dashboards, alerts, reports, and executive operational metrics.
- Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
- Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
- Perform root cause analysis for production incidents using observability platforms.
- Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
- Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible.
- Participate in incident, problem, change, and release management processes.
- Drive platform upgrades, patching, security compliance, and operational governance.
- Improve platform reliability through automation, self-healing, and AI-assisted operations.
Required Technical Skills
Observability Platforms
- Dynatrace Administration
- Splunk Enterprise Administration
- OpenSearch Administration
- Elasticsearch Administration
- Grafana
- Prometheus
- Kibana
- Jaeger
- OpenTelemetry
- Kafka (preferred)
Infrastructure
- Linux Administration
- Kubernetes
- Docker
- OpenShift or Rancher
- Networking (TCP/IP, DNS, Load Balancers, Firewalls)
- System Administration
Cloud & DevOps
- AWS, Azure, or Google Cloud Platform
- CI/CD pipelines
- Git
- Terraform
- Ansible
- REST APIs
Scripting
- Python
- Bash/Shell
- PowerShell (preferred)
AI & Automation Skills (Preferred)
- Experience using Generative AI (ChatGPT, GitHub Copilot, Amazon Q, Microsoft Copilot, or similar) to improve operational efficiency.
- Knowledge of AIOps platforms and AI-driven observability.
- Experience with Dynatrace Davis AI for anomaly detection and root cause analysis.
- Understanding of machine learning concepts for predictive monitoring and intelligent alerting.
- Experience building AI-assisted operational runbooks and troubleshooting workflows.
- Knowledge of Retrieval-Augmented Generation (RAG), vector databases, embeddings, and AI-powered knowledge search is a plus.
- Experience integrating AI with observability platforms using APIs.
- Familiarity with LLMs, prompt engineering, and AI-assisted automation.
- Experience using Python with AI frameworks (LangChain, LangGraph, OpenAI APIs, or similar) is desirable.
- Exposure to AI-driven incident summarization, log analysis, and automated ticket enrichment.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- 6 10+ years of IT infrastructure or observability operations experience.
- 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch.
- Strong Linux system administration experience.
- Experience supporting enterprise-scale production environments.
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder management skills.
Preferred Certifications
- Dynatrace Associate or Professional Certification
- Splunk Enterprise Certified Administrator
- Elastic Certified Engineer
- Kubernetes (CKA/CKAD)
- AWS/Azure/Google Cloud Platform Certification
- ITIL Foundation
- AI/ML or Generative AI certification (preferred)
Soft Skills
- Strong ownership and accountability
- Excellent problem-solving and analytical thinking
- Ability to work independently with minimal supervision
- Strong collaboration across cross-functional teams
- Continuous learning mindset
- Ability to thrive in fast-paced production environments
Thanks & Regards
Saravanan
DMinds Solutions Inc.
Skills
Docker
Shell
AWS
Machine Learning
Splunk
TCP/IP
Ansible
Azure
Bash
DNS
Generative AI
Git
Google Cloud
Grafana
Kafka
Kibana
Kubernetes
Phoenix
PowerShell
Prometheus
Python
REST
Stakeholder Management
Terraform
Similar jobs
Senior Microsoft Security Engineer
SDH Systems · Houston, United States
16 minutes agoOperating System Engineer
LTIMindtree · Irving, United States
16 minutes agoDesign Systems Engineer
Kforce Technology Staffing · Greenwood Village, United States
16 minutes agoIntegration Solution Architect
ISite Technologies Inc · United States
22 minutes agoNetwork Engineer Associate - DHA NE&S with Security Clearance
ASRC Federal · Albertville, United States
22 minutes agoSystems Engineer with Security Clearance
ASRC Federal · Brenham, United States
22 minutes ago