Haystack
← Back to Jobs
Remote
Other
HT

Hiring | Dynatrace Observability Lead | Remote | Contract

Healthcare Triangle IncUnited States🇺🇸United StatesPosted Oct 1, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
19 hours ago
AWSSplunkGoogle CloudGrafanaPythonServiceNowTerraformTriage

Job Description

Role : Dynatrace Observability Lead
Location : Remote
Duration : Long-term Contract
Job Description :
  • Primary/Must-have: Dynatrace, with strong hands-on monitoring experience
  • Good to have: Splunk
  • Soft skills: Leadership capabilities and effective client communication

Please find the detailed JD below for your reference.

Job Description: Observability & Automation Engineer

Required Experience

  • 7+ years of experience in Observability Engineering or related domains.
  • Strong experience in Observability and Automation Engineering, including AIOps.

Key Responsibilities

Observability Engineering

  • Design and implement end-to-end observability solutions across applications, infrastructure, and cloud environments, including:
    • Metrics
    • Logs
    • Distributed traces
    • Synthetic monitoring
    • Real User Monitoring (RUM)
  • Develop standardized dashboards, alerts, and telemetry frameworks to provide real-time visibility into system health.

Automation & Toil Reduction

  • Build and deploy automation solutions to eliminate repetitive operational tasks and improve efficiency.
  • Enable runbook automation, self-healing capabilities, and automated incident triage.

Reliability & Incident Optimization

  • Define and implement SLIs, SLOs, and alerting strategies to improve service reliability.
  • Drive improvements in MTTD and MTTR through actionable alerts and telemetry-driven insights.

Proactive Monitoring & AIOps

  • Implement proactive monitoring, anomaly detection, and predictive alerting to identify issues before customer impact.
  • Leverage AIOps capabilities for alert correlation and intelligent incident response.

Platform Integration & Enablement

  • Integrate observability platforms with CI/CD pipelines, cloud services, and ITSM tools such as ServiceNow.
  • Enable seamless workflows for alerting, escalation, and operational response.

Collaboration & Operational Readiness

  • Partner with engineering, product, and operations teams to standardize observability practices and establish standards and policies related to operational readiness.
  • Mentor teams and drive adoption of observability and operational best practices across the organization.

Technical Skills

Observability Tools

Hands-on experience with:

  • Dynatrace
  • Splunk
  • Grafana
  • OpenTelemetry preferred

Cloud & Infrastructure

  • Strong expertise in AWS and Google Cloud Platform.
  • Familiarity with cloud-native architectures.

Programming & Automation

  • Proficiency in Python.

Telemetry & Instrumentation

  • Experience implementing metrics, logs, and distributed tracing (MELT) across distributed systems.

Infrastructure as Code

  • Experience with Terraform.

SRE & Incident Management

Strong understanding of:

  • SLOs and SLIs
  • Alerting strategies
  • Incident response frameworks
  • Service reliability practices
  • MTTD and MTTR optimization

Preferred Background

  • Strong experience in Observability and Automation Engineering.
  • Experience working with AIOps technologies and practices.

AI Skills & Expectations

All contractor resources are expected to demonstrate baseline proficiency in enterprise-approved AI tools as part of their day-to-day responsibilities.

Consistent Use

  • Maintain a minimum of 90% weekly usage of AI tools such as GitHub Copilot, Microsoft 365 Copilot, and other enterprise-approved GenAI platforms.

Applied Productivity

  • Leverage AI tools to enhance:
    • Coding
    • Documentation
    • Data analysis
    • Decision-making workflows

Continuous Learning

  • Stay current with evolving AI capabilities and features.
  • Apply AI capabilities to improve delivery quality and velocity.

Similar jobs