Haystack
← Back to Jobs
Technology
DE

Senior Observability Platform Engineer

DTEL Engineering & Consultants IncUnited States🇺🇸United StatesPosted 11 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
3 weeks ago
Datadog

Job Description

  • Lead hands-on technical leadership for observability platform reliability and scalability across enterprise monitoring systems, including Dynatrace and IBM SevOne
  • Design and enforce observability patterns, standards, and data models to ensure alignment across multiple observability tools (Dynatrace, IBM SevOne, ServiceNow ITOM)
  • Drive AIOps enablement initiatives, including Davis AI implementation and causal analysis capabilities to improve operational decision-making
  • Architect and scale observability platforms across Application Performance Monitoring, Network Performance Monitoring, and Ingest (Dashboard & Visibility) tiers
  • Prevent uncontrolled log growth, reduce alert noise, and implement cost optimization strategies across the observability ecosystem
  • Lead root cause analysis initiatives and provide critical support during incident response "war room" sessions
  • Manage system health monitoring for servers, infrastructure, and applications across the enterprise
  • Collaborate with cross-functional teams to implement intelligent automation and advance observability maturity across the organization

What You'll Need to Have:

  • 8+ years of engineering experience with demonstrated expertise in both engineering and architecture roles
  • Deep expertise in designing and scaling enterprise observability platforms such as Dynatrace, DataDog, IBM SevOne, or similar tools
  • Proven ability to define and enforce observability patterns, standards, and data models at scale
  • Strong experience leading intelligent automation and root cause analysis initiatives within observability environments
  • Hands-on experience with AIOps platforms and AI-driven analysis tools (e.g., Davis AI, causal analysis engines)
  • Demonstrated expertise in managing observability data (logs, metrics, alerts) at scale with a focus on cost optimization and governance
  • Solid understanding of platform reliability, scalability, and cross-tool integration in complex enterprise environments

Strong analytical and problem-solving skills with the ability to operate effectively during high-pressure incident response scenarios

Similar jobs