Haystack
← Back to Jobs
Technology
ES

Lead Site Reliability Engineer

EPAM SystemsUnited States🇺🇸United StatesPosted Oct 4, 2026

Why This Role Stands Out

This Lead Site Reliability Engineer role at EPAM Systems offers an exciting opportunity to drive innovation by integrating AI capabilities and ensuring system excellence, perfect for a collaborative engineer eager to shape scalable cloud solutions. You'll thrive in this hybrid environment, contributing to a reputable company with significant impact and continuous learning. Apply now to elevate your career and enhance seamless user experiences!

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
5 days ago
AWSSplunkCDKDatadogGoogle CloudPagerDutyTerraformTypeScript

Job Description

We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.

Req.# Responsibilities Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, Google Cloud Platform Pub/Sub sinks, and OpenTelemetry (OTel) AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing Requirements AWS CDK & TypeScript for infrastructure provisioning and automation Agent & Cluster Agent deployment/management Monitors, Dashboards, Synthetics, and SLOs managed via Terraform Advanced constructs: Composite monitors, cardinality management, and retention controls OP2 / Vector (Observability Pipelines Worker) Log routing, enrichment, and redaction Splunk & Google Cloud Platform Pub/Sub sinks; OpenTelemetry standards CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice) AWS Lambda runtime monitoring & performance tuning PagerDuty integration and alert-to-page wiring from Datadog

Similar jobs