Haystack
← Back to Jobs
Technology
ES

Lead Site Reliability Engineer

EPAM SystemsNew York, NY🇺🇸United StatesPosted 14 Sept 2026

Why This Role Stands Out

This hybrid Lead Site Reliability Engineer role offers an exciting opportunity to architect and implement critical observability and disaster recovery solutions for a global leader in digital transformation. You'll develop cutting-edge platform-as-code strategies using AWS CDK and TypeScript, making this an ideal position for experienced SREs passionate about scalable infrastructure and deep system visibility. Apply now to shape the future of EPAM's resilient cloud footprint and drive significant impact.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
New York, NY, United States
Posted
15 hours ago
AWSSplunkCDKDatadogGoogle CloudPagerDutyTerraformTypeScript

Job Description

We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.

Req.# Responsibilities Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, Google Cloud Platform Pub/Sub sinks, and OpenTelemetry (OTel) AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing Requirements AWS CDK & TypeScript for infrastructure provisioning and automation Agent & Cluster Agent deployment/management Monitors, Dashboards, Synthetics, and SLOs managed via Terraform Advanced constructs: Composite monitors, cardinality management, and retention controls OP2 / Vector (Observability Pipelines Worker) Log routing, enrichment, and redaction Splunk & Google Cloud Platform Pub/Sub sinks; OpenTelemetry standards CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice) AWS Lambda runtime monitoring & performance tuning PagerDuty integration and alert-to-page wiring from Datadog

Similar jobs