Why This Role Stands Out
This hybrid Lead Site Reliability Engineer role offers an exciting opportunity to architect and implement critical observability and disaster recovery solutions for a global leader in digital transformation. You'll develop cutting-edge platform-as-code strategies using AWS CDK and TypeScript, making this an ideal position for experienced SREs passionate about scalable infrastructure and deep system visibility. Apply now to shape the future of EPAM's resilient cloud footprint and drive significant impact.
Quick Overview
Job Description
We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.
Req.# Responsibilities Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, Google Cloud Platform Pub/Sub sinks, and OpenTelemetry (OTel) AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing Requirements AWS CDK & TypeScript for infrastructure provisioning and automation Agent & Cluster Agent deployment/management Monitors, Dashboards, Synthetics, and SLOs managed via Terraform Advanced constructs: Composite monitors, cardinality management, and retention controls OP2 / Vector (Observability Pipelines Worker) Log routing, enrichment, and redaction Splunk & Google Cloud Platform Pub/Sub sinks; OpenTelemetry standards CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice) AWS Lambda runtime monitoring & performance tuning PagerDuty integration and alert-to-page wiring from Datadog
Similar jobs
- DC
Databricks Platform Engineer
NewDaman Consulting
United States🇺🇸Hybrid15 hours agoSQLAzureDatabricks+4Technology - PL
Senior Virtualization Platform Engineer (W2 only)
NewPatton Labs Inc.
Dearborn, MI🇺🇸Hybrid15 hours agoAnsibleGoogle CloudKubernetes+3Technology - JM
Senior Lead Software Engineer- SRE
NewJ.P. Morgan
Dallas, Texas🇺🇸On-site2 hours agoAWSLoad BalancingSnowflake+7Technology - JM
DevOps - Software Engineer III
NewJ.P. Morgan
Tampa, Florida🇺🇸On-site2 hours agoSplunkAgileApache+4Technology - OR
Manager, Site Reliability Engineering
NewOracle
Reston, Virginia🇺🇸Hybrid2 hours agoOracleTechnology - MT
Data Platform Engineer
NewMilestone Technologies, Inc.
New York, NY🇺🇸$55 - $65/hrOn-site15 hours agoSQLAWSETL+11Technology