Why This Role Stands Out
This Lead Site Reliability Engineer role at EPAM Systems offers an exciting opportunity to drive innovation by integrating AI capabilities and ensuring system excellence, perfect for a collaborative engineer eager to shape scalable cloud solutions. You'll thrive in this hybrid environment, contributing to a reputable company with significant impact and continuous learning. Apply now to elevate your career and enhance seamless user experiences!
Quick Overview
Job Description
We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.
Req.# Responsibilities Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, Google Cloud Platform Pub/Sub sinks, and OpenTelemetry (OTel) AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing Requirements AWS CDK & TypeScript for infrastructure provisioning and automation Agent & Cluster Agent deployment/management Monitors, Dashboards, Synthetics, and SLOs managed via Terraform Advanced constructs: Composite monitors, cardinality management, and retention controls OP2 / Vector (Observability Pipelines Worker) Log routing, enrichment, and redaction Splunk & Google Cloud Platform Pub/Sub sinks; OpenTelemetry standards CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice) AWS Lambda runtime monitoring & performance tuning PagerDuty integration and alert-to-page wiring from Datadog
Similar jobs
- JO
Senior Site Reliability Engineer - Banking & Finance
NewJooble
United States🇺🇸Hybrid2 days agoPythonTechnology - TC
DevSecOps Engineer
NewTECHNEPTUNE CONSULTING INC
United States🇺🇸RemoteYesterdaySAFeAgileEngineering - TB
Site Reliability Engineer (SRE)
NewThe Brixton Group
Benton Harbor, MI🇺🇸RemoteYesterdayDockerDynamoDBNode.js+15Technology - MR
Staff Site Reliability Engineer/ Azure
NewMotion Recruitment Partners, LLC
Toronto, ON🇺🇸HybridYesterdayMachine LearningAzure.NETTechnology - ED
DevOps Engineer III
NewAuto ApplyEnable Dental
Austin, Texas🇺🇸Remote11 hours agoDockerAWSEncryption+17Technology - TH
DevOps Engineer
NewAuto ApplyTheIncLab
Colorado Springs, Colorado🇺🇸Hybrid5 hours agoDockerShellAWS+22Technology - SG
DevOps Engineer
NewAuto ApplyStafford Gray
Lansing, Michigan🇺🇸Hybrid10 hours agoShellAuditingAzure+7Technology - JG
Azure DevOps Engineer
NewJudge Group, Inc.
Berkeley Heights, NJ🇺🇸On-siteYesterdayEncryptionAgileAnsible+4Technology - JG
Sr DevOps Engineer
NewJudge Group, Inc.
Lone Tree, CO🇺🇸HybridYesterdayDockerMicroservicesAnsible+2Technology - ST
Senior Azure DevOps Engineer
NewStefanini
Fort Myers, FL🇺🇸RemoteYesterdayAzureGitHub ActionsTerraformTechnology - MR
Azure DevOps Engineer
NewMotion Recruitment Partners, LLC
Toronto, ON🇺🇸HybridYesterdayService MeshArgoCDAzure+6Technology - HP
MLOps certified - Cloud ML Ops/SRE
NewHPTech Inc.
Sunnyvale, CA🇺🇸On-siteYesterdayMLOpsKubernetesTechnology