Haystack
← Back to Jobs
Technology
RY

SRE Observability (Staff or Senior)

RyanBPMUnited States🇺🇸United StatesPosted Oct 5, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
18 hours ago
DynamoDBNode.jsAPI GatewayAWSSplunkAgileCDKDatadogJavaJiraPostgreSQLPythonTerraformTypeScript

Job Description

SRE Observability (Staff or Senior)

Location: US (Remote)

Full time: Yes 

Contract on W2 : 12 Months +

Target Skills:

  • Resiliency Testing (Level 300) - Experienced in leading Failure Mode Effects Analysis and resiliency testing game days using AWS Fault Injection Service (FIS). Experience with multi-AZ failure, region failure, and network partition simulations. Experience with AWS Resilience Hub for recovery readiness assessment.
  • Observability — CloudWatch & Open Telemetry (Level 400) - Deep experience with CloudWatch at enterprise scale (Metrics V2, Log Insights, Contributor Insights, Composite Alarms), AppSignals for SLO monitoring, and ADOT/X-Ray for distributed tracing. Proficiency in Open Telemetry SDK instrumentation (Java, Python, Node.js). Experience in third-party tools (Datadog, Splunk, Dynatrace) and CloudWatch-native observability.
  • Multi-Region DR Patterns (Level 300) - Understanding of active-active, active-passive, warm standby, and pilot light architectures. Experience with Aurora Global Database failover, Route 53 health checks, and cross-region replication patterns.
  • AWS CDK TypeScript (Level 200) - Ability to read, modify, and contribute to CDK TypeScript constructs for observability and resilience resources (alarms, dashboards, FIS experiment templates). All infrastructure at this customer is CDK TypeScript — no Terraform.
  • AWS Services (Level 300) - Expertise in ECS, Fargate, ELB, EC2, S3, IAM, VPC, SQS, RDS, DynamoDB, Aurora PostgreSQL, Lambda, API Gateway, MSK, EFS, CloudWatch, X-Ray, FIS, Resilience Hub
  • Programming Languages (Level 300) - Python (Boto3), TypeScript (for CDK);
  • Documentation (Level 300) - Must be able to clearly write, draw, and organize various types (runbooks, incident playbooks, DR plans, READMEs, technical diagrams) of documentation for specific problems and solutions on a weekly basis.
  • Agile/Jira (Level 300) - On-daily basis you must be able to manage and report own Jira User Stories and help desk tickets during the current sprint. Must be able to articulate your user stories during sprint planning.

Similar jobs