Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
18 hours ago
DynamoDBNode.jsAPI GatewayAWSSplunkAgileCDKDatadogJavaJiraPostgreSQLPythonTerraformTypeScript
Job Description
SRE Observability (Staff or Senior)
Location: US (Remote)
Full time: Yes
Contract on W2 : 12 Months +
Target Skills:
- Resiliency Testing (Level 300) - Experienced in leading Failure Mode Effects Analysis and resiliency testing game days using AWS Fault Injection Service (FIS). Experience with multi-AZ failure, region failure, and network partition simulations. Experience with AWS Resilience Hub for recovery readiness assessment.
- Observability — CloudWatch & Open Telemetry (Level 400) - Deep experience with CloudWatch at enterprise scale (Metrics V2, Log Insights, Contributor Insights, Composite Alarms), AppSignals for SLO monitoring, and ADOT/X-Ray for distributed tracing. Proficiency in Open Telemetry SDK instrumentation (Java, Python, Node.js). Experience in third-party tools (Datadog, Splunk, Dynatrace) and CloudWatch-native observability.
- Multi-Region DR Patterns (Level 300) - Understanding of active-active, active-passive, warm standby, and pilot light architectures. Experience with Aurora Global Database failover, Route 53 health checks, and cross-region replication patterns.
- AWS CDK TypeScript (Level 200) - Ability to read, modify, and contribute to CDK TypeScript constructs for observability and resilience resources (alarms, dashboards, FIS experiment templates). All infrastructure at this customer is CDK TypeScript — no Terraform.
- AWS Services (Level 300) - Expertise in ECS, Fargate, ELB, EC2, S3, IAM, VPC, SQS, RDS, DynamoDB, Aurora PostgreSQL, Lambda, API Gateway, MSK, EFS, CloudWatch, X-Ray, FIS, Resilience Hub
- Programming Languages (Level 300) - Python (Boto3), TypeScript (for CDK);
- Documentation (Level 300) - Must be able to clearly write, draw, and organize various types (runbooks, incident playbooks, DR plans, READMEs, technical diagrams) of documentation for specific problems and solutions on a weekly basis.
- Agile/Jira (Level 300) - On-daily basis you must be able to manage and report own Jira User Stories and help desk tickets during the current sprint. Must be able to articulate your user stories during sprint planning.
Similar jobs
- II
Devops Engineer
NewIntake IT Solutions
Irvine, CA🇺🇸Hybrid18 hours agoDockerAWSAnsible+9Technology - RI
DevOps Platform Engineer
NewRavh IT Solutions
Irvine, CA🇺🇸On-site18 hours agoDockerShellAWS+16Technology - IC
Cloud DevOps Engineer – AI/ML- 7+ yrs- New York, United States- Onsite
NewiMedhas Consulting Services
New York, NY🇺🇸Hybrid18 hours agoAWSAzureGoogle CloudTechnology - KA
Resiliency AI Architect with .NET (AI/Agentic AI, Reliability Engineering, SRE, Observability, Production Engineering.NET Development Background) : Local to Texas
NewK Anand Corporation
Dallas, TX🇺🇸On-site18 hours agoJiraRoot Cause AnalysisManufacturing - ST
REMOTE JOB: Sr. Azure DevOps Engineer
NewSAR TECH LLC
United States🇺🇸Remote18 hours agoSAFeNode.jsScrum+17Technology - IS
Sr. DevOps/Enablement Engineer (STRICTLY OUR W2 -- Do NOT respond if you're looking for C2C or Remote)
NewInfoweb Systems, Inc.
Urbandale, IA🇺🇸$72/hrRemote18 hours agoAWSAgileGitHub Actions+4Technology