Haystack
← Back to Jobs
Remote
Technology

Lead Site Reliability Engineer

Goldenpick Technologies LLCUnited States🇺🇸United StatesPosted 31 Jul 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Role - Lead Site Reliability Engineer
Location - Remote 
 
Skills Must have
  • Strong experience in observability and monitoring, including hands-on expertise with:
  • Dynatrace
  • OpenTelemetry (OTel)
  • Distributed tracing
  • Metrics collection and analysis
  • Centralized logging and log aggregation
  • Alerting and dashboard development
  • Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform.
  • Experience with CI/CD pipelines, deployment automation, and operational tooling.
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
  • Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.
  • Cloud & Platform Expertise
  • Strong experience with Microsoft Azure, including:
  • Azure App Services
  • Resource Groups
  • Azure networking concepts
  • Scaling and performance optimization
  • Deployment and release management
  • Application lifecycle management
  • Experience leveraging Azure-native operational tooling such as:
  • Azure Monitor
  • Application Insights
  • Log Analytics
  • Azure dashboards and alerting
  • Experience supporting cloud-native and hybrid infrastructure environments.
  • Reliability & Engineering Practices
  • Demonstrated experience implementing and operating SRE practices, including:
  • Service Level Objectives (SLOs)
  • Service Level Indicators (SLIs)
  • Error budgets
  • Incident management
  • Problem management
  • Root Cause Analysis (RCA)
  • Reliability automation
  • Ability to improve system reliability through:
  • Performance tuning
  • Capacity planning
  • Observability-driven insights
  • Proactive issue detection
  • Reliability engineering initiatives
  • Experience developing automated recovery mechanisms and self-healing solutions.
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.

Skills

Azure
Terraform

Similar jobs