Haystack
← Back to Jobs
Remote
Technology
DS

Sr. SRE Engineer (Azure Expert)

Dminds Solutions Inc.United States🇺🇸United StatesPosted 3 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
20 hours ago
Load BalancingNew RelicSplunkAzureBashDatadogGitLLMPowerShellPythonReactTerraform

Job Description

Job Title: Sr. SRE Engineer

Location: Remote (need to travel based on request)

Duration: 12+Months Contract

Job Description:

Core Skills & Experience

  • Microsoft Azure Expert
    • Azure Administration and Architecture
    • Networking and Connectivity
    • Virtual Machines and Platform Services
    • Azure Monitor and Log Analytics
    • Azure Identity and Access Management
    • Azure Networking, NSGs, Firewalls, Load Balancing
    • Azure Backup and Disaster Recovery
  • Strong expertise in Azure networking (VNets, routing, firewalls, private links, load balancing).
  • Hands-on proficiency with infrastructure-as-code and automated deployments. (Must have Terraform and Git Enterprise, orchestration engines)
  • Exposure and understanding of building, deploying and managing API Gateways
  • Strong understanding of Azure security controls, governance, and compliance frameworks.
  • Full stack observability e.g. MELTS principles golden signals, and automation response using DataDog, New Relic, Splunk or other leading tools.
  • Strong FinOps expertise
  • Scripting skills (PowerShell, Bash, Python, React).
  • Strong understanding of Devops practices, tooling, and SDLC methods
  • Strong exposure to Anthropic, Open AI, platforms and associated tools & practices e.g Harness, Token usage, Skills, LLM and SLM concepts, Orchestration engines, and agent cost management.
  • Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement.
  • Experience designing observability strategies across metrics, logs, traces, synthetic monitoring, alerting, dashboards, and operational telemetry.
  • Ability to define actionable alerts that identify customer-impacting symptoms, reduce noise, and support rapid incident triage.
  • Proven capability in incident response, root cause analysis, blameless post-incident reviews, corrective action tracking, and operational learning.
  • Experience reducing toil through automation, self-service tooling, runbook automation, self-healing patterns, and repeatable engineering solutions.
  • Strong knowledge of capacity planning, performance engineering, load testing, scalability modelling, saturation analysis, and demand forecasting.
  • Experience with resilience validation techniques including chaos engineering, game days, failover testing, disaster recovery exercises, and operational readiness testing.
  • Ability to establish production readiness standards, reliability acceptance criteria, operational runbooks, service ownership models, and support handover practices.
  • Working knowledge of deployment reliability practices such as canary releases, blue-green deployments, rollback strategies, feature flags, and release health monitoring.

Thanks & Regards

Saravanan

DMinds Solutions Inc.

Similar jobs