Haystack
← Back to Jobs
Remote
Technology
DS

SRE Engineer

Dminds Solutions Inc.United States🇺🇸United StatesPosted 2 Sept 2026

Why This Role Stands Out

You'll have the opportunity to shape and manage critical AI platforms in a growing global business, leveraging your full-stack SRE expertise to drive innovation and reliability. This remote role is perfect for a high-caliber engineer passionate about infrastructure, coding, and AI, offering significant career growth and the chance to make a substantial impact. Apply now to join a forward-thinking team and contribute to cutting-edge technology.

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
22 hours ago
Load BalancingNew RelicSplunkAzureBashDatadogGitKubernetesLLMPowerShellPythonReactTerraformVaultZero Trust

Job Description

Hi,

We are looking for a SRE Engineer @ Remote. Kindly send me your updated profile and I am looking forward to work with you in this position.

Job Title: SRE Engineer

Location: Remote (but need to travel based on request)

Duration: Long term contract

Reporting Line: Cloud and Infrastructure Lead

Key Skills:

  • Full stack SRE
  • Coding
  • AI platform built
  • AI partners
  • API gateways
  • Infrastructure
  • Code base-quality
  • High caliber
  • Manage of core ai platform
  • Assets built on it

Job Description:

  • Client is expanding it s Global business in 2026 and the Cloud and Infrastructure team will need to support this growth by designing, provisioning then supporting the platforms to enable this.
  • We are seeking an experienced Site Reliability Engineer (SRE) to join our Cloud & Infrastructure team. The successful candidate will be responsible for designing, operating, automating, and continuously improving enterprise-scale Azure platforms, ensuring high availability, resiliency, security, and performance.
  • The role combines software engineering, cloud architecture, infrastructure automation, AI, and operational excellence to improve service reliability and reduce operational overhead through automation and engineering best practices.
  • The ideal candidate will have strong experience with Microsoft Azure, DevOps tools & practices, Terraform, API, AI platforms & tools, disaster recovery planning, and enterprise-scale resilience engineering.

Key Responsibilities

Platform Reliability & Operations

  • Ensure the availability, performance, scalability, and reliability of Azure-hosted services.
  • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Proactively monitor platform health and performance using observability tooling.
  • Perform root cause analysis and implement permanent fixes for recurring incidents.
  • Participate in incident management and on-call support rotations where required.
  • Lead blameless post-incident reviews, capture lessons learned, and drive corrective actions through to completion.
  • Reduce operational toil by identifying repetitive manual tasks and replacing them with automated, reusable engineering solutions.
  • Develop reliability dashboards and actionable alerts that focus on customer-impacting symptoms rather than infrastructure noise.

Azure Cloud Engineering

  • Design, deploy, and manage Azure infrastructure services including:
    • Virtual Networks
    • Application Gateways
    • API Management
    • Azure Kubernetes Service (AKS)
    • Azure Firewall
    • Azure Storage
    • Key Vault
    • Azure Monitor
    • Azure AI Services
  • Implement cloud platform standards and best practices.
  • Support multi-region Azure deployments and platform modernisation initiatives.
  • Undertake capacity planning and performance engineering to ensure platforms can scale reliably in line with business growth and peak demand.

Infrastructure as Code (Terraform)

  • Develop and maintain Terraform modules and reusable infrastructure patterns.
  • Implement Infrastructure as Code (IaC) standards and governance controls.
  • Ensure infrastructure is version controlled, peer-reviewed, and fully automated.
  • Manage Terraform state securely and consistently across environments.

DevOps, Automation & AI

  • Build and maintain Azure DevOps CI/CD pipelines.
  • Automate infrastructure provisioning and application deployments using pipelines with automated delivery and testing routines.
  • Implement testing, security scanning, policy compliance, and release gates.
  • Support DevOps and platform engineering practices.
  • Create automation for operational runbooks, self-healing processes, deployment validation, and environment consistency checks.
  • Manage and Implement AI platforms and tools such as Claude & Open AI, to develop skills and support business adoption of agentic AI capabilities.
  • Implement and enable self-service approach to technology services.

Resilience, Disaster Recovery & Failover

  • Design and implement highly available Azure architectures.
  • Develop and maintain disaster recovery and business continuity capabilities.
  • Implement and test:
    • Regional failover strategies
    • Active/Passive architectures
    • Active/Active deployments
    • Traffic Manager and Front Door failover patterns
    • Database resiliency and replication
    • Backup and recovery solutions
  • Conduct regular resilience and recovery testing exercises.
  • Identify and reduce single points of failure across platforms.
  • Define and execute game days, chaos testing, and controlled failure scenarios to validate operational resilience.

Security & Governance

  • Ensure platforms are secure-by-design.
  • Work closely with Security and Architecture teams to implement:
    • Zero Trust principles
    • RBAC controls / Managed Identities
    • Network segmentation and Zone based architecture
    • Secrets management
  • Support compliance requirements and operational audits.
  • Help coordinate security updates, patches, maintenance routines, and upgrades of the underlying system across partners and vendors
  • Embed reliability, security, and compliance controls into build and release pipelines to support production readiness.

Continuous Improvement

  • Drive automation and reduction of manual operational tasks.
  • Improve deployment reliability and platform observability.
  • Contribute to architecture standards, runbooks, and operational documentation.
  • Partner with engineering, architecture, security, and service teams to define production readiness standards and reliability acceptance criteria.

About You

At Client we work in a fast paces evolving environment with a growth mindset and outcome focus.

Core Skills & Experience

  • Microsoft Azure Expert
    • Azure Administration and Architecture
    • Networking and Connectivity
    • Virtual Machines and Platform Services
    • Azure Monitor and Log Analytics
    • Azure Identity and Access Management
    • Azure Networking, NSGs, Firewalls, Load Balancing
    • Azure Backup and Disaster Recovery
  • Strong expertise in Azure networking (VNets, routing, firewalls, private links, load balancing).
  • Hands-on proficiency with infrastructure-as-code and automated deployments. (Must have Terraform and Git Enterprise, orchestration engines)
  • Exposure and understanding of building, deploying and managing API Gateways
  • Strong understanding of Azure security controls, governance, and compliance frameworks.
  • Full stack observability e.g. MELTS principles golden signals, and automation response using DataDog, New Relic, Splunk or other leading tools.
  • Strong FinOps expertise
  • Scripting skills (PowerShell, Bash, Python, React).
  • Strong understanding of Devops practices, tooling, and SDLC methods
  • Strong exposure to Anthropic, Open AI, platforms and associated tools & practices e.g Harness, Token usage, Skills, LLM and SLM concepts, Orchestration engines, and agent cost management.
  • Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement.
  • Experience designing observability strategies across metrics, logs, traces, synthetic monitoring, alerting, dashboards, and operational telemetry.
  • Ability to define actionable alerts that identify customer-impacting symptoms, reduce noise, and support rapid incident triage.
  • Proven capability in incident response, root cause analysis, blameless post-incident reviews, corrective action tracking, and operational learning.
  • Experience reducing toil through automation, self-service tooling, runbook automation, self-healing patterns, and repeatable engineering solutions.
  • Strong knowledge of capacity planning, performance engineering, load testing, scalability modelling, saturation analysis, and demand forecasting.
  • Experience with resilience validation techniques including chaos engineering, game days, failover testing, disaster recovery exercises, and operational readiness testing.
  • Ability to establish production readiness standards, reliability acceptance criteria, operational runbooks, service ownership models, and support handover practices.
  • Working knowledge of deployment reliability practices such as canary releases, blue-green deployments, rollback strategies, feature flags, and release health monitoring.

Thanks & Regards

Saravanan

DMinds Solutions Inc.

Similar jobs