Haystack
← Back to Jobs
Technology
IT

Senior SRE / AIOps Engineer – Observability

ITBMS Inc.Minneapolis, MN🇺🇸United StatesPosted Sep 29, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Minneapolis, MN, United States
Posted
17 hours ago
AWSSplunkAzureBashCloudFormationGoogle CloudKubernetesPagerDutyPythonTerraform

Job Description

Position Details

Job Title: Senior SRE / AIOps Engineer – Observability & Intelligent Operations

Location: Minneapolis, MN(Onsite)
Experience: 8+ Years
Focus Areas: SRE | AIOps | Observability | Cloud | Automation | AI-Enabled Operations

Key Responsibilities

  • Define and track SLIs/SLOs, manage error budgets, and drive continuous improvements in availability, latency, and resiliency.
  • Build and optimize observability and AIOps platforms, including monitoring, dashboards, alerting, and log/metric/trace correlation.
  • Work with tools such as Dynatrace and Splunk to reduce alert noise and accelerate incident detection and troubleshooting.
  • Lead incident response, on-call activities, war rooms, and root-cause analysis.
  • Conduct post-incident reviews and ensure corrective and preventive actions are completed.
  • Develop runbooks, scripts, and automated remediation workflows to reduce operational toil and improve MTTR.
  • Partner with Engineering and business stakeholders on architecture, release readiness, capacity planning, and operational standards.
  • Translate reliability metrics and risks into executive-level reporting, including MTTD, MTTR, error-budget burn, and recurring toil.

Required Skills & Experience

  • 5+ years of experience in SRE / Production Operations.
  • Strong understanding of SLOs, SLIs, error budgets, incident management, and automated remediation.
  • 3+ years of hands-on experience with observability tools such as Dynatrace and Splunk.
  • Strong experience with logs, metrics, traces, alert tuning, dashboard development, and noise reduction.
  • 3+ years operating cloud-based services on AWS, Azure, or Google Cloud Platform.
  • Strong knowledge of Linux, networking, containers, Kubernetes, and Infrastructure as Code.
  • Experience with Terraform, CloudFormation, or similar IaC technologies.
  • Strong scripting/automation skills using Python and/or Bash.
  • Experience with CI/CD pipelines and automated operational workflows.
  • Experience building runbooks, self-healing workflows, and remediation automation.
  • Hands-on experience with AI-assisted incident response, including auto-triage, incident summarization, pattern detection, anomaly detection, or predictive alerting.
  • Working knowledge of Security/DevSecOps, vulnerability management, secrets/certificate governance, and secure CI/CD practices.

Preferred Skills

  • Experience with canary and blue-green deployments.
  • Experience with chaos engineering, disaster recovery drills, performance testing, and capacity planning.
  • Experience with Well-Architected Reviews, Azure WARA, or Azure Advisor.
  • Hands-on integration with ServiceNow and PagerDuty.
  • Experience building low-friction incident escalation and notification workflows.
  • Experience with AI tools such as GitHub Copilot, Microsoft 365 Copilot, or other enterprise-approved AI platforms.

Please send your profile to .

Similar jobs