Haystack
← Back to Jobs
Technology

Senior Engineer, Agentic SRE / Platform DevOps Scripting

IntraedgeUnited States🇺🇸United StatesPosted 31 Jul 2026

Why This Role Stands Out

This hybrid role offers significant growth potential as you engineer self-healing infrastructure and contribute to a modern cloud ecosystem, making it ideal for a proactive Senior SRE/DevOps Automation Engineer who thrives on eliminating manual tasks. You'll build innovative automation solutions within a reputable company, so seize this opportunity to elevate your career.

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

The Mission
We are looking for a forward-thinking Senior SRE / DevOps Automation Engineer who hates manual toil and believes in engineering self-healing infrastructure. In this role within our modern cloud ecosystem (Health100 ), you won''''t just respond to alerts—you will build the automation that triages them. We are actively infusing Generative AI and Agentic workflows into our platform operations to summarize logs, enrich tickets, and accelerate incident response. If you want to pioneer the intersection of SRE and AI-assisted platform engineering, this is your role. 
What You Will Do
  • Kill the Toil: Design and build intelligent automation scripts (Python, Bash ) and agentic workflows to handle log analysis, automated failure summaries, and issue routing.
  • Pioneer AI Engineering: Evaluate and deploy LLM orchestration tools (LangChain, LlamaIndex, Vertex AI, Gemini ) to create smart Slack bots and ticket enrichment pipelines.
  • Maintain Platform Health: Enforce SRE fundamentals—service health monitoring, strict error budgets, and high-quality alerting—across production GKE clusters.
  • Secure the Guardrails: Build secure-by-default, auditable automation patterns that safely operate within strict healthcare compliance and PHI handling standards.
  • Cross-Functional Triage: Act as the technical nexus between platform, network, database, and security teams to diagnose complex infrastructure and container dependencies.
The Skillset You Bring
  • SRE & Production Depth: 5+ years of experience across Site Reliability Engineering, DevOps, or Platform Operations in complex, matrixed enterprise environments.
  • Automation Mastery: 4+ years of advanced engineering using Python and Bash to build API-driven tools and automation utilities.
  • Cloud-Native Stack: 3+ years managing Kubernetes (GKE) , CI/CD pipelines, and enterprise observability/telemetry suites.
  • AI Curiosity or Competency: Hands-on interest or practical experience building agent-based automation or integrating with LLM APIs.  
Bonus Points If You Have
  • Production experience building operational Slack bots or incident post-mortem automation.
  • Deep knowledge of modern data layers (Kafka, Redis, Oracle DB@Google Cloud Platform, Postgres ).
  • Active certifications like Google Certified Professional Cloud DevOps Engineer or CKA/CKAD

Skills

Oracle
Bash
Generative AI
Google Cloud
Kafka
Kubernetes
LLM
Python
Redis

Similar jobs