Why This Role Stands Out
This hybrid Site Reliability Engineer role offers exciting opportunities to design and scale enterprise infrastructure using cutting-edge AI tooling, perfect for those passionate about automation and operational excellence. You'll thrive here if you enjoy reducing toil through code and ensuring 24/7 system reliability. Apply today to join a forward-thinking team and grow your expertise in a dynamic environment.
Quick Overview
Job Description
Client is looking for an experienced Site Reliability Engineer (SRE) to join our Infrastructure Platform Engineering team. In this role, you will help design, scale, and secure our enterprise on-premises and cloud hybrid infrastructure, driving high availability, automation, and operational excellence.
You will work at the intersection of traditional infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms, and automated pipelines. If you are passionate about reducing toil through code, leveraging modern AI agent frameworks (like Model Context Protocol / MCP), and ensuring 24/7 system reliability across massive fleet environments, we want to hear from you.
Key Responsibilities
- Infrastructure Reliability & Server Management: Architect, manage, and scale robust on-premises infrastructure and server fleets, ensuring high availability, performance optimization, and rigorous incident management.
- Configuration Management & Automation: Drive configuration management across our environment using Chef (Cinc) and Infrastructure as Code (IaC) principles to ensure zero-drift and consistent deployments.
- CI/CD & GitOps Pipelines: Design and maintain secure, scalable CI/CD pipelines (GitLab CI/CD) and GitOps workflows for automated system configuration, package rollout, and patch management.
- AI Agent Tooling & Service Building: Build and integrate next-generation internal tools and services utilizing AI agent frameworks and LLM tooling (such as Claude Code, Codex CLI, and Model Context Protocol) to automate diagnostics, ticket triage, and operational remediation workflows.
- Observability & Telemetry: Implement comprehensive observability platforms (Datadog, Grafana, custom data pipelines) to monitor fleet health, track Chef/Cinc run metrics, and proactively surface system anomalies.
- Cross-Platform Support: Partner with Windows and Linux engineering teams to maintain secure, compliant server and client environments, enforcing security standards (CIS benchmarks) and automated patching.
Qualifications & Required Skills
- Experience: 5+ years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations within large-scale enterprise environments.
- Configuration Management: Deep expertise in Chef (or Cinc) cookbook development, serverless execution modes, and automated provisioning.
- Infrastructure & Server Operations: Strong mastery of on-premises infrastructure, server management, hardware provisioning, and operating systems architecture.
- CI/CD & Automation: Proven track record of building automated CI/CD pipelines and GitOps workflows using modern version control (Git).
- AI Agent & Tool Building Experience: Hands-on experience building internal microservices, tools, or workflows leveraging AI agent tooling, LLM orchestration, or agentic frameworks.
- Observability & Reliability: Expertise in configuring end-to-end monitoring, metrics collection, logging, and alerting (Datadog/Grafana) to ensure platform reliability.
- OS Familiarity: Experience managing and securing Windows infrastructure (alongside Linux) (Note: Windows experience is a key supporting requirement, but the core focus remains on SRE, Chef, and on-prem/hybrid infrastructure).
- Scripting & Development: Proficiency in languages such as Python, Go, PowerShell, or Bash for automation and tooling development.
Similar jobs
- CO
AWS Cloud DevOps Engineer
NewCompunnel Inc.
Berkeley Heights, NJ🇺🇸Hybrid18 hours agoAWSCDKCloudFormation+1Technology - SI
Senior Cloud Platform Engineer (VMware Cloud Foundation)
NewStellent IT LLC
Irving, TX🇺🇸Hybrid18 hours agoAWSTCP/IPAnsible+8Technology - CA
Senior AI Platform Engineer
NewCapgemini
Atlanta, Georgia🇺🇸Hybrid1 hour agoGCPNeo4jAWS+7Technology - AL
Apple Client Platform Engineer (Remote, US)
NewAlectrona
United States🇺🇸$95k - $165k/yrRemote1 hour agoShellSwiftSAML+9Technology - MG
Senior DevOps Engineer – AWS/Kubernetes/GitHub Actions
NewMedinext Global LLC
Chicago, IL🇺🇸On-site18 hours agoDockerAWSNginx+9Technology - LT
Site Reliability Engineer
NewLorven Technologies, Inc.
Minneapolis, MN🇺🇸Hybrid18 hours agoAWSELKLogstash+3Technology