Haystack
← Back to Jobs
Technology
CS

Senior Site Reliability Engineer - Infrastructure and Agentic Automation

Cynet SystemsSanta Clara, CA🇺🇸United StatesPosted Sep 24, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Santa Clara, CA, United States
Posted
17 hours ago
MicroservicesSOC 2AgileBashChefDatadogGitGitLab CIGrafanaLLMPowerShellPython

Job Description

We are looking for Senior Site Reliability Engineer - Infrastructure and Agentic Automation for our client in Santa Clara, CA

Job Title: Senior Site Reliability Engineer - Infrastructure and Agentic Automation

Job Location: Santa Clara, CA

Job Type: Contract

Pay Range: $55.00hr - $60.00hr

Job Overview:

Our client is looking for an experienced Site Reliability Engineer (SRE) to join the Infrastructure Platform Engineering team. In this role, the successful candidate will help design, scale, and secure enterprise on-premises and cloud hybrid infrastructure, driving high availability, automation, and operational excellence. The candidate will work at the intersection of traditional infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms, and automated pipelines. This role is ideal for a professional passionate about reducing toil through code, leveraging modern AI agent frameworks like Model Context Protocol (MCP), and ensuring 24/7 system reliability across massive fleet environments.

Key Responsibilities:

  • Architect, manage, and scale robust on-premises infrastructure and server fleets, ensuring high availability, performance optimization, and rigorous incident management.
  • Drive configuration management across the environment using Chef (Cinc) and Infrastructure as Code (IaC) principles to ensure zero-drift and consistent deployments.
  • Design and maintain secure, scalable CI/CD pipelines (GitLab CI/CD) and GitOps workflows for automated system configuration, package rollout, and patch management.
  • Build and integrate next-generation internal tools and services utilizing AI agent frameworks and LLM tooling (such as Claude Code, Codex CLI, and Model Context Protocol) to automate diagnostics, ticket triage, and operational remediation workflows.
  • Implement comprehensive observability platforms (Datadog, Grafana, custom data pipelines) to monitor fleet health, track Chef/Cinc run metrics, and proactively surface system anomalies.
  • Partner with Windows and Linux engineering teams to maintain secure, compliant server and client environments, enforcing security standards (CIS benchmarks) and automated patching.

Qualifications & Required Skills:

  • 5+ years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations within large-scale enterprise environments.
  • Deep expertise in Chef (or Cinc) cookbook development, serverless execution modes, and automated provisioning.
  • Strong mastery of on-premises infrastructure, server management, hardware provisioning, and operating systems architecture.
  • Proven track record of building automated CI/CD pipelines and GitOps workflows using modern version control (Git).
  • Hands-on experience building internal microservices, tools, or workflows leveraging AI agent tooling, LLM orchestration, or agentic frameworks.
  • Expertise in configuring end-to-end monitoring, metrics collection, logging, and alerting (Datadog/Grafana) to ensure platform reliability.
  • Experience managing and securing Windows infrastructure alongside Linux.
  • Proficiency in languages such as Python, Go, PowerShell, or Bash for automation and tooling development.

== ==

Benefits
Our Benefits Include:
  • Medical, Dental, and Vision Insurance
  • 401(k) Retirement Plan
  • Health Savings Account (HSA)
  • Disability Insurance (Short-Term and Long-Term)
  • Life and AD&D Insurance
  • Paid Sick Leave (where required by applicable state or local law)
  • Supplemental Insurance Plans
  • Identity Theft Protection
  • Pet Insurance
  • Employee Wellness Programs
  • Employee Assistance Program (EAP)
  • Career Growth and Professional Development Opportunities
Disclaimer: Benefits eligibility, accrual rates, and usage limits may vary based on employment status, length of service, and work location. Paid Sick Leave is provided in strict accordance with applicable state and municipal mandates. Cynet Systems Inc. reserves the right to modify, amend, or terminate any benefit plans at any time in accordance with applicable laws.

About Cynet Systems

Founded in 2010 and headquartered in the Washington, DC metro area, Cynet Systems Inc. is a leading technology staffing and workforce solutions company serving Fortune 500 companies, government agencies, and enterprise organizations across the United States and Canada. We deliver agile, scalable talent solutions across IT, engineering, life sciences, clinical, and professional staffing, powered by a high-performing recruitment engine operating across North America and Asia.
As a nationally and locally certified Minority Business Enterprise (MBE), Cynet Systems is committed to helping organizations build high-performing teams while empowering professionals to grow rewarding careers. Our organization is certified to ISO 9001, ISO 14001, ISO 27001, and SOC 2 Type II standards, reflecting our commitment to quality, security, operational excellence, and customer success.

Similar jobs