Haystack
← Back to Jobs
Technology

Site Reliability Engineer Kubernetes Platform

Cosmic-I LLC DBA Northern BaseUnited States🇺🇸United StatesPosted 29 Jul 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description


Senior Site Reliability Engineer Kubernetes Platform
Remote ,San Jose, CA
10-12 years
Job Description
Must Have Technical/Functional Skills:
•    10+ years of experience in SRE, DevOps, or infrastructure engineering 
•    Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream) 
•    Hands-on experience working in FedRAMP High and/or DoD IL5 environments 
•    Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals 
•    Experience with Infrastructure as Code (Terraform preferred) 
•    Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD) 
•    Proficiency in scripting or programming (Python, Go) 
•    Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK) 
•    Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF) 
Roles & Responsibilities:
•    Design, build, and operate production-grade Kubernetes platforms in regulated environments 
•    Improve system reliability through automation, thoughtful design, and continuous iteration 
•    Define and drive SLOs, SLIs, and error budgets to guide reliability decisions 
•    Build and evolve CI/CD pipelines that are secure, scalable, and easy to use 
•    Implement robust observability (metrics, logs, traces) to make systems understandable and actionable 
•    Reduce operational toil by automating repetitive processes and improving workflows 
•    Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without sacrificing developer velocity 
•    Support ATO processes, including documentation, controls implementation, and audit readiness Confidential 
•    Participate in on-call rotations supporting customer requests and paging alerts 
•    Participate in incident response, blameless postmortems, and continuous improvement efforts 
•    Help shape a platform that engineers enjoy using 
Nice to Have 
•    Experience with service mesh technologies (Istio, Linkerd) 
•    Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno) 
•    Experience with GitOps workflows 
•    Exposure to multi-cluster or hybrid cloud architectures 
•    Knowledge of FIPS-compliant systems or DoD Cloud SRG 
•    Relevant certifications (CKA, CKS, cloud provider certs, Security+) 

Skills

ELK
Service Mesh
ArgoCD
GitHub Actions
GitLab CI
Grafana
Istio
Jenkins
Kubernetes
Prometheus
Python
Terraform

Similar jobs