Only W2- Site Reliability Engineer - Kubernetes Platform (FedRAMP High / IL5)
Quick Overview
Job Description
Job Title: Senior Site Reliability Engineer (SRE) – Kubernetes Platform (FedRAMP High / IL5)
Location: Remote (PST hours preferred)
Employment Type: W2
Experience Required: 10+ Years
Position Summary
We are seeking a Senior Site Reliability Engineer (Technical Leader) to help design, operate, and scale a Kubernetes-based platform supporting highly regulated environments, including FedRAMP High and DoD IL5. This role sits at the intersection of software engineering and infrastructure, working closely with engineers across the stack to ensure the platform is resilient, observable, compliant, and developer-friendly without slowing teams down.
What You'll Do
- Design, build, and operate production-grade Kubernetes platforms in regulated environments.
- Improve system reliability through automation, thoughtful design, and continuous iteration.
- Define and drive SLOs, SLIs, and error budgets to guide reliability decisions.
- Build and evolve CI/CD pipelines that are secure, scalable, and easy to use.
- Implement robust observability (metrics, logs, traces) to make systems understandable and actionable.
- Reduce operational toil by automating repetitive processes and improving workflows.
- Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without sacrificing developer velocity.
- Support ATO processes, including documentation, controls implementation, and audit readiness.
- Participate in on-call rotations supporting customer requests and paging alerts.
- Participate in incident response, blameless postmortems, and continuous improvement efforts.
Required Qualifications
- 10+ years of experience in SRE, DevOps, or infrastructure engineering.
- Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream).
- Hands-on experience working in FedRAMP High and/or DoD IL5 environments.
- Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals.
- Experience with Infrastructure as Code (Terraform preferred).
- Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD).
- Proficiency in scripting or programming (Python, Go).
- Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK).
- Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF).
Preferred Qualifications
- Experience with service mesh technologies (Istio, Linkerd).
- Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno).
- Experience with GitOps workflows.
- Exposure to multi-cluster or hybrid cloud architectures.
- Knowledge of FIPS-compliant systems or DoD Cloud SRG.
- Relevant certifications (CKA, CKS, cloud provider certs, Security+).
Similar jobs
- SI
Senior AI Engineer / Senior AI Platform Engineer - Agentic AI
NewStratEdge It consulting INC
Chicago, IL🇺🇸$82/hrOn-site19 hours agoAzureKafkaLLM+1Technology - SC
F2F INTERVIEW||C2H ROLE||DATA PLATFORM ENGINEER||READING, PA (HYBRID ROLE)
NewShift Code Analytics
Reading, PA🇺🇸Hybrid19 hours agoDynamoDBMongoDBMySQL+17Technology - PE
Network Engineer with Security Clearance
Peraton
Fort Meade, MD🇺🇸$86k - $138k/yrHybrid6 weeks agoHTTPSiOSTechnology - SM
Lead SRE (AWS CloudWatch/AppSignals/ADOT/X-Ray) - 10+ Years - Dallas, TX Hybrid
NewSystems Management Group, Inc
Dallas, TX🇺🇸On-site19 hours agoAWSSplunkCloudFormation+3Technology - DE
Sr Mgr, Site Reliability Engineer (SRE)
NewDisney Experiences
Orlando, FL🇺🇸$175k - $215k/yrHybridYesterdayAWSAnsibleAzure+4Technology - AT
Software Developer with DevOps
NewAivanta Tech Inc
San Diego, CA🇺🇸On-site19 hours agoDockerAWSFlink+8Technology