Quick Overview
Job Description
About the Role
Join an early-stage team building infrastructure that makes production Kubernetes easier to deploy and manage for demanding AI workloads. You will help shape the platform layer and its multi-cluster capabilities, working across systems engineering, automation, and cloud-native infrastructure.
What You'll Do
Build platform features in Go, including custom Kubernetes operators and controllers.
Design GitOps workflows with ArgoCD to streamline continuous deployment.
Develop infrastructure-as-code patterns with Terraform and Helm to provision and manage clusters.
Work with distributed storage technologies such as Ceph and WEKA for scalable cluster storage.
Create observability systems with Prometheus and Grafana to surface cluster health and performance.
Build and optimize container networking with Cilium for security and observability.
Design federated Kubernetes architectures for multi-cluster management.
Develop automation that reduces operational work for teams running production workloads.
What We're Looking For
At least 2 years of software development experience, with approximately 5 years preferred.
Hands-on experience writing Go for Kubernetes environments and managing production-scale clusters.
Experience with distributed storage such as Ceph or WEKA, federated Kubernetes, and multi-cluster management.
Experience with GitOps and ArgoCD, plus infrastructure as code using Terraform, Helm, or comparable tools.
Familiarity with cloud platforms, managed Kubernetes, container networking, or service mesh technologies.
Experience with Prometheus and Grafana is useful, as are contributions to open-source infrastructure projects.
Compensation & Benefits
Annual salary of $180,000 to $210,000 USD. Visa sponsorship is available.
Location
This is a full-time, on-site role in San Francisco, California, United States.
Similar jobs
- ON
Site Reliability Engineer II
NewAuto ApplyOnapsis
Dallas🇺🇸Hybrid11 hours agoOracleAWSTDD+10Technology - NE
Site Reliability Engineer
NewAuto ApplyNebius
Remote - United States🇺🇸$130k - $180k/yrRemote4 hours agoBashPythonTechnology - EL
Senior DevOps Engineer (NOAA badge required)
NewAuto ApplyElement84
Alexandria HQ (remote)🇺🇸Remote11 hours agoDockerDynamoDBSQL+20Technology - CM
Senior AI Platform Engineer
NewAuto ApplyCode Metal
Boston Hub🇺🇸Hybrid2 hours agoAPI GatewayMLflowSalesforce+7Technology - WE
Site Reliability Engineer, Cloud Infrastructure
NewAuto ApplyWeave
Weave - Headquarters (Lehi🇺🇸Remote11 hours agoDockerGCPAnsible+11Technology - CO
Senior Site Reliability Engineer
NewAuto ApplyCoalition, Inc.
Any location🇺🇸Hybrid6 hours agoMicroservicesAWSCapacity Planning+7Technology - CI
DevSecOps Engineer
NewAuto ApplyCHAOS Industries
Washington🇺🇸Hybrid7 hours agoDockerAWSOWASP+14Engineering - SA
Senior Infrastructure & Reliability Engineer
NewAuto ApplyStuut Ai
New York City🇺🇸Hybrid8 hours agoSOC 2AuditingERP+2Technology - CA
Associate Site Reliability Engineer/Site Reliability Engineer
NewAuto ApplyC3 AI
Redwood City🇺🇸Hybrid13 hours agoGCPAWSAnsible+8Technology - AI
Contractor: DevOps Engineer
NewAuto ApplyAbacus Insights
United States🇺🇸Hybrid13 hours agoDockerAWSSplunk+10Technology - ED
DevOps Engineer III
NewAuto ApplyEnable Dental
Austin, Texas🇺🇸Remote13 hours agoDockerAWSEncryption+17Technology - TH
DevOps Engineer
NewAuto ApplyTheIncLab
Colorado Springs, Colorado🇺🇸Hybrid7 hours agoDockerShellAWS+22Technology