Why This Role Stands Out
This remote Infrastructure/Cloud DevOps-SRE role offers excellent growth potential within a reputable company, allowing you to build and automate cutting-edge Kubernetes platforms while enjoying a competitive hourly rate of $55-$65. You'll thrive here if you're passionate about SRE principles, automation, and solving complex distributed systems challenges. Apply today to advance your career in a flexible, impactful position!
Quick Overview
Job Description
Infrastructure/Cloud DevOps-SRE
W2 Contract
Pay Rate: $55 - $65 per hour
Location: Cupertino, CA - Remote Role
Job Summary:
We are looking for a highly motivated DevOps / Site Reliability Engineer to support large-scale Kubernetes-based infrastructure and platform operations. This role is focused on building, automating, and operating highly reliable systems that power critical engineering platforms and services.
Duties and Responsibilities:
- Design, build, automate, and support scalable Kubernetes-based platforms and services
- Operate and troubleshoot production environments running at scale
- Develop automation and tooling to improve operational efficiency and reliability
- Monitor platform health, performance, and availability using observability tooling
- Troubleshoot infrastructure, application, and networking issues across distributed systems
- Work closely with engineering teams to improve deployment, reliability, and scalability practices
- Participate in operational support, incident response, and root cause analysis
- Improve CI/CD workflows and deployment automation
- Drive operational excellence through documentation, automation, and process improvements
- Take ownership of projects and independently drive deliverables to completion
Requirements and Qualifications:
- Strong hands-on experience with Kubernetes platforms such as:
- EKS
- GKE
- AKS or similar
- Experience running and supporting applications on Kubernetes at scale
- Strong understanding of containerized infrastructure and distributed systems
- Experience with monitoring and observability tools, preferably:
- Grafana
- Prometheus
- Experience with CI/CD pipelines and deployment automation
- Experience with Splunk logging, log analysis, and troubleshooting
- Strong scripting and automation experience using Python and/or Golang
- Experience troubleshooting production systems under pressure
- Strong communication and collaboration skills
- Self-starter mentality with strong ownership and accountability
Preferred Qualifications
- Experience operating Ray clusters/services
- Strong networking and troubleshooting experience
- Experience with cloud infrastructure and platform services
- Experience with Infrastructure as Code and automation frameworks
- Experience supporting high-scale production systems
- Familiarity with SRE principles and operational best practices
Bayside Solutions, Inc. is not able to sponsor any candidates at this time. Additionally, candidates for this position must qualify as a W2 candidate.
Bayside Solutions, Inc. may collect your personal information during the position application process. Please reference Bayside Solutions, Inc.'s CCPA Privacy Policy at
Similar jobs
- SE
Director, Site Reliability Engineering & Service Enablement
NewServiceNow
United States🇺🇸$221.2k - $387.1k/yrRemote17 minutes agoGCPAWSAzure+1Technology - NV
Software DevOps Engineer, Networking
NewNVIDIA
Santa Clara, California🇺🇸Hybrid17 minutes agoDockerShellSonarQube+10Technology - AL
Principal AI Platform Engineer – Agentic AI & MCP
NewAlfvo, LLC
United States🇺🇸Hybrid11 hours agoAzureGitHub ActionsLLM+3Technology - MC
Databricks Platform Engineer
NewMCKESSON
Columbus, Ohio🇺🇸$106.5k - $177.5k/yrHybrid5 hours agoSQLSnowflakeAzure+4Technology - MA
Senior Site Reliability Engineer
NewMastercard
O Fallon, Missouri🇺🇸$96k - $163k/yrHybrid11 hours agoDockerGCPMicroservices+12Technology - NG
Sr Principal DevOps Engineer (Cloud) (26-297) with Security Clearance
Northrop Grumman
Colorado Springs, CO🇺🇸$129.3k - $193.9k/yrOn-site2 months agoScrumActive DirectoryAgile+7Technology