Why This Role Stands Out
This Embedded Platform Engineer role at Balin Technologies offers significant opportunities for professional growth and skill development within a reputable tech company. You'll thrive here if you are a proactive, collaborative individual eager to contribute to robust system reliability and enjoy a dynamic, customer-focused environment. Apply now to join their innovative team!
Quick Overview
Seniority
Mid Senior
Work mode
On Site
Location
San Jose, CA, United States
Posted
Yesterday
KubernetesStakeholder Management
Job Description
Job Title: Core Platform /System Engineer
Location: Sunnyvale, CA , San Jose (On-Site)
Duration: Long-Term Contract
Duration: Long-Term Contract
Job Description
- As a Core Platform Engineer, we will serve as the first responder for production incidents,
- orchestrate incident management, drive reliability improvements, and establish SRE best
- practices across Compute, Networking, Storage, and GPU infrastructure teams.
- Candidates should possess hands-on infrastructure experience and sufficient technical
- depth to identify affected systems, engage the right subject matter experts, and drive
- incident resolution processes using data and observability signals.
Responsibilities
- Act as first responder during infrastructure incidents.
- Lead incident bridges and coordinate cross-functional response efforts.
- Perform incident triage and identify impacted infrastructure domains.
Gather evidence and telemetry to route incidents to the correct SME team.
- Drive incident communications and stakeholder updates.
- Improve reliability processes across platform engineering teams.
- Define and promote SRE best practices and operational standards. (SLO,SLI)
- Identify observability gaps and implement improvements.
- Build automation for incident response workflows.
- Manage and optimize incident management tooling (e.g., ).
- Support change management and operational readiness processes.
- Assist foundation engineering teams in identifying reliability risks and trends.
- Participate in on-call activities and operational reviews.
Required Qualifications
- 5 to 10+ years of experience in Site Reliability Engineering, Platform Engineering,
- Infrastructure Operations, or Systems Engineering.
- Strong infrastructure troubleshooting experience.
- GPU infrastructure
- KVM/virtualization
- SDN (OVN/OVS)
- Storage (Lightbits/Pure Storage)
- Proven incident management and operational leadership experience.
- Experience running high-severity production incidents.
- Strong understanding of observability, monitoring, SLIs, and SLOs.
- Experience building operational automation.
- Ability to make data-driven decisions during outages and service disruptions.
- Preferred Qualifications
- Experience with or similar incident management platforms.
- Kubernetes production operations experience.
- Cloud-native infrastructure experience.
- Experience supporting large-scale AI or GPU environments.
- Strong communication and stakeholder management skills.
- What Success Looks Like
- Quickly identifies affected infrastructure domains during incidents.
- Effectively coordinates SMEs and engineering teams.
- Reduces incident response and recovery times.
- Improves observability and operational processes.
- Establishes reliability standards across Crusoe's infrastructure platform.
- Ideal Candidate Screening Criteria (For Both Roles)
Similar jobs
- MC
Lead DevOps Engineer
NewMaven Companies
Seattle, WA🇺🇸Hybrid22 hours agoDockerShellAWS+7Technology - BA
Network Engineer
NewBooz Allen Hamilton
Washington, DC🇺🇸$77.6k - $176k/yrOn-site22 hours agoVMwareiOSTechnology - FG
Database Site Reliability Engineer (Database Operations)
NewFusion Global Solutions
Alpharetta, GA🇺🇸Hybrid22 hours agoMongoDBMySQLOracle+8Technology - OT
DevSecOps Engineer
NewOscar Technology
Dallas, TX🇺🇸On-site22 hours agoEngineering - KT
SRE (AI/ML)
NewKforce Technology Staffing
Maryland Heights, MO🇺🇸Hybrid22 hours agoAWSMLOpsSnowflake+11Technology - TM
Senior Technical SRE
NewTech Mahindra (Americas) Inc.
Redmond, WA🇺🇸Hybrid22 hours agoMicroservicesSQLAzure+3Technology