← Back to Jobs
Technology
Embedded Platform Engineer
Balin Technologies LLCSan Jose, CA🇺🇸United StatesPosted 5 Aug 2026
Quick Overview
Work Type
On Site
Level
Mid Senior
Job Description
Job Title: Platform Engineer
Location: We have two locations: Sunnyvale, CA AND San Jose CA(Onsite)
- You will engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones.
- Each morning starts with a structured review of overnight alerts and system performance metrics - identifying any anomalies, triaging what needs attention.
- You will collaborate with your team in a morning stand-up meeting to discuss ongoing projects, recent incidents, and priorities for the day..
- Your tasks will include automating routine processes, analyzing system logs, and developing tools to enhance our monitoring capabilities.
- You\''ll spend part of your day working closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment.
- Throughout the day, your focus is on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers.
- By days end, you will document your work, share insights with your team, and plan for the next days challenges, always with a customer-centric mindset.
What You’ll Bring to the Team:
- Strong experience with architecture, design patterns, reliability and scaling of new and current systems.
- Compute: HPC environments, AMD/NVIDIA GPUs, Linux Kernel/OS, NUMA awareness ,CPU pinning, huge pages, and performance optimization
- Storage: NVMe storage and distributed storage systems, High performance Block, file, and Object Storage Experience
- Software-Defined Networking (SDN): Design, deployment, and operational support of SDN infrastructure
- Experience leading and commanding incidents, including driving root cause analysis, coordinating cross-functional teams, and ensuring follow-through on corrective actions.
- Experience building observability from the ground up — defining SLOs/SLIs, closing monitoring gaps, and implementing alerting strategies that catch failures before customers do.
- Proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystem.
- Experience writing high quality code with at least one programming language (Python, Go, or similar).
- Experience with system-level debugging, including kdump, and kernel panic analysis..
- Proficiency in Infrastructure as Code tooling (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure..
- Experience with TCP/IP and network programming.
- Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms.
Primary skills:
- GPU
- Embedded Platform Engineering
- Network - OVS/OCS
- Storage- Lightbits / Pure storage
- Compute-GKN
- Infra
Skills
TCP/IP
Ansible
GitLab CI
Kubernetes
Python
Terraform
Similar jobs
Network Engineer
TEKsystems c/o Allegis Group · Indian Head, United States
21 minutes ago$45 - $50/hrPrincipal Site Reliability Engineer
Navy Federal Credit Union · Pensacola, United States
1 hour agoLead Site Reliability Engineer(only W2)
Analytics Solutions · Orlando, United States
1 hour agoSenior Cloud Platform Engineer
Avanciers LLC · United States
1 hour agoPrincipal Site Reliability Engineer
Navy Federal Credit Union · Vienna, United States
1 hour agoGPU Infrastructure Site Reliability Engineer(On-site, L2- Face to Face)
Balin Technologies LLC · Sunnyvale, United States
2 hours ago