Haystack
← Back to Jobs
Technology
BT

Core Platform Engineer/ SRE (incident management , Storage, Network, GPU)

Balin Technologies LLCSunnyvale, CA🇺🇸United StatesPosted 4 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Sunnyvale, CA, United States
Posted
2 days ago
TCP/IPAnsibleGitLab CIKubernetesPythonTerraform

Job Description

Role 1 — Core Platform Engineer, Level 1
The engineering first line of defense: incident response, triage, reliability, and automation across the full infrastructure stack.
Day-to-day:
  • Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence.
  • Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies.
  • Participates in team stand-ups on projects, incidents, and daily priorities.
  • Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring.
  • Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment.
  • Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset.
Must have:
  • Architecture, design patterns, reliability, and scaling of new and existing systems.
  • Incident command experience — driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through.
  • Observability built from the ground up — defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do.
  • Linux kernel internals — scheduler, memory allocation, driver subsystems.
  • High-quality code in at least one language (Python, Go, or similar).
  • System-level debugging — kdump, kernel panic analysis.
  • IaC (Ansible, Terraform, Kubernetes) and CI/CD (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.
  • TCP/IP and network programming.
  • Distributed storage systems — object, block, and/or file storage paradigms.
  • Strong communication skills.
Sourcing note: This is a deep SRE profile, not a pure generalist. The kernel-internals and system-level debugging bar is real and higher than a typical "L1" label implies — screen for genuine engineering depth, not helpdesk/NOC-tier breadth.

Similar jobs