Haystack
← Back to Jobs
Technology
FT

Core Platform Engineer – Infrastructure Reliability & Incident Management

FSTONE TechnologiesSunnyvale, CA🇺🇸United StatesPosted 26 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Sunnyvale, CA, United States
Posted
23 hours ago
KubernetesPagerDuty

Job Description

Core Platform Engineer – Infrastructure Reliability & Incident Management

We are seeking a Core Platform Engineer / SRE with strong infrastructure troubleshooting and production incident-management experience.

Key Responsibilities

  • Act as first responder and lead Sev-1/Sev-2 production incidents.
  • Lead incident bridges and coordinate cross-functional SMEs.
  • Perform infrastructure triage using logs, metrics, and telemetry.
  • Drive incident communication, resolution, and post-incident improvements.
  • Implement SRE best practices, SLIs/SLOs, observability, and automation.
  • Support on-call, change management, and operational readiness.

Required Skills

  • 5–10+ years in SRE, Platform Engineering, Infrastructure, or Systems Engineering.
  • Strong hands-on infrastructure troubleshooting.
  • Deep expertise in at least one: KVM/QEMU, OVN/OVS/SDN, GPU infrastructure, or Storage (Lightbits/Pure Storage).
  • Production Kubernetes/GKE experience preferred.
  • Strong knowledge of monitoring, observability, SLI/SLO, and incident management.
  • Experience with incident.io, PagerDuty, Opsgenie, or similar tools.
  • Strong communication and stakeholder-management skills.

Ideal Candidate

A senior infrastructure/SRE engineer who can quickly triage production incidents, identify the affected infrastructure domain, coordinate SMEs, and drive incidents to resolution.

Similar jobs