Haystack
← Back to Jobs
Manufacturing

Senior Production Engineer

Neural Strategic Solutions, Inc.Sunnyvale, CA🇺🇸United StatesPosted 4 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Senior Production Engineer

6 Month CTH+ extension

Onsite in Sunnyvale, CA or SFO, CA

 

JOB DESCRIPTION

 

Key priorities in candidate evaluation:

Strong hands-on experience in at least one of these areas:

  • GPU-based hardware (management, troubleshooting, at-scale operations)
  • SDN (OVN, OVS)
  • Storage (Lightbits, VAST, Pure Storage)

 

Candidates with depth across all three will be rare — what matters is genuine expertise in at least one of these domains rather than surface-level familiarity across all of them.

 

What You''ll Be Working On:

  • As a CORE PE at Crusoe, you will engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones.
  • Each morning starts with a structured review of overnight alerts and system performance metrics - identifying any anomalies, triaging what needs attention
  • You will collaborate with your team in a morning stand-up meeting to discuss ongoing projects, recent incidents, and priorities for the day.
  • Your tasks will include automating routine processes, analyzing system logs, and developing tools to enhance our monitoring capabilities.
  • You''ll spend part of your day working closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment
  • Throughout the day, your focus is on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers.
  • By day''s end, you will document your work, share insights with your team, and plan for the next day''s challenges, always with a customer-centric mindset.

 

What You’ll Bring to the Team:

  • Strong experience with architecture, design patterns, reliability and scaling of new and current systems
  • Experience leading and commanding incidents, including driving root cause analysis, coordinating cross-functional teams, and ensuring follow-through on corrective actions
  • Experience building observability from the ground up — defining SLOs/SLIs, closing monitoring gaps, and implementing alerting strategies that catch failures before customers do
  • Proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystem
  • Experience writing high quality code with at least one programming language (Python, Go, or similar)
  • Experience with system-level debugging, including kdump, and kernel panic analysis.
  • Proficiency in Infrastructure as Code tooling (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.
  • Experience with TCP/IP and network programming
  • Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms.
  • Hardware and GPU troubleshooting experience (nice to have)
  • Exposure to OVN/OVS-based networking stack (nice to have)
  • Strong communication skills

Skills

Root Cause Analysis

Similar jobs