Haystack
← Back to Jobs
Technology
WC

Staff Engineer - AI Workload Benchmarking

West Coast Consulting LLCSan Jose, CA🇺🇸United StatesPosted Oct 5, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
San Jose, CA, United States
Posted
Yesterday
C++LLMPyTorchPython

Job Description

Job Description
Onsite in San Jose, CA
Positioning:
Align benchmarking insights to our leadership in memory (HBM, DRAM, CXL) and storage (SSD/NAND) to inform product roadmaps. Enable next-generation AI infrastructure solutions through performance-driven system design, validation, and standards engagement. Requirements
Bachelor's/ Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field; advanced degree preferred. Professional experience in systems performance engineering, storage performance, or AI infrastructure benchmarking.

Demonstrated expertise with NVMe SSDs and storage stack performance analysis (block layer, page cache, file systems, asynchronous I/O). Hands-on experience with AI/ML workloads - LLM training and inference frameworks (PyTorch, vLLM, TensorRT-LLM, or equivalent), embedding pipelines, or vector databases (FAISS, Milvus, DiskANN, HNSW). Strong proficiency with Linux performance and tracing tools: blktrace, perf, eBPF/bpftrace, ftrace, BCC, iostat, fio.

Working knowledge of GPU systems and accelerator I/O paths
Experience designing and executing benchmarks against industry standards (MLPerf Storage, or equivalent). Proficiency in Python for benchmarking automation, data analysis, and visualization; comfort with C/C++ for systems-level work. Proven ability to deliver structured technical reports, characterization studies, and reproducible benchmark artifacts to a senior engineering audience.

Requirements

Bachelor's/ Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field; advanced degree preferred. Professional experience in systems performance engineering, storage performance, or AI infrastructure benchmarking. Demonstrated expertise with NVMe SSDs and storage stack performance analysis (block layer, page cache, file systems, asynchronous I/O).

Hands-on experience with AI/ML workloads LLM training and inference frameworks (PyTorch, vLLM, TensorRT-LLM, or equivalent), embedding pipelines, or vector databases (FAISS, Milvus, DiskANN, HNSW). Strong proficiency with Linux performance and tracing tools: blktrace, perf, eBPF/bpftrace, ftrace, BCC, iostat, fio. Working knowledge of GPU systems and accelerator I/O paths
Experience designing and executing benchmarks against industry standards (MLPerf Storage, or equivalent).

Proficiency in Python for benchmarking automation, data analysis, and visualization; comfort with C/C++ for systems-level work. Proven ability to deliver structured technical reports, characterization studies, and reproducible benchmark artifacts to a senior engineering audience.

Similar jobs