Quick Overview
Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
1 week ago
Job Description
Job Title: Performance Engineer
Preferred Location: Remote
Duration: Long term Contract
What you'll do
- Set the performance architecture agenda. Bring deep, current expertise across file, block, and object storage protocols and translate it into concrete performance requirements for //EXA's data path (NFSv3 direct-to-DN, pNFS layouts, S3 via MDN) informed by how HPC and hyperscale environments actually push storage systems (checkpointing, small-file metadata storms, GPU-starved read patterns, mixed-tenant burst I/O).
- Track and act on the NeoCloud / sovereign-cloud shift. Maintain a living view of where //EXA's highest-value deployments are heading GPU-cloud and sovereign-cloud operators (CoreWeave, Crusoe, Nscale, and similar) and make sure //EXA's performance roadmap, reference architectures, and sizing guidance map to how these operators actually buy and operate infrastructure (multi-tenant GPU clusters, bursty training/inference mixes, strict SLAs to end customers).
- Own competitive performance positioning. Build and maintain deep, technically substantiated comparisons against VAST Data, DDN, and WEKA not marketing bullet points, but real architectural analysis (metadata scaling model, erasure coding/durability tradeoffs, protocol support, GPU-direct paths, cost/performance at scale) that engineering and field teams can use to win technical evaluations and POCs.
- Drive performance tuning for multi-tenant HPC/AI workloads. Lead tuning and validation work spanning the full stack a GPU cluster touches storage (MDN/DN geometry, pack groups, erasure coding layout), networking (RDMA, RoCE/InfiniBand fabric behavior, NIC/queue tuning), and compute (GPU-side I/O patterns, checkpoint/restore, data loader behavior) with particular focus on how these interact when multiple tenants/workloads share the same //EXA fleet.
- Build and run the benchmark suite. Own //EXA's benchmark framework and result credibility: MLPerf Storage (v2/v3), elbencho, IO500, fio/vdbench-class synthetic tests, and workload-representative benchmarks for AI training/inference and traditional HPC. Ensure results are reproducible, defensible in public disclosure, and directly comparable to published competitor numbers.
- Define QoS, limits, and workload segmentation. Drive the technical requirements and validation for quality-of-service guarantees, per-tenant/per-workload throughput and IOPS limits, and workload isolation the mechanisms that let //EXA make hard SLA commitments in shared, multi-tenant NeoCloud deployments rather than best-effort performance.
What makes you competitive for this role
- Deep, hands-on background in storage performance engineering across file, block, and object protocols, ideally with direct HPC or hyperscale exposure (parallel filesystems, pNFS/NFS at scale, S3-scale object stores).
- Working knowledge of GPU cluster architecture RDMA fabrics, GPUDirect Storage, checkpoint/restore patterns for large model training and how storage bottlenecks manifest in mixed compute/network/storage systems.
- Fluency with industry benchmark standards (MLPerf Storage, IO500) and load-generation tooling (elbencho, fio, vdbench), plus the judgment to design workload-representative tests beyond canned benchmarks.
- Demonstrated ability to build rigorous, technically credible competitive analysis (not slideware) against systems like VAST, DDN, and WEKA architecture-level understanding, not just spec-sheet comparison.
- Experience with multi-tenant resource management concepts (QoS, rate limiting, workload isolation) in a distributed systems context.
- Comfortable operating across the stack and across audiences deep enough to debug an RDMA queue-pair stall or a metadata hot-partition, articulate enough to brief a NeoCloud customer's technical evaluation team.
Similar jobs
- SI
Software Developer AirFlow Dev
NewSaxon Infotech
United States🇺🇸HybridYesterdayAWSAirflowApache+3Technology - VT
Senior Staff Engineer
NewVailexa Technology LLC
United States🇺🇸HybridYesterdayAWSEnvoyAzure+7Technology - NS
Network Engineer II
NewNovaLink Solutions
Madison, WI🇺🇸RemoteYesterdayTechnology - DT
AI-Native Software Engineer
NewDanta Technologies
Durham, NC🇺🇸$60/hrHybridYesterdayDynamoDBMongoDBOracle+9Technology - AI
Software Engineer, Geospatial Platform (Terra) with Security Clearance
NewAnduril Industries
Seattle, WA🇺🇸$166k - $220k/yrHybridYesterdayRustAWSAzure+9Technology - SS
C++ Developer /Software Engineer
NewSimple Solutions
Boston, MA🇺🇸HybridYesterdayMongoDBEmbedded SystemsAgile+2Technology