Haystack
← Back to Jobs
Technology
ET

Senior Network Engineer

eTeam, Inc.Santa Clara, CA🇺🇸United StatesPosted 1 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Santa Clara, CA, United States
Posted
21 hours ago

Job Description

ENGAGEMENT SUMMARY

The Candidate will provide senior network engineering services for Ethernet-based AI data center fabrics

supporting GPU compute clusters, storage connectivity, and management plane services. The role requires

strong operational judgment, hands-on troubleshooting, and the ability to resolve failures across spine-leaf

architectures under production pressure.

WHAT THIS CANDIDATE WILL BE DOING

• Deploy, validate, and support Ethernet fabrics used for AI and HPC cluster environments.

• Troubleshoot L1 through L4 issues involving optics, transceivers, cabling, link bring-up, VLANs, MLAG,

ECMP, BGP, underlay/overlay reachability, congestion, and packet loss.

• Validate network readiness for distributed training and large east-west traffic patterns.

• Diagnose performance issues related to buffer pressure, microbursts, PFC behavior, QoS policy, MTU

mismatch, routing instability, and oversubscription.

• Work closely with Linux, storage, and cluster deployment teams to isolate host-versus-network fault

domains.

• Review and execute change plans for switch provisioning, firmware upgrades, topology expansion, and

maintenance events.

• Capture packet-level and counter-based evidence to drive root cause analysis.

• Build operational standards for cable mapping, port policy consistency, and fabric health validation.

WHAT WE NEE D TO SEE

• 7+ years in large-scale data center networking, including high-bandwidth Ethernet fabrics.

• Strong experience with spine-leaf design, routing, switching, and production troubleshooting.

• Hands-on skill with BGP, EVPN/VXLAN, MLAG, ECMP, QoS, PFC, RoCE considerations, and telemetry

interpretation.

• Experience validating optics, breakout configurations, cable plant integrity, and port-level consistency.

• Proven ability to troubleshoot distributed application impact caused by network behavior.

• Comfort using switch CLI, automation tooling, and packet/counter analysis workflows.

• Strong documentation habits for topology, incident timelines, and remediation plans.

PREFERRED EXPERIENCE

• Direct experience with AI fabrics carrying large-scale GPU collective traffic.

• Familiarity with SONiC, Cumulus Linux, or vendor NOS platforms used in AI data centers.

• Experience with streaming telemetry, PrometheGrafana, and network SRE operating models

 

Similar jobs