Quick Overview
Job Description
ENGAGEMENT SUMMARY
The Candidate will provide senior network engineering services for Ethernet-based AI data center fabrics
supporting GPU compute clusters, storage connectivity, and management plane services. The role requires
strong operational judgment, hands-on troubleshooting, and the ability to resolve failures across spine-leaf
architectures under production pressure.
WHAT THIS CANDIDATE WILL BE DOING
• Deploy, validate, and support Ethernet fabrics used for AI and HPC cluster environments.
• Troubleshoot L1 through L4 issues involving optics, transceivers, cabling, link bring-up, VLANs, MLAG,
ECMP, BGP, underlay/overlay reachability, congestion, and packet loss.
• Validate network readiness for distributed training and large east-west traffic patterns.
• Diagnose performance issues related to buffer pressure, microbursts, PFC behavior, QoS policy, MTU
mismatch, routing instability, and oversubscription.
• Work closely with Linux, storage, and cluster deployment teams to isolate host-versus-network fault
domains.
• Review and execute change plans for switch provisioning, firmware upgrades, topology expansion, and
maintenance events.
• Capture packet-level and counter-based evidence to drive root cause analysis.
• Build operational standards for cable mapping, port policy consistency, and fabric health validation.
WHAT WE NEE D TO SEE
• 7+ years in large-scale data center networking, including high-bandwidth Ethernet fabrics.
• Strong experience with spine-leaf design, routing, switching, and production troubleshooting.
• Hands-on skill with BGP, EVPN/VXLAN, MLAG, ECMP, QoS, PFC, RoCE considerations, and telemetry
interpretation.
• Experience validating optics, breakout configurations, cable plant integrity, and port-level consistency.
• Proven ability to troubleshoot distributed application impact caused by network behavior.
• Comfort using switch CLI, automation tooling, and packet/counter analysis workflows.
• Strong documentation habits for topology, incident timelines, and remediation plans.
PREFERRED EXPERIENCE
• Direct experience with AI fabrics carrying large-scale GPU collective traffic.
• Familiarity with SONiC, Cumulus Linux, or vendor NOS platforms used in AI data centers.
• Experience with streaming telemetry, PrometheGrafana, and network SRE operating models
Similar jobs
- RT
Platform Architect / Cloud Architect
NewRequest Technology, LLC
Chicago, IL🇺🇸Hybrid21 hours agoAWSFlinkSplunk+12Technology - MS
Senior AWS Cloud Engineer (6805) with Security Clearance
NewMetroStar Systems Inc.
Reston, VA🇺🇸$174k - $220k/yrHybrid21 hours agoDockerMicroservicesAWS+11Technology - EV
Senior Cloud Native Architect
NewEverpure
Remote🇺🇸RemoteYesterdayMicroservicesAnsibleHTTP+2 - SI
Cloud Engineer
NewSolomons International
Raleigh, NC🇺🇸Hybrid21 hours agoAWSAzureGoogle CloudTechnology - IN
Cloud Infrastructure Engineer (AWS/Azure/Google Cloud Platform)
NewInfoPeople Corp
White Plains, NY🇺🇸Hybrid21 hours agoDockerAWSAnsible+6Technology - NY
Salesforce Service Cloud / Contact Center Architect
NewNew York Technology Partners
United States🇺🇸Hybrid21 hours agoSalesforceTechnology