Haystack
← Back to Jobs
Remote
Technology
CD

Sr Ethernet AI Network Engineer

Cloud Destinations LLCUnited States🇺🇸United StatesPosted 14 Aug 2026

Why This Role Stands Out

This remote Sr Ethernet AI Network Engineer role offers the chance to tackle complex challenges in cutting-edge AI data center fabrics, developing your expertise in high-bandwidth networking. You'll thrive here if you have a strong background in Ethernet fabrics and advanced troubleshooting skills, eager to contribute to critical infrastructure. Don't miss this opportunity to make a significant impact and advance your career.

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Position: Sr Ethernet AI Network Engineer

Location: Remote

Duration: 12 Months Contract

 

Sr Ethernet AI Network Engineer to provide senior network engineering services for Ethernet based AI data center fabrics supporting GPU compute clusters, storage connectivity, and management plane services. This role requires strong operational judgment, advanced troubleshooting capabilities, and the ability to resolve complex network issues within large scale production environments.
This is a 12 month remote contract opportunity within the United States.

Responsibilities:

  • Deploy, validate, and support Ethernet fabrics used for AI and HPC cluster environments.
  • Troubleshoot Layer 1 through Layer 4 issues involving optics, transceivers, cabling, link bring up, VLANs, MLAG, ECMP, BGP, underlay and overlay reachability, congestion, and packet loss.
  • Validate network readiness for distributed training workloads and large east west traffic patterns.
  • Diagnose performance issues related to buffer pressure, microbursts, PFC behavior, Quality of Service policy, MTU mismatches, routing instability, and oversubscription.
  • Partner with Linux, storage, and cluster deployment teams to isolate host versus network fault domains.
  • Review and execute change plans for switch provisioning, firmware upgrades, topology expansion, and maintenance activities.
  • Capture packet level and counter based evidence to support root cause analysis.
  • Develop operational standards for cable mapping, port policy consistency, and fabric health validation. 

Qualifications:
Required Qualifications:

  • 7 or more years of experience in large scale data center networking, including high bandwidth Ethernet fabrics.
  • Strong experience with spine leaf architectures, routing, switching, and production troubleshooting.
  • Hands on experience with BGP, EVPN, VXLAN, MLAG, ECMP, Quality of Service, Priority Flow Control, RoCE considerations, and telemetry analysis.
  • Experience validating optics, breakout configurations, cable plant integrity, and port level consistency.
  • Proven ability to troubleshoot distributed application impacts caused by network behavior.
  • Experience using switch command line interfaces, automation tools, and packet and counter analysis workflows.
  • Strong documentation skills for topology diagrams, incident timelines, and remediation planning. 

Preferred Qualifications:

  • Direct experience supporting AI fabrics carrying large scale GPU collective traffic.
  • Familiarity with SONiC, Cumulus Linux, or similar network operating systems used in AI data centers.
  • Experience with streaming telemetry, Prometheus, Grafana, and network site reliability engineering operating models. 

Tools and Technologies:

  • Ethernet Data Center Fabrics
  • Spine Leaf Network Architectures
  • BGP
  • EVPN
  • VXLAN
  • MLAG
  • ECMP
  • Quality of Service
  • Priority Flow Control
  • RoCE
  • SONiC
  • Cumulus Linux
  • Prometheus
  • Grafana
  • Network Telemetry Platforms
  • Linux
  • Packet Analysis Tools
  • Network Automation Tooling

Skills

Grafana
Prometheus

Similar jobs