Haystack
← Back to Jobs
Technology
VT

Ethernet AI Network Engineer

Vailexa Technology LLCUnited States🇺🇸United StatesPosted 20 Aug 2026

Why This Role Stands Out

Leverage your expertise in Ethernet AI networking to contribute to cutting-edge AI and HPC cluster environments with a 12-month remote contract opportunity, offering flexibility and the chance to hone your troubleshooting and validation skills. This role is perfect for a mid-senior engineer who thrives on complex network challenges and enjoys collaborating with diverse technical teams to ensure optimal performance. Apply today to make a significant impact in this dynamic field!

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
6 days ago
GrafanaPrometheus

Job Description

Role- Sr Ethernet AI Network Engineer
Location- Remote US 
Duration - 12-month remote contract opportunity within the United States
 
Responsibilities:
  • Deploy, validate, and support Ethernet fabrics used for AI and HPC cluster environments.
  • Troubleshoot Layer 1 through Layer 4 issues involving optics, transceivers, cabling, link bring up, VLANs, MLAG, ECMP, BGP, underlay and overlay reachability, congestion, and packet loss.
  • Validate network readiness for distributed training workloads and large east west traffic patterns.
  • Diagnose performance issues related to buffer pressure, microbursts, PFC behavior, Quality of Service policy, MTU mismatches, routing instability, and oversubscription.
  • Partner with Linux, storage, and cluster deployment teams to isolate host versus network fault domains.
  • Review and execute change plans for switch provisioning, firmware upgrades, topology expansion, and maintenance activities.
  • Capture packet level and counter based evidence to support root cause analysis.
  • Develop operational standards for cable mapping, port policy consistency, and fabric health validation.
     
  • Required Qualifications
  • 7 or more years of experience in large scale data center networking, including high bandwidth Ethernet fabrics.
  • Strong experience with spine leaf architectures, routing, switching, and production troubleshooting.
  • Hands on experience with BGP, EVPN, VXLAN, MLAG, ECMP, Quality of Service, Priority Flow Control, RoCE considerations, and telemetry analysis.
  • Experience validating optics, breakout configurations, cable plant integrity, and port level consistency.
  • Proven ability to troubleshoot distributed application impacts caused by network behavior.
  • Experience using switch command line interfaces, automation tools, and packet and counter analysis workflows.
  • Strong documentation skills for topology diagrams, incident timelines, and remediation planning.
 
Preferred Qualifications
  • Direct experience supporting AI fabrics carrying large scale GPU collective traffic.
  • Familiarity with SONiC, Cumulus Linux, or similar network operating systems used in AI data centers.
  • Experience with streaming telemetry, Prometheus, Grafana, and network site reliability engineering operating models. Tools and Technologies:
  • Ethernet Data Center Fabrics
  • Spine Leaf Network Architectures
  • BGP
  • EVPN
  • VXLAN
  • MLAG
  • ECMP
  • Quality of Service
  • Priority Flow Control
  • RoCE
  • SONiC
  • Cumulus Linux
  • Prometheus
  • Grafana
  • Network Telemetry Platforms
  • Linux
  • Packet Analysis Tools

Similar jobs