Haystack
← Back to Jobs
Other
MU

AI Cluster Validation Engineer (Tester & Programmer)

MphasiS Corporation USAMorrisville, NC🇺🇸United StatesPosted 14 Sept 2026

Why This Role Stands Out

This hybrid role offers a fantastic opportunity to hone your programming and validation skills within cutting-edge AI cluster infrastructure at a reputable company. You'll thrive here if you're a proactive problem-solver with a knack for automation and a passion for ensuring system stability and performance. Apply now to contribute to impactful projects and grow your expertise in a collaborative environment.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Morrisville, NC, United States
Posted
Yesterday
BashKubernetesPythonRoot Cause Analysis

Job Description

Key Responsibilities

  • AI Cluster Infrastructure Validation: Build, deploy, and maintain hardware and software test environments for AI cluster validation, qualification, and performance benchmarking.
  • Execute test plans, reproduce complex issues, collect and analyze logs, and perform first-level root cause analysis across hardware, software, networking, and system components.
  • Develop and enhance automated test frameworks, scripts, and tools to improve validation efficiency, coverage, and repeatability.
  • Collaborate with software, hardware, networking, and system engineering teams to investigate issues, validate fixes, and improve overall cluster stability, scalability, and performance.
  • Conduct functional, performance, stress, and reliability testing for AI/HPC cluster solutions.
  • Document test methodologies, configurations, results, troubleshooting procedures, and operational best practices.

Basic Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
  • 1+ years of experience in deploying, testing, validating, or supporting data center hardware and software systems, with expertise in one or more of the following areas:

Server systems, Networking, Storage

  • Strong understanding of Linux operating systems, system administration, and troubleshooting.
  • Familiarity with data center cluster management platforms, distributed computing environments, and infrastructure validation methodologies.
  • Programming or scripting experience with Python, Bash, or similar languages.
  • Strong analytical, debugging, problem-solving, and troubleshooting skills.
  • Proven ability to quickly learn new technologies and adapt in a fast-paced engineering environment.
  • Self-motivated team player with strong communication and collaboration skills.

Preferred Qualifications

  • Experience operating, maintaining, or supporting data center, cloud, or laboratory environments.
  • Familiarity with AI/HPC clusters, GPU-based systems, and high-speed interconnect technologies such as InfiniBand, RoCE, NVLink, and Ethernet fabrics.
  • Experience developing test automation tools, validation frameworks, or CI/CD pipelines.
  • Knowledge of cluster orchestration and management technologies, including Kubernetes, Slurm, virtualization platforms, or cloud infrastructure.
  • Experience with performance analysis, benchmarking, workload characterization, and system optimization.
  • Understanding of storage technologies, distributed file systems, and AI workload deployment environments.

Similar jobs