Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Morrisville, NC, United States
Posted
2 days ago
BashKubernetesPythonRoot Cause Analysis
Job Description
JD:
Key Responsibilities
- AI Cluster Infrastructure Validation: Build, deploy, and maintain hardware and software test environments for AI cluster validation, qualification, and performance benchmarking.
- Execute test plans, reproduce complex issues, collect and analyze logs, and perform first-level root cause analysis across hardware, software, networking, and system components.
- Develop and enhance automated test frameworks, scripts, and tools to improve validation efficiency, coverage, and repeatability.
- Collaborate with software, hardware, networking, and system engineering teams to investigate issues, validate fixes, and improve overall cluster stability, scalability, and performance.
- Conduct functional, performance, stress, and reliability testing for AI/HPC cluster solutions.
- Document test methodologies, configurations, results, troubleshooting procedures, and operational best practices.
Basic Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
- 1+ years of experience in deploying, testing, validating, or supporting data center hardware and software systems, with expertise in one or more of the following areas:
Server systems, Networking, Storage
- Strong understanding of Linux operating systems, system administration, and troubleshooting.
- Familiarity with data center cluster management platforms, distributed computing environments, and infrastructure validation methodologies.
- Programming or scripting experience with Python, Bash, or similar languages.
- Strong analytical, debugging, problem-solving, and troubleshooting skills.
- Proven ability to quickly learn new technologies and adapt in a fast-paced engineering environment.
- Self-motivated team player with strong communication and collaboration skills.
Preferred Qualifications
- Experience operating, maintaining, or supporting data center, cloud, or laboratory environments.
- Familiarity with AI/HPC clusters, GPU-based systems, and high-speed interconnect technologies such as InfiniBand, RoCE, NVLink, and Ethernet fabrics.
- Experience developing test automation tools, validation frameworks, or CI/CD pipelines.
- Knowledge of cluster orchestration and management technologies, including Kubernetes, Slurm, virtualization platforms, or cloud infrastructure.
- Experience with performance analysis, benchmarking, workload characterization, and system optimization.
- Understanding of storage technologies, distributed file systems, and AI workload deployment environments.
Similar jobs
- JO
Manager, Analytics
NewJoveo Inc
Pinecrest, FL🇺🇸$105k - $131.3k/yrHybrid11 hours agoSQLTableauHTML+1 - JO
Manager, Analytics
NewJoveo Inc
Hollywood, FL🇺🇸$105k - $131.3k/yrHybrid11 hours agoSQLTableauHTML+1 - JO
Manager, Analytics
NewJoveo Inc
Ives Estates, FL🇺🇸$105k - $131.3k/yrHybrid11 hours agoSQLTableauHTML+1 - JO
Manager, Analytics
NewJoveo Inc
Virginia Gardens, FL🇺🇸$105k - $131.3k/yrHybrid11 hours agoSQLTableauHTML+1 - JO
Manager, Analytics
NewJoveo Inc
Sweetwater, FL🇺🇸$105k - $131.3k/yrHybrid11 hours agoSQLTableauHTML+1 - JO
Manager, Analytics
NewJoveo Inc
Country Walk, FL🇺🇸$105k - $131.3k/yrHybrid11 hours agoSQLTableauHTML+1