Quick Overview
Job Description
Role: Senior AI Validation Engineer
Location: Santa Clara, CA (Onsite)
Job Type : Contract -W2
Job Description :
The Candidate will provide AI infrastructure validation services focused on proving cluster readiness, identifying failure modes early, and accelerating root cause isolation before production impact. This role bridges systems, networks, and workload behaviour, and is ideal for a senior engineer who treats validation as an engineering discipline rather than a checklist.
Responsibilities :
• Design and execute validation plans for AI infrastructure spanning compute nodes, GPU communication, fabric health, storage access, orchestration, and workload readiness.
• Run structured bring-up, soak, regression, and qualification tests on new or changed AI cluster environments.
• Reproduce and isolate failures involving distributed training, node instability, communication libraries, container stacks, storage paths, or network transport behavior.
• Build validation coverage for Ethernet and InfiniBand environments, including host readiness and end-to- end workload verification.
• Correlate test failures with system logs, telemetry, firmware state, and application symptoms to accelerate defect isolation.
• Partner with deployment, Linux, network, and platform teams to close validation gaps before operational handoff.
• Create defect signatures, pass-fail criteria, readiness reports, and release recommendations.
• Improve automation for cluster certification, health scoring, and post-change validation.
Required Skills :
• 10+ years in systems validation, performance engineering, QA for infrastructure, or AI/HPC environment certification.
• Strong troubleshooting ability across Linux hosts, GPU systems, network fabrics, containers, and distributed workload behavior.
• Experience designing validation strategies rather than only executing scripted test cases.
• Familiarity with AI workload dependencies such as NCCL, RDMA paths, storage throughput, and multi-node orchestration behavior.
• Ability to distinguish infrastructure defects from workload, framework, or configuration issues.
• Strong scripting and automation capability for test execution and evidence collection.
• Clear written communication for readiness assessments and defect reports.
Preferred Skills :
• Experience validating GPU clusters, large training environments, or pre-production AI factories.
• Familiarity with telemetry analysis, burn-in workflows, and hardware-firmware-software compatibility testing.
• Experience building qualification suites for both deployment gates and steady-state operations.
Similar jobs
- QU
Oracle OIC+ B2B
NewQualis1 Inc.
Durham, NC🇺🇸Hybrid22 hours agoOracle - VS
Google Cloud Platform Cloud Migration Architect
NewVigna Solutions Inc.
Manor, TX🇺🇸On-site22 hours agoSQLAWSGoogle Cloud+1 - DL
Siemens Opcenter MOM platform
NewDiverse Lynx Llc
Seattle, WA🇺🇸Hybrid22 hours agoAgileAssemblyERP+3 - TB
Strategic Growth Partner – Technology Consulting
NewThink Big Solutions, Inc
United States🇺🇸Hybrid22 hours agoOracleAWSMachine Learning+10 - SI
Study Facilitation Support Executive - User Research
NewStratEdge It consulting INC
Cupertino, CA🇺🇸$25/hrOn-site22 hours agoComplianceSchedulingiOS - DW
Head of Data & Analytics
NewDale Workforce Solutions
Berwyn, PA🇺🇸Hybrid22 hours agoOracleAzureCRM+3