Lead end-to-end system validation strategies for AI and HPC hardware platforms, including AI accelerators, GPU clusters, and high-bandwidth memory subsystems in data center environments
Drive hands-on bring-up, characterization, and validation of AI server systems and associated components such as PCIe, NVLink, DRAM, and high-speed networking fabrics
Develop and maintain test specifications, validation procedures, and debug guides tailored to AI infrastructure NPI programs
Investigate and root-cause complex system failures spanning silicon, firmware, software, and hardware layers in collaboration with cross-functional engineering teams
Triage and track hardware and firmware defects through resolution while maintaining forward progress on NPI program milestones
Identify gaps in test coverage and drive improvements to test methodologies, tooling, and automation frameworks across the NPI lifecycle
Partner with AI platform and capacity engineering teams to define acceptance criteria and deployment readiness standards for new AI hardware systems
Guide data collection, analysis, and reporting efforts to surface systemic hardware quality trends and inform go/no-go decisions for production deployment
Communicate validation status, risk assessments, and technical findings to internal engineering teams and external hardware vendors
Collaborate with firmware and software teams to define hardware-software interface requirements for telemetry, diagnostics, and remote management of AI infrastructure
Minimum Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
6+ years of experience in hardware systems engineering, silicon validation, firmware validation, or system-level bring-up for AI servers, GPUs, TPUs, or AI accelerator platforms
Experience in one or more of the following domains: ASIC bring-up and characterization, board-level debug, firmware validation, or large-scale system validation in data center environments
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
Experience developing test specifications, validation procedures, and debug methodologies for complex hardware systems
Experience leading root-cause analysis and troubleshooting of system-level failures across hardware, firmware, and software stacks
Experience with high-speed interconnects or memory subsystems such as PCIe, NVLink, DDR5, or HBM in the context of AI or HPC system validation
Experience analyzing system telemetry and fleet health data to identify reliability trends and drive engineering improvements
Preferred Qualifications
Proficiency in scripting or programming languages such as Python for automation of infrastructure workflows and data analysis
Familiarity with Linux-based server environments and data center management tooling used in large-scale production operations
Experience defining hardware-software interface requirements for telemetry, out-of-band management, or remote diagnostics in data center AI systems
Experience with high-speed interconnects and memory subsystems such as PCIe, NVLink, InfiniBand, DDR5, or HBM in the context of AI or HPC infrastructure operations