Perform firmware upgrades, hardware validation, and storage setup
Configure and administer physical and logical resources, including M IG partitioning and BlueField platforms
Install and configure operating systems, cluster software, drivers, containers (Docker), and NGC CLI
Manage and orchestrate clusters using NVIDIA Base Command Manager, Slurm, Pyxis, Enroot, and Run: Ai
Perform stress, benchmarking, and burn-in tests using HPL, NCCL, NVIDIA Nemo, and ClusterKit
Verify cabling, firmware/software versions, and network signal quality
Troubleshoot and resolve hardware, software, storage, and performance faults
Replace faulty components and optimize systems for AMD/Intel platforms
Monitor, document, and report on cluster health, resource usage, and job performance
Ensure secure, efficient, and scalable operation of NVIDIA AI infrastructure, including user access and workload management
Requirements
Qualified candidates must hold an active NVIDIA Professional Certification in either AI Networking, AI Infrastructure, or AI Operations
Prior direct, hands-on professional experience administering NVIDIA GPU and data processing unit (DPU) technologies, AI software stacks, and data center environments for high-performance AI workloads
Comprehensive expertise in deploying and maintaining AI compute platforms, requiring proficiency in containerization and workload orchestration using Docker, Kubernetes, Slurm, NVIDIA Base Command Manager, and Run:Ai
Must be capable of configuring physical and logical resources, including Multi-Instance GPU (MIG) partitioning and BlueField platforms, while overseeing critical facility elements such as power, cooling, and storage solutions
The ability to demonstrate advanced skills in AI networking, specifically configuring and optimizing high-performance InfiniBand and Ethernet fabrics to ensure maximum throughput and minimal latency
Current active TS/SCI clearance with a CI Polygraph
About MARFORCYBER
Verified company details for this employer are not available yet.