Base Career helps you apply smarter for this job.
Key skills for this role
We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance. Responsibilities
Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
Build and support production AI infrastructure running hundreds to thousands of GPUs.
Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
Configure and tune distributed AI software stacks including: PyTorch NCCL CUDA UCX MPI Slurm Pyxis/Enroot
PyTorch
NCCL
CUDA
UCX
MPI
Slurm
Pyxis/Enroot
Optimize GPU scheduling and resource allocation for both training and inference environments.
Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
Identify performance regressions and troubleshoot distributed training issues at scale.
Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
Work closely with ML engineers to improve training scalability and inference efficiency.
Create automation to deploy, validate, benchmark, and monitor GPU clusters.
Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture. Required Qualifications
7+ years designing or operating large-scale Linux infrastructure.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
5+ years supporting production GPU clusters for AI or HPC workloads.
Demonstrated experience building multi-node GPU training environments from the ground up.
Deep expertise with distributed PyTorch training.
Extensive experience troubleshooting and optimizing NCCL communications.
Strong understanding of distributed AI communication patterns, including: AllReduce ReduceScatter AllGather Broadcast Point-to-point communications
AllReduce
ReduceScatter
AllGather
Broadcast
Point-to-point communications
Experience benchmarking distributed training using tools such as: nccl-tests NVIDIA DCGM Nsight Systems MLPerf (preferred)
nccl-tests
NVIDIA DCGM
Nsight Systems
MLPerf (preferred)
Strong understanding of GPU memory management, including: KV Cache Activation checkpointing Tensor Parallelism Pipeline Parallelism Data Parallelism
KV Cache
Activation checkpointing
Tensor Parallelism
Pipeline Parallelism
Data Parallelism
Experience optimizing LLM inference throughput, including: Tokens/sec optimization Batch sizing Continuous batching KV cache tuning Memory bandwidth optimization
Tokens/sec optimization
Batch sizing
Continuous batching
KV cache tuning
Memory bandwidth optimization
Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
Expert-level Linux systems administration skills.
Experience with Slurm workload manager.
Experience using Pyxis and Enroot for containerized GPU workloads.
Strong scripting skills using Python and Bash.
Technical Expertise AI Frameworks
Triton (preferred)
TensorRT-LLM (preferred)
We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance. Responsibilities
Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
Build and support production AI infrastructure running hundreds to thousands of GPUs.
Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
Configure and tune distributed AI software stacks including: PyTorch NCCL CUDA UCX MPI Slurm Pyxis/Enroot
PyTorch
NCCL
CUDA
UCX
MPI
Slurm
Pyxis/Enroot
Optimize GPU scheduling and resource allocation for both training and inference environments.
Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
Identify performance regressions and troubleshoot distributed training issues at scale.
Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
Work closely with ML engineers to improve training scalability and inference efficiency.
Create automation to deploy, validate, benchmark, and monitor GPU clusters.
Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture. Required Qualifications
7+ years designing or operating large-scale Linux infrastructure.
5+ years supporting production GPU clusters for AI or HPC workloads.
Demonstrated experience building multi-node GPU training environments from the ground up.
Deep expertise with distributed PyTorch training.
Extensive experience troubleshooting and optimizing NCCL communications.
Strong understanding of distributed AI communication patterns, including: AllReduce ReduceScatter AllGather Broadcast Point-to-point communications
AllReduce
ReduceScatter
AllGather
Broadcast
Point-to-point communications
Experience benchmarking distributed training using tools such as: nccl-tests NVIDIA DCGM Nsight Systems MLPerf (preferred)
nccl-tests
NVIDIA DCGM
Nsight Systems
MLPerf (preferred)
Strong understanding of GPU memory management, including: KV Cache Activation checkpointing Tensor Parallelism Pipeline Parallelism Data Parallelism
KV Cache
Activation checkpointing
Tensor Parallelism
Pipeline Parallelism
Data Parallelism
Experience optimizing LLM inference throughput, including: Tokens/sec optimization Batch sizing Continuous batching KV cache tuning Memory bandwidth optimization
Tokens/sec optimization
Batch sizing
Continuous batching
KV cache tuning
Memory bandwidth optimization
Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
Expert-level Linux systems administration skills.
Experience with Slurm workload manager.
Experience using Pyxis and Enroot for containerized GPU workloads.
Strong scripting skills using Python and Bash.
Technical Expertise AI Frameworks
Triton (preferred)
TensorRT-LLM (preferred)
Slurm
Pyxis
Enroot
GPU Networking Strong understanding of:
InfiniBand
RoCE v2
RDMA
GPUDirect RDMA
GPUDirect Storage
UCX
MPI
Network topology optimization
Congestion control
QoS
ECN/PFC
High-speed Ethernet (200/400/800 GbE)
Storage Experience designing or tuning storage for AI workloads, including:
Parallel file systems
Distributed storage
Object storage
NVMe
Checkpoint optimization
Dataset staging
Storage bandwidth optimization
Metadata performance
Performance Engineering Experience with:
NCCL benchmarking
Multi-node scaling analysis
GPU utilization optimization
Communication/computation overlap
NUMA optimization
CPU affinity
PCIe topology
GPU topology (NVLink/NVSwitch)
Memory bandwidth analysis
End-to-end performance profiling
Preferred Qualifications
Experience deploying AI workloads on Kubernetes.
Experience with NVIDIA GPU Operator.
Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).
Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.
U.S.-based IT infrastructure provider delivering managed cloud, cybersecurity, and GPU compute services to enterprises and AI teams.
Visit company websiteJobs and hiring trendsFull-time
Senior · 7+ years experience
Remote
Apply faster on company sites with our extension.
Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.
Familiarity with MLPerf benchmarking.
Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.
Experience automating infrastructure using Ansible, Terraform, or similar tools.
Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.