{bc}
linkedin

Technical Lead - GPU Infrastructure

Jobgether
Full-time
Mid-Senior
Onsite
Discovered Today
GPU infrastructureSlurmKubernetesBare-metal infrastructureNVIDIA GPUsCUDA
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

GPU infrastructureSlurmKubernetes
Smart Apply

Full Job Posting

Position Overview

Technical leadership role responsible for architecting and delivering a large-scale GPU infrastructure platform.

The platform supports research, model training, and managed inference workloads requiring reliable, scalable, and observable GPU compute.

The role is fully remote and based in Saudi Arabia.

The role combines systems expertise, engineering leadership, team management, and direct ownership of architecture and delivery.

Accountabilities

  • Lead the evolution from managed Kubernetes workloads toward bare-metal GPU infrastructure, including Slurm-based research computing and Kubernetes-powered inference.
  • Oversee a distributed team spanning backend, frontend, DevOps, QA, and documentation.
  • Own platform architecture, engineering delivery, infrastructure operations, technical decisions, and partner relationships.
  • Design and operate managed Slurm services and bare-metal NVIDIA GPU infrastructure.
  • Own Kubernetes lifecycle, GPU isolation, managed inference architecture, observability, incident response, and on-call practices.
  • Work with research, model-training, and product teams to translate workloads into platform requirements and manage capacity constraints.
  • Hire and develop platform team members while maintaining a high technical bar.
  • Contribute to decisions involving distributed systems, high-performance computing, networking, storage, virtualization, and GPU workloads.

Requirements

  • At least 8 years of hands-on engineering experience, including at least 3 years leading infrastructure platform teams.
  • Bachelor's or Master's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • Extensive production experience operating Slurm, including controllers, accounting, partitions, quality of service, node health, and upgrades.
  • Experience operating HPC or GPU training clusters for research or model-development users.
  • Strong experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycles, Fabric Manager, NVSwitch, DCGM, MIG, burn-in, and acceptance.
  • Ability to operate complex production systems, make architecture decisions, lead distributed teams, and remain close to code and infrastructure.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Jobgether