Technical Lead - GPU Infrastructure (100% Remote - Worldwide)
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Cosmic AC is a GPU compute and managed inference platform providing GPU containers, managed inference endpoints, and platform observability on Kubernetes.
The platform is expanding to bare-metal GPU infrastructure with a managed Slurm layer for research and model-training teams and a Kubernetes control plane for inference tenancy.
The Technical Lead owns architecture, implementation, delivery plans, and engineering leadership for a distributed team of about twelve engineers.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months.
Key Skills for This Role
Full Job Posting
About the Job
Cosmic AC is a GPU compute and managed inference platform providing GPU containers, managed inference endpoints, and platform observability on Kubernetes.
The platform is expanding to bare-metal GPU infrastructure with a managed Slurm layer for research and model-training teams and a Kubernetes control plane for inference tenancy.
The Technical Lead owns architecture, implementation, delivery plans, and engineering leadership for a distributed team of about twelve engineers.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months.
Responsibilities
- Own end-to-end platform architecture through proposals, high-level and low-level designs, reviews, and maintenance of the architecture baseline.
- Lead and line-manage backend, frontend, DevOps, QA, and documentation engineers across Europe and India.
- Design, build, and operate a managed Slurm service covering controllers, accounting, partitions, login nodes, onboarding, upgrades, health detection, storage, identity, and isolation.
- Own Kubernetes bootstrap and lifecycle on partner bare metal, GPU operators, VM-based GPU isolation, upgrades, backup, recovery, and node replacement.
- Own managed inference serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
- Develop metrics, logging, alerting, SLOs, incident response, post-incident reviews, and a sustainable on-call model.
- Interface with infrastructure partners and vendors on specifications, acceptance tests, escalations, capacity planning, and hardware sourcing.
- Translate research, model-training, and product workloads into platform requirements and broker capacity when necessary.
- Complete the platform team and establish technical standards for new engineers.
Must Have
- Eight or more years of hands-on engineering, including at least three years leading infrastructure platform teams.
- Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
- Hands-on Slurm operations for real users, including scheduling, accounting, node health, and upgrades.
- Bare-metal NVIDIA GPU fleet operations covering drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG, and node acceptance.
- InfiniBand, RDMA, SR-IOV, and multi-node NCCL troubleshooting experience.
- Deep Linux systems experience with drivers, passthrough, vfio-pci, cgroups, namespaces, and performance tuning.
- Production Kubernetes operations covering control planes, upgrades, networking, storage, operators, controllers, and multi-tenancy.
- HPC storage and data movement experience with shared filesystems, local NVMe caching, and distributed models and datasets.
- Observability and operations experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident reviews.
- Working fluency in JavaScript and Node.js for reviewing control-plane, CLI, and worker services.
- A shipped multi-tenant IaaS, PaaS, or research computing service with isolation, quotas, metering, and user-facing APIs or CLIs.
- Excellent written and spoken English, remote availability between UTC and UTC+5:30, and occasional travel.
Desirable
- Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers such as Kueue, Volcano, KAI, or Kubeflow Trainer.
- Experience with modern serving stacks such as vLLM, SGLang, or TensorRT-LLM.
- Experience with GPU workload isolation using KubeVirt, Kata Containers, QEMU, KVM, Firecracker, or confidential-computing technologies.
- Experience with Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, or GitOps.
- Experience in a GPU cloud, national or university HPC centre, AI lab platform team, peer-to-peer systems, or distributed systems.
Workplace
- The role is fully remote and based between UTC and UTC+5:30 so the working day overlaps with Europe and India.
- Occasional travel to partner sites and team events is expected.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Tether Operations Limited
Technical Lead - GPU Infrastructure
, IND
USAT Institutional Associate
, USA
USAT Institutional Associate
, USA
USAT Institutional Associate
, USA
Technical Lead - GPU Infrastructure
, GBR
Backend Engineer - Wallets (100% Remote)
, GBR
Backend Engineer - Wallets (100% Remote)
, UAE
Backend Engineer - Wallets (100% Remote)
, UAE