Base Career helps you apply smarter for this job.
Key skills for this role
We are seeking an exceptional Machine Learning Engineer who has made training and AI workload scheduling a specialty. This is a senior-level role for someone who has significant experience managing distributed machine learning workloads at scale using Slurm and/or Kubernetes.
As a technical visionary and hands-on expert, you will lead the evolution of our managed Slurm and Kubernetes offerings, as well as internal health checking and cluster automation.
At TensorWave, we’re leading the charge in AI compute, building a versatile cloud platform that’s driving the next generation of AI innovation. We’re focused on creating a foundation that empowers cutting-edge advancements in intelligent computing, pushing the boundaries of what’s possible in the AI landscape.
We are seeking an exceptional Machine Learning Engineer who has made training and AI workload scheduling a specialty. This is a senior-level role for someone who has significant experience managing distributed machine learning workloads at scale using Slurm and/or Kubernetes.
As a technical visionary and hands-on expert, you will lead the evolution of our managed Slurm and Kubernetes offerings, as well as internal health checking and cluster automation.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Las Vegas, USA
, USA
Las Vegas, USA
, USA
Las Vegas, USA
Tucson, USA
, USA
, USA
Las Vegas, USA
3+ years of hands-on Kubernetes experience , including deep knowledge of the Kubernetes API, internals, networking, and storage.
Proficiency in writing Kubernetes manifests, Helm charts, and managing releases.
Experience with DAGs using K8s native tools such as Argo Workflows.
Foundation in networking, especially as it pertains to RDMA, RoCE, and Infiniband.
Experience with low level kernel libraries, such as CUDA and Composable Kernel.
Contributions to open-source projects or ML/AI tooling.
A production-grade integrated Slurm platform that can support thousands of GPUs , with self-healing, scaling, and strong observability.
Infrastructure is resilient, secure, resource-optimized, and compliant.
Best practices and tooling are well-documented, standardized, and continuously improved across the company.
Make GPUs go Brrrrrrr
Stock Options
100% paid Medical, Dental, and Vision insurance
Life and Voluntary Supplemental Insurance
Short Term Disability Insurance
Flexible Spending Account
401(k)
Flexible PTO
Paid Holidays
Parental Leave
Mental Health Benefits through Spring Health
AMD-exclusive AI cloud provider serving teams with bare-metal GPU infrastructure for training, fine-tuning, and inference.
Visit company websiteJobs and hiring trendsFull-time
Senior · 5+ years experience
Onsite
Apply faster on company sites with our extension.