Base Career helps you apply smarter for this job.
Key skills for this role
Nscale is hiring a Staff Software Engineer to build Fleet Manager — the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale.
This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You'll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation. The problems are challenging and the stakes are high — the software you design and build determines how quickly and how reliably Nscale scales its GPU fleet to meet demand.
This is an opportunity to shape a foundational platform early, setting the patterns and standards that engineers across Fleet Manager build on.
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
Nscale is hiring a Staff Software Engineer to build Fleet Manager — the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale.
This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You'll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation. The problems are challenging and the stakes are high — the software you design and build determines how quickly and how reliably Nscale scales its GPU fleet to meet demand.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
New York City, USA
, USA
, USA
New York City, USA
New York City, USA
New York City, USA
New York City, USA
New York City, USA
This is an opportunity to shape a foundational platform early, setting the patterns and standards that engineers across Fleet Manager build on.
Device provisioning and enrollment: automation that takes bare-metal GPU nodes and network switches from first power-on to production-ready — BMC configuration, DHCP reservations, and provisioning state machines.
Burn-in and validation: multi-day testing workflows that qualify hardware before it enters, and re-enters, the fleet.
Workflow orchestration: durable, event-driven state machines that span multiple days, survive crashes, resume from checkpoints, support human-in-the-loop approval gates, and let thousands of concurrent idempotent workflows run without stepping on each other.
GPU health monitoring and self-healing: detection, diagnosis, and automated remediation workflows that keep nodes healthy in production.
Network configuration: switch lifecycle automation and network state management across the fleet.
Integrations: keeping Fleet Manager consistent with datacenter inventory tooling (DCIM, NetBox), bare-metal provisioning systems (MAAS, Ironic, IPMI), credential stores, and monitoring infrastructure.
Observability: structured logging, metrics, distributed tracing, and tooling that lets operators troubleshoot effectively.
Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar
Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems
Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation
Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows
GPU infrastructure experience: health monitoring, burn-in testing, or cluster management
HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE)
Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP
Open-source contributions in infrastructure automation or cloud-native tooling
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Builds scalable AI cloud infrastructure, offering GPU compute and platform services for machine learning workloads.
Visit company websiteJobs and hiring trendsSenior
Apply faster on company sites with our extension.