Base Career helps you apply smarter for this job.
Key skills for this role
Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role.
Responsibilities
Define Fuse's inference serving strategy and architecture from first principles.
Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.
Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
Act as a direct technical owner of inference performance and reliability.
Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
Set the standards, tooling, and benchmarks this function will run on as it grows.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
London, GBR
London, GBR
London, GBR
London, GBR
London, GBR
London, GBR
London, GBR
San Francisco, USA
Experience with Triton or custom ML inference/training frameworks.
Experience with autoscaling or capacity planning for large-scale inference workloads.
Exposure to multi-tenant serving or SLA-driven infrastructure.
Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
Familiarity with Kubernetes/Slurm for cluster orchestration.
Interest or experience in energy markets, grid systems, or sustainability-focused compute.
Building a full stack energy company to lower the cost of energy and accelerate energy abundance.
Visit company websiteJobs and hiring trendsFull-time
Mid · 4+ years experience
Remote
Apply faster on company sites with our extension.