Base Career helps you apply smarter for this job.
Key skills for this role
Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features and serving runtime algorithms.
Profile and optimize the inference framework (vLLM) with methods like speculative decoding, 5D Parallelism, and prefill-decode disaggregation.
Architect novel frameworks and runtime optimizations for inference infrastructure, benchmarking, and kernels.
Conduct and publish original research that advances the Pareto frontier in ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into production-grade, open-source software.
Develop, optimize, and benchmark GPU kernels (both hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization.
Bachelor’s, Master’s, or PhD degree in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE).
5+ years of industry experience in software engineering or equivalent research experience.
Strong programming skills in Python and one of C/C++, Go, or Rust. Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, software engineering, distributed systems, deep learning theories.
Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).
Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).
Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
, USA
, USA
Experience developing major features and optimizations for LLM inference engines (e.g., vLLM, SGLang).
Hands-on work with LLM inference and training runtimes (deploying LLMs to production, large-scale LLM pre-training and RL), ML compilers and DSLs (e.g., Triton, CuTe, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores).
Experience with speculative decoding training and runtime features: tree-structured drafting, parallel drafting, diffusion LLMs, DFlash, EAGLE.
Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.
At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work. Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential.
Our team consists of experts in AI, systems and performance optimization. Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. If you’re excited to build systems, kernels, and tools that make large-scale AI faster, more efficient, and easier to deploy, we’d love to hear from you.
#LI-Hybrid
You will also be eligible for equity and benefits .
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
Computing platform company for AI and accelerated graphics.
Visit company websiteJobs and hiring trendsCAD 135000-220000 yearly / year
Full-time
Senior · 5+ years experience
Onsite
Apply faster on company sites with our extension.