Base Career helps you apply smarter for this job.
Key skills for this role
Lead and grow a team building the on-device AI inference platform for RTX, RTX Pro, and DGX GPUs, with accountability for execution, technical direction, delivery quality, and roadmap alignment.
Drive cross-functional alignment with NVIDIA’s software, research, architecture, and product teams, along with industry partners and open-source communities, to build strategy and strengthen the AI ecosystem across RTX and DGX platforms.
Provide technical leadership for the architecture and evolution of modern inference runtimes and execution stacks across frameworks such as Llama.cpp, vLLM , PyTorch , WinML , DXCGC, and TensorRT-RTX, spanning workloads including LLMs, vision-language models, TTS, ASR, and diffusion models.
Mentor engineers, develop technical leaders, and foster a high-perfor man ce team culture centred on innovation, collaboration, and operational excellence.
Coordinate end-to-end optimization of AI models, data pipelines, and inference runtimes to improve perfor man ce across current and next-generation GPU architectures.
Drive adoption of model optimization techniques such as quantization, pruning, sparsity, and distillation to enable efficient deployment of large models on local and edge devices.
Establish team processes for system-level debugging, perfor man ce optimization, and perfor man ce-accuracy trade-off analysis, including infrastructure for perfor man ce and accuracy sweeps, gap analysis, and production-readiness improvements.
5+ overall years of industry experience and 2+ years of engineering leadership experience, combined with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or a related field.
Proven experience leading high-performing engineering teams in systems software, AI infrastructure, inference runtimes, or related domains.
Strong technical foundation in C++ software development, debugging, data structures, algorithms, and machine learning systems.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
, USA
, USA
Extensive background in AI inference pipelines and Deep Learning frameworks like Llama.cpp, vLLM , PyTorch , WinML , DXCGC, and TensorRT.
Deep understanding of inference backends and runtime internals, including scheduling, memory man a g e ment, KV-cache behavior , graph execution, quantization, and hardware-aware optimization.
Strong analytical and problem-solving skills, with the ability to balance technical depth, execution speed, and organizational priorities in a fast-paced environment.
Excellent written and verbal communication skills, with proven ability to collaborate across engineering, product, research, and executive collaborators.
Strong understanding of modern machine learning, deep neural networks, and generative AI, along with contributions to notable open-source projects.
Demonstrated success building teams, setting technical vision, and scaling execution through periods of rapid growth.
Track record of delivering end-to-end products with geographically distributed teams in multinational product organizations.
Experience in lower-level systems or GPU programming, including CUDA and high-perfor man ce systems development.
Contributions to open-source inference runtimes, model tooling, or perfor man ce infrastructure as well as p ractical experience working with frameworks and APIs including Llama.cpp, PyTorch , TensorRT, Vulkan, DirectX, and vLLM .
We're a top employer known for innovation and growth. We are an equal-opportunity employer and value diversity at our company. With competitive salaries and a generous benefits package, we are widely considered to be one of the world’s most desirable employers of technology. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our best-in-class engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we would like to hear from you.
Computing platform company for AI and accelerated graphics.
Visit company websiteJobs and hiring trendsFull-time
Senior · 5+ years experience
Onsite
Apply faster on company sites with our extension.