{bc}
linkedin

System Software Engineer - Local AI

NVIDIA AI
Pune, IND
Full-time
Entry
Onsite
Discovered 1 weeks ago
C++Data structures and algorithmsMachine learningAI inference pipelinesMachine learning and deep learning frameworksInference runtimes
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

C++Data structures and algorithmsMachine learning
Smart Apply

Full Job Posting

Role Overview

The LocalAI team builds efficient on-device AI software for RTX and DGX-class systems.

The role focuses on high-performance local inference, low latency, efficient memory use, infrastructure, and deployment on resource-constrained platforms.

What You'll Be Doing

  • Partner with software, research, architecture, and product teams to support the AI ecosystem on RTX and DGX PCs.
  • Build and optimize local AI inference stacks for RTX, RTX Pro, and DGX GPUs across hardware architectures.
  • Architect and develop inference runtimes and execution stacks for LLM, vision-language, TTS, ASR, and diffusion workloads.
  • Work with Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX across inference workloads.
  • Optimize models, data pipelines, and inference runtimes for current and next-generation GPU architectures.
  • Apply quantization, pruning, sparsity, and distillation for efficient deployment on local and edge devices.
  • Debug systems, optimize performance, analyze performance-accuracy trade-offs, and develop performance and accuracy sweep infrastructure.
  • Establish engineering guidelines that accelerate bring-up and production readiness of models and inference backends.

What We Need To See

  • At least 2 years of experience and a Bachelor's, Master's, or PhD in a related technical field, or equivalent experience.
  • Excellent C++ programming and debugging skills with strong knowledge of data structures, algorithms, and machine learning.
  • Proven experience with AI inference pipelines and applications using machine learning or deep learning frameworks.
  • Interest in inference backend and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
  • Strong analytical, problem-solving, multitasking, written communication, and oral communication abilities.

Ways To Stand Out From The Crowd

  • Understanding of machine learning, deep neural networks, and generative AI techniques, with contributions to major open-source projects.
  • Experience delivering end-to-end products with geographically distributed teams in multinational product companies.
  • Proficiency in lower-level system or GPU programming, CUDA, and high-performance systems.
  • Contributions to open-source inference runtimes, model tooling, or performance infrastructure.
  • Hands-on experience with Llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, or vLLM APIs and frameworks.

Employment

  • The posting lists the time type as full time.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at NVIDIA AI