System Software Engineer - Local AI
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
The LocalAI team builds efficient on-device AI software for RTX and DGX-class systems.
The role focuses on high-performance local inference, low latency, efficient memory use, infrastructure, and deployment on resource-constrained platforms.
Key Skills for This Role
Full Job Posting
Role Overview
The LocalAI team builds efficient on-device AI software for RTX and DGX-class systems.
The role focuses on high-performance local inference, low latency, efficient memory use, infrastructure, and deployment on resource-constrained platforms.
What You'll Be Doing
- Partner with software, research, architecture, and product teams to support the AI ecosystem on RTX and DGX PCs.
- Build and optimize local AI inference stacks for RTX, RTX Pro, and DGX GPUs across hardware architectures.
- Architect and develop inference runtimes and execution stacks for LLM, vision-language, TTS, ASR, and diffusion workloads.
- Work with Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX across inference workloads.
- Optimize models, data pipelines, and inference runtimes for current and next-generation GPU architectures.
- Apply quantization, pruning, sparsity, and distillation for efficient deployment on local and edge devices.
- Debug systems, optimize performance, analyze performance-accuracy trade-offs, and develop performance and accuracy sweep infrastructure.
- Establish engineering guidelines that accelerate bring-up and production readiness of models and inference backends.
What We Need To See
- At least 2 years of experience and a Bachelor's, Master's, or PhD in a related technical field, or equivalent experience.
- Excellent C++ programming and debugging skills with strong knowledge of data structures, algorithms, and machine learning.
- Proven experience with AI inference pipelines and applications using machine learning or deep learning frameworks.
- Interest in inference backend and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
- Strong analytical, problem-solving, multitasking, written communication, and oral communication abilities.
Ways To Stand Out From The Crowd
- Understanding of machine learning, deep neural networks, and generative AI techniques, with contributions to major open-source projects.
- Experience delivering end-to-end products with geographically distributed teams in multinational product companies.
- Proficiency in lower-level system or GPU programming, CUDA, and high-performance systems.
- Contributions to open-source inference runtimes, model tooling, or performance infrastructure.
- Hands-on experience with Llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, or vLLM APIs and frameworks.
Employment
- The posting lists the time type as full time.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at NVIDIA AI
Solutions Architect - NVIDIA Cloud Partners and Datacentre Infrastructure
Dubai, UAE
NVIDIA is seeking a Solutions Architect to guide cloud partners and strategic customers through the design, deployment, and operation of AI and HPC GPU infrastructure. The role focuses on datacentre power, cooling, MEP i
Senior Sales Account Manager, Government and Smart Spaces UAE
Dubai, UAE
NVIDIA is seeking a Senior Sales Account Manager to grow government and smart-spaces business in the UAE by developing strategic relationships, pipelines, and AI transformation opportunities. The role requires 10+ years
Senior Solutions Architect, InfiniBand and Networking Ethernet - NVIS
Sydney, AUS
NVIDIA is seeking a Senior Networking Solutions Architect to deliver large-scale Ethernet and InfiniBand projects for AI and high-performance computing customers. The role combines customer engagement, network and system
Senior Staff Site Reliability Engineer
Bengaluru, IND
NVIDIA is seeking a Senior Staff Site Reliability Engineer to transform and operate globally scaled on-premises and cloud infrastructure. The role requires deep experience in compute platform engineering, automation, dis
PhD Intern, AI ML in Wireless L1/L2 - Fall 2026
Bengaluru, IND
NVIDIA is seeking a PhD intern to develop and optimize AI/ML functions for wireless Layer 1 and Layer 2 software in its Aerial RAN team. The work includes signal-processing research, model architecture selection, trainin
Senior Account Manager Telecommunications
Dubai, UAE
The employer is seeking a senior account manager to grow NVIDIA’s telecommunications business through market strategies, executive relationships, strategic partnerships, and adoption of AI and accelerated-computing solut
Enterprise Account Manager
Dubai, UAE
NVIDIA AI is seeking a Senior Enterprise Account Manager to grow revenue and market share by managing strategic customers and ecosystem partners. The role requires enterprise and software sales experience, success agains
Senior System Software Engineer
Dubai, UAE
NVIDIA is seeking a Senior System Software Engineer to improve globally distributed authentication and storage services. The role covers backend microservices, distributed systems architecture, testing, API support, and
Solutions Architect - NVIDIA Cloud Partners and Datacentre Infrastructure
Dubai, UAE
Senior Sales Account Manager, Government and Smart Spaces UAE
Dubai, UAE
Senior Solutions Architect, InfiniBand and Networking Ethernet - NVIS
Sydney, AUS
Senior Staff Site Reliability Engineer
Bengaluru, IND
PhD Intern, AI ML in Wireless L1/L2 - Fall 2026
Bengaluru, IND
Senior Account Manager Telecommunications
Dubai, UAE
Enterprise Account Manager
Dubai, UAE
Senior System Software Engineer
Dubai, UAE