Base Career helps you apply smarter for this job.
Key skills for this role
Build and execute a technical Developer Relations strategy to grow NVIDIA platform adoption across AI Labs in India.
Develop trusted relationships with founders, CTOs, ML researchers, ML infrastructure teams, platform leaders, and developer communities.
Identify and accelerate high-value workloads such as foundation model training, fine-tuning, speech AI, retrieval augmented generation, multimodal AI, inference optimization, and production model serving.
Assess which AI Lab workloads are a strong fit for GPU acceleration by profiling bottlenecks and distinguishing compute-bound problems from memory-bound, IO-bound, network-bound, or orchestration-bound issues.
Build and adapt technical demos, sample code, notebooks, benchmark plans, reference architectures, and performance guides.
Run deep technical workshops, code labs, architecture reviews, office hours, developer sessions, technical webinars, and executive briefings.
Work hands-on with developers to debug integration issues, profile workloads, improve inference performance, and identify the right NVIDIA software stack for each use case.
Explain why similar model workloads may perform differently across labs due to model architecture, data pipeline design, batch size, latency targets, storage/network behavior, software stack, or deployment environment.
Capture developer feedback, technical blockers, competitive insights, and product requirements for NVIDIA product and engineering teams.
Bachelor's degree in engineering, computer science, data science, or a related technical discipline, or equivalent experience and 8 + years of experience in technical Developer Relations, ML engineering, AI platforms, model infrastructure, cloud platforms, solution architecture, or startup technical engagement.
Strong understanding of generative AI, deep learning, machine learning, model lifecycle, inference serving, model optimization, distributed training, and MLOps.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, GBR
Reading, GBR
Bengaluru, IND
Reading, GBR
, GBR
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Hands-on programming experience in Python, with working familiarity in C++, CUDA, or other performance-oriented development environments preferred.
Experience with modern AI frameworks and deployment tooling such as PyTorch, TensorFlow, JAX, Hugging Face, model serving systems, containers, Kubernetes, APIs, CI/CD, and distributed application architecture.
Ability to profile and reason about AI workloads, including GPU suitability, compute intensity, memory bandwidth, IO bottlenecks, network constraints, latency, throughput, batching, and cost/performance tradeoffs.
Ability to engage both deeply technical developers and senior business stakeholders, with credibility in whiteboarding, live demos, technical troubleshooting, and architecture reviews.
Hands-on exposure to core NVIDIA AI software and SDKs relevant to model builders, including NeMo Framework, NeMo Curator, Nemotron model families, TensorRT-LLM, NVIDIA NIM, Triton Inference Server, CUDA, NCCL, Nsight Systems, Nsight Compute, NGC containers, and NVIDIA AI Enterprise.
Ability to demonstrate how NeMo Framework supports model training, customization, alignment, evaluation, and deployment workflows for LLMs, multimodal models, and speech AI.
Ability to reason about when NeMo Curator or related data curation workflows can improve training data quality through filtering, deduplication, formatting, synthetic data workflows, and multimodal data preparation.
Ability to build or adapt demos that show how Nemotron models, NIM microservices, TensorRT-LLM, and Triton can move a model from experimentation to optimized inference.
Hands-on experience with large-scale model training, fine-tuning, inference optimization, or AI platform engineering.
Hands-on experience with inference optimization and model efficiency techniques such as quantization, distillation, pruning, speculative decoding, batching, KV cache optimization, RLHF, RLAIF, DPO, or other post-training and alignment methods.
Experience creating workload qualification frameworks, benchmark plans, or technical decision guides that help developers decide when GPU acceleration is the right fit.
Track record as a technical DevRel practitioner, including public talks, workshops, GitHub samples, blogs, tutorials, reference architectures, or developer community programs.
Strong point of view on India's sovereign AI, Indic language AI, and foundation model ecosystem.
Computing platform company for AI and accelerated graphics.
Visit company websiteJobs and hiring trendsFull-time
Senior · 8+ years experience
Onsite
Apply faster on company sites with our extension.