Staff Engineer - Inference /AI
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Key Skills for This Role
Full Job Posting
Your Role
- Architect, design, and develop horizontally scalable, containerized, fault-tolerant services on Kubernetes for enterprise AI and LLM workloads.
- Build and operate high-performance inference and platform services that deliver low-latency, high-throughput experiences for Generative AI and Agentic AI applications.
- Design and optimize critical system components across the stack, including distributed systems, storage, networking, and low-level infrastructure layers.
- Develop and enhance multi-tenant platform services supporting on-premises, hybrid, and cloud-based AI deployments.
- Design and implement scalable observability architectures using technologies such as Prometheus, Grafana, Datadog, OpenTelemetry, and related cloud-native monitoring frameworks.
- Debug complex production issues, perform root-cause analysis, and improve reliability, resiliency, and operational efficiency of platform services.
- Build and maintain CI/CD pipelines and deployment automation to accelerate delivery of production-grade services.
- Design and implement foundational LLM serving capabilities including request routing, rate limiting, token streaming, load balancing, quota management, and usage budgeting.
- Collaborate closely with globally distributed product management, AI, and software engineering teams to deliver high-quality products in a fast-paced environment.
- Contribute to all stages of the product lifecycle, including architecture, design, development, testing, experimentation, performance analysis, deployment, and operations.
- Leverage and contribute to relevant open-source cloud-native and AI ecosystem projects.
- Review code and design documents, provide feedback on product requirements, and champion engineering excellence across the team.
- Continuously evaluate emerging technologies and help shape the technical direction of Nutanix's Enterprise AI Platform.
- What You Will Bring
- Required Qualifications
- 8+ years of experience developing maintainable, modular, resilient, fail-safe, and long-lived software products within a product development organization.
- Strong computer science fundamentals including data structures, algorithms, operating systems, networking, and distributed systems.
- Hands-on experience with Docker, Kubernetes, and cloud-native architectures.
- Production experience developing backend systems using Go, Python, C++, or Rust.
- Experience building, owning, and maintaining CI/CD pipelines and release automation end-to-end.
- Strong understanding of datacenter architecture including compute, storage, networking, and virtualization.
- Experience designing and deploying software across on-premises, cloud, and hybrid environments.
- Demonstrated experience designing and tuning high-performance, performance-sensitive system software.
- Solid understanding of distributed computing, distributed data stores, and large-scale service architectures.
- Experience diagnosing and resolving production performance issues using observability and monitoring platforms such as Prometheus, Grafana, Datadog, Open Telemetry, or similar tools.
- Familiarity with LLM serving concepts including rate limiting, token streaming, request scheduling, load balancing, quota management, and usage budgeting.
- Familiarity with modern LLM concepts including reasoning workflows, tool calling, prompt templates, and agent.
- Experience building multi-tenant services running on virtualized or containerized infrastructure.
- Strong communication, collaboration, and problem-solving skills with the ability to work effectively across globally distributed teams.
- Master's degree in Computer Science or equivalent practical experience.
- Bonus Points If You Have Experience With
- Machine learning frameworks such as PyTorch or TensorFlow.
- GPU-based systems and acceleration technologies.
- Modern model-serving platforms such as vLLM, DeepSpeed, Hugging Face TGI, or Triton.
- Retrieval-Augmented Generation (RAG), vector databases, and AI orchestration frameworks.
- Open-source contributions or experience working in large distributed codebases.
- Production AI platforms, LLM APIs, agentic systems, or inference infrastructure.
- Building or scaling production LLM APIs, including streaming via SSE/WebSockets, prompt guardrails, rate limiting, and usage budgeting.
- Learn More About the Technology: https://www.nutanixbible.com/ [nutanixbible.com]
- #NAI
- Highlighted Benefits (Vancouver, Canada)
- Retirement: RRSP with dollar-for-dollar matching up to 7% of base salary
- Mental Health: Dedicated mental health coverage plus top-tier paramedical benefits
- Family: Fully paid maternity and parental leave and generous bereavement leave, including time for the loss of a pet
- Equity: RSUs and Employee Stock Purchase Plan at a 15% discount
- Work Arrangement Hybrid: This role operates in a hybrid capacity, blending the benefits of remote work with the advantages of in-person collaboration. In locations where our workplace policy applies (i.e. San Jose, Durham, Mexico City, Vancouver, Bangalore, Pune, Hoofddorp, Belgrade, Barcelona, Singapore, Sydney and Tokyo), employees are expected to work onsite a minimum of 3 days per week to foster collaboration, team alignment, and access to in-office resources. Workplace type may vary based on location and team requirements. Please speak with your recruiter for details. Additional team-specific guidance and norms will be provided by your manager.
- Pay Transparency - Role Location The pay range for this position at commencement of employment is expected to be between CAD $171,000 and CAD $257,000 per annual.
- However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. The total compensation package for this position may also include other elements, including a sign-on bonus, restricted stock units, and discretionary awards in addition to a full range of medical, financial and/or other benefits (including 401(k) eligibility and various paid time off benefits, such as vacation, sick time, and parental leave), dependent on the position offered. Details of participation in these benefit plans will be provided if an employee receives an offer of employment.
- If hired, employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors. Our application deadline is 40 days from the date of posting. In good faith, the posting may be removed prior to this date if the position is filled or extended in good faith.
About Nutanix
Nutanix develops cloud software and a unified platform for running applications, data, and AI across data centers, edge locations, and public and private clouds. Its hybrid multicloud technology helps organizations simplify infrastructure management and modernize IT operations.
Visit company websiteJobs and hiring trendsApply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Nutanix
Member of Technical Staff 2 - AI
Vancouver, CAN
Hungry, Humble, Honest, with Heart. The Opportunity When people talk about generative AI and other ML-powered solutions in today's conversation, they often refer to generative pre-trained transformers like ChatGPT that c
Manager, Business Intelligence & Systems Operations
San Jose, USA
Physical Security and Safety Manager - APAC
Bengaluru, IND
Member of Technical Staff 2 - AI
Vancouver, CAN
Physical Security and Safety Manager - APAC
Bengaluru, IND
Manager, Business Intelligence & Systems Operations
San Jose, USA
Advisory Enterprise Account Manager
, USA
Global Account Manager
, USA
Global Account Manager
Bentonville, USA
Resource Coordinator
San Jose, USA
