Staff Engineer - AI/Inference /Gateway
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Nutanix is building an enterprise AI platform for Generative AI, large language models, and Agentic AI applications.
The Enterprise AI team develops foundational technologies including LLM inference, the AI Gateway, and the Agentic AI Platform.
The role works at the intersection of large-scale distributed systems and machine learning infrastructure.
Key Skills for This Role
Full Job Posting
Opportunity
Nutanix is building an enterprise AI platform for Generative AI, large language models, and Agentic AI applications.
The Enterprise AI team develops foundational technologies including LLM inference, the AI Gateway, and the Agentic AI Platform.
The role works at the intersection of large-scale distributed systems and machine learning infrastructure.
Work Arrangement
- The team operates in a hybrid model combining office collaboration and remote work flexibility.
- New hires are expected to work in the office three days per week.
Role Responsibilities
- Architect and develop scalable, containerized, fault-tolerant Kubernetes services for enterprise AI workloads.
- Build high-performance inference and platform services for Generative AI and Agentic AI applications.
- Optimize distributed systems, storage, networking, and low-level infrastructure components.
- Develop multi-tenant services for on-premises, hybrid, and cloud AI deployments.
- Implement observability architectures using Prometheus, Grafana, Datadog, OpenTelemetry, or similar tools.
- Diagnose production issues and improve reliability, resiliency, and operational efficiency.
- Maintain CI/CD pipelines and deployment automation.
- Implement LLM serving features such as routing, rate limiting, token streaming, load balancing, quotas, and usage budgets.
- Collaborate with globally distributed product, AI, and software engineering teams.
- Participate in the full product lifecycle, from architecture and development through deployment and operations.
- Review code and designs and champion engineering excellence.
- Evaluate emerging technologies and influence the Enterprise AI Platform direction.
Required Qualifications
- At least 8 years of experience developing resilient software products in a product development organization.
- Strong fundamentals in data structures, algorithms, operating systems, networking, and distributed systems.
- Hands-on experience with Docker, Kubernetes, and cloud-native architectures.
- Production backend development experience using Go, Python, C++, or Rust.
- Experience owning CI/CD pipelines and release automation end to end.
- Strong datacenter architecture knowledge covering compute, storage, networking, and virtualization.
- Experience deploying software across on-premises, cloud, and hybrid environments.
- Experience designing and tuning high-performance system software.
- Experience resolving production performance issues with observability and monitoring platforms.
- Familiarity with LLM serving and modern LLM concepts.
- Master's degree in Computer Science or equivalent practical experience.
Preferred Experience
- Experience with PyTorch or TensorFlow, GPU systems, or model-serving platforms such as vLLM, DeepSpeed, Hugging Face TGI, or Triton.
- Experience with RAG, vector databases, AI orchestration frameworks, open-source contributions, production AI platforms, LLM APIs, or agentic systems.
- Experience scaling production LLM APIs with streaming, prompt guardrails, rate limiting, and usage budgeting.
Compensation
- The expected base pay range is CAD $171,000 to CAD $257,000 annually at commencement of employment.
- Total compensation may also include a sign-on bonus, restricted stock units, discretionary awards, and benefits.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Nutanix
Member of Technical Staff 2 - AI
Vancouver, CAN
Hungry, Humble, Honest, with Heart. The Opportunity When people talk about generative AI and other ML-powered solutions in today's conversation, they often refer to generative pre-trained transformers like ChatGPT that c
Manager, Business Intelligence & Systems Operations
San Jose, USA
Physical Security and Safety Manager - APAC
Bengaluru, IND
Member of Technical Staff 2 - AI
Vancouver, CAN
Physical Security and Safety Manager - APAC
Bengaluru, IND
Manager, Business Intelligence & Systems Operations
San Jose, USA
Advisory Enterprise Account Manager
, USA
Global Account Manager
, USA
Global Account Manager
Bentonville, USA
Resource Coordinator
San Jose, USA