{bc}
indeed

Staff Engineer - AI/Inference /Gateway

Nutanix
Vancouver, CAN
Full-time
Hybrid
$171,000
Discovered Yesterday
Distributed systemsKubernetesDockerCloud-native architecturesGoPython
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Distributed systemsKubernetesDocker
Smart Apply

Full Job Posting

Opportunity

Nutanix is building an enterprise AI platform for Generative AI, large language models, and Agentic AI applications.

The Enterprise AI team develops foundational technologies including LLM inference, the AI Gateway, and the Agentic AI Platform.

The role works at the intersection of large-scale distributed systems and machine learning infrastructure.

Work Arrangement

  • The team operates in a hybrid model combining office collaboration and remote work flexibility.
  • New hires are expected to work in the office three days per week.

Role Responsibilities

  • Architect and develop scalable, containerized, fault-tolerant Kubernetes services for enterprise AI workloads.
  • Build high-performance inference and platform services for Generative AI and Agentic AI applications.
  • Optimize distributed systems, storage, networking, and low-level infrastructure components.
  • Develop multi-tenant services for on-premises, hybrid, and cloud AI deployments.
  • Implement observability architectures using Prometheus, Grafana, Datadog, OpenTelemetry, or similar tools.
  • Diagnose production issues and improve reliability, resiliency, and operational efficiency.
  • Maintain CI/CD pipelines and deployment automation.
  • Implement LLM serving features such as routing, rate limiting, token streaming, load balancing, quotas, and usage budgets.
  • Collaborate with globally distributed product, AI, and software engineering teams.
  • Participate in the full product lifecycle, from architecture and development through deployment and operations.
  • Review code and designs and champion engineering excellence.
  • Evaluate emerging technologies and influence the Enterprise AI Platform direction.

Required Qualifications

  • At least 8 years of experience developing resilient software products in a product development organization.
  • Strong fundamentals in data structures, algorithms, operating systems, networking, and distributed systems.
  • Hands-on experience with Docker, Kubernetes, and cloud-native architectures.
  • Production backend development experience using Go, Python, C++, or Rust.
  • Experience owning CI/CD pipelines and release automation end to end.
  • Strong datacenter architecture knowledge covering compute, storage, networking, and virtualization.
  • Experience deploying software across on-premises, cloud, and hybrid environments.
  • Experience designing and tuning high-performance system software.
  • Experience resolving production performance issues with observability and monitoring platforms.
  • Familiarity with LLM serving and modern LLM concepts.
  • Master's degree in Computer Science or equivalent practical experience.

Preferred Experience

  • Experience with PyTorch or TensorFlow, GPU systems, or model-serving platforms such as vLLM, DeepSpeed, Hugging Face TGI, or Triton.
  • Experience with RAG, vector databases, AI orchestration frameworks, open-source contributions, production AI platforms, LLM APIs, or agentic systems.
  • Experience scaling production LLM APIs with streaming, prompt guardrails, rate limiting, and usage budgeting.

Compensation

  • The expected base pay range is CAD $171,000 to CAD $257,000 annually at commencement of employment.
  • Total compensation may also include a sign-on bonus, restricted stock units, discretionary awards, and benefits.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Nutanix