Base Career helps you apply smarter for this job.
Key skills for this role
We are looking for an ML Systems Engineer to help build and optimize the core serving infrastructure behind Agent Cloud.
This role focuses on high-performance inference across different accelerators.
You will work on model serving performance, accelerator utilization, long-context inference, batching, scheduling, KV cache management, runtime efficiency, and cost reduction.
This is a deeply technical role at the intersection of ML systems, infrastructure, and product.
Direct TPU experience is a strong plus, but not required.
We care most about strong ML systems fundamentals, performance intuition, and the ability to ship reliable systems quickly.
Location: Onsite in Palo Alto
Compensation: Competitive Salary + Equity
Model AI is building the infrastructure and application stack for the next generation of agentic AI systems .
We believe token usage will grow exponentially over the coming years, but routing all inference through closed model providers will remain too expensive for many users and enterprises. Our thesis is that agentic applications require a vertically integrated stack: high-throughput, cost-efficient serving infrastructure paired with an application layer designed for long-running, agentic workloads.
Model AI is building the Agent Cloud, a serving and training infrastructure platform purpose-built for agentic workloads, long-context inference, and large-scale open-source model deployment. By combining infrastructure and application design, we aim to make open-source models significantly more performant, practical, and competitive.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
We are looking for an ML Systems Engineer to help build and optimize the core serving infrastructure behind Agent Cloud. This role focuses on high-performance inference across different accelerators.
You will work on model serving performance, accelerator utilization, long-context inference, batching, scheduling, KV cache management, runtime efficiency, and cost reduction. This is a deeply technical role at the intersection of ML systems, infrastructure, and product.
Direct TPU experience is a strong plus, but not required. We care most about strong ML systems fundamentals, performance intuition, and the ability to ship reliable systems quickly.
Optimize large-scale LLM inference and serving systems.
Improve total tokens per second, decode tokens per second, latency, throughput, and cost efficiency.
Work on serving infrastructure for open-source models across different types of accelerators.
Improve batching, scheduling, KV cache management, memory usage, and accelerator utilization.
Support long-context inference, including workloads targeting up to 1M context.
Debug performance bottlenecks across model execution, runtime, networking, and infrastructure.
Work with frameworks such as JAX/XLA, PyTorch, vLLM, SGLang, TensorRT-LLM, or related systems.
Collaborate closely with the application team to ensure infrastructure is optimized for agentic workloads, not just generic chatbot inference.
Help turn research prototypes into reliable, high-performance production systems.
Hands-on technical excellence and strong engineering judgment.
End-to-end ownership, from design to implementation to production outcomes.
Bias for action: ship quickly, learn from failures, and iterate.
High intensity during critical milestones, with a focus on real customer impact.
Ability to do deep, focused work and sustain execution.
Clear communication with teammates, customers, and stakeholders.
Comfort with ambiguity, rapid change, and wearing multiple hats.
Low ego, high integrity, high accountability, and strong collaboration.
Continuous learning and a belief that judgment, intelligence, and capability compound over time.
If you are excited to build the infrastructure and agent systems behind the next generation of AI applications, push open-source models to production-grade performance, and turn ambitious research ideas into real-world impact, Model AI is the place for you.
Building high-performance infrastructure for agentic AI systems.
Visit company websiteJobs and hiring trendsFull-time
Senior
Onsite
Apply faster on company sites with our extension.