Base Career helps you apply smarter for this job.
Key skills for this role
FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud.
As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on.
Inference is an unforgiving workload for Kubernetes.
Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load.
This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further.
FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on.
Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud.
Cilium and eBPF, including kube-proxy replacement or upstream contributions.
Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling.
GPU orchestration: NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA).
High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning.
Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations.
Contributions to Kubernetes, Cilium, Istio, or other CNCF projects.
FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling.
We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.
FriendliAI is a private AI inference platform for enterprises deploying, scaling, and monitoring large language and multimodal models.
Visit company websiteJobs and hiring trendsFull-time
Senior · 5+ years experience
Hybrid
Apply faster on company sites with our extension.