Design, build, and operate infrastructure and backend services that power production AI features.
Build and improve model training infrastructure, including systems that support training jobs, experimentation, compute management, data workflows, and model artifacts.
Build infrastructure that supports the full model lifecycle, from training and experimentation through deployment and production serving.
Improve the scalability, performance, reliability, and cost efficiency of our AI platform.
Build and maintain cloud-based services and containerized workloads for AI and ML applications.
Develop systems that make it easier for ML engineers to train, evaluate, deploy, and iterate on models.
Build reusable platform capabilities that allow engineering teams to launch new AI-powered features quickly and safely.
Improve deployment, rollout, monitoring, and operational workflows for AI models and services.
Diagnose and resolve reliability and performance issues across application, infrastructure, compute, and ML system layers.
Support large-scale AI workloads across text, image, video, and multimodal applications.
Evaluate and adopt emerging AI infrastructure technologies when they provide meaningful improvements in productivity, performance, reliability, or cost.
Improve software development and operational workflows through effective use of modern AI coding agents and agentic engineering tools.
Work closely with AI scientiests, and product teams to translate rapidly changing requirements into practical technical solutions.
Own systems end-to-end, from architecture and implementation through deployment, monitoring, operation, and continuous improvement.
3+ years of professional software engineering experience building production backend systems, infrastructure, or distributed systems.
Strong programming skills in Python, Java, Go, or a comparable backend or systems language.
Strong understanding of distributed systems fundamentals, including concurrency, fault tolerance, messaging, backpressure, load balancing, and horizontal scaling.
Production experience with Kubernetes or Docker-based containerized workloads.
Experience operating production services on a major cloud platform such as AWS, GCP, or Azure.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
Experience designing and operating high-throughput or latency-sensitive production systems.
Familiarity with NoSQL databases such as DynamoDB, Cassandra, or comparable distributed data stores, including common data modeling and scalability considerations.
Strong debugging and problem-solving skills across application, infrastructure, and networking layers.
Experience with production observability, including metrics, logging, tracing, dashboards, and alerting.
Deep understanding of state-of-the-art AI coding agents, such as Claude Code, OpenAI Codex, or comparable agentic coding systems, with demonstrated ability to use them effectively in real software engineering workflows.
Ability to independently own complex systems from design through production operations.
Ability to move quickly, iterate with incomplete information, and adapt effectively as product requirements and technical priorities evolve.
Comfort operating in a fast-paced, high-pressure environment while maintaining sound engineering judgment and execution quality.
Understanding of LLM inference architecture and production model-serving systems.
Experience with inference frameworks such as SGLang, vLLM, TensorRT-LLM, TGI, or similar technologies.
Understanding of inference concepts such as prefill vs. decode, continuous batching, KV cache, prefix caching, and request scheduling.
Experience with NVIDIA GPUs such as A100 or H100.
Experience profiling or optimizing GPU workloads for throughput, latency, memory utilization, or cost.
Experience with Kubernetes GPU scheduling, GPU node pools, KEDA, or workload-aware autoscaling.
Experience with high-throughput image, video, or multimodal processing systems.