Senior Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Key Skills for This Role
Full Job Posting
What You'll Own
Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
Tune Linux performance deeply, at the OS and kernel level.
Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.
First 90 Days
One way the first 90 could unfold.
Days 1–30 — Immerse & Diagnose: Learn the current clusters across on-prem, AWS, and OCI, and where reliability and performance hurt most.
Days 30–60 — Ship & Validate: Take ownership of a production cluster and ship automation or tuning that measurably improves availability or performance.
Days 60–90 — Scale & Systemize: Contribute to the next-gen re-architecture and harden security and compliance practices.
What You Bring
5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
Working experience with Terraform, Airflow, and Ray.
Strong experience with AWS or OCI.
Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
Comfort in a less-structured, fast-paced environment.
Nice to Have
Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm).
Experience managing large-scale GPU clusters for AI/ML training or inference.
Familiarity with Kubernetes or orchestration frameworks like Ray.
Deep expertise in data pipelines and infrastructure.
About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
About Luma
Verified company details for this employer are not available yet.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Luma
Forward Deployed Creative [US]
Redwood City, USA
Account Executive – Entertainment
Los Angeles, USA
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Redwood City, USA
Research Scientist / Engineer – Performance Optimization
Redwood City, USA
Software Engineer, Inference
Redwood City, USA
Product Manager, Core Product
Redwood City, USA
Forward Deployed Engineer
Redwood City, USA
Staff AI Infrastructure Engineer
Redwood City, USA