Base Career helps you apply smarter for this job.
Key skills for this role
As part of the Cluster Orchestration team, you will play a key role in advancing CoreWeave’s orchestration platform including SUNK (Slurm on Kubernetes) and beyond our Kubernetes-native foundation that powers AI training and inference at scale. This is an opportunity to help shape one of the most critical layers of the AI cloud: ensuring workloads run seamlessly, reliably, and efficiently across massive GPU clusters. By building the systems that eliminate infrastructure bottlenecks and create new orchestration capabilities, you will directly empower customers to innovate faster and push the boundaries of what’s possible with AI.
As part of the Cluster Orchestration team, you will play a key role in advancing CoreWeave’s orchestration platform including SUNK (Slurm on Kubernetes) and beyond our Kubernetes-native foundation that powers AI training and inference at scale. This is an opportunity to help shape one of the most critical layers of the AI cloud: ensuring workloads run seamlessly, reliably, and efficiently across massive GPU clusters. By building the systems that eliminate infrastructure bottlenecks and create new orchestration capabilities, you will directly empower customers to innovate faster and push the boundaries of what’s possible with AI.
As a Senior Software Engineer I (IC3), you will own multiple services within the orchestration platform. You’ll lead design/code reviews, decompose projects into milestones, and drive measurable improvements in reliability and performance. You’ll define SLIs/SLOs for your services, strengthen operational practices, and mentor IC1/IC2 engineers. Your work will ensure customers see consistent improvements in throughput, latency, and system resilience.
~3–5 years of professional software engineering experience building distributed systems or cloud services.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bellevue, USA
Washington, USA
Bellevue, USA
Bellevue, USA
New York City, USA
Livingston, USA
Livingston, USA
Livingston, USA
Bellevue, USA
Strong coding in Go (Python or C++ a plus) with solid CS fundamentals.
Hands-on experience running Kubernetes at production scale.
Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry).
Proven ability to improve service reliability and performance using metrics (P95/P99 latency, throughput, error budgets).
Familiarity with orchestration and workflow technologies such as Ray, Kubeflow, Kueue, Istio, Knative, or Argo Workflows
Experience with distributed workloads, GPU-based applications, or ML pipelines.
Knowledge of scheduling concepts like quota enforcement, pre-emption, and scaling strategies.
Exposure to reliability practices including SLOs, alarms, and post-incident reviews.
The base salary range for this role is $139,000 to $204,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).
This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.
Specialized cloud provider for large-scale AI and machine learning.
Visit company websiteJobs and hiring trendsUSD 139000-204000 yearly / year
Full-time
Senior · 3+ years experience
Hybrid
Apply faster on company sites with our extension.