Base Career helps you apply smarter for this job.
Key skills for this role
As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries.
As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Memphis, USA
Memphis, USA
New York City, USA
, USA
Austin, USA
Memphis, USA
Dublin, USA
Palo Alto, USA
Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
Proven large-scale incident command experience and calm technical leadership on a bridge.
Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
Excellent problem-solving skills with a data-driven approach to reliability engineering.
Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.
Experience in AI/ML infrastructure or supercomputing environments.
Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
Experience running game days, dependency mapping, and closed-loop corrective action programs.
Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
Prior work in a fast-paced startup or tech company like SpaceXAI.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
xAI is an artificial intelligence company founded by Elon Musk that develops AI systems to accelerate scientific discovery. Its main products include the Grok AI chatbot and the Colossus supercomputer, one of the world's largest AI training clusters.
Visit company websiteJobs and hiring trendsMid · 5+ years experience
Apply faster on company sites with our extension.