Base Career helps you apply smarter for this job.
Key skills for this role
MatX's mission is to make the world’s best AI models run as efficiently as allowed by physics, bringing the world years ahead in AI quality and availability.
The role covers three areas of equal weight:
Branch and release methodology. Which branches exist, what each one is for, how a change reaches each one, and how that is verified.
CI performance and debugging. Keeping a large, EDA-heavy test suite fast, affordable, and reliable, and finding the cause when it is not.
Measurement and diagnostics. Instrumenting the above so that the state of the system is visible without anyone having to ask.
We expect every engineer to handle everyday collaboration: open a clean PR, review one, take feedback, and decline a change that is wrong. Where that is weak, we teach it. The specialized work is different: release branch structure, freeze enforcement, version pinning, merge-queue and runner behavior. That work is centralized because it requires specific expertise and consistency, not because other engineers are unable to do it. You would own the specialized work and improve the tooling and documentation that everyone else relies on. Success means people need your direct help less over time, not more.
Some common problems include:
Operating the long-lived branches. Protected release branches run alongside the main branch, each with its own rulesets, presubmit checks, and notifications. Each has its own failure modes; for example, a push-triggered workflow runs the copy of itself that exists on the pushed branch. You would own the branch structure and make "did this change land where it was supposed to?" answerable from a report rather than by inspection.
Making "what is the current version of X" mechanical. Several views of the same deliverable have to agree: a build from source, an archived release artifact, and whatever a downstream consumer is actually running. Drift detection should run continuously and its results should be visible without anyone asking.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, USA
, USA
, USA
, USA
, USA
, USA
, USA
, USA
, USA
Keeping CI fast and reliable as it grows. EDA-heavy test suites are slow and expensive by nature, which makes gradual degradation easy to miss. The work includes critical-path analysis, remote-execution behavior under real resource limits, runner-pool sizing and cost, and flake triage that ends in a fix rather than a re-run.
This section describes how the role operates day to day. It suits some people well and others poorly, so it is worth reading closely.
Ownership and completion. You would receive objectives rather than instructions. In return, a task is complete only when the change is pushed, CI is green, and the resulting state has been checked. Handing back work with the final verification still pending is the main thing that causes friction.
Verification. You will be asked how you know something, and the question is not a criticism. Asserting something unchecked, such as that a resource does not exist, a pool is too small, or a limit cannot be avoided, is what draws pushback. The expected answer to "how do you know?" is the command you ran and its output. "I have not verified that yet" is always an acceptable answer.
Durable fixes. The best outcome from a request is usually that the request does not need to be made again. Updating the runbook or adding the check that would have caught the problem is part of the work, not extra scope.
Deep experience in one of the three areas plus working knowledge of the other two is preferable to moderate experience in all three. Tell us which area is your strongest.
Version control internals. You can explain what a rebase does to commit identity, recover a branch that was deleted, verify by content that a cherry-pick landed, and describe what a squash merge queue does to history. You have built tooling on top of Git rather than only used it.
CI systems experience. Merge queues, required checks, event-trigger semantics, app-based authentication, concurrency controls, and self-hosted runners. Given a slow or flaky pipeline, you can identify the cause and quantify it. This includes reading a Bazel query and a build profile, and working in Python and shell against REST and GraphQL APIs.
Judgment under a schedule. You would be telling leadership what is and is not in a milestone. That requires saying "not yet" when it is true, and preferring a change that can be reverted to a migration that cannot.
Developer-infrastructure work under hard deadlines in another industry (silicon, aerospace, games)
Experience with repository structure at scale, whether monorepo or polyrepo
Familiarity with EDA flows.
This is not the SRE role for the compute fleet, and it is not a process role that hands the tooling to someone else.
AI semiconductor company designing high-throughput chips for large language models and frontier AI labs.
Visit company websiteJobs and hiring trendsUSD 160000-600000 yearly / year
Full-time
Senior
Hybrid
Apply faster on company sites with our extension.