Base Career helps you apply smarter for this job.
Key skills for this role
MatX is on a mission to be the compute platform for AGI. We are developing vertically integrated full-stack solutions from silicon to systems including hardware and software to train and run the largest ML workloads for AGI.
We are seeking a hands-on engineering TLM to lead the Kernel team. You will manage, mentor, and grow a team of kernel engineers while setting technical direction and executing. At the same time, you will be an active coding contributor by designing and implementing high-performance compute kernels for specialized AI hardware.
You will partner with the ML team to maximize output and improve diagnostic tools, and collaborate with hardware to co-design next-generation AI architectures.
Manage, mentor, and build a team of highly specialized kernel engineers and software developers
Design and optimize kernels that interface directly with our hardware
Work in partnership with our ML Research and Hardware Engineering teams
Provide expertise and guidance on hardware architecture from a programmer's perspective, ensuring seamless integration with the software stack
2+ years of engineering team management experience
7+ years of working directly within engineering teams experience
Bachelor of Computer Science or equivalent degree
Experience optimizing software for specialized hardware, employing techniques such as parallelism, SIMD programming, C, assembly-level optimization, or GPU/CUDA programming
Language: at least one of assembly, C++, C, Zig, or Rust
Ability to navigate ambiguity, manage escalations, and align stakeholders around infrastructure decisions.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, USA
, USA
, USA
, USA
, USA
, USA
, USA
, USA
, USA
Excellent communication and collaboration skills, with a track record of translating complex technical trade-offs with the founders and stakeholders
Experience implementing kernels for ML models such as Transformers
Experience using and implementing distributed parallelism techniques such as AllReduce, AllToAll, data parallelism, tensor parallelism.
Familiarity with how compilers work
AI semiconductor company designing high-throughput chips for large language models and frontier AI labs.
Visit company websiteJobs and hiring trendsUSD 120000-600000 yearly / year
Full-time
Senior · 7+ years experience
Hybrid
Apply faster on company sites with our extension.