Design, implement, and optimize high-performance compute and communication kernels for MTIA accelerators, taking ownership from architectural analysis through production deployment
Profile and root-cause performance across the full stack — instruction scheduling, memory hierarchy and DMA behavior, on-chip interconnect, multi-device collectives — and drive the fixes to the right layer, whether that is the kernel, the compiler, the runtime, or the hardware
Build and extend kernel authoring frameworks, templates, and libraries so that other engineers can reach high performance without deep architectural expertise; raise the ceiling and lower the floor at the same time
Deliver and maintain broad kernel coverage for PyTorch operators across recommendation, ranking, and generative AI workloads, in both eager and compiled execution paths
Partner with silicon architecture and design teams on hardware/software co-design: quantify the value of proposed features with real kernels, characterize rooflines pre-silicon, and advocate for the changes the software stack actually needs
Work with compiler, runtime, framework, and product-facing teams to land end-to-end wins on production models rather than isolated microbenchmark improvements
Investigate numerics and precision trade-offs, and design software mitigations that recover performance or accuracy lost to hardware limitations
Set technical direction for a kernel domain, write the design documents that align cross-functional partners, and mentor engineers on performance methodology and accelerator programming
Minimum Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
Bachelor's degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience
6+ years of professional experience in high-performance computing, accelerator kernel development, compiler backends, or systems performance engineering
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
Proficiency in C++ and Python, including low-level systems programming, templates and generic programming, and comfort reading and writing performance-critical code
Demonstrated experience writing and optimizing kernels for a parallel architecture — GPU (CUDA, ROCm/HIP, SYCL/OpenCL), TPU or other AI ASICs, or SIMD/vector CPU targets
Working knowledge of computer architecture as it applies to performance: memory hierarchies and bandwidth, latency hiding, occupancy and scheduling, vectorization, and synchronization
A rigorous, measurement-driven approach to performance: the ability to build a roofline or analytical model, profile against it, and explain the residual gap
Preferred Qualifications
Experience mentoring engineers and setting technical direction across teams
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Experience with distributed execution and collective communication (NCCL/RCCL-class primitives, tensor and expert parallelism, overlapping communication with compute)
Experience with low-precision numerics and quantization — FP8/E4M3/E5M2, MX and other block-scaled formats, INT8/INT4 — including error analysis and calibration
Experience with compiler and codegen technologies relevant to kernels: MLIR, LLVM, TVM, XLA, Halide, or polyhedral scheduling
8+ years of experience in accelerator software, HPC, or ML systems performance (or equivalent with an advanced degree)
Deep familiarity with transformer and attention kernel design: FlashAttention-class algorithms, KV-cache management, paged and chunked attention, linear and state-space attention variants, MoE routing and expert dispatch
Track record of open-source contribution in the kernel, compiler, or ML systems ecosystem
Experience with pre-silicon software development — architectural simulators, FPGA emulation, performance modeling — and with hardware/software co-design cycles
Familiarity with ML framework internals: PyTorch dispatch and eager execution, torch.compile / Inductor, custom operator integration, and inference serving stacks such as vLLM or SGLang
Experience building or contributing to high-performance kernel libraries or frameworks — CUTLASS, cuBLAS, cuDNN, CUTE, Triton, Helion, ThunderKittens, oneDNN, Composable Kernel, or comparable internal equivalents
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)