Base Career helps you apply smarter for this job.
Key skills for this role
Be part of the team creating the software foundation for next-generation AI compute platforms.
In this role, you'll work on compiler technologies, AI runtimes, graph optimization, and hardware-aware execution in close collaboration with inference engineers, ML scientists, and hardware specialists.
You'll leverage modern AI technologies and methodologies to accelerate software development, improve software-hardware co-design, and optimize the deployment and execution of AI workloads on next-generation compute platforms.
This position offers the opportunity to contribute to state-of-the-art AI infrastructure, optimize software for emerging AI hardware, and help define how modern machine learning workloads are represented, compiled, and executed at scale.
We are particularly interested in engineers who have applied AI technologies to solve complex systems, compiler, runtime, or hardware challenges, rather than solely using AI as a software productivity tool.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
Build and optimize compiler and runtime infrastructure for modern AI workloads
Enable efficient execution of machine learning models across GPUs, NPUs, TPUs, and custom AI accelerators
Apply modern AI technologies and methodologies to improve software development, system optimization, and software-hardware co-design processes
Collaborate with inference, systems, and hardware teams to improve software-hardware co-design
Investigate and resolve performance bottlenecks through profiling, benchmarking, and system-level analysis
AI compiler frameworks and infrastructure (e.g., Mojo/Max, MLIR, LLVM, XLA, OpenXLA, Triton, Gluon)
Compiler optimizations such as operator fusion, graph transformations, scheduling, code generation, and lowering pipelines
AI runtimes and execution frameworks (e.g., ONNX Runtime, TensorRT, TVM Runtime, IREE Runtime, XLA)
Deep understanding of AI systems, with experience applying AI technologies to solve engineering, performance, systems, or hardware-software optimization challenges
Performance optimization of machine learning workloads on GPUs, TPUs, NPUs, DSPs, or custom accelerators
Hardware-aware software development and AI accelerator enablement
Development of high-performance kernels and operators (e.g., GEMMs, convolutions, attention, normalization, quantization)
Distributed AI training or inference systems
Model execution frameworks such as Max, PyTorch, TensorFlow, JAX, or ONNX
Experience with Modular (Mojo/Max), OpenXLA, StableHLO, Torch-MLIR, Triton, TVM, or IREE
Experience developing software for AI accelerators or machine learning hardware platforms
Contributions to open-source projects such as LLVM, MLIR, PyTorch, OpenXLA, Triton, Gluon, or xDSL
Experience building software for AI accelerators or AI compute platforms
HTEC Group, Inc. is an equal opportunity employer.
Private AI-first digital engineering and technology consulting firm serving startups, enterprises, and Fortune 500 companies worldwide.
Visit company websiteUSD 200000-350000 yearly / year
Full-time
Senior
Hybrid
Apply faster on company sites with our extension.