{bc}
ashby

Research Engineer - Post-Training

Voltai
Palo Alto, USA
Full-time
Mid
Onsite
Discovered Yesterday
Reinforcement Learning
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Reinforcement Learning
Smart Apply

Full Job Posting

About the Team

Backed by Silicon Valley’s top investors, Stanford University, and CEOs/Presidents of Google, AMD, Broadcom, Marvell, etc. We are a team of previous Stanford professors, SAIL researchers, Olympiad medalists (IPhO, IOI, etc.), CTOs of Synopsys & GlobalFoundries, Head of Sales & CRO of Cadence, former US Secretary of Defense, National Security Advisor, and Senior Foreign-Policy Advisor to four US presidents.

Post-Training

In this role, you will post-train frontier models to autonomously perform complex tasks across the semiconductor design and verification pipeline. Models you train will propose and optimize chip architectures, generate and refine RTL code, run simulations, identify verification gaps, and iteratively improve designs — accelerating the pace of semiconductor innovation. You will collaborate with leading experts in hardware design, verification, and computer architecture to design rich reinforcement learning environments that capture the intricacies of chip design workflows. You’ll develop structured reward functions, scaling strategies, and evaluation frameworks that push models toward higher reliability, efficiency, and creativity in semiconductor reasoning. Your work will directly advance the goal of creating AI systems capable of reasoning about, designing, and verifying next-generation silicon systems.

You might thrive in this role if you have experience with

Creating and scaling RL environments for LLMs or multimodal agents

Building high-quality evaluation datasets and benchmarks for complex reasoning or design tasks

Working closely with domain experts in hardware and verification to define evaluation metrics, constraints, and simulation conditions

Designing reward functions and feedback pipelines that balance correctness, performance, and design efficiency

Running large-scale RL fine-tuning or post-training experiments for frontier models

Applying reinforcement learning or curriculum learning to structured reasoning or symbolic domains

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Voltai