More from this employer
San Francisco, USA
About the Role You'll be the second engineering hire at an early-stage B2B SaaS startup building AI agents for the industrial ingredient manufacturing industry, a sector worth hundreds of billions of dollars that underpi
San Francisco, USA
About the Role This is the first engineering hire at an early-stage B2B SaaS company building an agentic integration platform with paying enterprise customers. You will own core systems end-to-end and set the engineering
San Francisco, USA
About the Role This is a founding engineer role at an early-stage hardware-meets-software startup building AI-assisted PCB design tooling and EDA automation for the full electronics development lifecycle. You will own co
New York City, USA
About the Role This is a high-ownership engineering role at an early-stage healthtech startup building AI-powered payer compliance tools for surgical and procedural specialties. You'll sit between the founding team and c
San Francisco, USA
About the Role This is a full-stack product engineering role at a consumer fintech company helping everyday people manage cash flow, access earned wages early, and build credit. You'll work closely with product, engineer
, USA
, USA
, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
New York City, USA
San Francisco, USA
Base Career helps you apply smarter for this job.
This is a senior, hands-on technical leadership role owning the strategy and systems that measure, improve, and scale training data for frontier AI agents. You will sit at the intersection of research and engineering, leading a team that defines what high-quality agent training data looks like and building the infrastructure to enforce that bar at scale. The work directly shapes the post-training data that aligns AI models to real-world tasks.
This is a senior, hands-on technical leadership role owning the strategy and systems that measure, improve, and scale training data for frontier AI agents. You will sit at the intersection of research and engineering, leading a team that defines what high-quality agent training data looks like and building the infrastructure to enforce that bar at scale. The work directly shapes the post-training data that aligns AI models to real-world tasks.
Lead the data quality team in building evaluation systems across RL environments, synthetic data, benchmarks, and domain-specific workflows.
Define data quality strategy by building QC systems, enforcing standards, and designing experiments to grade agent outputs.
Develop methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing.
Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows.
Translate qualitative research insights into production systems: validation pipelines, dashboards, internal tools, and feedback loops.
Help build internal research taste around what makes agent training data realistic, learnable, diverse, reliable, and genuinely useful.
Mentor research engineers to maintain a high bar for technical rigor, clarity, and execution speed.
5+ years of experience in research or data quality engineering, specifically building systems for AI/ML data evaluation.
Demonstrated experience leading technical projects or teams in data quality or AI/ML evaluation, ideally on ambiguous, open-ended problems.
Advanced proficiency in Python, Docker, and Linux environments.
Deep, research-oriented understanding of AI evals and post-training, beyond surface-level agent frameworks.
Experience building QC systems, benchmarks, synthetic data pipelines, validation workflows, or model evaluation infrastructure.
Ability to reason carefully about what makes training data high-quality for AI agents, not just technically valid.
Experience translating research insights into production pipelines and internal tooling.
Ability to collaborate with domain experts and data vendors, capturing expert judgment and converting it into scalable review or generation systems.
Strong written communication skills, with the ability to explain methodology clearly to researchers, engineers, and external stakeholders.
Comfort designing metrics, experiments, and QA/QC processes independently.
Early-stage startup experience and the ability to move quickly in fast-paced, resource-constrained environments.
Detail-oriented mindset with a sharp eye for subtle inconsistencies and edge cases in data.
Salary range: $150,000 to $180,000 USD annually. Visa sponsorship is available.
AI-powered talent agent matching candidates to startup roles.
Visit company websiteJobs and hiring trendsSkip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career