Base Career helps you apply smarter for this job.
Key skills for this role
How do we make advanced AI/ML not just powerful but trustworthy enough to run in a regulated healthcare environment? We're looking for a Senior Staff engineer who can own that question end to end: the data pipelines that feed our models and evaluations, and the evaluation infrastructure — judges, scorers, statistical regression gates, red-team testing — that decides whether AI models are effective and safe to ship to patients.
This is a leadership role for someone who works comfortably across disciplines. You'll set technical direction for how we build, version, and trust the data and judgments that our AI products are evaluated against. Your scope will expand from raw data ingestion and feature/dataset pipelines, through evaluation methodology and statistical rigor, to the reporting surfaces that let clinical and product teams act on what we learn.
You'll spend most of your time on problems that don't have an existing playbook: ambiguous, cross-team, and genuinely hard to reason about. The job is to bring clarity to that ambiguity, chart a path the rest of the team and organization can follow, and see it through from idea to production, building the relationships and buy-in along the way to make it stick.
How do we make advanced AI/ML not just powerful but trustworthy enough to run in a regulated healthcare environment? We're looking for a Senior Staff engineer who can own that question end to end: the data pipelines that feed our models and evaluations, and the evaluation infrastructure — judges, scorers, statistical regression gates, red-team testing — that decides whether AI models are effective and safe to ship to patients.
This is a leadership role for someone who works comfortably across disciplines. You'll set technical direction for how we build, version, and trust the data and judgments that our AI products are evaluated against. Your scope will expand from raw data ingestion and feature/dataset pipelines, through evaluation methodology and statistical rigor, to the reporting surfaces that let clinical and product teams act on what we learn.
You'll spend most of your time on problems that don't have an existing playbook: ambiguous, cross-team, and genuinely hard to reason about. The job is to bring clarity to that ambiguity, chart a path the rest of the team and organization can follow, and see it through from idea to production, building the relationships and buy-in along the way to make it stick.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
New Albany, USA
Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to c
Menlo Park, USA
Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to c
Gilbert, USA
Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to c
New Albany, USA
Menlo Park, USA
Gilbert, USA
London, GBR
Gilbert, USA
Gilbert, USA
, USA
, USA
Set the technical direction for our evaluation systems; metric, judge and scorer design, the statistical methodology behind regression decisions, and the infrastructure that tracks and categorizes failures over time.
Design and scale the data pipelines — ingestion, transformation, dataset versioning, labeling and calibration workflows — that both evaluation and downstream data science work depend on.
Proactively address challenges in scaling and complexity AI evaluation.
Define how we evaluate any new AI service from scratch. Drive multi-team initiatives like replacing manual, inconsistent review processes with statistically sound, automated gates.
Own our approach to adversarial and red-team evaluation as a risk-reduction program, designing the test suites and failure taxonomies that catch safety and edge-case issues before they reach patients.
Work through complex, cross-team technical disagreements and drive alignment across engineering, product and AI leaders.
Originate new approaches and methodology that becomes a reusable standard rather than a one-off fix.
Take vague, cross-team pain points ("we do this manually and it's inconsistent") all the way from a rough idea to a fully-specified, shipped system, without needing to hand off any part of the journey.
Lead major platform improvements; re-architecting core systems, removing brittle logic, modernizing how things run with impact that's felt org-wide, not just on your own team.
Build real working relationships across ML engineering, data science, platform engineering, clinical, legal, and product, the kind of trust that gets you looped in early, before decisions are locked in.
Become a go-to voice on evaluation methodology and data pipeline design: share what you've learned in internal talks, write things up so other teams can use them, and expect your ideas to shape how others approach similar problems.
Mentor other engineers, including experienced ones, and help raise the technical and statistical bar of the teams you work with.
10+ years of experience in ML infrastructure, data engineering, or evaluation/testing systems, with a track record of impact that reaches beyond a single team or project.
Hands-on depth in evaluation systems: designing and calibrating LLM judges/scorers, building statistically sound regression-testing methodology (e.g., paired significance testing with proper correction for multiple comparisons), measuring agreement against human labels, and designing adversarial/red-team evaluation approaches.
Hands-on depth in data pipeline engineering: dataset versioning, feature and benchmark pipelines, labeling and calibration workflows, and high-throughput ingestion and transformation systems.
A history of building things that became the standard approach for others — not just solving your own problem, but changing how a broader group of people tackle a category of problem.
Experience leading multi-team projects to completion, including navigating and resolving genuine technical disagreement along the way.
A track record of mentoring other engineers, including senior ones, and visibly raising the bar for the teams around you.
Excellent communication — comfortable adapting the same idea for different audiences, and confident building support for it well before launch.
Strong Python, and enough statistical fluency to design and defend a testing framework that real production decisions ride on.
Experience with Databricks, MLflow, Unity Catalog, or similar data/eval platforms.
Experience building reporting tools for people without direct engineering access (e.g., automated Slack digests, spreadsheet reports for non-technical teams).
Prior experience in a regulated industry (healthcare, fintech, life sciences).
A track record of company-wide talks or write-ups that changed how other teams approached a problem.
Publicly traded U.S. telehealth platform connecting consumers with licensed providers and personalized prescription and wellness care.
Visit company websiteJobs and hiring trendsUSD 240000-265000 yearly / year
Full-time
Senior · 10+ years experience
Remote
Apply faster on company sites with our extension.