Own episode acceptance criteria for every task family — translate research and engineering data needs into operational, auditable standards, and version them as model training needs evolve.
Take over the annotation and episode-review queue in your first weeks, then build the team and workflows that scale it far beyond yourself.
Stand up and manage a remote annotation team (likely Philippines-based; direct hires or vendor/BPO) with an overnight turnaround SLA — episodes collected today are reviewed and scored before the next shift starts.
Design and run a sampling-based QA audit program with explicit coverage targets, including inter-rater reliability checks that keep annotators and auditors calibrated.
Run the operator quality feedback loop: operator-level quality scorecards delivered to site supervisors within 24 hours. You own the standard and the signal; supervisors own the coaching and people decisions.
Instrument your function: define the quality metrics (episode acceptance rate, audit coverage, feedback latency, operator quality distribution, annotation throughput) and build the operational dashboards your team runs on, partnering with our analytics function, which independently owns org-wide reporting.
Own the certification bar for new operator onboarding — no one collects production data without meeting it — while site teams run the day-to-day training reps.
Drive a standing weekly loop with research and engineering on failure modes, task-spec drift, and what "good" needs to mean next.
Qualifications
3+ years in data operations, annotation/labeling operations, or data collection QA for ML systems — robotics, autonomy, or teleoperation data strongly preferred.
A track record of building quality standards from scratch — rubrics, SOPs, QC workflows — not just executing against existing ones. Be ready to walk us through one you built.
Experience managing annotation or review teams, including remote/offshore or vendor/BPO teams, and driving their performance against SLAs.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
About Nimble
Nimble is an AI robotics company building the autonomous supply chain to power fast, efficient and economical commerce. We’re training robot AGI to power a proprietary generalist supply chain superhumanoid,
About Nimble
Nimble is an AI robotics company building the autonomous supply chain to power fast, efficient and economical commerce. We’re training robot AGI to power a proprietary generalist supply chain superhumanoid,
Experience delivering direct, frequent quality feedback to operators or annotators — including the hard conversations when someone isn't meeting the bar.
Data fluency: able to build and own your own reporting and pressure-test the numbers — SQL, BI tools, or AI-assisted, we don't care how — rather than waiting on someone else.
Meticulous judgment on edge cases, paired with the pragmatism to ship a v1 rubric this week instead of a perfect one next quarter.
High agency and comfort with ambiguity in a fast-paced, high-growth environment — the org will triple around you this year.
Able to work in person out of our San Francisco HQ, with regular time at our collection sites.
Alignment with Nimble's values: relentlessly resourceful, humble, dependable, and committed to legendary impact.
Additional Requirements
Familiarity with imitation learning / robot learning data — what makes a demonstration usable for model training.