Base Career helps you apply smarter for this job.
Key skills for this role
Develop computer vision and VLM approaches for scene understanding, object detection and tracking, activity recognition, and understanding events across video.
Adapt, fine-tune, and evaluate vision and vision-language models for visual grounding, temporal reasoning, and structured prediction grounded in observable evidence.
Design training and adaptation strategies, including supervised fine-tuning, representation learning, distillation, and domain adaptation, based on measurable product needs.
Build representative image and video datasets, annotation workflows, and evaluation sets that capture difficult edge cases while protecting sensitive data.
Create rigorous experiments and benchmarks that measure perception quality, temporal consistency, hallucinations, robustness, latency, and cost across locations and operating conditions.
Diagnose failures caused by occlusion, lighting changes, camera placement, rare events, and domain shift; use those findings to improve data and models.
Partner with infrastructure and product engineers to deploy efficient inference pipelines, with monitoring, quality gates, staged rollouts, and rollback paths.
Translate advances in computer vision, VLMs, and embodied AI into practical product capabilities, and communicate the evidence and tradeoffs behind your decisions.
Raise research and engineering standards through reproducible experiments, thoughtful reviews, and clear documentation.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
A PhD or research-focused master’s degree in computer vision, machine learning, robotics, or a related field, or equivalent research experience.
Industry research or engineering experience in autonomous driving, robotics, embodied AI, or other applications of perception in the physical world.
Publications at venues such as CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, CoRL, ICRA, or RSS.
Experience with monocular video perception, spatial understanding, long-video reasoning, or learning from limited and noisy labels.
Experience shipping vision models under real-time constraints, including model compression, distillation, quantization, or inference optimization.
When applying, please include links to relevant publications, research projects, or code, and briefly describe your own contribution.
AI platform helping restaurants capture demand, convert revenue, and manage operations through voice, text, and visual agents.
Visit company websiteJobs and hiring trendsFull-time
Mid · 3+ years experience
Onsite
Apply faster on company sites with our extension.