Base Career helps you apply smarter for this job.
Key skills for this role
Industrial labor is incredibly dangerous work - almost 3 million people in the US per year are injured in the workplace for entirely preventable and at times, fatal or debilitating causes. Protecting these essential people who power our world is what motivates Voxelitos, and we'd love for you to join us. At Voxel, we're passionate about revolutionizing workplace safety and operations with groundbreaking, full-stack AI and computer vision technology.
Voxel’s site intelligence platform helps safety and operations leaders see the unseen risks, make strategic decisions, and prevent workplace incidents before they happen. Our customers include Fortune 500 companies across major grocers and retailers, manufacturers, food and beverage warehousers, supply chain and logistics service providers. Based in SF with team members sitting all over the globe, Voxel is backed by industry leading VC’s.
Voxel is looking for a Staff Machine-Learning Infrastructure Engineer to drive the next wave of our computer-vision platform for workplace safety. You will be the technical owner for three pillars of our ML lifecycle — ground-truth data & labeling workflows, large-scale training infrastructure, and continuous model lifecycle management . If you excel at designing cloud-native, distributed systems that turn raw video into production-ready, version-controlled models, we’d love to meet you.
Own data & labeling pipelines – architect scalable labeling services (storage, query, retrieval), design ontologies, automate annotation workflows, and build quality-tiered datasets that stay within cost constraints.
Build and operate training infrastructure – create multi-GPU / multi-node training frameworks (Ray, Spark, Kubernetes), optimize distributed jobs, and integrate accelerators (TensorRT, CUDA-graph, FP8, etc.).
Manage the full model lifecycle – stand up model registries, version control, evaluation suites, and continuous-learning loops that push updates from dev → staging → prod with zero-downtime rollbacks.
Provide technical leadership, mentorship, and lightweight project management to a small infra + research squad.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
San Francisco, USA
, USA
San Francisco, USA
, USA
, USA
San Francisco, USA
Chicago, USA
Establish DevOps-for-ML best practices (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely.
Partner with ML engineers on architecture decisions, from data schemas to inference optimizations, ensuring infra and research road-maps stay tightly aligned.
Experience running multi-instance / multi-GPU training jobs, mixed-precision optimizations, or TensorRT / Triton inference.
Familiarity with active-learning, continuous-training, or online distillation pipelines.
Background in model registry tooling (MLflow, BentoML, SageMaker Registry) and evaluation dashboards.
Prior work with computer-vision models (YOLO, DETR, Faster RCNN) or video understanding at scale.
Contributions to open-source ML infra projects or published talks/blogs on MLOps.
Exposure to edge-deployment or real-time inference systems.
Experience shipping high quality production code in Python
Join a visionary team revolutionizing safety and operations, directly impacting the well-being of millions of essential workers. This is your chance to build an extraordinary business and foster a vibrant company culture that demands your absolute best. Alongside AI experts, experienced entrepreneurs, and passionate problem-solvers, you'll play a pivotal role in shaping the company's growth trajectory and market position. Enjoy a competitive salary, benefits, and a dynamic work environment.
Industrial AI company providing a site-intelligence platform that helps enterprise safety and operations teams reduce workplace risk.
Visit company websiteJobs and hiring trendsUSD 200000-250000 yearly / year
Full-time
Senior · 5+ years experience
Hybrid
Apply faster on company sites with our extension.