Base Career helps you apply smarter for this job.
Key skills for this role
We’ve just released Fabric — a DiT-based speech-to-video model that turns spoken words into coherent, human-centric video. We’re now building the next version and a suite of video generation & editing models focused on controllability, quality, and speed.
You will work across the full model lifecycle for Fabric 2 and related models — from exploring new research and approaches to training and fine-tuning models. Your work will turn ideas from research into features creators use every day.
At VEED , we're a generative AI company building the AI video creation platform made for social. It's where millions of marketers imagine and generate branded videos, in minutes.
We're powered by frontier models and VEED Fabric, our own image-to-video model. Our APIs also run AI video inside leading creative products.
We're 150 people, $50M ARR, backed by Sequoia. We're hiring. Come help us make AI videos worth posting.
We’re a distributed team with hubs in London, Barcelona, and Amsterdam. Our ML team is mainly in London, so we’re open to candidates already based there, willing to relocate, or working remotely within 4–5 hours of London time (GMT).
You’ll be joining a team of experienced ML engineers and researchers (ex-Spiritme, ex-Pipio) working on scaling our generative video models. The team’s goal is to turn breakthroughs in research into fast, reliable, and accessible systems that power VEED’s future.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
London, GBR
London, GBR
London, GBR
London, GBR
London, GBR
London, GBR
We’ve just released Fabric — a DiT-based speech-to-video model that turns spoken words into coherent, human-centric video. We’re now building the next version and a suite of video generation & editing models focused on controllability, quality, and speed.
You will work across the full model lifecycle for Fabric 2 and related models — from exploring new research and approaches to training and fine-tuning models. Your work will turn ideas from research into features creators use every day.
Training and fine-tuning models to create best-in-class generation and editing people-centric video models
Running experiments to boost visual quality, lip sync, and controllability
Defining clear metrics and evaluation loops and using them to guide decisions
Work closely with infra and performance engs and collaborate with the product team
Model/Training: PyTorch, diffusion/DiT, diffusers, transformers, torch.compile, torch.distributed, torchrun, Accelerate, PEFT/LoRA.
Data: Python, Hugging Face datasets, WebDataset-style storage, ffmpeg, large-scale curation & filtering.
Serving: Optimised PyTorch/TensorRT runtimes in GPU-backed Python services; experiment tracking, A/B testing, eval pipelines.
Hardware: Access to latest-gen NVIDIA GPUs H100 and B200 for training and inference.
Experience training diffusion or transformer models for video, vision, or audio with results in production or papers
You care about shipping useful ML, not just benchmarks. You enjoy owning problems and collaborating across teams.
Strong data and experimentation skills and confidence designing practical metrics
Comfortable writing clean Python and reviewing research with a pragmatic eye
Clear communicator who can explain trade offs and make sound decisions
Motivated by impact and by making generative video better, faster, and more accessible
UK private AI video creation platform helping marketers, founders, and creators generate, edit, and publish branded videos.
Visit company websiteJobs and hiring trendsGBP 76000-119400 yearly / year
Full-time
Senior
Hybrid
Apply faster on company sites with our extension.