Base Career helps you apply smarter for this job.
Key skills for this role
Research and develop predictive world models that learn how the physical world evolves, forecasting the future state of a scene from large-scale multimodal driving and robotics data.
Develop high-quality multi-view future prediction and generation, supporting both action-conditioned rollouts and formulations that forecast the future without explicit action conditioning.
Work at the boundary between world modeling and policy learning: develop architectures in which a shared backbone both predicts the future and produces trajectories or actions, and apply predictive pre-training to improve Vision-Language-Action (VLA) driving performance.
Extend prediction beyond 2D pixel into a shared multimodal latent space that spans 3D scene representations such as Gaussian Splatting, together with occupancy and reward signals, so that a single model can support simulation, evaluation, and policy training.
Advance cross-embodiment generalization: design unified observation and action representations, together with embodiment-conditioning mechanisms, so that a single world model transfers across vehicles, robots, and sensor configurations with only few-shot data.
Define and build the evaluation methodology for predictive world models, spanning representation quality, prediction accuracy, generation fidelity, physical plausibility, long-horizon rollout consistency, and ultimately closed-loop policy performance, then feed the resulting models back into training as a source of synthetic data and corner-case simulation.
MS or PhD level education in Engineering or Computer Science with a focus on Deep Learning, Computer Vision, Generative Models, or a related field, or equivalent experience. Open to recent graduates.
Strong experience in applied deep learning including model architecture design, large-scale model training, data curation, and empirical analysis.
1-3 years + of experience working with DL frameworks such as PyTorch, including hands-on experience with distributed training (FSDP, DeepSpeed, or Megatron-style parallelism).
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Brazil, USA
Dubai, UAE
, CAN
, CAN
, CAN
, CAN
, CAN
, CAN
Strong Python programming experience with software design skills.
Solid understanding of data structures, algorithms, code optimization and large-scale data processing.
Excellent problem-solving skills, including the ability to design controlled experiments and draw sound conclusions from noisy training signals.
Hands on experience with generative models for video or 3D, such as diffusion, flow matching, autoregressive video prediction, or neural scene representations including NeRF and Gaussian Splatting.
Experience with world models or learned simulators for decision making, including model-based reinforcement learning and Vision-Language-Action (VLA) models.
Experience with multimodal foundation models and video tokenizers or VAEs, including pretraining or adapting large pretrained backbones.
Experience with large-scale training infrastructure and performance optimization, such as mixed precision, torch.compile, kernel-level optimization, and multi-node scaling.
A fun, supportive and engaging environment.
Infrastructures and computational resources to support your work.
Opportunity to work on cutting edge technologies with the top talents in the field.
Opportunity to make a significant impact on the transportation revolution by the means of advancing autonomous driving.
Competitive compensation package.
Snacks, lunches, dinners, and fun activities.
Founded in 2014, XPENG is a Chinese smart electric vehicle manufacturer integrating advanced AI and autonomous driving technologies into its cars, eVTOL aircraft, and robotics products.
Visit company websiteJobs and hiring trendsUSD 174720-295680 / year
Senior · 1–3 years experience
Apply faster on company sites with our extension.