VLA 全栈研究员 / VLA Full-Stack Researcher

岗位使命

负责通用人形机器人基座模型的 VLA 全链路研究,让机器人完成导航、灵巧手操作、移动操作(Loco-manipulation)等多种真实任务,并让模型能力随数据、模型规模和训练计算持续增长。

你将负责

  • 负责 VLA 的数据构建、Pre-training、Post-training 和评测。
  • 研究人形机器人导航、灵巧手操作和移动操作(Loco-manipulation),建设统一导航、灵巧操作和移动操作的通用 VLA 模型。
  • 研究和优化 VLA 模型架构,探索 Linear Attention、KDA 等高效序列建模方法,提升模型对长时序任务和复杂动作序列的建模与推理能力。
  • 建设多模态和机器人数据的清洗、配比、标注、合成与质量评估流程。
  • 提升模型在跨任务、跨场景、跨本体和长时程任务中的泛化能力。
  • 研究 VLA 之上的顶层 Agent,负责任务分解、高层规划与长时程任务决策。
  • 将模型部署到真实人形机器人,并根据真机结果持续迭代。

我们希望你

  • 深入理解 Transformer、VLM、VLA 或多模态生成模型。
  • 在导航、灵巧手操作、Loco-manipulation、机器人学习或具身智能中的一个或多个方向有研究或工程经验;无需同时精通所有方向,欢迎在一个方向有深度并愿意拓展到完整链路。
  • 至少完整参与过一次较大规模模型训练,熟悉 Pre-training、Post-training、RL 中至少一个方向。
  • 能够分析训练曲线、数据问题、模型退化、推理瓶颈和评测偏差,有很强的工程实现和实验能力。

VLA Full-Stack Researcher

Mission

Own the full-stack VLA research for our general-purpose Humanoid Foundation Model — enabling robots to perform real-world tasks across navigation, dexterous manipulation, and loco-manipulation, with capability that keeps scaling with data, model size, and training compute.

What You'll Do

  • Own VLA data construction, pre-training, post-training, and evaluation.
  • Research humanoid navigation, dexterous manipulation, and loco-manipulation — building a single general VLA model that unifies all three.
  • Research and optimize VLA architectures, exploring efficient sequence modeling methods such as Linear Attention and KDA to improve modeling and reasoning over long-horizon tasks and complex action sequences.
  • Build pipelines for cleaning, mixing, annotating, synthesizing, and quality-assessing multimodal and robot data.
  • Improve generalization across tasks, scenes, embodiments, and long-horizon tasks.
  • Research the top-level agent above the VLA — task decomposition, high-level planning, and long-horizon decision-making.
  • Deploy models on real humanoid robots and iterate continuously based on real-world results.

What We're Looking For

  • Deep understanding of Transformers, VLMs, VLAs, or multimodal generative models.
  • Research or engineering experience in one or more of: navigation, dexterous manipulation, loco-manipulation, robot learning, or embodied AI. You don't need to master every direction — depth in one, plus willingness to expand across the full stack, is welcome.
  • Full-cycle participation in at least one large-scale model training run; familiar with at least one of pre-training, post-training, or RL.
  • Able to analyze training curves, data issues, model degradation, inference bottlenecks, and evaluation bias — with strong engineering and experimental skills.