AI Infra 全栈工程师 / AI Infra Full-Stack Engineer(训练与推理 / Training & Inference)

岗位使命

建设高性能、高可靠、可扩展的模型训练和推理基础设施,让训练更快、推理延迟更低、算力利用率更高,并显著缩短从实验到真机验证的周期。

你将负责

  • 建设 VLM/VLA 的分布式训练、Post-training 和 RL 基础设施。
  • 优化训练吞吐、显存利用率、通信效率和计算资源利用率,负责并行策略、Checkpoint、容错恢复和任务调度。
  • 优化模型推理,包括 Serving、Batching、KV Cache、量化、编译和 Kernel 优化。
  • 建设从实验提交、资源调度、指标监控到模型产物管理的完整研发平台。
  • 优化模型在服务器和机器人端的部署效率、稳定性和延迟。

我们希望你

  • 熟悉 PyTorch、CUDA、Triton、NCCL 或相关训练与推理系统。
  • 理解数据并行、张量并行、流水线并行、序列并行或专家并行中的一种或多种。
  • 能够使用 Profiling 工具定位系统瓶颈,并通过实验验证优化效果。
  • 熟悉 Linux、容器、集群调度、分布式存储和可观测性系统,对系统性能敏感,愿意深入框架、Runtime 和 Kernel 解决问题。

AI Infra Full-Stack Engineer (Training & Inference)

Mission

Build high-performance, reliable, and scalable training and inference infrastructure — faster training, lower inference latency, higher compute utilization, and a much shorter cycle from experiment to real-robot validation.

What You'll Do

  • Build distributed training, post-training, and RL infrastructure for VLM/VLA models.
  • Optimize training throughput, memory usage, communication efficiency, and resource utilization; own parallelism strategies, checkpointing, fault tolerance, and job scheduling.
  • Optimize model inference: serving, batching, KV cache, quantization, compilation, and kernel optimization.
  • Build the end-to-end R&D platform: experiment submission, resource scheduling, metrics monitoring, and model artifact management.
  • Optimize deployment efficiency, stability, and latency on both servers and robots.

What We're Looking For

  • Familiar with PyTorch, CUDA, Triton, NCCL, or related training/inference systems.
  • Understanding of one or more parallelism strategies: data, tensor, pipeline, sequence, or expert parallelism.
  • Able to locate system bottlenecks with profiling tools and validate optimizations experimentally.
  • Familiar with Linux, containers, cluster scheduling, distributed storage, and observability; performance-sensitive and willing to go deep into frameworks, runtimes, and kernels.