AI Infra 全栈工程师 / AI Infra Full-Stack Engineer(训练与推理 / Training & Inference)
岗位使命
建设高性能、高可靠、可扩展的模型训练和推理基础设施,让训练更快、推理延迟更低、算力利用率更高,并显著缩短从实验到真机验证的周期。
你将负责
- 建设 VLM/VLA 的分布式训练、Post-training 和 RL 基础设施。
- 优化训练吞吐、显存利用率、通信效率和计算资源利用率,负责并行策略、Checkpoint、容错恢复和任务调度。
- 优化模型推理,包括 Serving、Batching、KV Cache、量化、编译和 Kernel 优化。
- 建设从实验提交、资源调度、指标监控到模型产物管理的完整研发平台。
- 优化模型在服务器和机器人端的部署效率、稳定性和延迟。
我们希望你
- 熟悉 PyTorch、CUDA、Triton、NCCL 或相关训练与推理系统。
- 理解数据并行、张量并行、流水线并行、序列并行或专家并行中的一种或多种。
- 能够使用 Profiling 工具定位系统瓶颈,并通过实验验证优化效果。
- 熟悉 Linux、容器、集群调度、分布式存储和可观测性系统,对系统性能敏感,愿意深入框架、Runtime 和 Kernel 解决问题。
AI Infra Full-Stack Engineer (Training & Inference)
Mission
Build high-performance, reliable, and scalable training and inference infrastructure — faster training, lower inference latency, higher compute utilization, and a much shorter cycle from experiment to real-robot validation.
What You'll Do
- Build distributed training, post-training, and RL infrastructure for VLM/VLA models.
- Optimize training throughput, memory usage, communication efficiency, and resource utilization; own parallelism strategies, checkpointing, fault tolerance, and job scheduling.
- Optimize model inference: serving, batching, KV cache, quantization, compilation, and kernel optimization.
- Build the end-to-end R&D platform: experiment submission, resource scheduling, metrics monitoring, and model artifact management.
- Optimize deployment efficiency, stability, and latency on both servers and robots.
What We're Looking For
- Familiar with PyTorch, CUDA, Triton, NCCL, or related training/inference systems.
- Understanding of one or more parallelism strategies: data, tensor, pipeline, sequence, or expert parallelism.
- Able to locate system bottlenecks with profiling tools and validate optimizations experimentally.
- Familiar with Linux, containers, cluster scheduling, distributed storage, and observability; performance-sensitive and willing to go deep into frameworks, runtimes, and kernels.
