Navigator, a web agentNavigator 网页智能体
Built the RL model for Navigator, a state-of-the-art AI agent that operates web browsers.
为 Navigator 训练了强化学习模型。Navigator 是一个领先的、能自主操作浏览器的 AI 智能体。
Freelance AI engineer and robot builder.
LifeRPG player.
独立 AI 工程师,学造机器人中。
LifeRPG 玩家。
I built multimodal foundation models and AI agents: Llama 3 Vision post-training at Meta and the reinforcement learning (RL) for the Navigator web agent at Yutori. Now I am a freelancer, learning/building robots. I hope everyone has a humanoid companion by the time I retire.
That intelligence correlates with is a law of this universe.
我曾做多模态基础模型与 AI 智能体:在 Meta 参与 Llama 3 Vision 后训练,在 Yutori 做强化学习(RL)训练 Navigator 网页智能体。现正在独立学习/制造机器人,希望退休前让所有人都有机器人伙伴。
智能和 正相关是本宇宙的一个定律。
Built the RL model for Navigator, a state-of-the-art AI agent that operates web browsers.
为 Navigator 训练了强化学习模型。Navigator 是一个领先的、能自主操作浏览器的 AI 智能体。
Co-led post-training for Llama 3 Vision — SFT, DPO, reward modeling, and rejection sampling. Worked on multimodal data, SFT, and reward modeling for Llama 4.
共同负责 Llama 3 Vision 的后训练(SFT、DPO、奖励模型、拒绝采样),并参与 Llama 4 的多模态数据、SFT 与奖励模型。
Led the production of Emu Edit, a diffusion model for free-form image editing from text instructions.
主导 Emu Edit(根据文字指令自由编辑图像的扩散模型)的产品化。
Shipped face tracking on Meta Quest Pro — system design, data collection and annotation, distributed training and evaluation.
在 Meta Quest Pro 上发布面部追踪。负责系统设计、数据采集与标注、分布式训练与评测。


CVPR 2017 · IEEE Conference on Computer Vision and Pattern RecognitionSpotlight
@inproceedings{xiaoli2016end,
title = {Joint Detection and Identification Feature Learning for Person Search},
author = {Tong Xiao and Shuang Li and Bochao Wang and Liang Lin and Xiaogang Wang},
booktitle = {IEEE Conference on Computer Vision and Pattern Recognition},
year = {2017}
}
CVPR 2017 · IEEE Conference on Computer Vision and Pattern Recognition
@inproceedings{li2017person,
title = {Person Search with Natural Language Description},
author = {Shuang Li and Tong Xiao and Hongsheng Li and Bolei Zhou and Dayu Yue and Xiaogang Wang},
booktitle = {IEEE Conference on Computer Vision and Pattern Recognition},
year = {2017}
}
CVPR 2015 · IEEE Conference on Computer Vision and Pattern Recognition
@inproceedings{xiao2015learning,
title = {Learning from Massive Noisy Labeled Data for Image Classification},
author = {Tong Xiao and Tian Xia and Yi Yang and Chang Huang and Xiaogang Wang},
booktitle = {IEEE Conference on Computer Vision and Pattern Recognition},
year = {2015}
}
CVPR 2014 · IEEE Conference on Computer Vision and Pattern Recognition
@inproceedings{li2014filter,
title = {DeepReID: Deep Filter Pairing Neural Network for Person Re-Identification},
author = {Wei Li and Rui Zhao and Tong Xiao and Xiaogang Wang},
booktitle = {IEEE Conference on Computer Vision and Pattern Recognition},
year = {2014}
}
There are too many robot companies. We need a hub to better understand them. Check out this Robot Overflow site if you have the following questions:
机器人公司太多了,我们需要一个地方来更好地了解它们。如果你有下面这些问题,来 Robot Overflow 看看:
What does Real2Sim mean for robots? Here, we break down the system and components for you and show the value for each level of capability, from L1 to L5.
Real2Sim 对机器人意味着什么?我们在这里为你拆解整个系统和各个组件,并展示从 L1 到 L5 每一级能力的价值。