$ whoami

Tong Xiao肖桐

Freelance AI engineer and robot builder. LifeRPG player.

独立 AI 工程师,学造机器人中。LifeRPG 玩家。

I built multimodal foundation models and AI agents: Llama 3 Vision post-training at Meta and the reinforcement learning (RL) for the Navigator web agent at Yutori. Now I am a freelancer, learning/building robots. I hope everyone has a humanoid companion by the time I retire.

That intelligence correlates with log⁡(effective compute) is a law of this universe.

我曾做多模态基础模型与 AI 智能体:在 Meta 参与 Llama 3 Vision 后训练,在 Yutori 做强化学习(RL)训练 Navigator 网页智能体。现正在独立学习/制造机器人,希望退休前让所有人都有机器人伙伴。

智能和 log⁡(有效算力) 正相关是本宇宙的一个定律。

CV简历Google Scholar谷歌学术tong.xiao.work [at] gmail.com
  • Robot Overflow is live: a hub where we can better understand the robot industry.Robot Overflow 上线了。它能帮助我们更好地了解机器人产业。
  • I made Robot Real2Sim to better understand what the Real2Sim system is and its potential value.我制作了Robot Real2Sim 来更好地了解什么是 Real2Sim 以及它的潜在价值。
  • I left Yutori to spend more time with family.我离开了 Yutori,和家人多一些时间相处。
  • We introduced Yutori Navigator, a state-of-the-art AI web agent. I built the RL model for it.我们推出了最先进的 AI 网页智能体 Yutori Navigator。我负责模型的强化学习。
  • I joined Yutori to build AI agents for the web.我加入了 Yutori,致力于构建网页 AI 智能体。

Selected work代表工作

Yutori2025–2026

Navigator, a web agentNavigator 网页智能体

Built the RL model for Navigator, a state-of-the-art AI agent that operates web browsers.

为 Navigator 训练了强化学习模型。Navigator 是一个领先的、能自主操作浏览器的 AI 智能体。

Meta2024

Llama 3 Vision post-trainingLlama 3 Vision 后训练

Co-led post-training for Llama 3 Vision — SFT, DPO, reward modeling, and rejection sampling. Worked on multimodal data, SFT, and reward modeling for Llama 4.

共同负责 Llama 3 Vision 的后训练(SFT、DPO、奖励模型、拒绝采样),并参与 Llama 4 的多模态数据、SFT 与奖励模型。

Meta2023

Emu Edit, image editing with textEmu Edit 文本驱动图像编辑

Led the production of Emu Edit, a diffusion model for free-form image editing from text instructions.

主导 Emu Edit(根据文字指令自由编辑图像的扩散模型)的产品化。

Meta2022

Face tracking on Quest ProQuest Pro 面部追踪

Shipped face tracking on Meta Quest Pro — system design, data collection and annotation, distributed training and evaluation.

在 Meta Quest Pro 上发布面部追踪。负责系统设计、数据采集与标注、分布式训练与评测。

The Llama 3 Herd of Models

Llama team, core contributor

arXiv 2024

Joint Detection and Identification Feature Learning for Person Search

Tong Xiao*, Shuang Li*, Bochao Wang, Liang Lin, Xiaogang Wang

CVPR 2017 · IEEE Conference on Computer Vision and Pattern RecognitionSpotlight

Person Search with Natural Language Description

Shuang Li, Tong Xiao, Hongsheng Li, Bolei Zhou, Dayu Yue, Xiaogang Wang

CVPR 2017 · IEEE Conference on Computer Vision and Pattern Recognition

Learning from Massive Noisy Labeled Data for Image Classification

Tong Xiao, Tian Xia, Yi Yang, Chang Huang, Xiaogang Wang

CVPR 2015 · IEEE Conference on Computer Vision and Pattern Recognition

DeepReID: Deep Filter Pairing Neural Network for Person Re-Identification

Wei Li, Rui Zhao, Tong Xiao, Xiaogang Wang

CVPR 2014 · IEEE Conference on Computer Vision and Pattern Recognition