Yanjun Chen

PhD Candidate, Department of Computing, The Hong Kong Polytechnic University.香港理工大学 计算学系 博士候选人。香港理工大学 コンピューティング学科 博士候補者。

prof_pic.jpg

Yanjun Chen

PhD Candidate, PolyU博士候选人 · 理大博士候補者 · PolyU

I am a PhD candidate in the Department of Computing at The Hong Kong Polytechnic University, advised by Prof. Wenjie Li (Maggie) and Prof. Wei Zhang, with joint doctoral training at the Eastern Institute of Technology (EIT), Ningbo. I am currently a research intern at Microsoft Research Asia, Singapore.

In reinforcement learning, models are trainable. The environments that train them are not. I want to make the environment trainable, the way models are, and with it to lift the ceiling of what AI can become.

Research

Making the pieces around a model trainable starts with knowing what they actually do in training, and which signal can say exactly what each step is worth.

Reward-model accuracy and training outcomes in RLHF. The reward model is the most typical piece of a training environment and is usually judged by its accuracy. Yet moderately accurate reward models train better language models than the most accurate ones, so a piece’s worth has to be judged inside training.

The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models (EMNLP 2024).

Exact credit for LLM agent teams. What each message in a team of LLM agents was worth has mostly been predicted, but when the team communicates through a shared context and everything a downstream agent reads is written into the trace, the trace is the state. One message can then be replaced and the run continued to its end, which makes its credit exact. For an LLM, much of the environment arrives as context (retrieved text, tool outputs, prompts), and that is where the same method points next: assigning credit to those pieces as well.

The Trace Is the State: Exact Credit Assignment for LLM Agent Teams (arXiv:2603.06859, in submission).

When a policy absorbs action shaping. Reward shaping has a theorem guaranteeing that a potential-based term can be removed; the same practice on the action channel, a training-time offset, has none. A trainable policy absorbs an offset its own output layer can reproduce exactly, and once absorbed, the offset can be removed with the return almost unchanged.

Action Shaping: Policies Absorb What They Can Express (arXiv:2609.32752, in submission).

Where I’m Going

The destination: an environment that learns alongside the model it trains, from language to embodied agents. The path runs through credit: credit should not stop at the agent’s boundary. Once trustworthy learning signals reach the pieces around the model, the environment can begin to learn.

With thanks to Xiaoyu Shen and Dawei Zhu, whose ongoing mentorship and guidance have shaped much of how I think about research.

我是香港理工大学 计算学系的博士候选人,师从 Wenjie Li (Maggie) 教授与 Wei Zhang 教授,并在东方理工(EIT,宁波)联合培养。目前在微软亚洲研究院(新加坡)做研究实习。

在强化学习里,模型是可以训练的。训练模型的环境,还不行。我想让环境也变得可训练,像模型一样,并以此把 AI 的上限抬上去。

研究方向

要让模型周围的组件也能训练,先得知道它们在训练里究竟起什么作用,以及什么信号能精确告诉我们每一步值多少。

RLHF 中 reward model 的准确率与训练效果。 reward model 是训练环境里最典型的组件,通常按准确率评判。但中等准确率的 reward model 反而比最准的训出更好的语言模型,所以组件值多少,得放进训练里看。

The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models (EMNLP 2024).

LLM agent 团队的精确 credit。 LLM agent 团队里每条消息值多少,过去大多靠预测;但当团队通过共享上下文交流、下游 agent 读到的一切都写进记录(trace)时,记录就是状态。这时换掉一条消息、真实续跑到结束,它的 credit 就能精确算出。对 LLM 来说,环境大多以上下文的形式到达模型(检索到的文本、工具返回、提示词),同一个办法接下来指向的正是这些组件:给它们也分配 credit。

The Trace Is the State: Exact Credit Assignment for LLM Agent Teams (arXiv:2603.06859, in submission).

action shaping 何时被 policy 吸收。 reward shaping 有定理保证基于势函数的项可以拿掉,action channel 上的同类做法(训练时加的偏移)没有。可训练的 policy 会吸收它自己输出层能精确复现的偏移,吸收后拿掉,回报几乎不变。

Action Shaping: Policies Absorb What They Can Express (arXiv:2609.32752, in submission).

往哪里去

终局:环境与模型一起学习,从语言走向具身智能体。去那里的路,经过 credit assignment(信用分配):credit 不该停在 agent 的边界上。 当可信的学习信号能到达模型周围的组件,环境就能开始学习。

感谢 Xiaoyu Shen 老师与 Dawei Zhu 师兄一直以来的指导与帮助,他们在很多方面塑造了我做研究的方式。

香港理工大学 コンピューティング学科の博士候補者で、Wenjie Li (Maggie) 教授と Wei Zhang 教授の指導のもと、東方理工(EIT、寧波)との共同育成プログラムに参加しています。現在、Microsoft Research Asia(シンガポール)のリサーチインターンでもあります。

強化学習において、モデルは訓練できる。モデルを訓練する環境は、まだできない。私はその環境を、モデルと同じように訓練できるものにしたい。そしてそれによって、AI の到達点を引き上げたいのです。

研究内容

モデルの周りの構成要素まで訓練できるようにするには、まず、それらが訓練の中で実際に何をしているのか、そして一歩ごとの価値をどの信号なら正確に言えるのかを知る必要がある。

RLHF における reward model の精度と訓練結果。 reward model は訓練環境の中で最も典型的な構成要素で、ふつうは精度で評価される。しかし、中程度の精度の reward model のほうが最も精度の高いものより良い言語モデルを訓練するので、構成要素の価値は訓練の内側で判断する必要がある。

The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models (EMNLP 2024).

LLM agent チームの厳密な credit。 LLM agent のチームで各メッセージの価値はこれまで主に予測されてきたが、チームが共有コンテキストを通じてやり取りし、下流の agent が読むものがすべて記録(trace)に書き込まれるなら、記録がそのまま状態になる。このとき一つのメッセージを差し替えて最後まで実際に走らせれば、その credit は厳密に求まる。LLM にとって環境の多くはコンテキストとして届く(検索されたテキスト、ツールの返り値、プロンプト)ので、同じ方法が次に向かう先は、こうした構成要素にも credit を割り当てることだ。

The Trace Is the State: Exact Credit Assignment for LLM Agent Teams (arXiv:2603.06859, in submission).

action shaping が policy に吸収されるとき。 reward shaping には、ポテンシャルに基づく項を取り除けることを保証する定理があるが、action channel 上の同じ実践(訓練時に加えるオフセット)にはない。訓練可能な policy は、自分の出力層が正確に再現できるオフセットを吸収し、吸収された後に外してもリターンはほとんど変わらない。

Action Shaping: Policies Absorb What They Can Express (arXiv:2609.32752, in submission).

これから

目指す先は、モデルとともに学ぶ環境です。言語から身体性エージェントへ。そこへ至る道は credit assignment(貢献度の割り当て)にあります。credit は agent の境界で止まるべきではない。 信頼できる学習信号がモデルの周りの構成要素にまで届いたとき、環境は学び始めることができます。

Xiaoyu Shen 先生と Dawei Zhu 先輩から受けた継続的なご指導とお力添えに深く感謝いたします。私の研究との向き合い方の多くは、お二人からの影響によるものです。

News近况お知らせ

Sep 28, 20262026年9月28日2026年9月28日 Started a research internship at Microsoft Research Asia, Singapore.加入微软亚洲研究院(新加坡),担任研究实习生。Microsoft Research Asia(シンガポール)でリサーチインターンを開始。
Sep 26, 20262026年9月26日2026年9月26日 Released Action Shaping: Policies Absorb What They Can Express on arXiv, together with v3 of Exact Is Easier, now titled The Trace Is the State: Exact Credit Assignment for LLM Agent Teams (both in submission).Action Shaping: Policies Absorb What They Can Express 在 arXiv 首次发布;Exact Is Easier 同日更新至 v3,并更名为 The Trace Is the State: Exact Credit Assignment for LLM Agent Teams(两篇均在投稿中)。Action Shaping: Policies Absorb What They Can Express を arXiv に初回公開し、同日、Exact Is Easier を The Trace Is the State: Exact Credit Assignment for LLM Agent Teams に改題した v3 も公開(いずれも投稿中)。
Sep 24, 20262026年9月24日2026年9月24日 FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control accepted at NeurIPS 2026 (co-author).FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control 被 NeurIPS 2026 接收(共同作者)。FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control が NeurIPS 2026 に採択(共著)。
Aug 20, 20262026年8月20日2026年8月20日 Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning accepted at EMNLP 2026 Findings (co-author).Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning 被 EMNLP 2026 Findings 接收(共同作者)。Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning が EMNLP 2026 Findings に採択(共著)。
Jun 24, 20262026年6月24日2026年6月24日 Passed the confirmation of candidature for my PhD at The Hong Kong Polytechnic University, with the thesis Towards Efficient Reinforcement Learning via Environment Measurement and Shaping.通过了香港理工大学的博士候选人资格审核,论文题目为 Towards Efficient Reinforcement Learning via Environment Measurement and Shaping。香港理工大学の博士候補者資格審査に合格しました。論文題目は Towards Efficient Reinforcement Learning via Environment Measurement and Shaping です。
May 22, 20262026年5月22日2026年5月22日 Shortlisted for PolyU Micro Fund 2025/26 Cohort 2 (HK$20,000 cash prize), with a conditional offer to the HKSTP Ideation Programme.入围 PolyU Micro Fund 2025/26 第二轮(HK$20,000 现金奖励),并获得 HKSTP Ideation Programme 的有条件录取。PolyU Micro Fund 2025/26 Cohort 2(賞金 HK$20,000)にショートリスト入り、HKSTP Ideation Programme に条件付きで内定。
May 08, 20262026年5月8日2026年5月8日 Released v2 of Exact Is Easier: Credit Assignment for Cooperative LLM Agents on arXiv:2603.06859.Exact Is Easier: Credit Assignment for Cooperative LLM Agents v2 已发布至 arXiv:2603.06859。Exact Is Easier: Credit Assignment for Cooperative LLM Agents v2 を arXiv:2603.06859 で公開。

Selected Publications代表性论文主要論文

  1. arXiv
    Action Shaping: Policies Absorb What They Can Express
    Yanjun Chen, Jinghan Wang, Xiaoyu Shen, and 2 more authors
    arXiv preprint arXiv:2609.32752, 2026
    In submission.
  2. arXiv
    The Trace Is the State: Exact Credit Assignment for LLM Agent Teams
    Yanjun Chen, Yirong Sun, Hanlin Wang, and 5 more authors
    arXiv preprint arXiv:2603.06859, 2026
    In submission.
  3. EMNLP Findings
    Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
    Xinghao Chen, Anhao Zhao, Heming Xia, and 7 more authors
    In Findings of the Association for Computational Linguistics: EMNLP 2026, 2026
  4. ACL Findings
    Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning
    Xinghao Chen, Zhijing Sun, Wenjin Guo, and 8 more authors
    In Findings of the Association for Computational Linguistics: ACL 2025, 2025
  5. EMNLP
    The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models
    Yanjun Chen, Dawei Zhu, Yirong Sun, and 3 more authors
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024

Teaching & Service教学与服务教育・学会活動

Teaching教学教育 (Teaching Assistant, The Hong Kong Polytechnic University)(助教,香港理工大学)(ティーチング・アシスタント、香港理工大学)

2025/26 S2 COMP5221  Software Project Management
2025/26 S1 COMP1002  Computational Thinking and Problem Solving
2024/25 S2 COMP1411  Introduction to Computer Systems
2024/25 S1 COMP5567  Distributed Algorithms and Protocols for Blockchains

Reviewing审稿査読

2026 ACL Rolling Review (ARR)
2026 IEEE Transactions on Neural Networks and Learning Systems (TNNLS)