Yanjun Chen

PhD Candidate, Department of Computing, The Hong Kong Polytechnic University.香港理工大学 计算学系 博士候选人。香港理工大学 コンピューティング学科 博士候補者。

prof_pic.jpg

Yanjun Chen

PhD Candidate, PolyU博士候选人 · 理大博士候補者 · PolyU

I am a PhD candidate in the Department of Computing at The Hong Kong Polytechnic University, advised by Prof. Wenjie Li (Maggie) and Prof. Wei Zhang, with joint doctoral training at the Eastern Institute of Technology (EIT), Ningbo.

In reinforcement learning, models are trainable. The environments that train them are not. I want to make the environment trainable, the way models are, and with it to lift the ceiling of what AI can become.

Research

What does the environment actually do to the model it trains?

Reward-model evaluation in RLHF. A more accurate reward model does not always train a better policy.

The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models (EMNLP 2024).

Exact credit for cooperative LLM agents. When agents cooperate, a shared outcome hides what each decision contributed.

Exact Is Easier: Credit Assignment for Cooperative LLM Agents (arXiv:2603.06859, in submission).

Withdrawable shaping on the action interface. Shaping aids are added to help the agent learn, then kept forever.

Under review (2026).

Where I’m Going

The destination: an environment that learns alongside the model it trains, from language to embodied agents. The path runs through credit: credit should not stop at the agent’s boundary. Once trustworthy learning signals reach the pieces around the model, the environment can begin to learn.

With thanks to Xiaoyu Shen and Dawei Zhu, whose ongoing mentorship and guidance have shaped much of how I think about research.

我是香港理工大学 计算学系的博士候选人,师从 Wenjie Li (Maggie) 教授与 Wei Zhang 教授,并在东方理工(EIT,宁波)联合培养。

在强化学习里,模型是可以训练的。训练模型的环境,还不行。我想让环境也变得可训练,像模型一样,并以此把 AI 的上限抬上去。

研究方向

环境到底对它训练的模型做了什么?

RLHF 里 reward model 的评估。 更准的 reward model,不一定训出更好的 policy。

The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models (EMNLP 2024).

协作 LLM agent 的精确 credit。 多个 agent 协作时,共享的结果把每个决策的真实贡献藏了起来。

Exact Is Easier: Credit Assignment for Cooperative LLM Agents (arXiv:2603.06859, in submission).

action 接口上可撤回的 shaping。 Shaping 辅助是为了帮 agent 学习才加上的,却从此永远留了下来。

Under review (2026).

往哪里去

终局:环境与模型一起学习,从语言走向具身智能体。去那里的路,经过 credit assignment(信用分配):credit 不该停在 agent 的边界上。 当可信的学习信号能到达模型周围的组件,环境就能开始学习。

感谢 Xiaoyu Shen 老师与 Dawei Zhu 师兄一直以来的指导与帮助,他们在很多方面塑造了我做研究的方式。

香港理工大学 コンピューティング学科の博士候補者で、Wenjie Li (Maggie) 教授と Wei Zhang 教授の指導のもと、東方理工(EIT、寧波)との共同育成プログラムに参加しています。

強化学習において、モデルは訓練できる。モデルを訓練する環境は、まだできない。私はその環境を、モデルと同じように訓練できるものにしたい。そしてそれによって、AI の到達点を引き上げたいのです。

研究内容

環境は、訓練するモデルに実際のところ何をしているのか。

RLHF における reward model の評価。 より精度の高い reward model が、必ずしもより良い policy を訓練するわけではない。

The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models (EMNLP 2024).

協調的 LLM agent の厳密な credit。 Agent が協調するとき、共有された結果は各決定の実際の寄与を隠してしまう。

Exact Is Easier: Credit Assignment for Cooperative LLM Agents (arXiv:2603.06859, in submission).

action interface 上の撤回可能な shaping。 Shaping 補助は agent の学習を助けるために加えられ、そのまま永久に残される。

Under review (2026).

これから

目指す先は、モデルとともに学ぶ環境です。言語から身体性エージェントへ。そこへ至る道は credit assignment(貢献度の割り当て)にあります。credit は agent の境界で止まるべきではない。 信頼できる学習信号がモデルの周りの構成要素にまで届いたとき、環境は学び始めることができます。

Xiaoyu Shen 先生と Dawei Zhu 先輩から受けた継続的なご指導とお力添えに深く感謝いたします。私の研究との向き合い方の多くは、お二人からの影響によるものです。

News近况お知らせ

Jun 24, 20262026年6月24日2026年6月24日 Passed the confirmation of candidature for my PhD at The Hong Kong Polytechnic University, with the thesis Towards Efficient Reinforcement Learning via Environment Measurement and Shaping.通过了香港理工大学的博士候选人资格审核,论文题目为 Towards Efficient Reinforcement Learning via Environment Measurement and Shaping香港理工大学の博士候補者資格審査に合格しました。論文題目は Towards Efficient Reinforcement Learning via Environment Measurement and Shaping です。
May 22, 20262026年5月22日2026年5月22日 Shortlisted for PolyU Micro Fund 2025/26 Cohort 2 (HK$20,000 cash prize), with a conditional offer to the HKSTP Ideation Programme.入围 PolyU Micro Fund 2025/26 第二轮(HK$20,000 现金奖励),并获得 HKSTP Ideation Programme 的有条件录取。PolyU Micro Fund 2025/26 Cohort 2(賞金 HK$20,000)にショートリスト入り、HKSTP Ideation Programme に条件付きで内定。
May 08, 20262026年5月8日2026年5月8日 Released v2 of Exact Is Easier: Credit Assignment for Cooperative LLM Agents on arXiv:2603.06859.Exact Is Easier: Credit Assignment for Cooperative LLM Agents v2 已发布至 arXiv:2603.06859Exact Is Easier: Credit Assignment for Cooperative LLM Agents v2 を arXiv:2603.06859 で公開。
Mar 06, 20262026年3月6日2026年3月6日 First arXiv release of Exact Is Easier: Credit Assignment for Cooperative LLM Agents (in submission).Exact Is Easier: Credit Assignment for Cooperative LLM Agents 在 arXiv 首次发布(投稿中)。Exact Is Easier: Credit Assignment for Cooperative LLM Agents を arXiv に初回公開(投稿中)。
May 22, 20252025年5月22日2025年5月22日 Co-authored a comprehensive survey on latent chain-of-thought reasoning (arXiv:2505.16782).合著的隐式思维链推理(latent chain-of-thought reasoning)综述发布于 arXiv:2505.16782latent chain-of-thought reasoning に関する包括的なサーベイ論文を共著として発表(arXiv:2505.16782)。
May 15, 20252025年5月15日2025年5月15日 Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning accepted at ACL 2025 Findings (co-author).Unveiling the Key Factors for Distilling Chain-of-Thought ReasoningACL 2025 Findings 接收(共同作者)。Unveiling the Key Factors for Distilling Chain-of-Thought ReasoningACL 2025 Findings に採択(共著)。
Jan 15, 20252025年1月15日2025年1月15日 Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation accepted at the NAACL 2025 Student Research Workshop (co-author).Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine TranslationNAACL 2025 学生研究工作坊接收(合作论文)。Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine TranslationNAACL 2025 Student Research Workshop に採択(共著)。

Selected Publications代表性论文主要論文

  1. ACL Findings
    Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning
    Xinghao Chen, Zhijing Sun, Wenjin Guo, and 8 more authors
    In Findings of the Association for Computational Linguistics: ACL 2025, 2025
  2. EMNLP
    The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models
    Yanjun Chen, Dawei Zhu, Yirong Sun, and 3 more authors
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024
  3. arXiv
    Exact Is Easier: Credit Assignment for Cooperative LLM Agents
    Yanjun Chen, Yirong Sun, Hanlin Wang, and 5 more authors
    arXiv preprint arXiv:2603.06859, 2026
    In submission.
  4. arXiv
    Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
    Xinghao Chen, Anhao Zhao, Heming Xia, and 7 more authors
    arXiv preprint arXiv:2505.16782, 2025

Teaching & Service教学与服务教育・学会活動

Teaching教学教育 (Teaching Assistant, The Hong Kong Polytechnic University)(助教,香港理工大学)(ティーチング・アシスタント、香港理工大学)

2025/26 S2 COMP5221  Software Project Management
2025/26 S1 COMP1002  Computational Thinking and Problem Solving
2024/25 S2 COMP1411  Introduction to Computer Systems
2024/25 S1 COMP5567  Distributed Algorithms and Protocols for Blockchains

Reviewing审稿査読

2026 ACL ARR (March and May cycles)(3 月与 5 月周期)(3 月・5 月サイクル)
2026 IEEE Transactions on Neural Networks and Learning Systems (TNNLS), invited受邀招待