News 近况 お知らせ

Sep 28, 20262026年9月28日2026年9月28日 Started a research internship at Microsoft Research Asia, Singapore.加入微软亚洲研究院(新加坡),担任研究实习生。Microsoft Research Asia(シンガポール)でリサーチインターンを開始。
Sep 26, 20262026年9月26日2026年9月26日 Released Action Shaping: Policies Absorb What They Can Express on arXiv:2609.32752.Action Shaping: Policies Absorb What They Can Express 发布于 arXiv:2609.32752。Action Shaping: Policies Absorb What They Can Express を arXiv:2609.32752 で公開。
Sep 24, 20262026年9月24日2026年9月24日 FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control accepted at NeurIPS 2026 (co-author).FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control 被 NeurIPS 2026 接收(共同作者)。FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control が NeurIPS 2026 に採択(共著)。
Aug 20, 20262026年8月20日2026年8月20日 Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning accepted at EMNLP 2026 Findings (co-author).Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning 被 EMNLP 2026 Findings 接收(共同作者)。Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning が EMNLP 2026 Findings に採択(共著)。
Jun 24, 20262026年6月24日2026年6月24日 Passed the confirmation of candidature for my PhD at The Hong Kong Polytechnic University, with the thesis Towards Efficient Reinforcement Learning via Environment Measurement and Shaping.通过了香港理工大学的博士候选人资格审核,论文题目为 Towards Efficient Reinforcement Learning via Environment Measurement and Shaping。香港理工大学の博士候補者資格審査に合格しました。論文題目は Towards Efficient Reinforcement Learning via Environment Measurement and Shaping です。
May 22, 20262026年5月22日2026年5月22日 Shortlisted for PolyU Micro Fund 2025/26 Cohort 2 (HK$20,000 cash prize), with a conditional offer to the HKSTP Ideation Programme.入围 PolyU Micro Fund 2025/26 第二轮(HK$20,000 现金奖励),并获得 HKSTP Ideation Programme 的有条件录取。PolyU Micro Fund 2025/26 Cohort 2(賞金 HK$20,000)にショートリスト入り、HKSTP Ideation Programme に条件付きで内定。
Mar 06, 20262026年3月6日2026年3月6日 Released The Trace Is the State: Exact Credit Assignment for LLM Agent Teams on arXiv:2603.06859.The Trace Is the State: Exact Credit Assignment for LLM Agent Teams 发布于 arXiv:2603.06859。The Trace Is the State: Exact Credit Assignment for LLM Agent Teams を arXiv:2603.06859 で公開。
May 22, 20252025年5月22日2025年5月22日 Co-authored a comprehensive survey on latent chain-of-thought reasoning (arXiv:2505.16782).合著的隐式思维链推理(latent chain-of-thought reasoning)综述发布于 arXiv:2505.16782。latent chain-of-thought reasoning に関する包括的なサーベイ論文を共著として発表(arXiv:2505.16782)。
May 15, 20252025年5月15日2025年5月15日 Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning accepted at ACL 2025 Findings (co-author).Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning 被 ACL 2025 Findings 接收(共同作者)。Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning が ACL 2025 Findings に採択(共著)。
Jan 15, 20252025年1月15日2025年1月15日 Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation accepted at the NAACL 2025 Student Research Workshop (co-author).Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation 被 NAACL 2025 学生研究工作坊接收(合作论文)。Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation が NAACL 2025 Student Research Workshop に採択(共著)。
Oct 09, 20242024年10月9日2024年10月9日 The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models accepted at EMNLP 2024.The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models 被 EMNLP 2024 接收。The Accuracy Paradox in RLHF: When Better Reward Models Don’t Yield Better Language Models が EMNLP 2024 に採択。