About
Previously, I obtained master degree in Machine Learning at the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), advised by Prof. Kun Zhang. I work closely with Guangyi Chen and Yunlong Deng in the MBZUAI Causality Group.
From July 2025 to December 2025, I was a visiting student in the Department of Philosophy at Carnegie Mellon University (CMU), working with Prof. Peter Spirtes and Prof. Kun Zhang, and collaborating with the CMU CLeaR Group. I obtained a B.Eng. in Computer Science and Technology from Tsinghua University in June 2024, where I worked at the Research Institute of Human–Computer Interaction and Media Integration with Chun Yu, mainly on HCI and context-aware systems.
My research primarily focuses on representation learning, particularly representation learning for foundation models. My broader goal is to improve the reliability and robustness of large language models, for example by reducing hallucinations and enabling stronger generalization.
I have published extensively on causal representation learning, with much of this work examining conventional adaptation problems through a causal perspective. My collaborators and I have also constructed relevant datasets and authored position papers that aim to clarify important research directions in this area.
Since 2026, my research has increasingly shifted toward improving the capabilities of LLMs through the lens of representation learning. I developed MPM, a lossless approach for extending multimodal large language models with additional modalities. I am also collaborating with Prof. Sean Du on noise detection in LLMs, particularly in the reasoning trajectories produced by agents.
Another major direction of my current work is improving the long-horizon capabilities of agents. Instead of treating context as a single flat sequence of tokens, we organize it into structured blocks. During generation, the model can emit dedicated editing tokens from its vocabulary to perform atomic operations on specific blocks, including deletion, rewriting, and exact restoration by block identifier.
News
- 2026.01 - “Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens” accepted to ICLR 2026 (poster).
- 2025.12 — I will attend NeurIPS 2025 in San Diego and present two posters. Feel free to talk to me at the poster sessions.
- 2025.09 — “CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations” accepted to NeurIPS 2025 (spotlight).
- 2025.09 — “Towards Self-Refinement of Vision-Language Models with Triangular Consistency” accepted to NeurIPS 2025 (poster).
- 2025.04 — “A General Representation-Based Approach to Multi-Source Domain Adaptation” presented at ICML 2025 (poster).
- 2025.07–2025.12 — Visiting student at Carnegie Mellon University, Department of Philosophy.
- 2024.10 — Started the M.Sc. in Machine Learning at MBZUAI, advised by Kun Zhang.
- 2024.06 — Graduated from Tsinghua University (B.Eng. in Computer Science and Technology).
- 2023.06–2023.07 — Research intern at Mian Bi Intelligent Technology Co., Ltd., working on pre-training data preparation for large language models.
- 2022 — Third Prize, 40th Challenge Cup of Tsinghua University.
- 2020 — Second-Class Scholarship for Freshmen, Tsinghua University.
Publications
Equal contribution is indicated by * after author names.
Featured paper
|
CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations Guangyi Chen*, Yunlong Deng*, Peiyuan Zhu*, Yan Li*, Yifan Shen, Zijian Li, Kun Zhang. NeurIPS 2025 Datasets and Benchmarks Track, Spotlight. |
Full papers
-
A Dialogue between Causal and Traditional Representation Learning: Toward Mutual Benefits in a Unified Formulation
Yan Li, Yuewen Sun, Shaoan Xie, Gongxu Luo, Yunlong Deng, Kun Zhang, Guangyi Chen. -
Multimodal LLMs under Pairwise Modalities
Yan Li*, Yunlong Deng*, Yuewen Sun, Gongxu Luo, Kun Zhang, Guangyi Chen. -
Should Bias Always Be Eliminated? A Principled Framework to Use Data Bias for OOD Generation
Yan Li, Yunlong Deng, Zijian Li, Anpeng Wu, Zeyu Tang, Kun Zhang, Guangyi Chen. -
Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
Yunlong Deng, Boyang Sun, Yan Li, Zeyu Tang, Lingjing Kong, Kun Zhang, Guangyi Chen.
ICLR 2026, Poster. -
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
Yunlong Deng, Guangyi Chen, Tianpei Gu, Lingjing Kong, Yan Li, Zeyu Tang, Kun Zhang.
NeurIPS 2025, Poster. -
A General Representation-Based Approach to Multi-Source Domain Adaptation
Ignavier Ng*, Yan Li*, Zijian Li, Yujia Zheng, Guangyi Chen, Kun Zhang.
ICML 2025, Poster. -
ProtoGS: Efficient and High-Quality Rendering with 3D Gaussian Prototypes
Zhengqing Gao, Dongting Hu, Jia-Wang Bian, Huan Fu, Yan Li, Tongliang Liu, Mingming Gong, Kun Zhang. -
Revealing Personalized Positions in Indoor Spaces: A User-Centric Approach to Context-Aware Area Discovery
Zhaoheng Li, Chun Yu, Yan Li, Yuanchun Shi.
Undergraduate work on human–computer interaction and context-aware area discovery.
Industry experience
Mian Bi Intelligent Technology Co., Ltd. · Jun. 2023 – Jul. 2023
Supervisor: Jie Cai, Algorithm Engineer
- Worked on data preparation for pre-training large language models, including data-cleaning algorithms and large-scale batch cleaning.
- Participated in discussions on classical text LLMs, data pipelines, and preliminary design.
Education
-
M.Sc. in Machine Learning, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · Oct. 2024 – Jun. 2026
Supervisor: Kun Zhang -
Visiting student, Department of Philosophy, Carnegie Mellon University (CMU) · Jul. 2025 – Dec. 2025
Visiting tutors: Peter Spirtes, Kun Zhang -
B.Eng. in Computer Science and Technology, Tsinghua University · Sept. 2020 – Jun. 2024
Honors and awards
- MBZUAI Conference Travel Scholarship, 2025
- Tsinghua University Academic Progress Scholarship, 2023
- Social Practice Excellence Scholarship of Tsinghua University, 2022
- Third Prize, 40th Challenge Cup of Tsinghua University, 2022
- Second-Class Scholarship for Freshmen, Tsinghua University, 2020
Service
- Reviewer: ICLR, ICML, NeurIPS
Contact
- Email: lyan012010 (at) gmail (dot) com