About Me

Hello! I'm Zixuan Gong. I am a third-year Ph.D. student at the Gaoling School of Artificial Intelligence (GSAI), Renmin University of China, advised by Prof. Yong Liu and Prof. Jiaye Teng. My research broadly focuses on LLM theory, driven at its core by a simple curiosity: how do foundation models actually learn, and why do they work so well? My long-term goal is to uncover the fundamental principles governing their behavior. By bridging theory and practice, I hope to translate these insights into the design and development of more principled architectures, training paradigms, and optimization algorithms.

Research

LLM Theory

Understanding the architecture, optimization, and generalization principles behind large language models.

Looped Transformers MoE Training Dynamics LLM Optimizer In-Context Learning Scaling Law

LLM Safety

Studying intrinsic risks and emergent behaviors that affect the reliability of LLMs.

Intrinsic Bias Self-Awareness

Machine Learning Theory

Exploring the theoretical foundations of machine learning.

Generalization Federated Learning

Currently, I am interested in the design and development of next-generation optimizers. I am exploring optimization algorithms and paradigms such as Muon and the muP principles, with a clear and simple goal—making large-scale training fundamentally more efficient, stable, and scalable.

Publications

Preprint (2025) loop

What Makes Looped Transformers Perform Better Than Non-Recursive Ones

Zixuan Gong, Yong Liu, Jiaye Teng

Preprint

Paper
Preprint (2025) optimization

Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers

Zixuan Gong, Shijia Li, Yong Liu, Jiaye Teng

Preprint

Paper
ICLR 2025 AR-NTP

Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization

Zixuan Gong*, Xiaolin Hu*, Huayi Tang, Yong Liu

* Equal contribution

The 13th International Conference on Learning Representations (ICLR 2025)

Paper
AAAI 2025 (Oral) AAAI

Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology

Xiaolin Hu*, Zixuan Gong*, Gengze Xu, Wei Liu, Jian Luan, Bin Wang, Yong Liu

* Equal contribution

The 39th Annual AAAI Conference on Artificial Intelligence (AAAI 2025)

Paper
Preprint (2025) survey

Beyond the Black Box: A Survey on the Theory and Mechanism of Large Language Models

Zeyu Gan, Ruifeng Ren, Wei Yao, Xiaolin Hu, Gengze Xu, Chen Qian, Huayi Tang, Zixuan Gong, Xinhao Yao, Pengwei Tang, Zhenxing Dou, Yong Liu

Preprint

Paper
Preprint (2026) scaling

Effective Frontiers: A Unification of Neural Scaling Laws

Jiaxuan Zou*, Zixuan Gong*, Ye Su, Huayi Tang, Yong Liu

* Equal contribution

Preprint

Paper
Preprint (2026) moe

Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry

Ye Su, Huayi Tang, Zixuan Gong, Yong Liu

Preprint

Paper

Beyond

  • Keep exploring, discovering, and thinking!
  • Communication creates ideas! Feel free to reach out via email.