A curated, annotated list of papers on looped / recurrent-depth / weight-tied transformers and the latent-reasoning ideas around them, from Universal Transformers (2018) to the compute-matched looped MoE models of 2026 — the technique reported to sit behind GPT-6 Astra.
This repository started from the WeChat article 《刚刚,GPT-6 Astra 背后的 Loop 技术全被挖出来了》 (PaperAgent, after a thread by @nrehiew_). The five papers it threads together are marked with ⭐ below and are covered by full-length reports, layout-preserving Chinese translations, and a survey with ten trend insights. Everything else was collected through arXiv searches on 2026-09-06.
中文说明见 README_zh.md。 五篇核心论文的逐篇深度解读在 reports/zh/,保版式中译 PDF 在 papers/zh/,趋势综述在 survey/zh/,完整中文报告 PDF 见 reports/pdf/awesome_loop_transformer_report_zh.pdf,HTML 演示见 slides/html/index.html。
| Deliverable | Path |
|---|---|
| Five core papers, original PDFs + arXiv full text | papers/original/, papers/src/ |
| Layout-preserving Chinese translations with object-level QA reports | papers/zh/ (QA record) |
| Per-paper deep-dive reports, Chinese and English | reports/zh/, reports/en/ |
| Full report PDFs (survey + five reports), Chinese and English | reports/pdf/awesome_loop_transformer_report_zh.pdf, reports/pdf/awesome_loop_transformer_report_en.pdf |
| Survey: taxonomy, timeline, ~60-paper map, ten trend insights, open problems | survey/zh/survey_zh.md (PDF), survey/en/survey_en.md (PDF) |
| HTML slide deck (21 slides, Chinese) + PDF export | slides/html/index.html, slides/html/awesome_loop_transformer_slides_zh.pdf |
| Beamer slide deck (English, 16:9, with backup slides) | slides/beamer/loop_transformer_slides.pdf (source) |
| Editable PowerPoint deck (Chinese, 21 slides, native shapes via ppt-master) + LibreOffice preview | slides/pptx/awesome_loop_transformer_slides_zh.pptx (preview PDF, page SVGs) |
| Machine-readable paper list and abstracts | data/papers.csv, data/related_abstracts.json |
Online slides: asimfish.github.io/awesome_loop_transformer/slides/html/
| You have | Read |
|---|---|
| 5 minutes | the ten insights below, then Appendix A of the survey (five popular claims the papers do not support) |
| 30 minutes | the survey §2–4 and §6, or the 21-slide HTML deck |
| 2 hours | the five paper reports — each opens with an at-a-glance box, a diagram, and closes with a glossary |
| a talk to give | the Beamer deck (English) or the editable PPTX (Chinese) |
-
Recurrent depth passed the compute-matching gate in 2026, through a narrow door. Loopie (matched measured step time) and SMELT (matched FLOPs/params/KV) beat vanilla MoE transformers; the win needs MoE + layer/mid-stack looping + two loops, and comes from converting saved activation memory into width rather than from loops replacing parameters (one recurrence is worth
$r^{0.46}$ unique parameters). - "How many loops" splits into two regimes. Fixed pretraining budget: 2 (Loopie, LoopCoder-v2, SMELT, Nanbeige4.2). Variable inference compute: adaptive and deep (Huginn to 32–64). LoopCoder-v2's "three regress" is the positional tax of the Parallel Loop Transformer, not a law.
- Loop granularity is shrinking and converging on the middle of the stack (depth fraction ≈ 0.5); MoE requires layer-mode looping.
- Two philosophies: treat a pretrained mid-stack as an ODE flow (sub-step, keep the endpoint) versus train the loop to be a fixed point (STARS, Fixed-Point Reasoners, Attractor Models).
- Training stability went from folklore to theorems (residual scaling 1/N, DeepLoop exponent 1/2, LR transfer across loop counts).
- Inference cost decomposes into latency, KV memory, and quantization, each with 2025–2026 fixes (PLT, diffusion-style samplers, MELT, Looped Latent Attention 32×, LoopQ). Matched training compute is not matched inference cost.
- Latent loops and explicit CoT are complements: super-additive in LoopCoder-v2, gap-closing in LOTUS; standard GRPO mis-assigns credit to looped models (RLTT, LoopRPT).
- Loops manipulate knowledge, they do not store it — pair them with retrieval, long context, or MoE expert banks.
- Applications moved from puzzles to agents and embodiment (SWE-bench, tool calling, VLA).
- Productization has started; the toolchain has not caught up.
Full text with sources: survey/en/survey_en.md · survey/zh/survey_zh.md.
The five papers threaded together by the WeChat article / @nrehiew_ thread, in reading order. Each has a full report (ZH/EN) and a layout-preserving Chinese translation.
-
⭐Universal Transformers. ICLR, 2019. paper, code, report-zh, report-en, 中译PDF
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Łukasz Kaiser
-
⭐Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. arXiv, 2025. paper, code, report-zh, report-en, 中译PDF
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein
-
⭐Training-Free Looped Transformers. arXiv, 2026. paper, report-zh, report-en, 中译PDF
Lizhang Chen, Jonathan Li, Chen Liang, Ni Lao, Qiang Liu
-
⭐LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling. arXiv, 2026. paper, code, report-zh, report-en, 中译PDF
Jian Yang, Shawn Guo, Wei Zhang, Tianyu Zheng, Yaxin Du, Haau-Sing Li, Jiajun Wu, Yue Song, Yan Xing, Qingsong Cai, Zelong Huang, Chuan Hao, Ran Tao, Xianglong Liu, Wayne Xin Zhao, Mingjie Tang, Weifeng Lv, Ming Zhou, Bryan Dai
-
⭐Loop the Loopies! arXiv, 2026. paper, report-zh, report-en, 中译PDF
Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai
-
A Survey on Latent Reasoning. arXiv, 2025. paper, code
Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng, Xingwei Qu, Jinfa Huang, Dawei Zhu, Hao Wang, Kaiwen Xue, Xuanliang Zhang, Yong Shan, Tianle Cai, Taylor Kergan, Assel Kembay, Andrew Smith, Chenghua Lin, Binh Nguyen, Yuqi Pan, Yuhong Chou, Zefan Cai, Zhenhe Wu, Yongchi Zhao, Tianyu Liu, Jian Yang, Wangchunshu Zhou, Chujie Zheng, Chongxuan Li, Yuyin Zhou, Zhoujun Li, Zhaoxiang Zhang, Jiaheng Liu, Ge Zhang, Wenhao Huang, Jason Eshraghian
-
Neural GPUs Learn Algorithms. ICLR, 2016. paper
Łukasz Kaiser, Ilya Sutskever
-
Adaptive Computation Time for Recurrent Neural Networks. arXiv, 2016. paper
Alex Graves
-
⭐Universal Transformers. ICLR, 2019. paper, code, report-zh, report-en, 中译PDF
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Łukasz Kaiser
-
Deep Equilibrium Models. NeurIPS, 2019. paper
Shaojie Bai, J. Zico Kolter, Vladlen Koltun
-
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. ICLR, 2020. paper
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut
-
Lessons on Parameter Sharing across Layers in Transformers. arXiv, 2023. paper
Sho Takase, Shun Kiyono
-
Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent Networks. NeurIPS, 2021. paper
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, Tom Goldstein
-
PonderNet: Learning to Ponder. ICML Workshop, 2021. paper
Andrea Banino, Jan Balaguer, Charles Blundell
-
End-to-end Algorithm Synthesis with Recurrent Networks: Logical Extrapolation Without Overthinking. NeurIPS, 2022. paper
Arpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam, Furong Huang, Micah Goldblum, Tom Goldstein
-
Path Independent Equilibrium Models Can Better Exploit Test-Time Computation. NeurIPS, 2022. paper
Cem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein, Yuhuai Wu, Shaojie Bai, Zico Kolter, Roger Grosse
-
Sparse Universal Transformer. EMNLP, 2023. paper
Shawn Tan, Yikang Shen, Zhenfang Chen, Aaron Courville, Chuang Gan
-
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference. ICLR, 2025. paper
Amirkeivan Mohtashami, Matteo Pagliardini, Martin Jaggi
-
MoEUT: Mixture-of-Experts Universal Transformers. NeurIPS, 2024. paper
Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts, Christopher D. Manning
-
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA. ICLR, 2025. paper
Sangmin Bae, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Seungyeon Kim, Tal Schuster
-
Looped Transformers as Programmable Computers. ICML, 2023. paper
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D. Lee, Dimitris Papailiopoulos
-
Looped Transformers are Better at Learning Learning Algorithms. ICLR, 2024. paper
Liu Yang, Kangwook Lee, Robert Nowak, Dimitris Papailiopoulos
-
Simulation of Graph Algorithms with Looped Transformers. ICML, 2024. paper
Artur Back de Luca, Kimon Fountoulakis
-
Looped Transformers for Length Generalization. ICLR, 2025. paper
Ying Fan, Yilun Du, Kannan Ramchandran, Kangwook Lee
-
On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding. arXiv, 2024. paper
Kevin Xu, Issei Sato
-
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning? ICML, 2024. paper
Khashayar Gatmiry, Nikunj Saunshi, Sashank J. Reddi, Stefanie Jegelka, Sanjiv Kumar
-
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers. arXiv, 2024. paper
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, Yufa Zhou
-
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent. AISTATS, 2025. paper
Bo Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song
-
On the Role of Depth and Looping for In-Context Learning with Task Diversity. arXiv, 2024. paper
Khashayar Gatmiry, Nikunj Saunshi, Sashank J. Reddi, Stefanie Jegelka, Sanjiv Kumar
-
Reasoning with Latent Thoughts: On the Power of Looped Transformers. ICLR, 2025. paper
Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank J. Reddi
-
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers. arXiv, 2025. paper
Kevin Xu, Issei Sato
-
What Makes Looped Transformers Perform Better Than Non-Recursive Ones. arXiv, 2025. paper
Zixuan Gong, Yong Liu, Jiaye Teng
-
SpiralFormer: Looped Transformers Can Learn Hierarchical Dependencies via Multi-Resolution Recursion. arXiv, 2026. paper
Chengting Yu, Xiaobo Shu, Yadao Wang, Yizhen Zhang, Haoyi Wu, You Wu, Rujiao Long, Ziheng Chen, Yuchi Xu, Wenbo Su, Bo Zheng
-
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization. arXiv, 2026. paper
Hung-Hsuan Chen
-
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers. COLM, 2026. paper
Harsh Kohli, Srinivasan Parthasarathy, Huan Sun, Yuekun Yao
-
Stability and Generalization in Looped Transformers. arXiv, 2026. paper
Asher Labovich
-
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models. arXiv, 2026. paper
Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis
-
Universal Transformers Need Memory: Depth-State Trade-offs in Adaptive Recursive Reasoning. arXiv, 2026. paper
Grigory Sapunov
-
Length Generalization with Log-Depth Recurrent Units. arXiv, 2026. paper
Charles Pert, Dalal Alrajeh, Alessandra Russo
-
Chain-of-Thought and Compressed Looped Transformers: A Memory-Budget Separation. arXiv, 2026. paper
Haozhou Zhang
-
Looped Transformers with Layer Normalization Provably Learn the Power Method. arXiv, 2026. paper
Lyumin Wu, Chenyang Zhang, Yuan Cao
-
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning. arXiv, 2026. paper
Hengyu Fu, Tianyu Guo, Zixuan Wang, Hanlin Zhu, Jason D. Lee, Jiantao Jiao, Stuart Russell, Song Mei
-
When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers. arXiv, 2026. paper
Tong Zhang, Junhao Hu, Yun Peng, Tao Xie
-
Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers. arXiv, 2026. paper
Takuya Ito, Ruchir Puri, Murray Campbell, Parikshit Ram
-
MoEUT: Mixture-of-Experts Universal Transformers. NeurIPS, 2024. paper
Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts, Christopher D. Manning
-
⭐Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. arXiv, 2025. paper, code, report-zh, report-en, 中译PDF
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein
-
Parallel Loop Transformer for Efficient Test-Time Computation Scaling. arXiv, 2025. paper
Bohong Wu, Mengzhao Chen, Xiang Luo, Shen Yan, Qifan Yu, Fan Xia, Tianqi Zhang, Hongrui Zhan, Zheng Zhong, Xun Zhou, Siyuan Qiao, Xingyan Bin
-
Scaling Latent Reasoning via Looped Language Models. arXiv, 2025. paper, code
Rui-Jie Zhu, Zixuan Wang, Kai Hua, Tianyu Zhang, Ziniu Li, Haoran Que, Boyi Wei, Zixin Wen, Fan Yin, He Xing, Lu Li, Jiajun Shi, Kaijing Ma, Shanda Li, Taylor Kergan, Andrew Smith, Xingwei Qu, Mude Hui, Bohong Wu, Qiyang Min, Hongzhi Huang, Xun Zhou, Wei Ye, Jiaheng Liu, Jian Yang, Yunfeng Shi, Chenghua Lin, Enduo Zhao, Tianle Cai, Ge Zhang, Wenhao Huang, Yoshua Bengio, Jason Eshraghian
-
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves. arXiv, 2026. paper
Jonas Knupp, Jan Hendrik Metzen, Jeremias Bohn, Georg Groh, Kristian Kersting
-
Adaptive Loops and Memory in Transformers: Think Harder or Know More? ICLR Workshop, 2026. paper
Markus Frey, Behzad Shomali, Ali Hamza Bashir, David Berghaus, Joachim Koehler, Mehdi Ali
-
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models. arXiv, 2026. paper
Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis
-
Hyperloop Transformers. arXiv, 2026. paper
Abbas Zeitoun, Lucas Torroba-Hennigen, Yoon Kim
-
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior. arXiv, 2026. paper
Zeyi Huang, Xuehai He, LiLiang Ren, Yiping Wang, Baolin Peng, Hao Cheng, Shuohang Wang, Pengcheng He, Jianfeng Gao, Yong Jae Lee, Yelong Shen
-
A Dual-Path Architecture for Scaling Compute and Capacity in LLMs. arXiv, 2026. paper
Markus Frey, Behzad Shomali, Joachim Koehler, Mehdi Ali
-
⭐LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling. arXiv, 2026. paper, code, report-zh, report-en, 中译PDF
Jian Yang, Shawn Guo, Wei Zhang, Tianyu Zheng, Yaxin Du, Haau-Sing Li, Jiajun Wu, Yue Song, Yan Xing, Qingsong Cai, Zelong Huang, Chuan Hao, Ran Tao, Xianglong Liu, Wayne Xin Zhao, Mingjie Tang, Weifeng Lv, Ming Zhou, Bryan Dai
-
Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models. arXiv, 2026. paper — contrast: sequence-length scaling
Aiwei Liu, Cheng Shi, Chuhan Wu, Ci Lei, Di Lu, Donald He, Fan Zhang, Fanhao Kong, Feifei Zhang, Guan Wang, Haicheng Wang, Haoyu Liu, Houjin Yu, Jiachen Ding, Jiayi Feng, Jie Zhou, Jijun Chi, Jindi Shi, Jing Lei, Junjie Zhang, Laiyi Li, Le Tian, Linhao Zhang, Miao Fan, Sijun Zhang, Wei Jia, Weiwei Shi, Wenhan Li, Wentao Zhao, Wenteng Liang, Xiao Zhou, Xiaojin Zhou, Xihuai Wang, Xinyu Gao, Xuanliang Wang, Xuyang Ao, Yang Yu, Yangxiu You, Yinuo Zhao, Yufei Kuang, Yufei Wang, Yuan Liu, Yuan Liu, Yuwen Chen, Zhencong Tian, Zhongyin Zhao, Zilin Yu, Zitao Wang
-
⭐Loop the Loopies! arXiv, 2026. paper, report-zh, report-en, 中译PDF
Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai
-
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model. arXiv, 2026. paper
Nanbeige Lab, :, Chen Yang, Chengrui Huang, Fufeng Lan, Hanhui Chen, Hao Zhou, Huatong Song, Jiaqi Cao, Jiaying Zhu, Jinlin Niu, Kai Wang, Lisheng Huang, Qiliang Liang, Ran Le, Ruixiang Feng, Shuang Sun, Tao Gu, Tao Zhang, Tianyu Luo, Yang Song, Yun Xing, Yuntao Wen, Ziyao Xu, Zongchao Chen, Zongqiang Li
-
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation. arXiv, 2026. paper
Amr Hegazy, Amr Alanwar, Mostafa Elhoushi
-
Allocating Recurrent Compute in Looped Language Models. arXiv, 2026. paper
Ruhai Lin, Yiyang Guo, Rui-Jie Zhu, Hao Ye, Jason K. Eshraghian
-
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. arXiv, 2026. paper
Shaowen Wang, Ge Zhang, Kairong Luo, Yuhao Wu, Shaofan Liu, Jiaheng Liu, Wenhao Huang, Shen Yan, Jian Li
-
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation. NeurIPS, 2025. paper, code
Sangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim, Jiyoun Ha, Tal Schuster, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Aaron Courville, Se-Young Yun
-
Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models. arXiv, 2025. paper, code
Jonas Geiping, Xinyu Yang, Guinan Su
-
Parallel Loop Transformer for Efficient Test-Time Computation Scaling. arXiv, 2025. paper
Bohong Wu, Mengzhao Chen, Xiang Luo, Shen Yan, Qifan Yu, Fan Xia, Tianqi Zhang, Hongrui Zhan, Zheng Zhong, Xun Zhou, Siyuan Qiao, Xingyan Bin
-
Hyperloop Transformers. arXiv, 2026. paper
Abbas Zeitoun, Lucas Torroba-Hennigen, Yoon Kim
-
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models. arXiv, 2026. paper
Victor Conchello Vendrell, Arnau Padres Masdemont, Niccolò Grillo, Jordi Ros-Giralt, Arash Behboodi, Fabio Valerio Massoli
-
LoopQ: Quantization for Recursive Transformers. arXiv, 2026. paper
Rui Fang, Hsi-Wen Chen, Ming-Syan Chen
-
LT2: Linear-Time Looped Transformers. arXiv, 2026. paper
Chunyuan Deng, Yizhe Zhang, Rui-Jie Zhu, Yuanyuan Xu, Jiarui Liu, T. S. Eugene Ng, Hanjie Chen
-
CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models. arXiv, 2026. paper
Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
-
Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers. arXiv, 2026. paper
James O' Neill, Fergal Reid
-
Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead. arXiv, 2026. paper, code
John T. Halloran
-
SCORE: Replacing Layer Stacking with Contractive Recurrent Depth. arXiv, 2026. paper
Guillaume Godin
-
Simply Stabilizing the Loop via Fully Looped Transformer. arXiv, 2026. paper
Rao Fu, Zixuan Yang, Jiankun Zhang, Jing Ma, Hechang Chen, Yu Li, Yi Chang
-
Solve the Loop: Attractor Models for Language and Reasoning. arXiv, 2026. paper
Jacob Fein-Ashley, Paria Rashidinejad
-
Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models. ICML, 2026. paper
Xiao-Wen Yang, Ziyu Han, Xi-Hua Zhang, Wen-Da Wei, Jie-Jing Shao, Lan-Zhe Guo, Yu-Feng Li
-
On the Residual Scaling of Looped Transformers: Stability and Transferability. arXiv, 2026. paper
Shaowen Wang, Bingrui Li, Ge Zhang, Wenhao Huang, Shen Yan, Jian Li
-
Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers. arXiv, 2026. paper, code
Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin, Alexander Theus, Thomas Hofmann, Valentina Boeva, T. Konstantin Rusch, Antonio Orvieto
-
Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping. arXiv, 2026. paper
Hsun-Yu Kuo, El Mahdi Chayti, Patrik Reizinger, Wieland Brendel, Martin Jaggi
-
LayerNorm as Implicit Gain Control in Looped Transformers. arXiv, 2026. paper
Matthias M. M. Buehlmaier
-
DeepLoop: Depth Scaling for Looped Transformers. arXiv, 2026. paper
Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang
-
Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth. arXiv, 2026. paper
Ivan Viakhirev, Kirill Borodin, Amirah Almutairi, Serguei Barannikov, Maxim Abramov, Grach Mkrtchian
-
PonderNet: Learning to Ponder. ICML Workshop, 2021. paper
Andrea Banino, Jan Balaguer, Charles Blundell
-
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference. ICLR, 2025. paper
Amirkeivan Mohtashami, Matteo Pagliardini, Martin Jaggi
-
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation. NeurIPS, 2025. paper, code
Sangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim, Jiyoun Ha, Tal Schuster, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Aaron Courville, Se-Young Yun
-
Two-Scale Latent Dynamics for Recurrent-Depth Transformers. arXiv, 2025. paper
Francesco Pappone, Donato Crisostomi, Emanuele Rodolà
-
Think-at-Hard: Dynamic Looped Transformers for Improved Reasoning. ICML, 2026. paper, code
Tianyu Fu, Yichen You, Zekai Chen, Guohao Dai, Huazhong Yang, Yu Wang
-
LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation. ICLR, 2026. paper
Ahmadreza Jeddi, Marco Ciccone, Babak Taati
-
Adaptive Loops and Memory in Transformers: Think Harder or Know More? ICLR Workshop, 2026. paper
Markus Frey, Behzad Shomali, Ali Hamza Bashir, David Berghaus, Joachim Koehler, Mehdi Ali
-
Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping. arXiv, 2026. paper
Hsun-Yu Kuo, El Mahdi Chayti, Patrik Reizinger, Wieland Brendel, Martin Jaggi
-
Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts. arXiv, 2026. paper
Andrei Cristian Popescu, Haitz Sáez de Ocáriz Borde, Pietro Liò
-
Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers. arXiv, 2026. paper
Joe Logan
-
RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory. arXiv, 2026. paper
Yuxiang Wang, Kunyu Feng, Yingda Shen, Haoning Xu, Junyu Wang, Zhizheng Wu
-
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA. ICLR, 2025. paper
Sangmin Bae, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Seungyeon Kim, Tal Schuster
-
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence. arXiv, 2025. paper, code
Sean McLeish, Ang Li, John Kirchenbauer, Dayal Singh Kalra, Brian R. Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Jonas Geiping, Tom Goldstein, Micah Goldblum
-
Improving Recursive Transformers with Mixture of LoRAs. arXiv, 2025. paper
Mohammadmahdi Nouriborji, Morteza Rohanian, Omid Rohanian
-
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation. arXiv, 2026. paper
Jaber Jaber, Osama Jaber
-
⭐Training-Free Looped Transformers. arXiv, 2026. paper, report-zh, report-en, 中译PDF
Lizhang Chen, Jonathan Li, Chen Liang, Ni Lao, Qiang Liu
-
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs. ICML, 2026. paper
Ziyue Li, Yang Li, Tianyi Zhou
-
Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets. arXiv, 2026. paper
Mark Shapiro
-
Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning. arXiv, 2025. paper
Qifan Yu, Zhenyu He, Sijie Li, Xun Zhou, Jun Zhang, Jingjing Xu, Di He
-
Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Models. ICML, 2026. paper, code
Jonathan Williams, Esin Tureci
-
LoopRPT: Reinforcement Pre-Training for Looped Language Models. arXiv, 2026. paper
Guo Tang, Shixin Jiang, Heng Chang, Nuo Chen, Yuhan Li, Huiming Fan, Jia Li, Ming Liu, Bing Qin
-
Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning. arXiv, 2026. paper
Jiayu Yang, Chao Chen, Shengen Wu, Yinhong Liu, Yuxuan Fan, Lujundong Li, Songning Lai, Chengwei Qin, Zhijiang Guo
-
Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models. arXiv, 2026. paper
Rituraj Sharma, Tu Vu
-
Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers. arXiv, 2026. paper, code
Ying Fan, Anej Svete, Kangwook Lee
-
LoopMTP: A looped transformer guided by latent multi-token prediction. arXiv, 2026. paper
Behzad Shomali, Markus Frey, David Berghaus, Joachim Koehler, Mehdi Ali
-
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer. Workshop, 2025. paper, code
Wenquan Lu, Yuechuan Yang, Kyle Lee, Yanshu Li, Enqi Liu
-
Two-Scale Latent Dynamics for Recurrent-Depth Transformers. arXiv, 2025. paper
Francesco Pappone, Donato Crisostomi, Emanuele Rodolà
-
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs? arXiv, 2026. paper
Guanxu Chen, Dongrui Liu, Jing Shao
-
Step-resolved data attribution for looped transformers. arXiv, 2026. paper
Georgios Kaissis, David Mildenberger, Juan Felipe Gomez, Martin J. Menten, Eleni Triantafillou
-
Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models. arXiv, 2026. paper
Darpan Aswal, Thomas Palmeira Ferraz, Yongxin Zhou, Maxime Peyrard
-
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence? arXiv, 2026. paper
Wenlong Wang, Fergal Reid
-
Hierarchical Reasoning Model. arXiv, 2025. paper
Guan Wang, Jin Li, Yuhao Sun, Xing Chen, Changling Liu, Yue Wu, Meng Lu, Sen Song, Yasin Abbasi Yadkori
-
Hierarchical Reasoning Models: Perspectives and Misconceptions. arXiv, 2025. paper
Renee Ge, Qianli Liao, Tomaso Poggio
-
Less is More: Recursive Reasoning with Tiny Networks. arXiv, 2025. paper
Alexia Jolicoeur-Martineau
-
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator. arXiv, 2025. paper
Arip Asadulaev, Rayan Banerjee, Fakhri Karray, Martin Takac
-
Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models. arXiv, 2026. paper
Zirui Ren, Ziming Liu
-
Dynamical Systems Theory Behind a Hierarchical Reasoning Model. arXiv, 2026. paper
Vasiliy A. Es'kin, Mikhail E. Smorkalov
-
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models. arXiv, 2026. paper
Chris Cameron, Wangzheng Wang, Nikita Ivanov, Ashmita Bhattacharyya, Didier Chételat, Yingxue Zhang
-
Solve the Loop: Attractor Models for Language and Reasoning. arXiv, 2026. paper
Jacob Fein-Ashley, Paria Rashidinejad
-
Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers. arXiv, 2026. paper, code
Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin, Alexander Theus, Thomas Hofmann, Valentina Boeva, T. Konstantin Rusch, Antonio Orvieto
-
Block-Recurrent Dynamics in Vision Transformers. arXiv, 2025. paper
Mozes Jacobs, Thomas Fel, Richard Hakim, Alessandra Brondetta, Demba Ba, T. Andy Keller
-
LoopViT: Scaling Visual ARC with Looped Transformers. arXiv, 2026. paper, code
Wen-Jie Shu, Xuerui Qiu, Rui-Jie Zhu, Harold Haodong Chen, Yexin Liu, Harry Yang
-
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning. arXiv, 2026. paper, code
Yalcin Tur, Jalal Naghiyev, Haoquan Fang, Wei-Chuan Tsai, Jiafei Duan, Dieter Fox, Ranjay Krishna
-
Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models. arXiv, 2026. paper
Ruihan Xu, Yuting Gao, Lan Wang, Jianing Li, Weihao Chen, Qingpei Guo, Ming Yang, Shiliang Zhang
-
ELT: Elastic Looped Transformers for Visual Generation. arXiv, 2026. paper
Sahil Goyal, Swayam Agrawal, Gautham Govind Anil, Prateek Jain, Sujoy Paul, Aditya Kusupati
-
RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction. arXiv, 2026. paper
Renjie He
-
Déjà View: Looping Transformers for Multi-View 3D Reconstruction. arXiv, 2026. paper
Alessandro Burzio, Tobias Fischer, Sven Elflein, Qunjie Zhou, Riccardo de Lutio, Jiawei Ren, Jiahui Huang, Shengyu Huang, Marc Pollefeys, Laura Leal-Taixé, Zan Gojcic, Haithem Turki
-
Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers. arXiv, 2026. paper
Yacouba Kaloga, Shashi Kumar, Shakeel A. Sheikh, Driss Khalil, Petr Motlicek, Ina Kodrasi
-
LatentMT: Machine Translation with Latent Reasoning. arXiv, 2026. paper
Wei-Rui Chen, Samar M. Magdy, Chiyu Zhang, Wenhui Zhu, Zhipeng Wang, Muhammad Abdul-Mageed
-
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model. arXiv, 2026. paper
Nanbeige Lab, :, Chen Yang, Chengrui Huang, Fufeng Lan, Hanhui Chen, Hao Zhou, Huatong Song, Jiaqi Cao, Jiaying Zhu, Jinlin Niu, Kai Wang, Lisheng Huang, Qiliang Liang, Ran Le, Ruixiang Feng, Shuang Sun, Tao Gu, Tao Zhang, Tianyu Luo, Yang Song, Yun Xing, Yuntao Wen, Ziyao Xu, Zongchao Chen, Zongqiang Li
-
ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding. arXiv, 2026. paper
Shijie Wang, Xiangzhao Hao, Yueti Li, Guangyu Cao, Xinyu Tang, Haiyun Guo
-
Looped Language Models Improve Compositional Tool Calling. arXiv, 2026. paper
Andrei Cristian Popescu, Haitz Sáez de Ocáriz Borde, Pietro Liò
-
Think before you speak: Training Language Models With Pause Tokens. ICLR, 2024. paper
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, Vaishnavh Nagarajan
-
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models. arXiv, 2024. paper
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, Adam Santoro
-
Training Large Language Models to Reason in a Continuous Latent Space. COLM, 2025. paper
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, Yuandong Tian
-
Compressed Chain of Thought: Efficient Reasoning Through Dense Representations. arXiv, 2024. paper
Jeffrey Cheng, Benjamin Van Durme
-
PonderLM: Pretraining Language Models to Ponder in Continuous Space. arXiv, 2025. paper
Boyi Zeng, Shixiang Song, Siyuan Huang, Yixuan Wang, He Li, Ziwei He, Xinbing Wang, Zhiyu Li, Zhouhan Lin
-
AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling. arXiv, 2025. paper
Preslav Aleksandrov, Meghdad Kurmanji, Fernando Garcia Redondo, David O'Shea, William Shen, Alex Iacob, Lorenzo Sani, Xinchi Qiu, Nicola Cancedda, Nicholas D. Lane
-
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts. arXiv, 2025. paper
Yeskendir Koishekenov, Aldo Lipani, Nicola Cancedda
-
Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium. arXiv, 2025. paper
Akbar Anbar Jafari, Gholamreza Anbarjafari
- GPT-6 Astra. OpenAI, 2026-09. page — release page; no architectural details disclosed.
- Secret Technique Behind OpenAI's Astra Model Sparks Security Concerns. The Information, 2026-09-02. Stephanie Palazzolo — reports Astra uses "recurrent depth" (anonymously sourced; paywalled, not accessed for this repository).
- Thread: five papers behind recurrent depth. X, 2026-09. @nrehiew_ · @kimmonismus.
- 刚刚,GPT-6 Astra 背后的 Loop 技术全被挖出来了. 微信公众号 PaperAgent, 2026-09. article — the Chinese write-up this repository started from; a snapshot is kept in
source/.
| Model | Params | Tokens | Loops | Link |
|---|---|---|---|---|
| Huginn-0125 | 3.5B dense | 0.8T | random at train, free at test (r̄=32) | HF · code |
| Ouro 1.4B / 2.6B (+Thinking) | dense | 7.7T | 4 | site |
| LoopCoder-V2 | 7B dense PLT | 18T | 2 | HF |
| Loopie-20B-A2B / 6B-A0.6B | MoE | 3.5T | 2 (layer-loop) | see paper 2607.16051 (megatron-loopie, vllm-loopie) |
| Nanbeige4.2-3B | 3B | 28T | 2 | see paper 2607.22083 |
# 1. Paper list -> README
python3 src/fetch_metadata.py # refresh titles/authors from arXiv into data/papers.csv
python3 src/generator.py # regenerate README.md from data/header.md + papers.csv + footer.md
# 2. Chinese translations (SuperTranslate, https://github.com/asimfish/super_translate)
export PAPER_CHINA_DEEPSEEK_API_KEY=... # never on the command line
export SUPER_TRANSLATE_HOME=/path/to/super_translate # after `uv sync` there
bash scripts/translate_papers.sh # -> papers/zh/*.zh.pdf + *.inspect.json
# 3. Report PDFs (pandoc + XeLaTeX; macOS CJK fonts by default, override with CJK_FONT / CJK_SANS)
bash scripts/build_pdfs.sh
# 4. Slides
open slides/html/index.html # arrow keys / space to navigate, P to print
python3 scripts/export_slides.py # HTML -> PDF + PNG previews with headless Chromium (set CHROME_BIN if needed)
cd slides/beamer && xelatex loop_transformer_slides.tex && xelatex loop_transformer_slides.tex
# 5. Editable PPTX through ppt-master (https://github.com/hugohe3/ppt-master; python>=3.10 + its requirements.txt)
export PPT_MASTER_HOME=/path/to/ppt-master
bash scripts/build_pptx.sh # page SVGs -> quality gate -> SVG-to-DrawingML export -> slides/pptx/- Reading. Each core paper was read in full from the arXiv PDF and HTML text; every number in the reports carries its table or section reference, and an appendix in the survey lists five claims from secondary coverage that the papers do not support.
- Translation. SuperTranslate native engine (layout-preserving, formulas and figures frozen, DeepSeek backend) followed by its object-level
inspectQA. Three of five PDFs pass with zero issues; the remaining flags are documented inpapers/zh/QA_SUMMARY.md. - Literature search. arXiv API queries on 2026-09-06 for looped / recurrent-depth / depth-recurrent transformers, latent reasoning with loops or recursion, and named model families; abstracts of the 60 most relevant works are kept in
data/related_abstracts.json. - Writing style. Claims first, necessary limits stated once, numbers checkable in place — following anti-defensive-writing and 说人话 shuorenhua.
- Layout. Report PDFs via pandoc + XeLaTeX with a shared header (
scripts/pandoc_header.tex); the Beamer deck follows the beamer-skill rules (16:9, 10pt, no overlays, semantic color-blind-safe palette, references and backup slides); the HTML deck and the PPTX follow thetechnical-deepdivestyle contract of ppt-master (problem → mechanism → trade-off → implication, blueprint palette, every number with its conditions); the PPTX itself is produced by ppt-master's Quick Generate route —scripts/build_pptx_svgs.pyemits contract-conforming page SVGs (root bounds per module, one text frame per paragraph block, NBSP frame slack so renderers that ignorewrap="none"do not re-wrap), which pass its final quality gate and are exported to native DrawingML shapes. Paper-writing conventions were checked against PaperOrchestra. - Readability. Each report opens with an at-a-glance box (problem / method / result / limits / plain words) and one explanatory diagram, and closes with a glossary; the survey has a reading guide for 5-minute, 30-minute and 2-hour readers. Diagrams are authored in
scripts/build_diagrams.py(SVG) and rendered to PNG byscripts/render_diagrams.sh. - List format. Follows awesome-ml4co:
data/papers.csvis the source of truth andsrc/generator.pyrenders the README.
Add a row to data/papers.csv (or an entry to CURATED in src/fetch_metadata.py and re-run it), then run python3 src/generator.py and open a pull request. Use ; to list a paper under several categories. Please keep the venue field to the accepted venue or arXiv, and add a code link when one exists.
@misc{awesome_loop_transformer_2026,
title = {Awesome Loop Transformer: Looped Transformers, Recurrent Depth and Latent Reasoning (2018--2026)},
author = {{awesome\_loop\_transformer contributors}},
year = {2026},
url = {https://github.com/asimfish/awesome_loop_transformer}
}Text, reports, and slides in this repository are released under CC BY 4.0; code under src/ and scripts/ under MIT. Original papers remain under their authors' licenses; PDFs in papers/original/ are redistributed from arXiv, and the Chinese translations in papers/zh/ are derived works provided for study.

