
Daily Paper Cast
byJingwen Liang, Gengyu Wang
ScienceTechnology
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.com Creator: Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/ Gengyu Wang, LLM ML, http://wanggengyu.com Listen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236 Cover Image by Kawen Kuang https://kawen.art
Episodes(40 episodes)
Episode 2087
AREX: Towards a Recursively Self-Improving Agent for Deep Research
π€ Upvotes: 118 | cs.AI
Authors:
Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Zheng Liu, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang, Zheng Liu
Title:
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Arxiv:
http://arxiv.org/abs/2607.21461v1
Abstract:
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas ver...
Published: Jul 25, 2026Duration: 19m 51s
Episode 2086
ReferTrack: Referring Then Tracking for Embodied Visual Tracking
π€ Upvotes: 45 | cs.RO
Authors:
Hanjing Ye, Tianle Zeng, Jiazhao Zhang, Shaoan Wang, Zibo Zhang, Weisi Situ, Yuchen Zhou, Yonggen Ling, Hong Zhang
Title:
ReferTrack: Referring Then Tracking for Embodied Visual Tracking
Arxiv:
http://arxiv.org/abs/2607.20061v1
Abstract:
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and wea...
Published: Jul 25, 2026Duration: 20m 48s
Episode 2085
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
π€ Upvotes: 42 | cs.CL
Authors:
Hao Liang, Qihan Lin, Zhaoyang Han, Xiaochen Ma, Zhen Hao Wong, Meiyi Qiang, Linzhuang Sun, Wentao Zhang
Title:
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Arxiv:
http://arxiv.org/abs/2605.09635v3
Abstract:
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual gro...
Published: Jul 25, 2026Duration: 22m 45s
Episode 2084
Visual Contrastive Self-Distillation
π€ Upvotes: 41 | cs.CV, cs.AI
Authors:
Yijun Liang, Yunjie Tian, Yijiang Li, Yuqi Jia, Furong Huang, Tianyi Zhou, Di Fu
Title:
Visual Contrastive Self-Distillation
Arxiv:
http://arxiv.org/abs/2607.21556v1
Abstract:
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be...
Published: Jul 25, 2026Duration: 21m 24s
Episode 2083
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
π€ Upvotes: 34 | cs.CV
Authors:
Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang
Title:
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
Arxiv:
http://arxiv.org/abs/2607.21072v1
Abstract:
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete tex...
Published: Jul 25, 2026Duration: 20m 39s
Episode 2082
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
π€ Upvotes: 51 | cs.CL, cs.AI
Authors:
Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tol...
Published: Jul 24, 2026Duration: 19m 57s
Episode 2081
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
π€ Upvotes: 170 | cs.CV, cs.AI, cs.LG
Authors:
Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
Title:
ABot-World-0: Infinite Interactive Wor...
Published: Jul 23, 2026Duration: 21m 15s
Episode 2080
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
π€ Upvotes: 122 | cs.SE, cs.AI
Authors:
Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang
Title:
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Arxiv:
http://arxiv.org/abs/2607.16617v1
Abstract:
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the \textit{NL2Pipeline gap}. To bridge it, we introduce \textsc{DataFlow-Harness}, a platform that guides an...
Published: Jul 23, 2026Duration: 19m 21s
Episode 2079
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
π€ Upvotes: 66 | cs.CV
Authors:
Maohua Li, Qirui Li, Yanke Zhou, Yiduo Li, Zhaosheng Chi, Chao Xu, Cuifeng Shen, Yixuan Xu, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang
Title:
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
Arxiv:
http://arxiv.org/abs/2607.19139v1
Abstract:
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention decomposition with targeted interventions across token spans, hea...
Published: Jul 23, 2026Duration: 22m 47s
Episode 2078
Generative World Renderer at the Speed of Play
π€ Upvotes: 65 | cs.CV
Authors:
Guixu Lin, Zheng-Hui Huang, Siqi Yang, Ming-Hsuan Yang, Kaipeng Zhang, Zhixiang Wang
Title:
Generative World Renderer at the Speed of Play
Arxiv:
http://arxiv.org/abs/2607.18703v1
Abstract:
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves scene structure without altering the underlying world dynamics. This demonstrates an alternative path toward interactive world modeling and user-controllable play. However, the original AlayaRenderer is too computationally expensive for...
Published: Jul 23, 2026Duration: 19m 30s
Episode 2077
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
π€ Upvotes: 56 | cs.CV, cs.AI, cs.LG, cs.MM, eess.IV
Authors:
Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu
Title:
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Arxiv:
http://arxiv.org/abs/2607.19064v2
Abstract:
Large-scale visual generators are increasingly capable but costly to...
Published: Jul 23, 2026Duration: 20m 13s
Episode 2076
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
π€ Upvotes: 47 | cs.AI
Authors:
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
Title:
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
Arxiv:
http://arxiv.org/abs/2607.18367v1
Abstract:
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to...
Published: Jul 23, 2026Duration: 21m 14s
Episode 2075
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
π€ Upvotes: 21 | cs.AI, cs.CL
Authors:
Kunlun Zhu, Xuyan Ye, Zhiguang Han, Yuchen Zhao, Bingxuan Li, Weijia Zhang, Muxin Tian, Xiangru Tang, Pan Lu, James Zou, Jiaxuan You, Heng Ji
Title:
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
Arxiv:
http://arxiv.org/abs/2607.18754v1
Abstract:
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cau...
Published: Jul 23, 2026Duration: 21m 0s
Episode 2074
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
π€ Upvotes: 144 | cs.CV
Authors:
Yuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, Zhiqiu Zhang, Songze Li, Jun Zhang, Tianxiang Jiang, Yuandong Yang, Ziang Yan, Zikang Wang, Xinyu Chen, Haoran Chen, Shaowei Zhang, Limin Wang
Title:
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Arxiv:
http://arxiv.org/abs/2607.17423v1
Abstract:
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence int...
Published: Jul 22, 2026Duration: 17m 21s
Episode 2073
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
π€ Upvotes: 76 | cs.CL
Authors:
Xinyu Geng, Xuanhua He, Sixiang Chen, Yanjing Xiao, Fan Zhang, Shijue Huang, Haitao Mi, Zhenwen Liang, Tianqing Fang, Yi R. Fung
Title:
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
Arxiv:
http://arxiv.org/abs/2607.07820v2
Abstract:
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and ver...
Published: Jul 22, 2026Duration: 19m 42s
Episode 2072
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
π€ Upvotes: 74 | cs.CL
Authors:
Qing Zong, Yue Guo, Mengxin Yang, Yiwen Guo, Yangqiu Song
Title:
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
Arxiv:
http://arxiv.org/abs/2607.17250v1
Abstract:
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture how characters and worlds evolve together over time. To address this, EvolvingWorld models literary simulation as...
Published: Jul 22, 2026Duration: 20m 46s
Episode 2071
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
π€ Upvotes: 65 | cs.CL, cs.SE
Authors:
Yuhang Wang, Yuling Shi, Shaoqiu Zhang, Jialiang Liang, Shilin He, Siyu Ye, Yuting Chen, Kai Cai, Xiaodong Gu
Title:
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
Arxiv:
http://arxiv.org/abs/2607.18213v1
Abstract:
Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations indicating the relevance of code context whe...
Published: Jul 22, 2026Duration: 19m 58s
Episode 2070
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
π€ Upvotes: 77 | cs.RO
Authors:
Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang, Jianwei Gao, Shijian Lu, Chengju Liu, Jianfei Yang, Mingxiu Chen, Deli Zhao
Title:
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
Arxiv:
http://arxiv.org/abs/2607.17977v1
Abstract:
We present RynnBrain 1.1, a family of emb...
Published: Jul 22, 2026Duration: 22m 34s
Episode 2069
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
π€ Upvotes: 47 | cs.CV
Authors:
Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo
Title:
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
Arxiv:
http://arxiv.org/abs/2607.18217v2
Abstract:
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and...
Published: Jul 22, 2026Duration: 19m 11s
Episode 2068
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
π€ Upvotes: 49 | cs.RO, cs.CV
Authors:
Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li
Title:
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Arxiv:
http://arxiv.org/abs/2607.14183v2
Abs...
Published: Jul 22, 2026Duration: 21m 19s