Daily Paper Cast

Daily Paper Cast

byJingwen Liang, Gengyu Wang

ScienceTechnology

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.com Creator: Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/ Gengyu Wang, LLM ML, http://wanggengyu.com Listen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236 Cover Image by Kawen Kuang https://kawen.art

Episodes(40 episodes)

Episode 2087
AREX: Towards a Recursively Self-Improving Agent for Deep Research
πŸ€— Upvotes: 118 | cs.AI Authors: Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Zheng Liu, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang, Zheng Liu Title: AREX: Towards a Recursively Self-Improving Agent for Deep Research Arxiv: http://arxiv.org/abs/2607.21461v1 Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas ver...
Published: Jul 25, 2026Duration: 19m 51s
Episode 2086
ReferTrack: Referring Then Tracking for Embodied Visual Tracking
πŸ€— Upvotes: 45 | cs.RO Authors: Hanjing Ye, Tianle Zeng, Jiazhao Zhang, Shaoan Wang, Zibo Zhang, Weisi Situ, Yuchen Zhou, Yonggen Ling, Hong Zhang Title: ReferTrack: Referring Then Tracking for Embodied Visual Tracking Arxiv: http://arxiv.org/abs/2607.20061v1 Abstract: Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and wea...
Published: Jul 25, 2026Duration: 20m 48s
Episode 2085
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
πŸ€— Upvotes: 42 | cs.CL Authors: Hao Liang, Qihan Lin, Zhaoyang Han, Xiaochen Ma, Zhen Hao Wong, Meiyi Qiang, Linzhuang Sun, Wentao Zhang Title: K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Arxiv: http://arxiv.org/abs/2605.09635v3 Abstract: Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual gro...
Published: Jul 25, 2026Duration: 22m 45s
Episode 2084
Visual Contrastive Self-Distillation
πŸ€— Upvotes: 41 | cs.CV, cs.AI Authors: Yijun Liang, Yunjie Tian, Yijiang Li, Yuqi Jia, Furong Huang, Tianyi Zhou, Di Fu Title: Visual Contrastive Self-Distillation Arxiv: http://arxiv.org/abs/2607.21556v1 Abstract: On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be...
Published: Jul 25, 2026Duration: 21m 24s
Episode 2083
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
πŸ€— Upvotes: 34 | cs.CV Authors: Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang Title: Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Arxiv: http://arxiv.org/abs/2607.21072v1 Abstract: Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete tex...
Published: Jul 25, 2026Duration: 20m 39s
Episode 2082
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
πŸ€— Upvotes: 51 | cs.CL, cs.AI Authors: Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tol...
Published: Jul 24, 2026Duration: 19m 57s
Episode 2081
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
πŸ€— Upvotes: 170 | cs.CV, cs.AI, cs.LG Authors: Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo Title: ABot-World-0: Infinite Interactive Wor...
Published: Jul 23, 2026Duration: 21m 15s
Episode 2080
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
πŸ€— Upvotes: 122 | cs.SE, cs.AI Authors: Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang Title: DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Arxiv: http://arxiv.org/abs/2607.16617v1 Abstract: Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the \textit{NL2Pipeline gap}. To bridge it, we introduce \textsc{DataFlow-Harness}, a platform that guides an...
Published: Jul 23, 2026Duration: 19m 21s
Episode 2079
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
πŸ€— Upvotes: 66 | cs.CV Authors: Maohua Li, Qirui Li, Yanke Zhou, Yiduo Li, Zhaosheng Chi, Chao Xu, Cuifeng Shen, Yixuan Xu, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang Title: Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Arxiv: http://arxiv.org/abs/2607.19139v1 Abstract: Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention decomposition with targeted interventions across token spans, hea...
Published: Jul 23, 2026Duration: 22m 47s
Episode 2078
Generative World Renderer at the Speed of Play
πŸ€— Upvotes: 65 | cs.CV Authors: Guixu Lin, Zheng-Hui Huang, Siqi Yang, Ming-Hsuan Yang, Kaipeng Zhang, Zhixiang Wang Title: Generative World Renderer at the Speed of Play Arxiv: http://arxiv.org/abs/2607.18703v1 Abstract: Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves scene structure without altering the underlying world dynamics. This demonstrates an alternative path toward interactive world modeling and user-controllable play. However, the original AlayaRenderer is too computationally expensive for...
Published: Jul 23, 2026Duration: 19m 30s
Episode 2077
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
πŸ€— Upvotes: 56 | cs.CV, cs.AI, cs.LG, cs.MM, eess.IV Authors: Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu Title: Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Arxiv: http://arxiv.org/abs/2607.19064v2 Abstract: Large-scale visual generators are increasingly capable but costly to...
Published: Jul 23, 2026Duration: 20m 13s
Episode 2076
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
πŸ€— Upvotes: 47 | cs.AI Authors: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao Title: AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Arxiv: http://arxiv.org/abs/2607.18367v1 Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to...
Published: Jul 23, 2026Duration: 21m 14s
Episode 2075
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
πŸ€— Upvotes: 21 | cs.AI, cs.CL Authors: Kunlun Zhu, Xuyan Ye, Zhiguang Han, Yuchen Zhao, Bingxuan Li, Weijia Zhang, Muxin Tian, Xiangru Tang, Pan Lu, James Zou, Jiaxuan You, Heng Ji Title: AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents Arxiv: http://arxiv.org/abs/2607.18754v1 Abstract: LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cau...
Published: Jul 23, 2026Duration: 21m 0s
Episode 2074
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
πŸ€— Upvotes: 144 | cs.CV Authors: Yuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, Zhiqiu Zhang, Songze Li, Jun Zhang, Tianxiang Jiang, Yuandong Yang, Ziang Yan, Zikang Wang, Xinyu Chen, Haoran Chen, Shaowei Zhang, Limin Wang Title: TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Arxiv: http://arxiv.org/abs/2607.17423v1 Abstract: Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence int...
Published: Jul 22, 2026Duration: 17m 21s
Episode 2073
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
πŸ€— Upvotes: 76 | cs.CL Authors: Xinyu Geng, Xuanhua He, Sixiang Chen, Yanjing Xiao, Fan Zhang, Shijue Huang, Haitao Mi, Zhenwen Liang, Tianqing Fang, Yi R. Fung Title: DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Arxiv: http://arxiv.org/abs/2607.07820v2 Abstract: Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and ver...
Published: Jul 22, 2026Duration: 19m 42s
Episode 2072
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
πŸ€— Upvotes: 74 | cs.CL Authors: Qing Zong, Yue Guo, Mengxin Yang, Yiwen Guo, Yangqiu Song Title: EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Arxiv: http://arxiv.org/abs/2607.17250v1 Abstract: This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture how characters and worlds evolve together over time. To address this, EvolvingWorld models literary simulation as...
Published: Jul 22, 2026Duration: 20m 46s
Episode 2071
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
πŸ€— Upvotes: 65 | cs.CL, cs.SE Authors: Yuhang Wang, Yuling Shi, Shaoqiu Zhang, Jialiang Liang, Shilin He, Siyu Ye, Yuting Chen, Kai Cai, Xiaodong Gu Title: SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Arxiv: http://arxiv.org/abs/2607.18213v1 Abstract: Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations indicating the relevance of code context whe...
Published: Jul 22, 2026Duration: 19m 58s
Episode 2070
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
πŸ€— Upvotes: 77 | cs.RO Authors: Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang, Jianwei Gao, Shijian Lu, Chengju Liu, Jianfei Yang, Mingxiu Chen, Deli Zhao Title: RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Arxiv: http://arxiv.org/abs/2607.17977v1 Abstract: We present RynnBrain 1.1, a family of emb...
Published: Jul 22, 2026Duration: 22m 34s
Episode 2069
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
πŸ€— Upvotes: 47 | cs.CV Authors: Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo Title: HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement Arxiv: http://arxiv.org/abs/2607.18217v2 Abstract: Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and...
Published: Jul 22, 2026Duration: 19m 11s
Episode 2068
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
πŸ€— Upvotes: 49 | cs.RO, cs.CV Authors: Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li Title: Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning Arxiv: http://arxiv.org/abs/2607.14183v2 Abs...
Published: Jul 22, 2026Duration: 21m 19s