
Daily Paper Cast
byJingwen Liang, Gengyu Wang
ScienceTechnology
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.com Creator: Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/ Gengyu Wang, LLM ML, http://wanggengyu.com Listen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236 Cover Image by Kawen Kuang https://kawen.art
Episodes(40 episodes)
Episode 2067
Apple-$Ο$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
π€ Upvotes: 39 | cs.CV
Authors:
Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu
Title:
Apple-$Ο$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Arxiv:
http://arxiv.org/abs/2607.16401v1
Abstract:
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a fa...
Published: Jul 22, 2026Duration: 22m 3s
Episode 2066
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
π€ Upvotes: 119 | cs.SE, cs.AI
Authors:
Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yang, Mingxi Cheng, Qi Dai, Bei Liu, Kai Qiu, Yue Dong, Ji Li, Chong Luo
Title:
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Arxiv:
http://arxiv.org/abs/2606.29538v4
Abstract:
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial videos and other multimodal human resources largely underused. We pre...
Published: Jul 21, 2026Duration: 19m 13s
Episode 2065
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
π€ Upvotes: 115 | cs.CL, cs.AI
Authors:
Mikhail Komarov, Ivan Bondarenko, Stanislav Shtuka, Oleg Sedukhin, Roman Shuvalov, Yana Dementyeva, Matvey Solovyov, Nikolay O. Nikitin
Title:
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
Arxiv:
http://arxiv.org/abs/2607.11683v1
Abstract:
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval. RAGU, an open-source modular GraphRAG engine, addresses this by separating extraction from consolidation: entities and relations pass through two...
Published: Jul 21, 2026Duration: 19m 44s
Episode 2064
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
π€ Upvotes: 57 | cs.RO, cs.CV
Authors:
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou
Title:
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Arxiv:
http://arxiv.org...
Published: Jul 21, 2026Duration: 20m 54s
Episode 2063
Loop the Loopies!
π€ Upvotes: 56 | cs.CL, cs.AI
Authors:
Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai
Title:
Loop the Loopies!
Arxiv:
http://arxiv.org/abs/2607.16051v2
Abstract:
We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N times increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N tim...
Published: Jul 21, 2026Duration: 19m 35s
Episode 2062
xHC: Expanded Hyper-Connections
π€ Upvotes: 47 | cs.LG, cs.CL
Authors:
Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan
Title:
xHC: Expanded Hyper-Connections
Arxiv:
http://arxiv.org/abs/2607.14530v1
Abstract:
Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising sca...
Published: Jul 21, 2026Duration: 20m 38s
Episode 2061
Cura 1T: Specialized Model for Agentic Healthcare
π€ Upvotes: 43 | cs.AI
Authors:
actAVA AI, :, Haolin Chen, Leon Qi, Steve Brown, Deon Metelski, Tao Xia, Joonyul Lee, Qixuan Wang, Kevin Riley, Frank Wang, Weiran Yao
Title:
Cura 1T: Specialized Model for Agentic Healthcare
Arxiv:
http://arxiv.org/abs/2607.15314v1
Abstract:
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in dif...
Published: Jul 21, 2026Duration: 20m 49s
Episode 2060
On-Policy Delta Distillation
π€ Upvotes: 28 | cs.LG, cs.CL
Authors:
Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
Title:
On-Policy Delta Distillation
Arxiv:
http://arxiv.org/abs/2607.15161v1
Abstract:
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output dis...
Published: Jul 21, 2026Duration: 19m 41s
Episode 2059
RecGPT-V3 Technical Report
π€ Upvotes: 26 | cs.IR
Authors:
Bowen Zheng, Chao Yi, Dian Chen, Gaoyang Guo, Han Zhu, Jiakai Tang, Jian Wu, Mao Zhang, Wen Chen, Yifan Lu, Yujie Luo, Yuning Jiang, Zhujin Gao, Bo Zheng, Dixuan Wang, Hao Fang, Jiancai Liu, Jing Yu, Ke Chen, Kewei Zhu, Mingke Xu, Wenjun Yang, Xunke Xi, Zile Zhou
Title:
RecGPT-V3 Technical Report
Arxiv:
http://arxiv.org/abs/2607.15591v1
Abstract:
Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pio...
Published: Jul 21, 2026Duration: 22m 0s
Episode 2058
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
π€ Upvotes: 109 | cs.CV
Authors:
Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
Title:
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Arxiv:
http://arxiv.org/abs/2607.14935v1
Abstract:
Recent advances in video understanding have spanned motion, long video, and...
Published: Jul 18, 2026Duration: 24m 41s
Episode 2057
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
π€ Upvotes: 71 | cs.CL
Authors:
Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao
Title:
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Arxiv:
http://arxiv.org/abs/2607.14777v1
Abstract:
Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between epi...
Published: Jul 18, 2026Duration: 19m 16s
Episode 2056
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
π€ Upvotes: 83 | cs.LG, cs.DC
Authors:
Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Steven Chiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, Cheng Jin
Title:
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Arxiv:
http://arxiv.org/abs/2607.14952v1
Abstract:
A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K t...
Published: Jul 18, 2026Duration: 19m 36s
Episode 2055
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
π€ Upvotes: 49 | cs.AI, cs.IR
Authors:
Yuyao Zhang, Junjie Gao, Zhengxian Wu, Jiaming Fan, Jin Zhang, Shihan Ma, Yao Yao, Weiran Qi, Chuyan Jin, Guiyu Ma, Xingzhong Xu, Kai Yang, Ji-Rong Wen, Zhicheng Dou
Title:
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Arxiv:
http://arxiv.org/abs/2607.15257v1
Abstract:
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current sin...
Published: Jul 18, 2026Duration: 15m 57s
Episode 2054
BadWAM: When World-Action Models Dream Right but Act Wrong
π€ Upvotes: 36 | cs.LG, cs.RO
Authors:
Qi Li, Xingyi Yang, Xinchao Wang
Title:
BadWAM: When World-Action Models Dream Right but Act Wrong
Arxiv:
http://arxiv.org/abs/2607.15207v1
Abstract:
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we sho...
Published: Jul 18, 2026Duration: 21m 3s
Episode 2053
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
π€ Upvotes: 30 | cs.CV
Authors:
Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang
Title:
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Arxiv:
http://arxiv.org/abs/2607.14202v1
Abstract:
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it rem...
Published: Jul 18, 2026Duration: 22m 20s
Episode 2052
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
π€ Upvotes: 29 | cs.CV, cs.SD
Authors:
Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li
Title:
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Arxiv:
http://arxiv.org/abs/2607.14189v1
Abstract:
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these set...
Published: Jul 18, 2026Duration: 20m 51s
Episode 2051
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
π€ Upvotes: 23 | cs.LG
Authors:
Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro VΓ©lez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
Title:
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
Arxiv:
http://arxiv.org/abs/2607.13188v1
Abstract:
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws $\textit{together}$, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet...
Published: Jul 18, 2026Duration: 21m 40s
Episode 2050
From Pixels to States: Rethinking Interactive World Models as Game Engines
π€ Upvotes: 23 | cs.CV
Authors:
Zhen Li, Zian Meng, Shuwei Shi, Mingliang Zhai, Jiaming Tan, Chuanhao Li, Kaipeng Zhang
Title:
From Pixels to States: Rethinking Interactive World Models as Game Engines
Arxiv:
http://arxiv.org/abs/2607.14076v1
Abstract:
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by predicting future observations conditioned on user actions, and are increasingly regarded as potential next-generation game engines. Rea...
Published: Jul 18, 2026Duration: 17m 56s
Episode 2049
UniVR: Thinking in Visual Space for Unified Visual Reasoning
π€ Upvotes: 22 | cs.CV
Authors:
Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin
Title:
UniVR: Thinking in Visual Space for Unified Visual Reasoning
Arxiv:
http://arxiv.org/abs/2607.12800v1
Abstract:
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and ste...
Published: Jul 18, 2026Duration: 18m 17s
Episode 2048
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
π€ Upvotes: 172 | cs.AI, cs.SE
Authors:
Ruhan Wang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Yue Yu, Junyao Yang, Kishan Panaganti, Haitao Mi, Dongruo Zhou, Leoweiliang
Title:
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Arxiv:
http://arxiv.org/abs/2607.13285v1
Abstract:
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a c...
Published: Jul 17, 2026Duration: 19m 12s