
Sign up to save your podcasts
Or


【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 15 篇论文如下:
[00:27] 🎮 AgentGarten: Code Worlds for Evolving Agents(AgentGarten:面向进化智能体的代码世界)
[01:03] 🎮 Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?(Learn2Play Bench:LLM 智能体在陌生环境中从经验中学习的效果如何?)
[01:43] 🚦 TokenRouter: Efficient Serving System for Token-Level LLM Routing(TokenRouter:面向 Token 级 LLM 路由的高效服务系统)
[02:29] 🌍 From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation(从轨迹到智能体世界:用于交互式环境模拟的智能体语言世界模型)
[03:26] 🤖 SuperNav: An Agentic Navigation System for Any Task in Any Scene(SuperNav:面向任意场景中任意任务的智能体导航系统)
[04:11] 🚀 MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement(MiMo-V2.6:将强化学习扩展至自我改进)
[04:55] 🤖 In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks(让机器人上下文学习变简单:面向操作任务的普惠方案)
[05:39] 🤖 Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction(面向细粒度具身交互的多智能体自我中心世界模型)
[06:27] 🤖 DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training(DreamTrue:基于反事实后训练的动作忠实机器人世界模型)
[07:05] 🌀 OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs(OuroWorld:将任意3D世界变为多样且无限循环的3D动态影像)
[07:50] ⚡ MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers(MC-Sparse:解构并弥合扩散 Transformer 中的稠密-稀疏注意力差距)
[08:34] 🔗 Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching(超越时空先验:面向稠密对应匹配的可泛化方法)
[09:16] 🧪 TestPrism: Rethinking Test Evaluation Beyond a Single Reference(TestPrism:重新思考超越单一参考的测试评估)
[10:02] 🎨 Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards(通过组合偏好奖励与评分规则奖励对前沿文本到图像模型进行后训练)
[10:51] ✂ SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference(SparseDecoding:面向准确高效大语言模型推理的解码感知剪枝)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 15 篇论文如下:
[00:29] 🗜 STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization(STEPQuant:Delta-Rule 递归状态量化中的误差何时与何处重要)
[01:12] 🤖 Long-WAM: Scaling the Context of World-Action Models(Long-WAM:扩展世界-动作模型的上下文)
[01:53] 🤖 nanoMuse: An Open-Source Personal Agent for Every Device You Own(nanoMuse:面向你拥有的每台设备的开源个人智能体)
[02:33] 🎮 Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness(递归游戏创作者:面向体验的智能体产品级游戏开发框架)
[03:17] 🧠 Questioning the Questions: Sustaining Self-Evolution in Reasoning Models(质疑问题本身:维持推理模型的自我进化)
[04:09] ⚡ GRACE: Generation-aware latent compression for efficient video generation(GRACE:面向高效视频生成的生成感知潜在压缩)
[04:48] 🎭 DecepEval: A Benchmark for Evaluating Deception in LLM Agents(DecepEval:用于评估LLM智能体欺骗行为的基准)
[05:33] 🔮 VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction(VepAgent:通过工具增强强化学习桥接因果转移以实现视频事件预测)
[06:14] 🎬 SGF+: Decoupling Gradient Flows for Autoregressive Video Generation(SGF+:解耦自回归视频生成的梯度流)
[06:58] 🤖 UniWAM: Unified World-Action Model(UniWAM:统一世界-动作模型)
[07:36] 🧠 Semifactual Credit-Augmented Policy Optimization(半事实信用增强策略优化)
[08:15] 🗂 RunningTab: Direct Workspace Interaction with Environment-Side Tabs(RunningTab:基于环境侧标签的直接工作区交互)
[08:54] 🔊 WorldSonus: Bringing Sound to Worlds(WorldSonus:为世界带来声音)
[09:42] 🧩 Tetris3D: 3D Scene Generation With Objects That Fit Together(Tetris3D:生成物体相互契合的3D场景)
[10:25] 🤖 ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation(ReSAIL:缓解迭代式智能体自蒸馏中的崩溃)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 5 篇论文如下:
[00:50] TOP1(🔥609) | 🕵 False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents(虚假前沿:诊断与缓解自演化搜索智能体中的共同作弊)
[02:49] TOP2(🔥563) | 🤖 Raven: The Harness of Harnesses for Composable Agentic Intelligence(Raven:面向可组合智能体智能的“框架之框架”)
[05:15] TOP3(🔥532) | 🧭 The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation(教师是方向,而非终点:在在线策略蒸馏中外推强化学习诱导的表征残差)
[07:52] TOP4(🔥467) | 🔁 LoopVL: Recurrent Visual Intelligence(LoopVL:循环视觉智能)
[09:59] TOP5(🔥411) | 🎨 MaLiang-Harness: A Programmable Path to Image and Video Generation(MaLiang-Harness:通往图像与视频生成的可编程路径)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 15 篇论文如下:
[00:29] 🎥 OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction(OneStreamer:统一流式视频交互中的感知、记忆与主动响应)
[01:11] 🔀 Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL(自适应奖励路由:通过前向过程强化学习实现联合音视频扩散的动态多奖励优化)
[01:55] 🧠 Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States(超越记忆:利用显式信念状态驾驭长时程智能体)
[02:38] 🤖 Agent Priors-guided Policy Learning(智能体先验引导的策略学习)
[03:25] 🌀 Hierarchical Continuous Diffusion Language Models(分层连续扩散语言模型)
[04:03] 👁 World Observer: Joint Actor-Observer Generation for Persistent World Modeling(World Observer:面向持久世界建模的联合行动者-观察者生成)
[04:54] 📉 Sharpening Tax in Post-Training(后训练中的锐化税)
[05:39] 🎯 ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization(ActiveSaddler:面向智能体执行框架优化的自动化课程学习)
[06:22] 🤖 A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review(可信AI审稿人缺失的一环:从修辞鲁棒性基准测试到SciCore评审)
[07:08] 🎮 ROWBench: Do Video Models Render What the Program Specifies?(ROWBench:视频模型能否渲染程序所指定的内容?)
[07:55] 🤖 AutoGUIWorld: Image Generators as Visual World Models for GUI Agent(AutoGUIWorld:将图像生成器用作 GUI 智能体的视觉世界模型)
[08:39] 🔍 Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation(基于跨执行框架适配的检索增强技能优化)
[09:23] ⚖ Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL(让稀疏奖励算数:多奖励强化学习中的密度感知奖励聚合)
[10:05] 🤖 Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding(去中心化 Master-Mind:多智能体路径规划中通过迭代意图去噪的联合动作精炼)
[11:02] 🧩 E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models(E-MoE:面向非因子化扩散语言模型的增强混合专家)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 15 篇论文如下:
[00:27] 🧭 The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation(教师是方向,而非终点:在在线策略蒸馏中外推强化学习诱导的表征残差)
[01:11] 🪞 UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement(UniEvo-VL:面向多模态模型自我改进的在线策略自蒸馏训练方案)
[01:54] 🕵 False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents(虚假前沿:诊断与缓解自演化搜索智能体中的共同作弊)
[02:42] 🤖 AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks(AREX-2:通过长时程反思任务推进自我改进智能体)
[03:29] 🖥 Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents(Mid-Harness:在模型与执行框架之间扩展终端智能体的动作)
[04:09] 🧬 EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery(EvoDuet:面向科学发现的网络搜索与任务求解双层协同演化)
[05:01] 🕵 WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents(WorldAuditBench:使用多模态智能体进行交互式3D世界审计)
[05:47] 🛠 Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI(测试时 AI4AI 中面向智能体执行框架设计的元技能学习)
[06:26] 🧠 EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making(EVOKE:激发智能体中的世界知识以实现可迁移决策)
[07:12] 🎮 RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement(RSIGame:具备递归自我改进能力的自主智能体游戏开发)
[07:53] 📉 More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models(更多选择,更少决策:类JEV直接决策模型中的序数尺度偏差)
[08:45] 🧠 Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering(Imagine3D-LLM:教会多模态大语言模型在回答前想象3D场景)
[09:31] 🛠 Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training(智能体错误数据集:面向失败分析与错误感知后训练,规模化构建5万条错误—诊断配对)
[10:21] 🏮 LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models(LANTERN:照亮语言模型中的隐藏数学知识)
[11:05] 🖼 It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them(并非图像所示:无关上下文会扰乱 VLM 评判模型却不为其提供信息)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 10 篇论文如下:
[00:38] TOP1(🔥810) | 🧠 LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence(LimiX-2:面向通用结构化数据智能的上下文机制网络)
[02:36] TOP2(🔥704) | 🎬 Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation(Vidu S2:实时交互、可编辑与空间视频生成)
[04:41] TOP3(🔥494) | 🤖 Raven: The Harness of Harnesses for Composable Agentic Intelligence(Raven:面向可组合智能体智能的“框架之框架”)
[06:43] TOP4(🔥493) | 🎓 StudentSim: Training LLM-based Student Simulators(StudentSim:训练基于大语言模型的学生模拟器)
[08:51] TOP5(🔥484) | 🤖 Scaling Automatic Research Agents via World Models(通过世界模型扩展自动研究智能体)
[11:10] TOP6(🔥431) | 🤖 Atria Dawn: The Dawn of Agentic Superintelligence(Atria Dawn:智能体超级智能的黎明)
[13:35] TOP7(🔥396) | 🤖 Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills(从仓库到技能:将GitHub代码库蒸馏为AI4AI技能)
[15:38] TOP8(🔥382) | 🎨 MaLiang-Harness: A Programmable Path to Image and Video Generation(MaLiang-Harness:通往图像与视频生成的可编程路径)
[17:42] TOP9(🔥376) | 🧠 Continual Learning Mechanisms Compose for Long-Horizon Memorization(持续学习机制组合用于长时程记忆)
[19:55] TOP10(🔥358) | 🤖 In-Context Learning for Robots: Methods and Applications(面向机器人的上下文学习:方法与应用)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 5 篇论文如下:
[00:48] TOP1(🔥236) | 🧠 Training Object Permanence in World Models(在世界模型中训练客体永久性)
[03:10] TOP2(🔥235) | 🎓 OmniEdu: Open Foundation Models for Learning and Teaching(OmniEdu:面向学习与教学的开放基础模型)
[05:22] TOP3(🔥218) | 🧬 RRSI: Regularized Recursive Self-Improvement of Agent Harnesses(RRSI:智能体运行框架的正则化递归自我改进)
[07:44] TOP4(🔥217) | 🎙 Realtime-Venus: A full-duplex interaction system with asynchronous delegation(Realtime-Venus:一种支持异步委派的全双工交互系统)
[09:47] TOP5(🔥160) | 🧭 The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks(有品味的智能体:长时程任务中的品味测量与提升)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 15 篇论文如下:
[00:28] 🧩 FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders(FuseReg:正则化层融合缓解表征自编码器中的重建-生成差距)
[01:12] 🔗 RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation(RayOrch:面向基础模型数据准备的受血缘控制多粒度数据流编程与执行)
[01:58] ⚡ Block Sparse Attention with Log-Linear Complexity(对数线性复杂度的块稀疏注意力)
[02:40] 🤖 InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data(InternW0-Δ:一个连接预测动态与动作、基于20K+小时开放数据的世界动作模型)
[03:24] 🖐 Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors(Tactile-JEPA:面向分布式触觉传感器的拓扑感知自监督表示学习)
[04:11] 🛰 Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning(利用预训练扩散模型与多模态条件增强摄影测量数字表面模型)
[04:58] 👁 FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance(FoMo:生成轨迹中的分叉时刻作为感知距离)
[05:42] 📊 Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem(真实环境中的 Jev:Jev 模型功能、应用与生态的数据驱动分析)
[06:30] 🎯 TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations(TrackEverything:通过去重3D场景表示实现长时程密集跟踪)
[07:21] 🧩 SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL(SLCA-GRPO:解决工具调用强化学习中的跨段信用误归因)
[08:08] 🎯 CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation(CARD:面向个性化文本生成的聚类级适配与奖励引导解码)
[08:54] ⚖ Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs(隐式个性化与显式风格会冲突吗?PsPLUG:用于平衡定制化 LLM 中个性化与风格的轻量级插件)
[09:41] 🤝 AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs(AgentWorld:多智能体大语言模型长时程协作基准测试)
[10:24] 🎮 Game Arena: Strategic LLM Evaluation in Competitive Environments(游戏竞技场:竞争环境中的大语言模型策略评估)
[11:08] 🏦 IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking(IndicBankBench:评估印度零售银行中语言模型助手的安全性与可靠性)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 15 篇论文如下:
[00:27] 🧠 Training Object Permanence in World Models(在世界模型中训练客体永久性)
[01:12] 🧠 Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs(你的Transformer能同时容纳两个想法:LLM中线性叠加的证据)
[01:53] 🎬 WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation(WanPE:面向现代文本到视频生成的电影级提示增强)
[02:40] 🧭 OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents(OmniEcho:面向全模态具身智能体的音视频空间理解)
[03:18] 🤖 Agent-Editing World Model: Rethinking World Modeling for LLM Agents(智能体编辑世界模型:为LLM智能体重新思考世界建模)
[04:05] 🧩 Parts-of-Speech as Emergent Categories in SAE Latent Space(词性作为SAE潜空间中的涌现类别)
[04:46] 🤖 Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents(Qwen-Planner-Agent:面向真实世界移动规划智能体的闭环 AI-for-AI 框架)
[05:20] 🔍 IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis(IterSynth:通过角色解耦的迭代合成重新思考深度搜索智能体)
[06:04] 🤖 Coding Agents for Generalized Task and Motion Planning Problems(面向泛化任务与运动规划问题的编码智能体)
[06:49] 🧠 Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone(神经谱容量:仅凭网络规格测量与设计架构)
[07:28] 🛡 AgentKernel: The Trust-Native Agentic Operating System(AgentKernel:信任原生的智能体操作系统)
[08:15] 🧪 Rufus-Air: An Open LLM Post-Training Recipe(Rufus-Air:开放的大语言模型后训练配方)
[08:59] 🤖 PUBG Ally: A Conversational Embodied Agent as an AI Teammate(PUBG Ally:作为AI队友的对话式具身智能体)
[09:44] 🤖 World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal(世界动作智能体:利用视觉语言模型通过世界动作预演实现机器人操作)
[10:27] 🛸 ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds(ExplorationBench:测量 AI 系统在可验证异星世界中的探索能力)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43
【目录】
本期的 15 篇论文如下:
[00:28] 🧠 SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue(SpeakerMem-R1:以说话者为中心的多方对话双轨记忆)
[01:13] 🤖 Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World(Spatial-Interactor:通过与可观测物理世界交互学习空间推理)
[01:59] 🌍 HappyWorld-Bench(快乐世界基准(HappyWorld-Bench))
[02:40] 🧠 The Past Frames the Future: Memory for Autoregressive Video Generation(过往帧塑造未来:自回归视频生成中的记忆机制)
[03:29] 🧠 Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents(即时记忆:学习为LLM智能体整理任务自适应记忆)
[04:11] 🎬 RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling(RewardVerse:面向视频奖励建模的评分量规引导策略优化)
[05:00] 🎯 PACT: From Credit Assignment to Critic Alignment(PACT:从信用分配到评论家对齐)
[05:43] 🧪 Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?(薛定谔的代码仓库:LLM 是学会了 SWE-bench,还是记住了它?)
[06:24] 📐 GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression(GeoPair:用于免训练 Transformer 压缩的几何保持跨层因子分解)
[07:00] 📦 PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing(PackLab:在机器人装箱中开发、训练与评估多模态大语言模型的综合框架)
[07:53] 🧠 MemBodied: Recurrent Associative Memory for Vision-Language-Action Models(MemBodied:面向视觉-语言-动作模型的循环联想记忆)
[08:39] 🧪 WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents(WhatWorkedBench:基准测试AI智能体的实验理解能力)
[09:21] 🧠 Hunyuan-A13B Technical Report(混元-A13B 技术报告)
[10:03] 🎮 Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms(可验证隐藏动力学游戏:从已求解机制生成智能体强化学习环境)
[10:45] 🎥 All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation(所有模态都是平等的,但视频更平等:弥合联合视频生成中的交叉注意力差距)
【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
From the publisher's feed

294 Listeners

304 Listeners

12 Listeners

253 Listeners