每日简报

2026-06-15

← 历史归档

iptv-org/iptv

TypeScript · ★ 120,943 · 🍴 6,472 · 📈 1,528 stars today

Collection of publicly available IPTV channels from all over the world

中文介绍 该项目是一个全球公开可用的 IPTV 频道集合。它为各类 IPTV 播放器提供了大量可用的电视频道源,主要采用 M3U 格式组织频道列表,方便用户获取和观看网络电视节目,适合想免费收看全球电视内容的爱好者。

freeCodeCamp/freeCodeCamp

TypeScript · ★ 447,197 · 🍴 44,951 · 📈 146 stars today

freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.

中文介绍 这是 freeCodeCamp.org 的开源代码库与课程体系。它提供完全免费的交互式学习平台,涵盖数学、编程和计算机科学等课程。通过项目实战和认证,帮助零基础学习者系统掌握 Web 开发、数据科学等技能,适合所有编程初学者。

pytest-dev/pytest

Python · ★ 14,037 · 🍴 3,179 · 📈 14 stars today

The pytest framework makes it easy to write small tests, yet scales to support complex functional testing

中文介绍 pytest 是一个功能强大的 Python 测试框架。它让编写小型单元测试变得简单,同时也能很好地支持复杂的功能和集成测试。通过 fixture、插件和参数化等特性简化了测试代码的编写和维护,是 Python 开发者进行自动化测试的首选工具。

swc-project/swc

Rust · ★ 33,778 · 🍴 1,401 · 📈 163 stars today

Rust-based platform for the Web

中文介绍 SWC 是一个基于 Rust 编写的、面向 Web 的超快平台。它主要用作 JavaScript/TypeScript 的编译器和打包工具,凭借 Rust 的高性能,在编译和构建速度上远超传统工具如 Babel 或 Webpack,适用于追求极致构建效率的大型前端项目。

chatwoot/chatwoot

Ruby · ★ 31,220 · 🍴 7,586 · 📈 400 stars today

Open-source live-chat, email support, omni-channel desk. An alternative to Intercom, Zendesk, Salesforce Service Cloud etc. 🔥💬

中文介绍 Chatwoot 是一个开源的实时聊天与客服支持平台。它集成了在线聊天、邮件、社交媒体等多渠道客服功能,是 Intercom、Zendesk 等商业产品的替代方案。适合中小企业快速搭建统一的客户沟通界面,提升服务效率。

NVIDIA/SkillSpector

Python · ★ 5,296 · 🍴 405 · 📈 964 stars today

Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, and security risks.

中文介绍 SkillSpector 是由 NVIDIA 开发的 AI 智能体技能安全扫描器。它能自动检测 AI 技能(如 LLM 应用)中的安全漏洞、恶意模式及潜在风险。采用静态分析等方法,帮助 AI 开发者和运维团队在部署前发现并修复安全问题。

meshery/meshery

TypeScript · ★ 10,401 · 🍴 3,420 · 📈 20 stars today

Meshery, the cloud native manager

中文介绍 Meshery 是一个云原生管理平台。它作为一个统一的控制平面,用于管理 Kubernetes、服务网格等云原生基础设施,提供可视化界面来部署、配置和管理各种云原生应用。适合运维和 DevOps 团队简化复杂的云原生环境操作。

cypress-io/cypress

TypeScript · ★ 49,936 · 🍴 3,413 · 📈 39 stars today

Fast, easy and reliable testing for anything that runs in a browser.

中文介绍 Cypress 是一款专为现代 Web 应用打造的端到端测试工具。它提供快速、可靠且易于调试的测试体验,可直接在浏览器内运行。通过自动等待、时间旅行调试等功能,帮助前端开发者编写和维护高质量的 UI 与集成测试。

GorvGoyl/Clone-Wars

★ 35,461 · 🍴 3,185 · 📈 269 stars today

100+ open-source clones of popular sites like Airbnb, Amazon, Instagram, Netflix, Tiktok, Spotify, Whatsapp, Youtube etc. See source code, demo links, tech stack, github stars.

中文介绍 该项目收录了 100 多个流行网站(如 Airbnb、Netflix、Instagram 等)的开源克隆版本。它提供了这些项目的源代码、演示链接和技术栈信息,是开发者学习现代 Web 和移动端应用架构、获取项目灵感的绝佳资源库。

Introduction-to-Autonomous-Robots/Introduction-to-Autonomous-Robots

TeX · ★ 2,721 · 🍴 613 · 📈 293 stars today

Introduction to Autonomous Robots

中文介绍 这是一个关于自动驾驶机器人的开源课程或教材项目。它旨在系统介绍自主机器人的基本概念、关键技术和相关算法,内容可能涵盖传感器、导航、规划等,适合机器人领域的学生和爱好者进行入门学习。

shiyu-coder/Kronos

Python · ★ 29,909 · 🍴 5,151 · 📈 244 stars today

Kronos: A Foundation Model for the Language of Financial Markets

中文介绍 Kronos 是一个专为金融市场语言设计的基础模型(Foundation Model)。它旨在理解和生成金融领域的文本、数据或报告,可能应用于市场情绪分析、信息抽取或自动化报告生成,服务于金融分析师和研究机构。

music-assistant/server

Python · ★ 2,185 · 🍴 437 · 📈 197 stars today

Music Assistant is a free, opensource Media library manager that connects to your streaming services and a wide range of connected speakers. The server is the beating heart, the core of Music Assistant and must run on an always-on device like a Raspberry Pi, a NAS or an Intel NUC or alike.

中文介绍 Music Assistant 是一个开源媒体库管理器的核心服务器。它连接本地文件和各种流媒体服务(如 Spotify、Tidal),并能将音乐发送到众多无线音箱和播放器设备,实现统一管理和跨设备播放,适合拥有多源音乐库的爱好者。

Free-TV/IPTV

Python · ★ 16,936 · 🍴 2,543 · 📈 70 stars today

M3U Playlist for free TV channels

中文介绍 该项目主要提供用于观看免费电视节目的 M3U 播放列表。它为用户收集和整理了公开可用的免费电视频道流地址,方便在 IPTV 播放器软件中直接导入和使用,获取全球免费的电视内容。

puppeteer/puppeteer

TypeScript · ★ 94,645 · 🍴 9,439 · 📈 29 stars today

JavaScript API for Chrome and Firefox

中文介绍 Puppeteer 是一个提供高级 API 的 Node.js 库,用于控制 Chrome 或 Firefox 浏览器。它支持无头模式,广泛应用于网页抓取、生成截图/PDF、自动化测试和 UI 自动化等场景,是前端测试和爬虫开发者的常用工具。

andrewyng/aisuite

Python · ★ 14,392 · 🍴 1,504 · 📈 291 stars today

Simple, unified interface to multiple Generative AI providers

中文介绍 AISuite 由吴恩达团队开发,旨在为多种主流生成式 AI 提供商(如 OpenAI、Google 等)提供一个简单统一的接口。开发者可以用一套代码轻松切换和调用不同模型,降低了应用开发和比较不同 AI 服务的成本。

A frontier without an ecosystem is not stable

@satyanadella · 5.9M 粉丝 · 1.9M 阅 · 3.5K 赞 · 541 转

I’ve been thinking a lot about the future of the firm in an AI-driven economy. This transition is different than any previous platform shift. In the past, we used digital systems to enhance human

中文介绍 Satya Nadella 探讨 AI 驱动经济中“公司”的未来。他指出,当前的平台转变与以往不同,不再是用数字系统来增强人类,而是一场更深层的结构性变革。

Codex-maxxing: treating Codex like an operating loop

@BradGroux · 5.9K 粉丝 · 714.6K 阅 · 1.0K 赞 · 638 转

Most people still use coding agents like fancy autocomplete or a one-shot chat box. That leaves a lot of value on the table. The better pattern is to treat Codex like a durable operating loop:

中文介绍 分享将 Codex 用作“持久运行循环”的高级工作流。作者认为多数人只将编码代理当作单次聊天工具,而更好的模式是将其构建为一个包含目标设定、执行和反馈的持续操作循环,以挖掘更多价值。

Anthropic is losing the mandate of heaven

@haridigresses · 12.5K 粉丝 · 281.7K 阅 · 513 赞 · 36 转

Four months ago, in early February, Anthropic was the darling. OpenAI was the dominant behemoth to root against. Over the last 1-2 years, we'd seen the Sam ouster / return drama, Ilya and Mira had

中文介绍 对 Anthropic 公司声誉转变的行业观察。作者指出,在 OpenAI 发生动荡后,Anthropic 曾被视为更受青睐的“安全牌”,但近来其行业地位和公众观感已发生显著变化。

Autonomous Long-Running Coding Agents

@omarsar0 · 307.3K 粉丝 · 81.2K 阅 · 518 赞 · 66 转

Autonomous coding is moving from better prompting to better control systems. The important shift is that engineers are learning how to wrap agents in goals, evaluators, loops, and artifacts that let

中文介绍 分析自主编码代理的技术演进趋势。其核心是从优化提示词转向构建更完善的控制系统,工程师开始学习如何用目标、评估器、执行循环和产物来封装代理,使其能长期自主运行。

Anthropic's War on Opensource AI

@TheAhmadOsman · 61.0K 粉丝 · 74.9K 阅 · 507 赞 · 98 转

Anthropic wants the public to see one thing: the careful lab, the safety lab, the grown-up in the room trying to keep frontier AI from running off a cliff. However, the pattern around Anthropic does

中文介绍 对 Anthropic 公司策略的批评分析。作者认为 Anthropic 对外塑造“谨慎安全实验室”的形象,但其实际行为模式(如对待开源 AI)与之存在差距,引发了“表里不一”的讨论。

Mastering Codex (Mobile) for Engineering

@Dimillian · 51.8K 粉丝 · 38.2K 阅 · 558 赞 · 37 转

How to turn your phone into a Codex control center?? At first glance, it looks like a way to check on a Codex task from your phone. That is useful, but it misses the bigger idea. The power of Codex

中文介绍 分享在移动端高效使用 Codex 的技巧。核心观点是不应只将手机用于查看任务状态,而应将其转变为实时的“Codex 工程控制中心”,实现更灵活的项目管理和干预。

Building recursive agent systems

@leerob · 258.6K 粉丝 · 36.8K 阅 · 586 赞 · 40 转

At Cursor, we run thousands of agents to help us train the next version of Composer. We give them research tasks, and if they aren't succeeding or run into issues, they DM us on Slack or page us via

中文介绍 分享在 Cursor 内部构建递归代理系统的实践案例。公司运行数千个代理进行研究任务来训练下一代 Composer,当代理遇到问题时,会通过 Slack 或寻呼系统主动向工程师求助,形成高效反馈闭环。

Introducing the OpenAI Partner Network

OpenAI launches the Partner Network, investing $150M to help global partners accelerate enterprise AI adoption, deployment, and transformation.

中文介绍 OpenAI推出合作伙伴网络,将投资1.5亿美元,旨在帮助全球合作伙伴加速企业人工智能的采用、部署和转型。

New OpenAI Academy courses for the next era of work

OpenAI introduces three Academy courses that help people build practical AI skills, create repeatable workflows, and apply agents in everyday work.

中文介绍 OpenAI推出三门新的学院课程,旨在帮助人们在工作中构建实用的AI技能、创建可重复的工作流程并应用AI代理。

[AINews] Loopcraft: The Art of Stacking Loops

a quiet day lets us highlight a great concept from Peter Steinberger, Boris Cherny, and Andrej Karpathy

中文介绍 人工智能新闻报道了Peter Steinberger、Boris Cherny和Andrej Karpathy提出的「Loopcraft」概念。

How Preply combines AI and human tutors to personalize learning

Preply uses OpenAI to launch AI-generated lesson summaries, providing personalised feedback and language learning exercises.

中文介绍 语言学习平台Preply利用OpenAI技术推出AI生成的课程总结,提供个性化反馈和语言练习。

Google DeepMind is worried about what happens when millions of agents start to interact

Google DeepMind is funding research into the potential dangers of situations where millions of different AI agents interact with each other online. According to Rohin Shah, who directs the company’s AGI safety and alignment research, the mass-market arrival of agents that can carry out tasks without

中文介绍 谷歌DeepMind正资助研究,以评估数百万不同AI代理在线交互可能带来的潜在风险。

How an astrophysicist uses Codex to help simulate black holes

Discover how astrophysicist Chi-kwan Chan uses Codex to build black hole simulations, helping scientists study extreme physics and test Einstein’s theory of general relativity.

中文介绍 天体物理学家Chi-kwan Chan利用OpenAI的Codex工具构建黑洞模拟,以帮助研究极端物理现象和验证爱因斯坦的广义相对论。

BBVA puts AI at the core of banking with OpenAI

Learn how BBVA scaled ChatGPT Enterprise to 100,000 employees and partnered with OpenAI to accelerate AI-powered banking transformation worldwide.

中文介绍 西班牙对外银行(BBVA)将ChatGPT企业版扩展至10万名员工,并与OpenAI合作,以加速全球范围内由AI驱动的银行业转型。

OpenAI to acquire Ona

OpenAI plans to acquire Ona to expand Codex with secure, persistent cloud environments, enabling long-running AI agents across enterprise workflows.

中文介绍 OpenAI计划收购Ona,以将安全、持久的云环境集成到Codex中,从而支持企业工作流中长时间运行的AI代理。

Supporting Europe’s work in ensuring a trustworthy AI ecosystem

OpenAI supports the EU Code of Practice on AI content transparency, advancing provenance standards and tools to help people understand AI-generated content.

中文介绍 OpenAI支持欧盟在AI内容透明度方面的实践准则,以推进溯源标准和工具,帮助人们理解AI生成的内容。

Beyond the IT Checklist: Engineering a Reasonable Standard of Care for Cyber Safety

第一作者: Matthew E. Jablonski · 方向: 软件安全

Abstract:Current U.S. cyber policy, centered on security, often treats documentation of controls and incident reports as a proxy for safety in the built environment. This paper argues that such an approach is inadequate for cyber-physical systems, where digital failures can produce kinetic harm. We construct and code a corpus of critical infrastructure policy documents (N=292, 2000-2025) to examine how "reasonable care" is operationalized across the NIST SP 800-160 Vol.~2 resilience lifecycle. The resulting maps show that obligations are concentrated in the Anticipate phase and emphasize administrative compliance, while Withstand and Recover phases rely heavily on delegated references to IT-focused control catalogs that are poorly aligned with physics-based hazards. We identify three major disconnects: miscalibrated delegated standards, recovery defined as notification rather than...

论文介绍 本文探讨美国网络政策中基于文档的合规性不足问题,尤其针对网络物理系统。研究构建关键基础设施政策语料库,分析NIST SP 800-160 Vol. 2弹性生命周期中"合理关怀"的操作化。发现政策义务过度集中于预期阶段,而承受和恢复阶段依赖IT控制,与物理风险不匹配,揭示三个主要脱节,为改进网络安全标准提供参考。

Differentially Private Hierarchical Heavy Hitters

第一作者: Ari Biswas · 方向: 安全研究

Abstract:The task of finding _Hierarchical_ Heavy Hitters (HHH) was introduced by Cormode et al. [VLDB 2003] as a generalisation of the heavy hitter problem. While finding HHH in data streams has been studied extensively, the question of releasing HHH when the underlying data is private remains unexplored. In this paper, we study differentially private HHH release in both the streaming and non-streaming setting. In the non-streaming setting, we show the surprising result that the relative error in estimating the residual count for any prefix is independent of the height of the hierarchy and the number of heavy hitters in the stream. Meanwhile, in the streaming setting, although the exact version of HHH has low global sensitivity (as counting queries are 1-sensitive), the approximation functions due to streaming have high global sensitivity, linear in the available space. Despite this...

论文介绍 本文研究差分隐私约束下分层重击者的发布问题。在非流式设置中,估计残差计数的相对误差独立于层次高度和重击者数量;在流式设置中,近似函数因高全局敏感性带来挑战。提出算法解决隐私保护下的数据流分析,适用于统计查询和隐私数据分析应用。

Intent-Based Cryptographic API Design for Cryptographic Agility

第一作者: Navaneeth Rameshan · 方向: 密码学协议

Abstract:As organizations move toward post-quantum cryptography, they face the major challenge of updating cryptographic algorithms across large, complex software portfolios. However, most cryptographic APIs in use today were designed around specific algorithms. These APIs expect explicit use of specific algorithms, provide little or no support for policy-based algorithm selection, and offer no straightforward way to migrate existing keys to newer algorithms. This makes the transition to post-quantum cryptography challenging. The companion assessment framework identifies the barriers to cryptographic agility and explains why algorithm transition is largely a software engineering problem. To address the limitations of current cryptographic APIs, we identify the principles necessary to design a cryptographically agile API. The design principles are derived from five fundamental...

论文介绍 当前密码学API基于特定算法设计,阻碍后量子密码学过渡。本文指出其缺乏灵活性,提出基于意图的API设计原则,支持策略化算法选择和密钥迁移,旨在简化软件更新过程,为构建密码学灵活系统提供指导。

An Assessment Framework for Application-Level Cryptographic Agility

第一作者: Navaneeth Rameshan · 方向: 密码学协议

Abstract:The impending post-quantum transition to new cryptography will require complete replacement of algorithms within all software. The cryptographic APIs used today make this transition challenging because they were not designed with agility as a concern. There is no method for systematically assessing cryptographic agility as an overall ability. In addition to this, the term itself refers to multiple independent capabilities. Specifically, it includes replacing algorithms, selecting by policy, and substituting implementations. This lack of structured decomposition limits both the evaluation of systems and the development of cryptographically agile APIs. We introduce a component-based assessment framework that characterizes application-level cryptographic agility along seven orthogonal dimensions: three coupling dimensions that measure what the application code knows about...

论文介绍 密码学灵活性涵盖算法替换、策略选择和实现替代等多个能力。本文引入组件化评估框架,沿七个正交维度特征化应用层密码学灵活性,为系统评估和API设计提供结构化方法,有助于量化和提升密码学灵活性。

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

第一作者: Zihao Wang · 方向: AI 安全

Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences. This makes them vulnerable to prompt-injection attacks, in which seemingly benign content embeds adversarial instructions that manipulate agent behaviour. Existing security benchmarks adopt an \textit{attack-centric} perspective, focusing on the technical feasibility of injections while overlooking the nuanced distribution of resulting harms. In practice, however, prompt-injection risk is victim-dependent: a single exploit can produce asymmetric consequences for different stakeholders, and the same attack pattern may exhibit substantially different effectiveness depending on whom it targets. To capture these properties, we introduce \textbf{\sysname}, a \textit{stakeholder-centric} benchmark...

论文介绍 大型语言模型驱动的网络代理易受提示注入攻击。现有安全基准以攻击为中心,忽视不同利益相关者的受害差异。本文提出利益相关者为中心的基准,评估攻击对不同目标的影响,更全面衡量风险,为真实世界部署提供安全评估工具。

The Invisible Ink of the Android Malware World: A Longitudinal Study on the Usage of Covert Communication Channels

第一作者: Zeya Umayya · 方向: 软件安全

Proxies, VPNs and Tor have long helped the privacy community and users in censored regions to fight censorship. However, the same tools can be maliciously exploited by malware and botnets to conceal their communication to external command and control servers. Despite being a critical concern fueled by the proliferation of malware based attacks, no longitudinal studies have analyzed how malware applications use covert channels (CC) to evade detection. We fill this gap by performing the first study of the usage of covert channels in the Android malware ecosystem. To that end, we develop a multistage pipeline that combines static and dynamic analysis to investigate both system and network-level features. We applied this pipeline on a corpus of 3.5M Android malware spanning 2009 to July 2025. Our carefully crafted static validation rules uncovered 288K APKs that used CCs spanning 511...

论文介绍 Android恶意软件常使用代理、VPN等隐蔽通信渠道规避检测。本文进行首次纵向研究,开发多阶段管道分析2009-2025年的350万恶意软件样本,揭示隐蔽渠道的使用模式和演变趋势,为恶意软件检测和防御提供数据支持。

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

第一作者: Jiaqi Luo · 方向: AI 安全

Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross. Within this broader red-line scenario, autonomous penetration represents a core enabling capability and subtask: the ability of LLM-powered AI systems to independently conduct adversarial operations against a target server without human intervention, identify and exploit vulnerabilities, and obtain unauthorized access or control. A growing body of work has sought to assess the autonomous penetration capabilities of AI systems. However, existing evaluations often employ opaque methodologies, rely on unrealistic or overly simplified penetration-testing scenarios, or provide LLMs with excessive prior knowledge and task-specific guidance, and cannot accurately capture the extent to which modern AI...

论文介绍 自主渗透是大型语言模型AI系统的关键安全红线。现有评估方法存在方法不透明、场景简化等问题,无法准确捕捉能力。本文分析自主渗透能力的评估挑战,旨在定义AI安全边界和改进评估方法,推动安全AI系统发展。

DIG: Oracle-Guided Directed Input Generation for One-Day Vulnerabilities

第一作者: Andrew Bao · 方向: 软件安全

One-day vulnerabilities pose significant risks due to delayed or incomplete patch adoption. Generating proof-of-concept (PoC) inputs is therefore essential for assessing real-world impact. The key challenge is identifying necessary constraints for triggering the vulnerability and solving them effectively. Existing directed fuzzing approaches prioritize inputs toward target locations, but neither explicitly identify necessary constraints nor solve them effectively, relying instead on target-distance feedback and random mutation. Agentic approaches show strong potential through code reasoning and structured input generation, but goal drift in long-horizon reasoning limits their effectiveness. DIG addresses this challenge by exploiting a key property of one-day vulnerabilities: patches often reveal necessary preconditions for triggering. DIG uses an LLM to analyze the patch and synthesize...

论文介绍 一日漏洞因补丁延迟带来风险,生成PoC输入至关重要。DIG利用LLM分析补丁以合成必要约束,解决定向模糊测试中的目标漂移问题,提升漏洞触发输入生成的效率,适用于漏洞评估和补丁验证场景。

SoK: The Constant Time Model

第一作者: Billy Bob Brumley · 方向: 软件安全

Abstract:Constant time programming patterns is the primary defense against timing attacks on cryptographic implementations, yet what "constant time" means varies across academia and industry. This work systematizes constant time models and their evolution, identifies a recurring gap between what models protect and what specifications assume, and distills an offensive methodology for discovering timing vulnerabilities that originate outside the cryptographic primitive boundary. Applying this methodology, we locate a specification-level vulnerability related to private key loading, and confirm the leak in both OpenSSL and BoringSSL. Counterintuitively, BoringSSL's per-observation signal is several orders of magnitude stronger than OpenSSL's, despite an explicitly stricter threat model.

论文介绍 研究问题:密码学实现中常量时间编程模式的定义不一致,导致防御漏洞。核心方法:系统化常量时间模型及其演化,识别模型保护与规范假设间的差距,并提炼攻击方法以发现时间漏洞。可能应用:应用于私钥加载,在OpenSSL和BoringSSL中确认规范级漏洞,强调威胁模型对信号强度的影响。

ViPER: Vision-based Packing-Aware Encoder for Robust Malware Detection

第一作者: Fatima Qaiser · 方向: 软件安全

Abstract:Visualization-based malware detection maps raw binary bytes to grayscale images and applies learned visual classifiers, providing an evasion-resistant and disassembly-free alternative to conventional analysis pipelines. However, executable packing remains a critical failure mode: packed binaries produce high-entropy images that obscure the structural patterns these models rely on. Because packing is also prevalent in benign software (e.g., for compression or copy protection), packing state alone is not a reliable indicator of maliciousness, and existing approaches do not address this challenge within a unified supervised framework. We present ViPER, a Vision-based Packing-Aware Encoder for Robust malware detection. ViPER builds on a LoRA-adapted ViT-B/14 backbone with a dual-head architecture that jointly learns malware classification and packing detection. A packing-aware...

论文介绍 研究问题:现有基于可视化的恶意软件检测在打包可执行文件上失败,因为高熵图像掩盖结构模式。核心方法:提出ViPER,一种基于视觉的打包感知编码器,使用LoRA-adapted ViT-B/14骨干和双头架构联合学习恶意软件分类和打包检测。可能应用:提高恶意软件检测的鲁棒性,减少传统分析中打包的干扰。

MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

第一作者: Chejian Xu · 方向: 软件安全

Abstract:Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering. In these systems, safety and security are inherently distributed across role-specialized agents, significantly expanding the attack surface, particularly under coordinated adversarial behaviors such as privilege escalation and cross-agent collusion. Existing red-teaming approaches for MAS remain limited: they rely on heuristic selection of target agents and perturb isolated message streams, leaving critical questions unanswered as which agents are most responsible for system safety, and how compromised agents can coordinate to bypass defenses. We propose MAStrike, a closed-loop framework for collusive red-teaming in hierarchical MAS. We propose the first agent-level Shapley value analysis for MAS, quantifying each agent's marginal...

论文介绍 研究问题:层次化多智能体系统的安全漏洞,特别是智能体协调对抗行为如权限升级和跨代理共谋。核心方法:提出MAStrike闭环框架,首次引入智能体级Shapley值分析量化安全贡献,进行共谋红队测试。可能应用:识别关键智能体,评估系统防御,增强多智能体系统安全性。

LNTest: A Testbed for Evaluating Bitcoin Lightning Network-Based Botnets

第一作者: Thomas Bakaysa · 方向: 密码学协议

Abstract:Bitcoin's Lightning Network (LN) can be exploited as a covert, low-cost command-and-control (C&C) channel for botnets, as demonstrated by the LNBot and D-LNBot designs. However, both remain proof-of-concept prototypes evaluated only through simulation, leaving key questions about real-world topology formation, propagation complexity, and resilience to takedowns unanswered. We present LNTest, the first reusable testbed for LN-based botnets, built from Core Lightning nodes containerized with Docker over a shared Bitcoin Core regtest chain. LNTest supports three overlay topology modes (a deterministic chain, autonomous peer discovery, and user-supplied graphs), enabling controlled experiments across different botnet structures. Using LNTest, we report three main findings. First, D-LNBot's autonomous formation protocol does not produce the uniform chain from its design; instead...

论文介绍 研究问题:比特币闪电网络可被用作僵尸网络的隐蔽命令与控制通道,但现有评估仅基于模拟,缺乏真实世界测试。核心方法:开发LNTest可重用测试床,基于Docker容器化节点和regtest链,支持三种拓扑模式进行控制实验。可能应用:评估闪电网络僵尸网络的形成、传播和抗打击能力,为安全研究提供工具。

A Privacy-Preserving Framework Using Remote Data Science for Inter-Institutional Student Retention Prediction

第一作者: John Fields · 方向: AI 安全

This study explores privacy-preserving machine learning (PPML) techniques using the PySyft platform to enable collaborative prediction of student retention between institutions. We developed a remote data science (RDS) framework with a semi-air-gapped architecture consisting of high-side and low-side servers, allowing researchers from three universities to build predictive models on sensitive student data without direct data access. Using historical data from a small private university (N=720), we evaluated three synthetic data generation approaches and validated the framework through inter-institutional collaboration. The results demonstrate consistent classification performance across institutions (Macro F1: 0.690--0.695) while maintaining strict Family Educational Rights and Privacy Act (FERPA) compliance. We also propose Data-Type-Aware Templates, a novel synthetic data method that...

论文介绍 研究问题:机构间协作预测学生留存需保护敏感数据隐私。核心方法:开发基于PySyft的远程数据科学框架,采用半气隙架构,允许在不直接访问数据下构建模型。可能应用:实现跨机构协作,保持FERPA合规,分类性能稳定。

Semantic Identification of IoT Devices from Behavioral Primitives

第一作者: Samuel Witt · 方向: 密码学协议

Abstract:Accurate identification of IoT devices is important for security management and policy enforcement. Existing approaches typically learn device signatures from packets or flow records. These methods operate on low-level communication observations whose traffic patterns may vary across deployments, software versions, and user interactions. This paper studies device identification using Manufacturer Usage Description (MUD) profiles. MUD profiles describe device behavior using Access Control Entries (ACEs), where each ACE represents a behavioral primitive consisting of protocol, endpoint, direction, and port semantics derived from device communication policy. Our contributions are threefold. First, using 28 publicly available MUD profiles containing 1,023 ACE instances, we construct ACE-level semantic representations from compact behavioral text and analyze their geometric...

论文介绍 研究问题:物联网设备识别依赖低层通信观察,易受部署、软件版本和用户交互变化影响。核心方法:使用制造商使用描述(MUD)配置文件中的行为原语(如访问控制条目)构建语义表示进行设备识别。可能应用:提高设备识别准确性,支持安全管理和策略执行。

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

第一作者: Pengfei He · 方向: AI 安全

Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources. Existing defenses mainly focus on blocking malicious content at inference time, and current red-teaming methods primarily optimize attack success. As a result, developers have limited visibility into how latent prompt injections emerge and propagate through agents. We propose PI-Hunter, an automated agentic auditing framework for proactive vulnerability exposure in LLM agents. PI-Hunter constructs realistic source-aware test cases and iteratively evolves them through feedback-driven exploration to induce agents to retrieve and reveal latent malicious instructions embedded within external environments. Extensive experiments across multiple benchmarks, agent...

论文介绍 研究问题:大型语言模型代理面临间接提示注入攻击,现有红队方法聚焦攻击成功,缺乏漏洞暴露。核心方法:提出PI-Hunter自动化框架,构建源感知测试用例,通过反馈驱动探索诱导代理检索和揭示嵌入环境中的恶意指令。可能应用:主动暴露LLM代理安全漏洞,提高防御能力。

SMSR: Certified Defence Against Runtime Memory Poisoning in Persistent LLM Agent Systems

第一作者: Tarun Sharma · 方向: 软件安全

Abstract:Retrieval-augmented generation (RAG) agents increasingly run with persistent memory that accumulates across user sessions. This creates a new attack surface: an adversary interacting only through normal channels can inject crafted memories that, once retrieved, steer the agent's responses for future users, without touching model weights or code. We call this Multi-Session Memory Poisoning (MSMP) and show that no existing defence certifies against it; static-corpus defences (RobustRAG, ReliabilityRAG) assume a fixed knowledge base, and heuristic filters are bypassed by fluent enterprise-style text. We present Signed Memory with Smoothed Retrieval (SMSR), the first defence with a certified robustness bound for this setting. Component 1 adds HMAC-SHA256 provenance at write time, blocking unsigned injection. Component 2 applies randomised memory ablation with verdict-based...

论文介绍 研究问题:检索增强生成代理的持久记忆存在多会话记忆投毒攻击风险,现有防御缺乏认证鲁棒性。核心方法:提出SMSR防御,使用HMAC-SHA256签名阻止未签名注入,并采用随机记忆消融提供认证鲁棒性界限。可能应用:保护持久记忆LLM代理免受记忆投毒,确保系统安全。

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents

第一作者: Siyu Shen · 方向: 隐私保护

Screenshot-based mobile GUI agents can operate ordinary smartphone apps through the same visual interface as a human user, but this capability also turns every screen observation into a privacy boundary. During normal task execution, screenshots may expose contacts, messages, photos, files, recommendations, health cues, and other sensitive context that is unrelated to the user's request. We call this problem incidental visual privacy exposure. It is difficult to address with existing defenses: text anonymization misses many visual and inferential cues, while generic privacy masking can remove the evidence and controls that a GUI agent needs to complete the task. This paper presents CAPED, a context-aware pre-upload exposure control layer for mobile GUI agents. CAPED is designed as a phone-side protection layer: before screenshots are released to a remote multimodal agent, it extracts...

论文介绍 基于截图的移动GUI代理在执行任务时可能意外暴露联系人、消息等敏感信息,称为附带视觉隐私暴露。现有防御方法如文本匿名化或通用掩码效果有限。CAPED提出一种上下文感知的预上传曝光控制层,作为手机侧保护,在截图发送到远程多模态代理前进行提取和处理,以平衡任务完成与隐私保护。

Amnesia: A Stealthy Replay Attack on Continual Learning Dreams

第一作者: Ahmed Sharshar · 方向: 安全研究

Continual learning (CL) models often use experience replay to reduce catastrophic forgetting, but their robustness to replay sampling interference remains underexplored. Existing CL attacks alter inputs or training pipelines (poisoning/backdoors) and rarely include explicit auditable constraints, limiting realism. Here, auditability means a monitor can verify compliance from sampler-visible telemetry - e.g., logged replay index/label statistics - by checking that the realized replay class histogram stays close to a nominal baseline and that replay rate is unchanged per batch and/or over a rolling window. We study a limited-privilege insider who controls only replay index selection, not pixels, labels, or model parameters, while staying within auditable limits such as queue priorities. We introduce Amnesia, a replay composition attack that maximizes degradation under two budgets: a...

论文介绍 持续学习模型常使用经验回放减少灾难性遗忘,但其回放采样鲁棒性未充分研究。现有攻击多针对输入或训练流程,缺乏可审计性。Amnesia是一种回放组合攻击,在仅控制回放索引选择且遵守可审计限制(如队列优先级)下,最大化模型性能退化,揭示了持续学习系统在回放采样中的安全漏洞。

Beyond Attack Success Rate: Examining Trigger Leakage in Vision-Language Agentic Systems

第一作者: Jiamin Chang · 方向: 系统安全

Vision-Language Agentic Systems (VLAS) connect visual perception to planning, tool use, and physical actions. This means backdoor-type triggers can propagate through both decision pipelines and their connected interfaces, thus making visual backdoors a system-level threat. Current evaluations on such backdoors focus on clean accuracy and attack success rate (ASR), metrics that capture whether a trigger works, but not whether an attack is actually "precise" -- i.e. whether it triggers hidden behaviors only when intended. In this work, we formalize the failure of trigger precision as "trigger leakage": inputs that are visually or semantically close to the intended trigger and therefore inadvertently activate the attacker-specified behavior. To quantify this leakage, we introduce Neighbor Leakage Rate (NLR). Our experiments show that at a 3% poisoning ratio, icon and text triggers remain...

论文介绍 视觉-语言智能体系统将视觉感知与规划、工具使用等连接,使后门触发器可传播到决策管道和接口,构成系统级威胁。当前评估仅关注攻击成功率,忽略触发器精确性。本文形式化触发器泄漏问题,即视觉或语义上接近的输入意外激活攻击行为,并引入邻居泄漏率来量化,以改进安全评估。

From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model Merging

第一作者: Zhenqian Zhu · 方向: 安全研究

Abstract:Model merging (MM) has gained significant attention as a cost-effective approach to integrate multiple task-specific models into a unified model. However, recent work reveals that MM is highly susceptible to backdoor attacks. Existing defenses based on task arithmetic often fail to eliminate backdoors without substantially degrading clean-task performance, owing to their reliance on direct parameter-space editing. To address this gap, we propose Linear Feature Path Minimization (LFPM), a backdoor mitigation framework for model merging, which introduces an anti-backdoor task vector into the backdoored merged model. Unlike prior approaches, LFPM formulates the backdoor robustness of the merged model from a unified feature-space perspective under the Cross-Task Linearity (CTL) framework, which leverages the approximate linearity of features across tasks. This perspective guides...

论文介绍 模型合并是集成多任务模型的有效方法,但易受后门攻击。现有基于任务算术的防御依赖参数空间编辑,难以消除后门而不影响性能。本文提出线性特征路径最小化框架,从特征空间视角在跨任务线性性下引入反后门任务向量,以缓解后门并保持模型性能。

Influence Factors on RAG Poisoning

第一作者: Pedro Pereira · 方向: AI 安全

Abstract:Retrieval-Augmented Generation (RAG) systems enhance large language models by grounding responses in retrieved documents from external knowledge sources at inference time. However, this reliance on retrieved content introduces vulnerabilities to poisoning attacks, in which adversarial documents can manipulate both the retrieval process and the generated outputs. This paper investigates poisoning robustness in RAG through a full factorial experimental study covering 432 configurations. We analyze the impacts of dataset, retriever type, retrieval depth, database composition, chunking strategy, and generator model on retrieval-level and generation-level metrics. The results show that retriever architecture, dataset, and retrieval depth are the strongest factors affecting poisoning exposure, while generator choice and database composition have a major impact on downstream attack...

论文介绍 检索增强生成系统通过检索外部文档增强大语言模型,但依赖检索内容易受投毒攻击。本文通过全因子实验研究432种配置,分析数据集、检索器类型、检索深度等因素对投毒鲁棒性的影响。结果发现检索器架构、数据集和检索深度是主要影响因素,为系统安全设计提供依据。

Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks

第一作者: Achraf Hsain · 方向: AI 安全

Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata restricting an agent's actions. We argue this is the wrong product. The same automata-theoretic machinery -- specification compilation, product game construction, attractor computation, and winning-region extraction -- is better read as a design-time analytical instrument whose outputs are structural insights about a system rather than runtime constraints on a deployed agent. We instantiate this through a constrained two-player safety game for network defense. The two specifications are enforced asymmetrically: the defender specification defines the unsafe region of the game, whereas the attacker specification restricts the adversary's legal actions during attractor computation. Solving the game yields a defensibility verdict -- a formal...

论文介绍 屏蔽强化学习通常作为运行时安全机制,将时序逻辑规范编译为自动机以限制智能体动作。本文认为这非最佳应用,同一自动机理论工具更适合作为设计时分析工具。通过构建双人安全游戏,其中防御者和攻击者规范不对称执行,求解游戏生成可防御性判定,用于网络防御的结构化安全分析。

Split Tallies: A Discrete Certificate Calculus for Auditing Dynamic Ordered Sets in Constant Memory

第一作者: Faruk Alpay · 方向: 安全研究

Abstract:We study retrospective auditing for dynamic ordered sets maintained by an untrusted party. A passive auditor watches insert, delete, membership, predecessor, successor, min, and max operations, stores five machine words and a flag, and receives a constant-size public tally record per operation. At audit time the maintainer discloses the claimed live vacant intervals. The method represents order semantics by maximal gaps: gaps are born, cited, consumed, and timestamped, while two hidden field accumulators test equality of the birth and consumption ledgers. Honest executions are accepted with probability one. If any answer in a T-operation session is wrong, acceptance occurs with probability at most (4T+1)/p over one secret field element, against computationally unbounded maintainers. We prove that deterministic and visible-coin auditors require linear state, and that removing...

论文介绍 针对不受信任方维护的动态有序集,回溯审计需监控插入、删除等操作。本文提出Split Tallies方法,使用常数内存和公开记录进行审计,代表顺序语义通过最大间隙的生成、引用和消耗。诚实执行总是被接受,而错误操作导致接受概率有界,提供高效安全审计方案。

Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models

第一作者: Jianwei Fei · 方向: 安全研究

Abstract:Model fingerprinting, embedding user-specific identifiers (fingerprints) into generated outputs, has recently emerged as a popular solution to protect the intellectual property rights (IPR) of generative text-to-image (T2I) models and prevent unauthorized redistribution. In this work, we reveal a previously unexplored systematic vulnerability in existing generative model fingerprinting methods: they lack robustness against collusion attacks, where multiple attackers combine their models to remove or obscure the fingerprints. To address this issue, we take the first step towards a robust fingerprinting method for T2I models with anti-collusion capabilities. The proposed method encodes strings of bits, namely fingerprints, into the coefficients of a personalized normalization module (PNM) incorporated into T2I models, so that fingerprints can be reliably recovered from any...

论文介绍 生成模型指纹用于保护文本到图像模型的知识产权,但现有方法缺乏抗共谋攻击能力。本文提出将指纹编码到个性化归一化模块系数中,实现高效、鲁棒和抗共谋的指纹方法。该方法可从任何输出可靠恢复指纹,以应对多个攻击者结合模型去除指纹的威胁。

PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learning Agent

第一作者: Junfeng Guo Heng Huang · 方向: AI 安全

Abstract:While real-world applications of reinforcement learning (RL) are becoming increasingly popular, the security of RL systems deserve more attention and exploration. In particular, recent work has revealed that RL agents are vulnerable to backdoor attacks, where a victim agent behaves normally under standard conditions but executes malicious actions when a specific trigger is activated. Existing backdoor defenses for RL either require access to the agent's internal parameters, operate only at the model or trajectory level, or are limited to specific attack types. To ensure the security of RL agents, we propose \texttt{PolicyGuard}, a \textit{test-time step-level} backdoor defense which leverages Gaussian Process (GP) posterior variance and adapts pseudo trajectories to enable uncertainty computation for individual time step. Besides, we also provide theoretical foundations to...

论文介绍 该研究针对强化学习代理易受后门攻击的问题,提出PolicyGuard防御方法。该方法在测试时对每个时间步进行不确定性估计,利用高斯过程后验方差和伪轨迹实现步级防御。现有防御方法多限于模型或轨迹级别,PolicyGuard提供了更精细的机制,有助于提升RL系统在现实应用中的安全性。

Detecting Functional Memorization in Code Language Models

第一作者: Matthieu Meeus · 方向: 软件安全

Abstract:Large language models (LLMs) are increasingly used to generate code at scale. Meanwhile, prior work has investigated whether training data may be recoverable from model outputs, by auditing the textual overlap between training examples and model generations. Code, however, can be functionally equivalent while textually dissimilar. In this work, we study functional memorization: extraction of functional logic beyond what verbatim metrics detect. We construct a counterfactual setup for Olmo-3-32B, comparing a midtrained model (exposed to target code) against a pretrained reference (not exposed). We prompt both models with Python function signatures and measure both textual and functional similarity (i.e., LLM-as-a-judge, execution-based). Our results show clear evidence of functional memorization, highlighting the need for auditing metrics that go beyond textual overlap.

论文介绍 本文研究代码语言模型中的功能性记忆,即训练数据的功能逻辑提取。研究构建反事实设置,比较中训练模型和预训练模型,通过LLM作为裁判和执行测量来验证功能性记忆。结果发现功能性记忆现象,强调需要超越文本重叠的审计指标,以更好地保护代码生成模型的隐私。

Smarter Saboteurs, Better Fixers: Scaling & Security in Linear Multi-Agent Workflows

第一作者: Timothy McAllister · 方向: AI 安全

Abstract:As LLM-based multi-agent systems (MAS) are deployed in the wild, the resilience of their collaboration structures against adversarial compromise becomes a critical safety concern. Attackers may leverage prompt-injection or jailbreaking to sabotage individual agents within MAS workflows, but the interaction between model scaling and system-level resilience remains poorly understood. This paper investigates how model scale affects the security of linear multi-agent workflows. Our experiments across scales of two open-weight model families on the HumanEval benchmark reveal a compliance-correction symmetry: larger models are far more likely to faithfully execute malicious instructions, with the control-to-malicious performance drop reaching 53.7pp at 27B in uncorrected pipelines. However, appending a lightweight terminal Fixer stage collapses this to 0.6pp and restores statistical...

论文介绍 本文探究模型规模对基于LLM的多智能体系统安全性的影响。实验基于HumanEval基准,涉及两种开源模型家族,显示较大模型更易执行恶意指令,但在工作流末尾添加修复阶段可显著降低风险。研究揭示了「合规-纠正对称性」,为设计安全协作系统提供实践指导。

Fed-FBD: Federated Functional Block Diversification for Isolation, Privacy, and Surgical Unlearning

第一作者: Weijie Chen · 方向: AI 安全

Abstract:Federated learning (FL) enables collaborative model training without sharing raw patient data, but standard approaches such as FedAvg treat each client as a black box and provide no mechanism for isolating an adversarial contributor, auditing per-client influence, or honoring a departed participant's right to be forgotten. We present Fed-FBD (Federated Functional Block Diversification), a modular federated architecture that decomposes a ResNet backbone into six functional blocks (the stem, four residual groups, and the classification head) and maintains a warehouse of N color variants, each assembled from independently tracked and contributor-stamped blocks. Fed-FBD provides three capabilities absent in FedAvg: (i) architecturally guaranteed block-level isolation, so that an adversarial or mislabelled client cannot contaminate the clean colous; (ii) privacy-by-design, where...

论文介绍 联邦学习缺乏隔离和遗忘机制,Fed-FBD通过将ResNet分解为功能块并维护变体仓库来解决。它提供块级隔离、隐私设计和精确遗忘能力,允许隔离对抗客户端、审计每个客户端影响,并支持参与者遗忘权利,增强了联邦学习的安全性和可控性。

SAIGuard: Communication-State Simulation for Proactive Defense of LLM Multi-Agent Systems

第一作者: Ruxue Shi · 方向: AI 安全

Abstract:LLM-based multi-agent systems (MAS) solve complex tasks through inter-agent collaboration, but their communication-driven nature also allows security risks to spread across agents and trigger system-wide failures. Existing MAS defenses mainly follow a reactive paradigm after execution by detecting and isolating harmful agents, which may cause irreversible damage and degrade collaborative utility. To address this, we propose a proactive defense framework for MAS security, namely a Simulation-aware Interception Guard (SAIGuard). SAIGuard performs communication-state simulation over the MAS interaction graph, estimates the impact of incoming messages on local agent states and the global MAS state, and detects risky messages via reconstruction deviations from benign communication patterns. Instead of isolating agents, SAIGuard sanitizes or regenerates suspicious messages before it...

论文介绍 现有LLM多智能体系统防御多为反应式,SAIGuard提出主动防御框架。它通过模拟通信状态和检测重建偏差来预防风险,在消息传播前进行拦截,估计其对智能体状态的影响,从而在系统层面净化或重生成可疑消息,避免故障并保持协作效用。

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

第一作者: Daichi Azuma · 方向: 导航与运动 · 来源: cs.RO

Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change brings it closer to the goal. Navigation world models provide such visual foresight, but they remain prediction modules that require an external planner to convert predicted futures into closed-loop control. We propose Navigation World Action Model (NavWAM), a diffusion-transformer policy that turns navigation world-model prediction into executable action by representing future observations, goal-progress values, and action chunks in a shared latent sequence. By learning future prediction jointly with the action and value targets that determine closed-loop behavior, NavWAM makes visual foresight directly usable for robot control. We build NavWAM through simulation pretraining and real-robot adaptation, and...

论文介绍 目标条件视觉导航需将预测转换为控制,NavWAM提出扩散变换器策略,联合学习未来观察、目标进度值和动作块。通过模拟预训练和实际适应,使视觉预测直接驱动机器人闭环控制,无需外部规划器,简化导航系统并提高效率。

See Selectively, Act Adaptively: Dual-Level Structural Decomposition for Bimanual Robot Manipulation

第一作者: Yoon-Ji Choi · 方向: VLA 通用模型 · 来源: cs.RO

In bimanual robotic manipulation, task-relevant visual information varies with the task stage and context, while the interaction of the two arms shifts between independent and coordinated modes, making policy learning challenging. However, existing monolithic Vision-Language-Action (VLA) policies process diverse visual inputs and interaction patterns through a single shared representation and action generation pathway, often failing to separately account for visual relevance and bimanual interaction structure. To address this issue, we propose a bimanual manipulation VLA framework based on Dual-Level Structural Decomposition. The View-Selective Visual Router dynamically adjusts wrist-view contributions to emphasize relevant visual cues, while the Interaction-Aware Action Mixture-of-Experts (MoE) decomposes action generation into coordinated and arm-wise pathways to adapt to varying...

论文介绍 双臂机器人操作中视觉信息和交互模式多变,现有VLA模型处理不足。本文提出双级结构分解框架,包括视觉路由器和动作混合专家,动态调整视觉输入贡献,分解动作为协调和独立路径,以适应双臂操作的不同阶段,提高策略的自适应能力。

EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation

第一作者: Kailin Wang · 方向: 机器人操作 · 来源: cs.RO

Pretrained-feature world models provide a useful substrate for robot imagination, but visual or latent prediction alone does not determine whether an imagined future satisfies task-relevant events. Long-horizon manipulation requires progress signals that are relational, predicate-level, and physically grounded: whether an object has moved, whether a drawer or contact state has changed, whether a placement predicate is satisfied, and whether a candidate future is reliable enough for execution. We introduce EA-WM, an event-aware world-model framework that augments frozen visual-feature dynamics with task-specification-grounded event prediction and verification. EA-WM rolls out candidate futures in pretrained visual-feature space, decodes them into structured event states, and scores them using task-progress, semantic-consistency, physical-feasibility, and uncertainty terms. The verifier...

论文介绍 长 horizon 操作需任务进展信号,EA-WM增强预训练世界模型 with 任务规范接地的事件预测和验证。它滚动候选未来,解码为结构化事件状态,基于任务进度、语义一致性、物理可行性和不确定性评分,为机器人想象提供可靠性保障,确保操作安全。

GeoCFNet: Geometry-Aware Confidence Field Network for Robot-Assisted Endoscopic Submucosal Dissection

第一作者: Rui Tang · 方向: 具身智能 · 来源: cs.CV

Advanced surgical robotics has made robot-assisted endoscopic submucosal dissection (ESD) a promising approach for the en-bloc resection of large lesions, with the potential to reduce recurrence and improve long-term outcomes. However, the technical complexity and risk of complications in ESD demand stable and precise visual guidance to maintain an accurate dissection corridor and a safe tissue margin. Dense confidence fields provide an effective representation for this purpose by describing both the preferred dissection region and its spatial transition to surrounding tissue. However, reliable confidence field estimation remains challenging in dynamic endoscopic scenes due to smoke, specular highlights, tissue deformation, weak texture, and the thin geometric structure of the target region. To address these challenges, we formulate dissection guidance as a geometry-aware confidence...

论文介绍 本文研究了机器人辅助内镜黏膜下剥离中的视觉引导问题。在烟雾、反光、组织变形等复杂内镜场景下,准确估计描述解剖区域和安全边界的置信场极具挑战。作者提出了几何感知置信场网络 GeoCFNet,该方法通过结合几何信息来提升动态手术场景中的置信场估计可靠性,旨在为机器人手术提供稳定精确的视觉引导,以维持安全的解剖分离路径。

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

第一作者: Yangcen Liu · 方向: 机器人操作 · 来源: cs.RO

Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but directly using them for robot learning requires bridging two gaps: the visual gap between human and robot observations, and the action gap between human motion and robot-executable action. We propose EgoEngine, a scalable framework for transforming egocentric human manipulation videos into high-fidelity robot data. Given an egocentric RGB video, EgoEngine produces: (i) a high-fidelity robot observation video replacing human with robot while preserving scene context and temporal alignment, and (ii) a task-aligned, executable robot action trajectory under feasibility constraints. Experiments in simulation and on real robots show that EgoEngine enables scalable conversion of human videos into robot data and, to...

论文介绍 灵巧操作机器人学习受限于大规模机器人演示数据的收集成本。本文提出了 EgoEngine 框架,旨在将低成本、可扩展的自我中心人类操作视频转换为高保真的机器人演示数据。该方法通过生成包含机器人视角的观测视频和对齐任务的可执行动作轨迹,来弥合人类与机器人之间的视觉和动作域差距,为机器人策略学习提供了新的数据来源。

Learning to Assist: Collaborative VLAs for Implicit Human-Robot Collaboration

第一作者: Leo Xu · 方向: VLA 通用模型 · 来源: cs.RO

Human-robot collaboration (HRC) combines the complementary strengths of humans and robots to improve task efficiency. However, many existing collaborative systems rely on hand-engineered pipelines, limiting their scalability and flexibility for new tasks. In this work, we show that models trained end-to-end with imitation learning, specifically vision-language-action (VLA) models, can support collaborative manipulation, and characterize the key factors affecting their real-world performance. We evaluate two state-of-the-art models and identify a failure mode of action-chunking policies in implicit HRC, where demonstration action leakage (i.e., action chunks crossing latent task transitions) can cause premature assistive behavior. We find that this issue increases with longer execution horizons and occurs in real-world collaborative VLA systems, such as when a robot attempts to hand...

论文介绍 现有人机协作系统常依赖手工设计的流程,限制了其可扩展性。本文探索了通过模仿学习训练的端到端视觉语言动作模型来支持隐式人机协作操作。研究评估了当前先进模型,发现动作块策略在隐式协作中存在「演示动作泄漏」问题,即动作块跨越了潜在的任务转换点,导致机器人过早执行辅助行为。该工作分析了影响此类系统实际性能的关键因素。

Mana: Dexterous Manipulation of Articulated Tools

第一作者: Zhao-Heng Yin · 方向: 机器人操作 · 来源: cs.RO

Abstract:Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal degrees of freedom and contact-rich interactions. While prior work has largely focused on rigid objects, articulated tool use remains underexplored because of its physical complexity and the difficulty of learning functional grasping and manipulation policies. We present Mana (Manipulation Animator), a general sim-to-real framework that reinterprets dexterous manipulation as an animation problem. Inspired by computer animation, Mana employs a coarse-to-fine pipeline that transforms procedurally-generated grasp keyframes into manipulation trajectories through motion planning and reinforcement learning. The data generation process is largely automatic, requiring only a few mouse clicks to specify functional affordances (<1 minute per tool). Across four articulated...

论文介绍 铰接工具因其内部自由度与丰富的接触交互,对灵巧机器人操作构成挑战。本文提出了 Mana 框架,将灵巧操作重新定义为动画问题。该框架采用从粗到精的流程:首先自动生成抓取关键帧,随后通过运动规划和强化学习将其转化为操作轨迹。这种方法实现了数据的自动化生成,只需少量交互即可为不同工具指定功能,为解决铰接工具操作提供了通用仿真到真实的路径。

Improving Robotic Generalist Policies via Flow Reversal Steering

第一作者: Andy Tang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Generalist policies can learn a wide range of skills from diverse robot datasets. In order to solve or improve on challenging new tasks, we need a way to infer and invoke the appropriate actions from the policy's rich behavioral prior, especially when directly commanding the policy fails. We focus on flow matching generalists and propose Flow Reversal Steering (FRS): a method that takes suboptimal but ``reasonable'' actions, finds their latent noises by passing them through the flow policy in reverse, and maps them to nearby generalist action modes. We evaluate FRS across many simulated and real-world manipulation settings. First, FRS can turn coarse semantic guidance from humans or vision-language models (VLMs) into corresponding good robot actions, improving zero-shot control. These gains can be distilled with behavioral cloning by training an auxiliary policy to output...

论文介绍 通用机器人策略能从多样数据中学习广泛技能,但在面对新任务时可能无法直接输出理想动作。本文针对流匹配类泛化策略,提出了流程逆转引导方法。该方法能从人类或视觉语言模型提供的粗糙语义引导或次优但合理的动作中,通过反向流程找到其潜在噪声,并映射回策略的良好动作模式,从而改进零样本控制性能,并可通过行为克隆将改进效果蒸馏到新策略中。

$\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

第一作者: Arnav Kumar Jain · 方向: 机器人操作 · 来源: cs.RO

Abstract:The potential impacts of world models (WMs, i.e., learned simulators) on robotics are far-reaching -- policy evaluation, policy improvement, and test-time planning -- all with limited real-world interaction. To unlock these downstream capabilities, a WM needs to jointly satisfy three desiderata: $\textit{(i)}$ fidelity (i.e., producing simulated trajectories that correlate with reality), $\textit{(ii)}$ consistency (i.e., producing simulated trajectories that are coherent over long horizons), and $\textit{(iii)}$ efficiency (i.e., producing simulated trajectories quickly). We propose $\texttt{WEAVER}$ (World Estimation Across Views for Embodied Reasoning): a WM architecture that simultaneously achieves all three desiderata, providing state-of-the-art results on robotic manipulation tasks. $\texttt{WEAVER}$ is a multi-view WM trained to predict future latents and reward values...

论文介绍 高质量的世界模型对机器人操作中的策略评估、改进和规划至关重要。本文提出了 WEAVER,一种多视角世界模型架构,旨在同时满足高保真度、长时一致性和高效率三个关键要求。该模型通过预测未来的潜在状态和奖励值,在机器人操作任务上达到了先进性能,为以有限真实世界交互来解锁世界模型的下游应用潜力提供了有效工具。

MCR-Bionic Hand: Anatomical Structural Priors for Dexterous Manipulation

第一作者: Haosen Yang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dexterous robotic hands are usually formulated as high dimensional active control systems governed by degrees of freedom, actuation, and algorithms. Human hand dexterity, however, is partly encoded in the physical architecture of bones, ligaments, tendons, aponeuroses, and intrinsic muscles. This work describes that contribution as two linked forms of structural intelligence: structural prior generation, in which wrist to finger tenodesis, FDS/FDP routing, and the dorsal extensor hood transform low dimensional posture inputs into default grasp configurations and PIP to DIP coordination; and muscle mediated modulation, in which extrinsic muscles, lumbricals, and interossei regulate MCP posture, distal stability, fingertip force paths, and contact states around that default state. Based on this framework, MCR-Bionic Hand is developed as a 1:1 musculoskeletal biomimetic hand...

论文介绍 人类手部灵巧性部分源于其骨骼、韧带和肌腱等物理结构编码的「结构智能」。本文基于此认识,提出了 MCR 仿生手设计。该设计利用了两种结构先验:一是由肌腱路由等结构生成默认抓取构型,二是由肌肉系统对姿态和接触状态进行调制。通过1:1 模拟肌肉骨骼系统,该仿生手旨在实现更自然、更高效的灵巧操作。

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale

第一作者: Nils Blank · 方向: 机器人操作 · 来源: cs.RO

Abstract:This work introduces Spatial Annotations from Robot Demonstrations with Reliability Calibration (SPARC), a risk-aware framework that automatically labels robot demonstrations with structured spatial annotations and assigns each annotation a reliability score. Structured spatial annotations, such as bounding boxes, object trajectories, and manipulation phase labels, benefit a broad range of robotics applications from training grounded robot policies and embodied foundation models to motion planning and hierarchical task composition. Existing automated pipelines generate such annotations at scale but provide no reliable quality signal: detector confidence is poorly calibrated for annotation correctness, forcing a choice between accepting noisy labels or discarding useful samples. In contrast to existing automated pipelines, SPARC leverages the spatio-temporal structure inherent...

论文介绍 为机器人演示自动生成结构化空间标注对多种应用有益。现有自动化流程缺乏可靠的质量信号。本文提出了 SPARC 框架,它利用演示数据固有的时空结构,自动为机器人演示生成如边界框、物体轨迹和操作阶段标签等结构化空间标注,并为每个标注分配一个经校准的可靠性评分。这种风险感知的标注方式为下游机器人学习和任务组合提供了更可靠的数据基础。

GIVE: Grounding Human Gestures in Vision-Language-Action Models

第一作者: Pengfei Liu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Human communication is inherently multimodal, where language is often accompanied by non-verbal cues such as gestures to convey intentions. However, current Vision-Language-Action (VLA) models treat robotic manipulation as a pure text-driven task, overlooking the important role of gestures in Human-Robot Interaction (HRI). This often leads to inaccurate intent grounding and unreliable manipulation when language instructions are ambiguous or underspecified. To address this challenge, we propose GIVE (Gesture Intent via Visual-Semantic Enhancement), an effective approach that enhances pre-trained VLA models with human gesture understanding without architectural modifications. Specifically, GIVE incorporates gesture information through two complementary pathways: a visual pathway that overlays hand skeletons and fingertip rays onto robot observations for explicit object...

论文介绍 本文针对当前视觉-语言-动作(VLA)模型在机器人操作中忽略手势等非语言线索的问题,提出GIVE方法。该方法通过视觉路径(叠加手部骨架和指尖射线)和语义路径,在不修改模型架构的情况下增强对手势的理解,从而在语言指令模糊时提升意图接地和操作可靠性,适用于人机交互场景。

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation

第一作者: Xiangyu Zhu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Whole-body mobile manipulation requires coordinating mobile base and manipulator under shifting viewpoints, posing challenges in geometric perception and action generation. Current policies either rely on 2D features or sparse 3D representations that lack dense spatial structure, and typically encode arm and base within one action vector that ignores their distinct control demands. Moreover, existing dense fusion strategies risk corrupting pretrained representations under noisy depth while incurring heavy computational overhead. We present GeoHAT, an end-to-end diffusion-based framework built on a simple principle: geometry should be injected only where reliable and attended to only where needed. GeoHAT employs a lightweight Fourier spatial encoder that maps dense per-pixel 3D coordinates into geometric tokens without an additional 3D vision backbone. These tokens are then...

论文介绍 本文解决全身移动操作中几何感知与动作生成的协调难题。提出GeoHAT端到端扩散框架,采用轻量傅里叶空间编码器将密集3D坐标注入为几何token,并运用混合注意力机制,在可靠时注入几何信息,在需要时关注,以提升机器人在变化视点下的操作性能。

Real-Time Execution with Autoregressive Policies

第一作者: Sangkyu Lee · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Real-time execution, enabled by asynchronous inference that ensures both smooth action trajectories and fast reactivity, is critical for realistic deployments of large-scale Vision-Language-Action models. However, recent work on real-time execution primarily focuses on variants of diffusion policies, even though it is more critical for autoregressive policies given their slower rollout speed in synchronous inference. In contrast, we demonstrate that autoregressive policies can achieve real-time execution by adjusting the tokenization horizon and applying constrained decoding, thereby guaranteeing strict latency bounds that enable multi-trajectory decoding to maximize performance. Across simulated and real-world environments, we find that the autoregressive policy consistently outperforms its equivalent-level flow-matching policy counterpart while achieving significantly...

论文介绍 本文探讨大规模视觉-语言-动作模型的实时执行问题,提出针对自回归策略的方法。通过调整tokenization horizon和应用约束解码,保证严格延迟界限,实现多轨迹解码以最大化性能,在模拟和真实环境中表现优于流匹配策略,促进机器人部署。

Low cost, easily manufactured, highly flexible strain and touch sensitive fiber for robotics applications

第一作者: Christian Diaz Herrera · 方向: 机器人操作 · 来源: cs.RO

Abstract:Existing stretch and touch sensors for robots are generally expensive with respect to at least one of material costs, required manufacturing equipment, or manufacturing time. We present and experimentally characterize a conductive fiber made using only inexpensive commercial off-the-shelf parts (conductive thread at $0.07/ft, silicone tubing at $0.94/ft) and tools (loop-style needle threader at $2), which can be manufactured quickly (20 cm length in 2 minutes.) We demonstrate its use as a resistive strain sensor with three applications: Triggering a grasp in a pneumatically actuated assistive finger, sensing the pose of a pneumatically actuated robotic strap, and estimating the pose of a flexible solid. We also demonstrate that it can be used as a capacitive sensor with two applications: First, as a touch sensor which triggers a commercial robot arm to move, and second, as a...

论文介绍 本文针对机器人传感器成本高、制造复杂的问题,提出一种低成本、易制造的柔性导电纤维。该纤维使用廉价商业部件制成,可作为电阻应变传感器和电容触觉传感器,应用于气动辅助手指抓取、机器人带姿态估计和柔性固体姿态估计等场景。

EMG-Based Adaptation of Anisotropic Virtual Fixtures for Robot-Assisted Surgical Resection and Dissection

第一作者: Dario Onfiani · 方向: 具身智能 · 来源: cs.RO

Abstract:In this paper, we address the development of an adaptive assistance system for robot-assisted laparoscopic surgery, specifically for delicate tasks such as Resection and Dissection. Even if Virtual Fixtures offer significant advantages for guiding a surgeon's movements, conventional Virtual Fixtures are often defined by fixed geometries, lacking the flexibility to adapt to the surgical workflow or the surgeon's immediate intent. To address these limitations, we propose a novel framework for an adaptive and anisotropic virtual fixture. In addition, we introduce an intuitive control interface that modulates the fixture's geometry in real-time based on the surgeon's intent, inferred from EMG signals. This approach allows the surgeon to dynamically expand or disengage the constraint by contracting their forearm muscles, enabling seamless transitions between precise guided motion...

论文介绍 本文开发用于机器人辅助腹腔镜手术的自适应辅助系统,针对切除和解剖任务。提出基于肌电信号(EMG)的自适应各向异性虚拟夹具框架,允许外科医生通过前臂肌肉收缩动态调整夹具几何,实现精确引导运动与自由操作的无缝切换,提升手术灵活性。

Humor Style Drives Laughter, Topic Shapes Acceptability: Evaluating Bilingual Personal and Political Robot-Delivered AI Jokes

第一作者: Anna-Maria Velentza · 方向: 具身智能 · 来源: cs.RO

Abstract:Humor plays a central role in human social relationships, and recent advances in computational humor create new opportunities for integrating humor into human-robot interaction (HRI). While large language models (LLMs) can generate diverse forms of humor, it remains unclear how humor style, joke content, and language preference shape perceptions of robot-delivered humor in group settings. In this exploratory study, we employed a mixed factorial design in which participants evaluated AI-generated jokes delivered by a robot in a university classroom. We examined the effects of humor type (Affiliative, Self-Enhancing, Aggressive, Self-Defeating) and joke content (person-related vs. political) on perceived funniness and appropriateness, as well as preferred language. Results show that humor type significantly influences funniness, with Aggressive and Affiliative humor rated...

论文介绍 本文探索机器人讲笑话时,幽默风格、笑话内容和语言偏好如何影响感知。通过混合因子实验,评估AI生成的笑话在教室环境中的表现,发现幽默类型显著影响趣味性,而笑话内容影响可接受性,为人机交互中的幽默集成提供参考。

WT-UMI: Tactile-based Whole-Body Manipulation via Force-Supervised Contact-Aware Planning

第一作者: Jaehwi Jang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Whole-body humanoid manipulation of bulky, deformable, and shared-load objects requires distributed contact sensing and explicit force regulation, yet most imitation policies treat contact force only implicitly. On the other hand, different demonstration sources provide complementary modalities with inherent trade-offs: human demonstrations capture natural contact forces but not robot-executable actions, while teleoperation directly records robot actions but with less natural force regulation. This paper presents \textbf{WT-UMI}, a wearable whole-body tactile interface worn by human operators or mounted on humanoids, providing accurate observations of tactile images, contact forces, and end-effector poses across both human demonstration and humanoid teleoperation modes. We introduce a force-conditioned target-pose correction module that converts measured human poses into...

论文介绍 本文针对人形机器人全身操作中分布式接触感知和力调节的需求,提出WT-UMI穿戴式触觉接口。该接口在人类示范和机器人遥操作模式下提供准确的触觉观测,并引入力条件目标姿态校正模块,实现基于力监督的接触感知规划,适用于笨重、可变形物体的操作。

Proprioceptive-visual correspondence enables self-other distinction in humanoid robots

第一作者: Yurun Chen · 方向: 具身智能 · 来源: cs.RO

Abstract:Distinguishing self from others is a prerequisite for social intelligence, yet humanoid robots that increasingly share workspaces with humans still lack this ability. Here we show that a humanoid robot can learn self-other distinction from proprioceptive-visual correspondence, without any identity labels or kinematic models. Once established, this distinction bootstraps a predictive self-model that maps joint configurations to three-dimensional body occupancy, capturing how the robot's body changes with action. In multi-agent scenes involving humans or morphologically identical robots, the system reliably identifies itself, learns a 3D self-model, and supports downstream tasks including target reaching, collision-aware motion planning, and human-to-robot motion retargeting. Together, these results outline a route toward bodily self-representation in robots that act and...

论文介绍 本文解决人形机器人区分自我与他人的难题,提出基于本体感觉-视觉对应的无监督方法。该方法无需身份标签或运动学模型,学习预测自我模型,支持目标到达、碰撞感知运动规划和动作重定向等下游任务,推动机器人社交智能发展。

Embedding ISO 10218 Safety Compliance in Robots via Control Barrier Functions for Human-Robot Collaboration

第一作者: Federico Parma · 方向: 导航与运动 · 来源: cs.RO

Abstract:Human-Robot Collaboration (HRC) requires strict adherence to safety standards, such as ISO 10218, to prevent harmful interactions. Standard Speed and Separation Monitoring (SSM) filters calculate safe robotic speeds based on conservative assumptions, such as constant human velocity, which prevents accurate predictions of minimum separation distances and causes unnecessary operational halts. This paper proposes a Control Barrier Function (CBF) that explicitly incorporates human acceleration data to analytically forward-predict the minimum human-robot separation distance during a worst-case robotic stopping trajectory. To guarantee safety at the control level, this predictive CBF is integrated as an inequality constraint within a Sequential Quadratic Programming (SQP) framework. Specifically, two methods are proposed: Method I, a CBF-constrained PD safety filter; and Method II...

论文介绍 研究问题:人机协作需严格遵守ISO 10218安全标准,但现有安全速度与分离监控方法基于保守假设如常数人类速度,导致不必要的操作停机。核心方法:提出一种控制障碍函数,整合人类加速度数据以分析性地前向预测最小人机分离距离,并集成到顺序二次规划框架中作为不等式约束,提供了两种安全滤波方法。可能应用:提高人机协作的安全性和效率。

Multi-Modal Multi-Agent Robotic Cognitive Alignment enabled by Non-Invasive Consumer Brain Computer Interfaces: A Proof of Concept Exploration

第一作者: Nataliya Kosmyna · 方向: 具身智能 · 来源: cs.RO

Abstract:While non-verbal behaviors and expressive movements are essential for natural human-robot interaction, existing methods often overlook a crucial element: the human's internal cognitive state. Frequently, proactive multi-agent systems can interrupt humans at inopportune moments, leading to cognitive overload and decreased task performance. This paper introduces a framework for generating "cognitively aligned" multi-agent interactions, enhancing the ability of robotic systems to contextually defer communications to the user of an agent system during moments of high human mental workload and engagement. We present the design and implementation of a closed-loop architecture that explores the interplay between autonomous task execution and real-time neurophysiological focus. Using a consumer-grade Brain-Computer Interface (BCI), our approach continuously monitors...

论文介绍 研究问题:现有多智能体系统可能在不当时机中断人类,导致认知超负荷和任务性能下降。核心方法:提出一个闭环架构,使用消费级脑机接口实时监测人类神经生理焦点,生成「认知对齐」的多智能体交互,在人类高心理负荷时延迟通信。可能应用:改善人机交互的自然性和任务效率。

FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation

第一作者: Chengbo Yuan · 方向: 机器人操作 · 来源: cs.RO

Abstract:Despite the success of vision-based generalist robotic policies, existing tactile-based policies remain tied to fixed embodiments and sensor setups. This is because tactile signals are highly heterogeneous across hardware, making cross-sensor generalization difficult. We present FTP-1,the first generalist foundation tactile policy pretrained to acquire transferable tactile manipulation abilities across diverse sensors and embodiments. FTP-1 supports varied tactile inputs, including image-, array-, and state-based signals, by using heterogeneous encoders to project them into unified morphology-aware latent tokens that are jointly modeled by a shared tactile Transformer expert. Pretrained on around 3,000 hours of tactile manipulation data aggregated from 26 data sources, spanning human and robot demonstrations across 21 sensors, FTP-1 learns tactile skills that transfer beyond...

论文介绍 研究问题:现有触觉策略绑定于固定硬件和传感器设置,跨传感器泛化困难,限制触觉操作技能的转移。核心方法:提出FTP-1,首个通用基础触觉策略,预训练于约3000小时触觉操作数据,使用异构编码器将多样触觉输入统一为形态感知潜在令牌,并通过共享Transformer专家建模。可能应用:实现触觉技能在不同传感器和具身间的转移。

Y-BotFrame: An Extensible Embodied Agent Framework for Quadruped Robot Assistants

第一作者: Luyao Zhang · 方向: 导航与运动 · 来源: cs.RO

Abstract:Quadruped robots are capable of traversing a wide range of complex terrains with high flexibility. As highly mobile ground-based intelligent platforms, they can be equipped with modules for navigation control, environmental perception, and intelligent interaction, thereby serving as real-world mobile deployment platforms for various algorithms. In this paper, we introduce Y-BotFrame, an extensible embodied platform that turns a robot into an intelligent ground assistant. Y-BotFrame integrates multimodal perception capabilities, including speech, vision, and LiDAR, and employs a large language model as the cognitive core for environmental understanding, contextual reasoning, and task planning. The system maps user natural-language instructions into executable embodied task units that can be carried out by the robot. Y-BotFrame supports natural interaction through voice commands...

论文介绍 研究问题:四足机器人作为高移动性平台,需要集成感知和交互能力以服务各种算法。核心方法:提出Y-BotFrame框架,集成语音、视觉和LiDAR多模态感知,使用大语言模型作为认知核心进行环境理解、推理和任务规划,将自然语言指令映射为可执行任务单元。可能应用:作为现实世界移动部署平台,支持语音命令和自然交互。

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

第一作者: Dayu Xia · 方向: 机器人操作 · 来源: cs.RO

Abstract:Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models to judge not only final task success, but also how a manipulation execution is physically and temporally progressing. However, existing evaluations fail to test whether VLMs possess fine-grained process understanding. To address this gap, we present RoboProcessBench, a benchmark for process-aware understanding in vision-language robotic manipulation. RoboProcessBench decomposes such capability into two complementary dimensions, \emph{static monitoring} and \emph{dynamic reasoning}, instantiated as 12 diagnostic question families covering phase, contact, motion, coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions. Built from physically grounded execution...

论文介绍 研究问题:视觉语言模型在机器人操作中作为视觉评论家等角色,但现有评估缺乏对操作执行过程精细理解的测试。核心方法:提出RoboProcessBench基准,将能力分解为静态监控和动态推理两个维度,包含12个诊断问题族,覆盖阶段、接触、运动等。可能应用:评估和改进视觉语言模型对操作过程的感知能力。

GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training

第一作者: Zhihai Bi · 方向: 导航与运动 · 来源: cs.RO

Abstract:Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated...

论文介绍 研究问题:人形机器人物体交互因动态平衡和稳定交互的耦合而具有挑战性,现有方法需要任务特定训练或依赖轨迹回放。核心方法:提出GenHOI框架,通过重建场景和生成交互视频,使人形机器人零样本模仿单个生成视频执行多样物体交互任务。可能应用:实现无需任务特定训练的灵活物体交互。

Trajectory-Level Redirection Attacks on Vision-Language-Action Models

第一作者: Gokul Puthumanaillam · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) policies bring natural language into closed-loop robot control, enabling robots to execute manipulation tasks directly from text instructions. The same interface gives text a recurring role in control because the prompt is reused at every replanning step, and each prompt-conditioned action changes the future observations on which the policy acts. Existing VLA attacks study adversarial prompts that elicit targeted low-level actions or make such actions persist across changing images. We identify a stronger trajectory-level failure mode: a prompt that still $\textit{appears}$ to specify the intended task but redirects the final physical outcome. We mathematically formalize this setting as $\textit{command-preserving trajectory redirection}$, a prompt-only threat model in which the attacker chooses one prompt before the episode, all policy and...

论文介绍 研究问题:视觉语言动作模型中,现有攻击研究对抗提示引发低级动作,但存在更强的轨迹级失败模式,其中提示仍看似指定任务但重定向最终物理结果。核心方法:形式化命令保留轨迹重定向威胁模型,攻击者选择提示后,策略在重新规划步骤中重用提示,导致轨迹重定向。可能应用:揭示VLA模型的安全性漏洞,提高鲁棒性。

EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment

第一作者: Shihefeng Wang · 方向: 模仿学习 · 来源: cs.RO

Abstract:Scalable robot imitation learning relies on large-scale heterogeneous data from diverse robots or body-free data, making Cartesian end-effector actions a key interface for embodiment-agnostic policy learning. However, end-effector-only abstraction leaves Cartesian policies unaware of the deployed robot body, making them brittle under robot-specific constraints such as whole-body collision avoidance. To overcome this limitation, we present EmbodiSteer, a training-free framework that steers embodiment-agnostic visuomotor policies toward zero-shot, embodiment-aware deployment. EmbodiSteer keeps policy learning in Cartesian space while efficiently lifting inference-time diffusion sampling into the target robot's joint space via forward kinematics and Jacobian-based updates. With whole-body collision-aware guidance over joint trajectories after each denoising step, the arm can be...

论文介绍 研究问题:具身无关视觉运动策略在笛卡尔空间学习,但部署时忽略机器人特定约束如全身避碰,导致策略脆弱。核心方法:提出EmbodiSteer框架,保持策略在笛卡尔空间学习,但在推理时通过运动学和Jacobian更新将扩散采样提升到目标机器人的关节空间,提供全身碰撞感知引导。可能应用:实现具身无关策略的零样本跨具身部署。

SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation

第一作者: Sunghwan Kim · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Long-horizon robot mobile manipulation requires continual reasoning about localization, environment changes, and task progress, all of which are challenging to infer from image observations alone. In this paper, we show that conditioning a mobile manipulation policy on a spatiotemporal feature map improves reasoning over long horizons. The map represents the environment and the articulated robot body as neural points in a shared latent space and is updated online from egocentric observations and proprioceptive state. We update the environment neural points using object-level rigid tracking and the robot neural points using forward kinematics. We use our spatiotemporal environment and robot feature (SERF) map as a state input to a vision-language-action (VLA) model by extracting map tokens from multiple reference frames and spatial scales, providing the policy with both local...

论文介绍 本文研究长期机器人移动操作中从图像观察进行持续推理的挑战,提出SERF地图方法。该方法将环境和机器人身体表示为共享潜空间中的神经点,通过对象级刚体追踪和正向运动学在线更新,并作为视觉-语言-动作模型的状态输入,以提升策略在长时程任务中的推理能力,可能应用于复杂移动操作场景。

Towards Reliable Sequential Object Picking in Clutter: The Runner-up Solution to RGMC 2025

第一作者: Wei Yu · 方向: 机器人操作 · 来源: cs.RO

Abstract:As a long-standing challenge in robotic manipulation, stable and efficient grasping in cluttered environments is of great importance in industrial settings. While recent studies have achieved relatively high success rates in grasping from clutter, there remain few mature solutions for more demanding tasks such as sequential object search and sorting. This work addresses sequential object picking in cluttered environments based on the Cluttered Environment Picking Benchmark (CEPB) and presents our solution to the Pick-in-Clutter track of the 10th Robotic Grasping and Manipulation Competition (RGMC) at ICRA 2025. The task poses several key challenges. First, it requires robust and collision-aware grasping with high success rates across a diverse set of objects, including both rigid and deformable ones. Second, it demands efficient search for target objects, which places...

论文介绍 本文针对杂乱环境中顺序物体抓取的稳定性与效率挑战,基于CEPB基准和RGMC 2025竞赛,提出解决方案。核心方法包括实现鲁棒的碰撞感知抓取和高效的目标物体搜索,以处理刚性及可变形物体,旨在提升工业设置中的物体分拣可靠性。

An Embodied Simulation Platform, Benchmark, and Data-Efficient Augmentation Framework for Wet-Lab Robotics

第一作者: Zhe Liu · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Wet-lab robots can improve the reproducibility, throughput, and safety of biomedical experiments, but scaling their learning requires customizable simulators for safe and reproducible task generation, open editable laboratory assets, and efficient pipelines that turn limited demonstrations into usable training data. We present Pipette, an embodied simulation platform, benchmark, and data-efficient augmentation framework for wet-lab robot learning. Pipette releases over 43 open-source and re-editable wet-lab assets, together with an extensible asset-building pipeline. A key component of Pipette is its simulation-based data augmentation pipeline, replaying human demonstrations in simulation, applies lighting, camera, speed, and action perturbations, and filters generated episodes with automatic task success checks, rapidly expanding usable training data from limited manual...

论文介绍 本文介绍Pipette平台,一个用于湿实验室机器人学习的具身仿真平台、基准和数据高效增强框架。它提供开源可编辑的实验室资产,通过模拟数据增强管道从有限示例生成训练数据,以支持安全可复现的任务生成,加速机器人在生物医学实验中的应用。

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

第一作者: Allison Andreyev · 方向: 机器人操作 · 来源: cs.RO

Abstract:For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time. Although Vision-Language Models (VLMs) have enabled zero-shot generalization in robot task and motion planning (TAMP), current state-of-the-art approaches often remain computationally "heavyweight" or require extensive training on thousands of demonstrations. We present GRASP (Grounded Reasoning and Symbolic Planning), a framework designed as a step toward open-vocabulary tabletop manipulation. Our approach leverages a pretrained VLM to translate natural-language queries into neuro-symbolic goal states, grounded in the physical world via a bounding-box detection pipeline. Unlike methods that rely on fixed color lists or hard-coded coordinates, GRASP enables robots to interpret abstract spatial concepts such as "top shelf" and...

论文介绍 本文提出GRASP框架,用于语言条件下的开放词汇桌面操作。该方法利用预训练视觉-语言模型将自然语言查询转换为神经符号目标状态,并通过边界框检测实现物理世界接地,使机器人能解释抽象空间概念,适用于家庭或工业环境中的实时任务规划。

AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

第一作者: Jianli Sun · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-level Unmanned Aerial Vehicle (UAV) movement and end-effector-level arm manipulation differ substantially in action scale, dynamics, and control objectives. In this paper, we propose AIR-VLA+, a flow matching action generation architecture specifically designed for aerial manipulation, featuring cascaded dual-action decoders and an asymmetric feature-level Mixture of Experts (MoE). We construct cascaded manipulation and movement decoders, allowing the UAV to unidirectionally observe the manipulator's intent during movement to achieve workflow coordination, while isolating the impact of UAV movement information backpropagation on arm manipulation stability. Addressing the characteristic that UAV movement is highly dependent on high-level semantics and responsible for...

论文介绍 本文针对空中机器人操作中平台运动与末端操作器控制的表示耦合问题,提出AIR-VLA+架构。该方法采用级联双动作解码器和非对称混合专家模块,解耦两者以协调工作流,并通过动作生成架构提升控制稳定性,可能用于复杂空中操纵任务。

Sparse2Act: Learning Action-Aligned Sparse 3D Representations for Cross-Domain Robot Manipulation

第一作者: Yu Guo · 方向: 机器人操作 · 来源: cs.RO

Abstract:Explicit 3D representations are attractive for manipulation because they expose object shape, workspace geometry, and robot-object relations in metric coordinates. However, sparse 3D encoders are often learned through downstream task objectives, tying the representation to a particular data distribution, policy architecture, and action parameterization. We introduce Sparse2Act, an observation-action alignment framework for pretraining sparse point-cloud encoders. The key idea is to use task-space end-effector actions as geometric supervision: masked sparse 3D tokens are trained to organize scene features around the workspace motion paired with the observation. After pretraining, only the encoder initialization is reused by downstream policies, allowing them to retain their own architectures and action spaces, including joint-space commands. On the LIBERO-10 benchmark, our...

论文介绍 本文引入Sparse2Act框架,用于预训练稀疏点云编码器以支持跨域机器人操作。核心是通过任务空间末端执行器动作作为几何监督进行观察-动作对齐训练,使编码器能解耦下游策略,提高泛化能力,已在基准测试中验证有效性。

EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence

第一作者: Xin Zhou · 方向: 模仿学习 · 来源: cs.RO

Abstract:In this paper, we propose the Enhanced World Action Model (EWAM), a closed-loop online adaptation architecture built upon a pretrained and fully frozen Cosmos3 backbone network. Evaluated entirely under a zero-shot task protocol, EWAM is centrally focused on reducing the amount of additional deployment data required to adapt to new task layouts. Notably, no extra task-specific demonstration sets were introduced in any of the evaluations, and no fine-tuning was performed on the backbone network. Its performance gains stem entirely from an inference-time co-reasoning mechanism composed of four inserted lightweight neural layers: the Neural Experience Memory Layer located in the intermediate layers of the Diffusion Transformer (DiT) provides task-relevant execution context; the Neural Anomaly Detection Layer after the state prediction head monitors the divergence between...

论文介绍 本文提出EWAM模型,一种基于预训练骨干网络的闭环在线适应架构,用于具身智能。该方法通过推理时轻量神经层进行协同推理,无需额外任务数据或微调,以减少对新任务布局的部署数据需求,实现零样本任务协议下的高效适应。

DARRMS -- An Efficient Algorithm for Dynamic Attention Radius in Resource-Constrained Multi-Agent Systems

第一作者: Benjamin Alcorn · 方向: 具身智能 · 来源: cs.RO

Abstract:Multi-agent systems are integral tools for various domains such as robotics, cybersecurity, and autonomous vehicle planning. These types of systems often have constraints on the computational resources, leading to a need for efficient lightweight algorithms. Traditional decision making frameworks often assume ideal conditions, such as full observability and unlimited computational capacity, which do not align with real-world challenges. In this paper, we introduce a new algorithm that allows for reduced demand on computational resources without a large cost of other performance metrics. Agents will limit their observability to some attention radius, which intentionally allows them to ignore parts of the environment that might be unnecessary for action planning. By optimizing both the attention radius and decision-making, our approach enhances coordination and scalability in...

论文介绍 本文介绍DARRMS算法,用于资源受限多智能体系统的动态注意力半径优化。该算法允许智能体限制观察范围以减少计算需求,通过优化注意力半径和决策来增强协调与可扩展性,适用于机器人、自动驾驶等领域中的高效规划。

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation

第一作者: Honglin He · 方向: 导航与运动 · 来源: cs.RO

Abstract:Autonomous long-horizon sidewalk navigation is essential for micro-mobility applications such as robotic food delivery and assistive electronic wheelchairs. Unlike autonomous driving on the road, long-horizon sidewalk navigation requires precise maneuvering through unpredictable sidewalk terrains and pedestrians, with a lightweight perception stack as minimal as a single monocular RGB camera. While imitation learning (IL) from demonstrations offers a practical solution, the resulting autopilot policy often suffers from compounding errors, a lack of social compliance on sidewalks, and deficiencies in counterfactual reasoning to handle complex situations. To address these challenges, we introduce FlowPilot, a mapless navigation policy that achieves robust and efficient long-horizon navigation performance using only a monocular RGB camera. We first propose to use anchored flow...

论文介绍 本文针对单目视觉下的长程人行道导航问题,提出了FlowPilot策略。该策略旨在解决传统模仿学习在复杂人行道环境中因误差累积、缺乏社会合规性和反事实推理能力不足而导致的挑战。其核心方法是使用「锚定流」与人类偏好策略相结合,仅依靠单一RGB摄像头实现无地图导航。该系统有望应用于机器人送餐、辅助电子轮椅等微型出行场景。

G-MAPP: GPU-accelerated Multi-Agent Planning and Perception for Reactive Motion Generation

第一作者: Tanmay Bishnoi · 方向: 导航与运动 · 来源: cs.RO

Abstract:Reactive motion generation in unstructured environments remains an open challenge in robotics. Due to the computational complexity of collision-free motion generation, existing methods either generate global trajectories for static scenarios, or employ models that make conservative assumptions about the environment. This paper identifies the primary bottleneck as the runtime performance demand of planning on high-fidelity environments, and the temporal integration between the perception and planning modules. Therefore, we propose a framework that does not compromise on runtime performance and world representations for perception and planning by accelerating world modeling and vector-field based planning using the GPU. This allows us to achieve faster parallel state exploration for quasi-global trajectory planning, and tighter coupling of the perception-action loop in real-time...

论文介绍 本文旨在解决非结构化环境中机器人反应式运动生成的实时性与高保真世界建模之间的矛盾。作者提出了G-MAPP框架,其核心是利用GPU并行加速世界建模和基于矢量场的规划过程。这使得系统能够在高保真环境中实现更快的并行状态探索,并实时耦合感知与规划循环,从而生成准全局无碰撞轨迹。

Action-Effect Memory Pretraining for Robot Manipulation

第一作者: Yijing Zhou · 方向: 机器人操作 · 来源: cs.RO

Abstract:We present AEM, an Action-Effect Memory pretraining framework for robot manipulation that learns compact temporal representations from vision-action history. Unlike prior robot representation pretraining methods that mainly focus on single-frame visual encoding, AEM targets the temporal nature of manipulation, where the current observation alone is often insufficient under partial observability. AEM models manipulation as an action-driven interaction process by interleaving visual and action features and applying masked modeling to recover missing content from incomplete histories, thereby learning action-conditioned state evolution. The Mamba-encoded output of the final vision token is used as a compact history representation, serving as the global context for decoding and downstream control. This design preserves a single-vector temporal bottleneck while keeping inference...

论文介绍 针对机器人操作任务中的部分可观测性问题,本文提出了动作效果记忆预训练框架AEM。不同于专注于单帧视觉编码的方法,AEM通过交错视觉与动作特征,并应用掩码建模来恢复不完整历史,从而学习动作驱动的状态演化。其使用Mamba编码器生成紧凑的历史表示作为全局上下文,以支持下游控制任务。

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

第一作者: Baochang Ren · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach. AI can help read literature, generate hypotheses, and plan protocols, yet the execution of those protocols at the bench still requires a human operator. Vision-Language-Action (VLA) models provide one possible interface between written protocols and robot execution, but existing policies are trained mostly on household and tabletop demonstrations and rarely encounter the instruments, transparent liquids, or fixed protocol workflows found in scientific laboratories. Closing this gap requires both laboratory-specific supervision and a unified learning framework that can accommodate the diverse robot embodiments used to execute experimental protocols. We therefore identify data and embodiment as central bottlenecks...

论文介绍 针对自动驾驶VLA模型中思维链与生成轨迹之间关系缺乏评估的问题,本文提出了VLADriveBench评估框架。该框架结合了观测指标和干预协议,以从互补视角分析思维链-动作关系。研究应用该框架评估模型后发现,思维链的观测对齐度与其对动作的因果影响可能显著分离,揭示了现有评估的不足。

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

第一作者: Hanyang Yu · 方向: 多模态具身 · 来源: cs.RO

Abstract:World Action Models (WAMs) present a promising paradigm for robotic control via video prediction. However, current WAMs suffer from fundamental spatial bottlenecks: standard text inputs introduce referential ambiguity in cluttered scenes, while unstructured RGB predictions lack semantic grounding and remain biased by task-irrelevant backgrounds. To overcome these limitations, we introduce MaskWAM, an object-centric world-action model. By jointly integrating masks as both explicit inputs and predictions via a unified Mixture of Transformers (MoT), MaskWAM unlocks robust policy generalization. This design provides two key benefits: (1) predicting future masks yields object-centric semantic supervision that suppresses visual noise, significantly enhancing even standard text-conditioned WAMs; and (2) coupling this predictive supervision with first-frame visual prompts, such as...

论文介绍 现有世界动作模型存在文本输入歧义和预测缺乏语义锚定的问题。本文提出MaskWAM,一个对象中心的世界动作模型。其核心是通过统一的混合Transformer架构,将遮罩同时作为显式输入和预测目标。这种设计提供了两方面优势:预测未来遮罩提供以对象为中心的语义监督,抑制视觉噪声;并结合首帧视觉提示,显著增强了策略的泛化能力。

$μ$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models

第一作者: Egor Cherepanov · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) models predict chunks of future actions from the current observation, an assumption that fails under partial observability, where decisions depend on information no longer visible. Existing memory-augmented VLAs simultaneously introduce recurrence, retrieval, compression modules, auxiliary objectives, hierarchical memory, or task-specific architectural changes, so the contribution of recurrence itself remains entangled with surrounding machinery. We present a controlled isolation study of recurrence in a strong pretrained VLA backbone. Our formulation augments the transformer with a small set of learnable memory tokens carried across timesteps and updated through self-attention, trained end to end with truncated backpropagation through time, with no auxiliary losses and no architectural changes. We instantiate this as $\mu$VLA, a family of...

论文介绍 VLA模型在部分可观测条件下表现不佳,因其决策依赖于不可见的信息。现有记忆增强方法常将循环机制与其他复杂组件混合,导致贡献难以分离。本文对预训练VLA骨干中的循环机制进行受控隔离研究,提出μVLA。它仅通过一组可学习的记忆令牌跨时间步携带并更新状态,以截断时间反向传播端到端训练,无需辅助损失或架构修改。

VLADriveBench: Evaluating CoT-Action Relationship in VLA for Autonomous Driving

第一作者: Thach Nguyen · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Vision-language-action (VLA) models generate chain-of-thought (CoT) reasoning alongside driving trajectories, but existing benchmarks evaluate only trajectory quality and do not assess whether the CoT is relevant, consistent, or causally connected to the driving action. We introduce VLADriveBench, a framework that combines observational metrics (mentioning, hallucination, contradiction, action alignment) with a CoT intervention protocol to provide complementary views of the CoT-action relationship. Applying VLADriveBench to three models across two architectures, we find that the two analyses can diverge sharply: ORION scores highest on observational alignment yet its CoT is epiphenomenal, while Alpamayo v1.5 scores lower yet its CoT is strongly causal, with visual salience gating the extent of CoT influence.

论文介绍 本研究针对自动驾驶中视觉-语言-动作(VLA)模型的链式思维(CoT)推理评估问题。现有基准仅评估轨迹质量,而作者提出VLADriveBench框架,结合观察性指标和CoT干预协议,以互补视角分析CoT与驾驶动作的相关性和因果性。实验应用于多个模型,发现观察性对齐与因果影响可能分离,例如ORION模型观察性对齐高但CoT为附带现象,而Alpamayo v1.5的CoT因果性强。该工作为VLA模型的CoT分析提供了新方法。

Rarity-Gated Context Conditioning for Offline Imitation Learning-Based Maritime Anomaly Detection

第一作者: Yongmin Kim · 方向: 模仿学习 · 来源: cs.LG

Abstract:Contextual anomaly detection aims to identify abnormal behavior conditional on context variables, but practical deployments often face highly imbalanced context distributions where rare regimes can be critical information. Under such frequency bias, context-conditioned models can produce unstable decisions and excessive false alarms in rare contexts. We propose Rarity-Gated Feature-wise Linear Modulation (RGFiLM), a rarity-aware conditioning module that combines feature-wise modulation (i.e., context-conditioned scaling and shifting of hidden features) with a gate controlled by a data-driven rarity score. The rarity score is estimated from the empirical distribution of context variables and regulates how strongly context modulates intermediate representations: the gate becomes more decisive under rare contexts while remaining conservative under frequent contexts. We evaluate...

论文介绍 本文关注在上下文分布严重不平衡场景下的上下文条件异常检测问题,例如海事监控中罕见但关键的工况。现有模型在稀有上下文中可能产生不稳定决策。作者提出RGFiLM模块,结合特征调制和数据驱动的稀有度门控机制。该门控根据上下文稀有度调节其对隐藏特征的调制强度,在稀有情境下更果断,在常见情境下更保守,以减少误报。

PersonaDrive: Human-Style Retrieval-Augmented VLA Agents for Closed-Loop Driving Simulation

第一作者: Mahmoud Srewa · 方向: VLA 通用模型 · 来源: cs.AI

Abstract:Closed-loop driving simulators typically populate their environments with non-ego traffic agents that behave largely the same way, produced either by rule-based traffic managers or by learned models trained toward a single behavioral mode. Recent work introduces style variation through post-hoc labels on observational data or LLM-inferred reward weights, but these signals act as proxies for what a style should reward rather than demonstrations of humans explicitly asked to drive in that style. We introduce PersonaDrive, a pipeline that conditions a vision-language-action (VLA) driving agent on retrieved demonstrations from a style-instructed human driving dataset, in which participants drive CARLA leaderboard routes under aggressive, neutral, and conservative instructions on a driver-in-the-loop rig. The pipeline has three stages: (i) offline triplet mining over per-style...

论文介绍 在闭环驾驶模拟中,非自车交通代理行为模式单一,缺乏人类驾驶风格的多样性,现有方法多依赖代理信号。本文提出 PersonaDrive 管道,通过检索增强技术从风格指导的人类驾驶数据集中获取演示,条件化视觉-语言-动作(VLA)代理以模拟不同驾驶风格。数据集包括攻击性、中性和保守性指令下的驾驶记录。该管道包含离线三元组挖掘等阶段,旨在提升模拟环境交通行为的真实性和多样性,对自动驾驶系统测试具有应用潜力。

市场总览

美股方面,SPY与QQQ呈多头排列,RSI中性,但个股分化明显:AAPL、MSFT、META等RSI偏低或超卖,趋势偏空,显示承压。加密市场整体技术面矛盾,BTC、ETH、SOL虽出现MACD金叉,但趋势仍bearish且空头排列;恐慌贪婪指数20反映极度恐慌,总市值2.31T,BTC主导率56.6%,市场情绪谨慎。中概股如BABA、PDD、JD等RSI处于超卖或偏低区域,空头排列显著,下行压力持续。商品外汇中,黄金期货中性,原油期货下跌且RSI偏低;美元指数偏上行,多头排列;人民币偏下行,虽MACD金叉但趋势bearish。整体技术面呈现美股分化、加密短期反弹但趋势未改、中概承压、商品外汇混合的态势。

今日关注

BABA 阿里巴巴 (BABA)
偏下行

RSI14为29.6,处于超卖状态,趋势为bearish,信号包括「RSI 超卖」和「空头排列」。当前价格112.82,低于20日均线125.81、50日均线130.33和200日均线149.57,均呈空头排列。MACD为-4.9544,低于信号线-3.3542,显示下行动量持续。

^TNX 10Y 美债收益率 (%)
偏上行

趋势为bullish,信号有「多头排列」和「接近 52 周低」。当前价格4.49,高于50日均线4.42和200日均线4.21,但低于20日均线4.52。RSI14为50.8,中性略偏上。MACD为0.0203,高于信号线0.0289,但差值小,显示温和上行倾向。

BTC-USD Bitcoin
中性

信号包括「MACD 金叉」和「空头排列」。MACD为-3217.09但金叉可能预示短期反弹,RSI14为40.6,未超买超卖。趋势bearish,价格65422.73低于所有均线(20日66920.75、50日73914.25、200日77644.41),但近期5日涨幅6.13%,指标矛盾显示中性状态。

DX-Y.NYB 美元指数 DXY
偏上行

趋势为bullish,信号有「接近 52 周高」和「多头排列」。当前价格99.58,高于20日均线99.44、50日均线98.89和200日均线98.65。RSI14为55.1,中性偏上。MACD为0.2806,高于信号线0.265,显示上行动量。

全部资产

^VIX

VIX 恐慌指数

$17.68 -9.05%
5 日
-17.81%
距 52w 高
-49.9%
RSI(14)
48.6
趋势
空头
SMA 20 / 50 / 200
17.53 / 18.25 / 18.54
MACD / 信号
0.309 / -0.098
死叉(SMA50↓SMA200) (2 天前)空头排列

^TNX

10Y 美债收益率 (%)

$4.49 +0.54%
5 日
-1.08%
距 52w 高
-10.2%
RSI(14)
50.8
趋势
多头
SMA 20 / 50 / 200
4.52 / 4.42 / 4.21
MACD / 信号
0.020 / 0.029
接近 52 周低多头排列

DX-Y.NYB

美元指数 DXY

$99.58 -0.17%
5 日
-0.47%
距 52w 高
-1.1%
RSI(14)
55.1
趋势
多头
SMA 20 / 50 / 200
99.44 / 98.89 / 98.65
MACD / 信号
0.281 / 0.265
接近 52 周高多头排列

SPY

S&P 500 ETF

$741.75 +0.54%
5 日
+0.57%
距 52w 高
-2.5%
RSI(14)
52.9
趋势
多头
SMA 20 / 50 / 200
745.07 / 722.80 / 686.30
MACD / 信号
3.772 / 7.393
接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$721.34 +0.59%
5 日
+2.31%
距 52w 高
-3.6%
RSI(14)
55.2
趋势
多头
SMA 20 / 50 / 200
721.50 / 681.81 / 625.38
MACD / 信号
8.315 / 13.604
多头排列

AAPL

Apple

$291.13 -1.52%
5 日
-5.27%
距 52w 高
-8.3%
RSI(14)
44.0
趋势
多头
SMA 20 / 50 / 200
303.88 / 285.49 / 266.87
MACD / 信号
2.261 / 5.760
多头排列

MSFT

Microsoft

$390.74 +0.10%
5 日
-6.22%
距 52w 高
-29.7%
RSI(14)
37.1
趋势
空头
SMA 20 / 50 / 200
419.75 / 411.80 / 453.73
MACD / 信号
-4.135 / 1.271
空头排列

NVDA

Nvidia

$205.19 +0.16%
5 日
+0.04%
距 52w 高
-13.3%
RSI(14)
45.2
趋势
中性
SMA 20 / 50 / 200
214.62 / 206.91 / 189.26
MACD / 信号
-1.101 / 1.257

GOOGL

Alphabet

$359.68 +0.53%
5 日
-2.40%
距 52w 高
-12.0%
RSI(14)
42.4
趋势
中性
SMA 20 / 50 / 200
376.42 / 362.26 / 307.94
MACD / 信号
-2.903 / 1.104

TSLA

Tesla

$406.43 +1.82%
5 日
+3.95%
距 52w 高
-18.5%
RSI(14)
48.9
趋势
中性
SMA 20 / 50 / 200
415.74 / 398.30 / 415.69
MACD / 信号
-2.173 / 2.196

META

Meta

$566.98 -0.26%
5 日
-4.39%
距 52w 高
-28.8%
RSI(14)
34.7
趋势
空头
SMA 20 / 50 / 200
604.21 / 621.83 / 658.09
MACD / 信号
-12.809 / -7.969
空头排列
加密恐慌贪婪
20
极度恐慌
加密总市值
$2.31 T
+1.25% / 24h
BTC 主导率
56.6%
ETH 8.9%
24h 成交量
$60.2 B
活跃币 17,446

BTC-USD

Bitcoin

$65,422.73 +1.55%
5 日
+6.13%
距 52w 高
-48.2%
RSI(14)
40.6
趋势
空头
SMA 20 / 50 / 200
66,920.75 / 73,914.25 / 77,644.41
MACD / 信号
-3,217.093 / -3,495.419
MACD 金叉 (1 天前)空头排列

ETH-USD

Ethereum

$1,714.77 +2.06%
5 日
+4.71%
距 52w 高
-65.4%
RSI(14)
36.8
趋势
空头
SMA 20 / 50 / 200
1,804.75 / 2,068.44 / 2,407.85
MACD / 信号
-122.934 / -129.239
MACD 金叉 (今天)空头排列

SOL-USD

Solana

$70.78 +2.77%
5 日
+8.96%
距 52w 高
-72.0%
RSI(14)
43.6
趋势
空头
SMA 20 / 50 / 200
72.52 / 81.45 / 99.68
MACD / 信号
-4.580 / -4.938
MACD 金叉 (今天)空头排列

BABA

阿里巴巴 (BABA)

$112.82 +0.12%
5 日
-6.81%
距 52w 高
-41.4%
RSI(14)
29.6
趋势
空头
SMA 20 / 50 / 200
125.81 / 130.33 / 149.57
MACD / 信号
-4.954 / -3.354
RSI 超卖空头排列

PDD

拼多多 (PDD)

$81.56 +0.32%
5 日
-4.13%
距 52w 高
-41.5%
RSI(14)
32.4
趋势
空头
SMA 20 / 50 / 200
88.52 / 95.37 / 111.40
MACD / 信号
-4.262 / -3.794
空头排列

JD

京东 (JD)

$28.56 +1.78%
5 日
-1.11%
距 52w 高
-22.5%
RSI(14)
41.6
趋势
空头
SMA 20 / 50 / 200
29.87 / 30.11 / 30.31
MACD / 信号
-0.573 / -0.378
空头排列

0700.HK

腾讯控股 (0700.HK)

HK$480.00 +3.54%
5 日
+7.53%
距 52w 高
-29.7%
RSI(14)
57.5
趋势
中性
SMA 20 / 50 / 200
451.63 / 472.49 / 568.38
MACD / 信号
-1.125 / -5.718

GC=F

黄金期货

$4,316.00 +2.40%
5 日
-0.46%
距 52w 高
-22.7%
RSI(14)
42.9
趋势
中性
SMA 20 / 50 / 200
4,409.89 / 4,581.05 / 4,424.06
MACD / 信号
-109.622 / -92.054

CL=F

WTI 原油期货

$80.92 -4.67%
5 日
-11.37%
距 52w 高
-32.3%
RSI(14)
34.1
趋势
中性
SMA 20 / 50 / 200
92.75 / 96.13 / 73.49
MACD / 信号
-3.246 / -2.143

USDCNY=X

美元 / 人民币

¥6.76 -0.20%
5 日
-0.05%
距 52w 高
-6.2%
RSI(14)
35.4
趋势
空头
SMA 20 / 50 / 200
6.78 / 6.81 / 6.97
MACD / 信号
-0.012 / -0.014
MACD 金叉 (3 天前)接近 52 周低空头排列
风险提示

本报告基于公开行情数据的技术指标解读,仅供参考。过去走势不代表未来表现,技术分析不能保证未来收益。投资者应结合自身情况谨慎决策。

Iran War Live Updates: U.S. and Iran Reach Cease-Fire Agreement

The deal was expected to halt fighting for 60 days, open the Strait of Hormuz and lift the U.S. naval blockade on Iranian ports. But it would leave the thorniest nuclear issues for another day.

中文摘要 美国与伊朗达成停火协议,预计停止战斗60天,开放霍尔木兹海峡并解除美国对伊朗港口的封锁,但最棘手的核问题留待日后解决。

Here’s the latest.

中文摘要 标题为「最新消息」,但无具体事件描述。

Japan Is Running Out of Royals. Are More Men the Answer?

Japan’s legislature is drafting a plan to allow the imperial family to adopt distant male relatives. But some in Japan would prefer a female emperor.

中文摘要 日本立法机构正起草计划,允许皇室收养远亲男性以增加继承人,但部分人希望出现女天皇。

Australia news live: Wong welcomes US-Iran peace deal; Joyce says fundraised millions to pay for One Nation ads

Follow today’s news live Get our breaking news email, free app or daily news podcast The prime minister and foreign minister have issued a lengthier statement welcoming the agreement made by the US and Iran, and called for continued restraint to avoid further escalation. President Donald Trump made

中文摘要 澳大利亚总理和外长欢迎美伊和平协议,呼吁继续克制以避免进一步升级。Joyce筹集数百万美元支付One Nation广告费用。

Stock markets soar, oil falls as US, Iran confirm deal to end war

Asian stock markets surge as Washington and Tehran announce agreement to end hostiles and reopen Strait of Hormuz.

中文摘要 美伊确认结束战争协议后,亚洲股市大涨,油价下跌。华盛顿和德黑兰宣布协议结束敌对行动并重开霍尔木兹海峡。

US-Iran peace deal could see strait of Hormuz open on Friday ‘under Iranian arrangements’, says state media – Middle East crisis live

The US president confirmed the agreement as Pakistan’s prime minister said the official signing will be in Geneva on 19 June Full report: peace deal between US and Iran, Pakistan says, with strait of Hormuz to reopen Iranian hardliners in vociferous push to reject US peace deal Israel says it has st

中文摘要 美伊和平协议可能使霍尔木兹海峡于周五在伊朗安排下开放。巴基斯坦总理表示官方签署将于6月19日在日内瓦举行,但伊朗强硬派强烈反对。

Claims Israel’s Beirut strike pushed Trump on Iran announcement

US diplomat Alan Eyre says despite the US-Iran ceasefire announcement, there is no deal until it has been formalised.

中文摘要 美国外交官艾伦·艾尔表示,尽管美伊宣布停火,但协议尚未正式化。有说法称以色列对贝鲁特的袭击推动了特朗普的伊朗声明。

In Israel, Broad Discontent Even Before Deal’s Details Are Known

Israelis across the political spectrum have said the agreement appears to leave fundamental security threats posed by Iran unaddressed.

中文摘要 在以色列,协议细节未知前就广泛不满。以色列各政治派别认为协议未能解决伊朗构成的根本安全威胁。

Peace deal between US and Iran announced, with strait of Hormuz expected to reopen

Agreement was struck despite ⁠an Israeli strike on Lebanon on Sunday that drew criticism from both Iran and Trump Middle East crisis – live updates A peace deal between the US and Iran has been reached following nearly four months of fighting in the region, Donald Trump and senior Iranian officials

中文摘要 美伊宣布和平协议,霍尔木兹海峡预计重开。协议在近四个月战斗后达成,尽管以色列周日袭击黎巴嫩引发批评。

Trump Claims Strait Will be ‘Permanently Toll Free’ Under Agreement With Iran

In a call to The New York Times, President Trump praised Russia and China’s leaders and described Israel’s prime minister as “a very difficult guy.”

中文摘要 特朗普称在美伊协议下,霍尔木兹海峡将永久免费通行。特朗普在电话中赞扬俄罗斯和中国领导人,称以色列总理为「非常难搞的家伙」。

Anger among Iranian hardliners at terms of deal agreed with US

Those in favour forced to defend themselves against claims the terms amount to capitulation Middle East crisis – live updates Iranian hardliners have mounted a rearguard rejection of a deal with the US as as they say it does not guarantee sanctions relief, compensation or control of the strait of Ho

中文摘要 伊朗强硬派愤怒反对与美国协议的条款,认为其相当于投降,不保证制裁解除、赔偿或控制。

The U.S. and Iran announce a deal to end the war

President Trump said the U.S. would remove its blockade of the Strait of Hormuz.

中文摘要 美伊宣布结束战争协议。特朗普表示美国将解除对霍尔木兹海峡的封锁。

Oil prices slide after Pakistan announces deal between US and Iran

Under the agreement, the key Strait of Hormuz waterway will be reopened, US President Donald Trump said.

中文摘要 在巴基斯坦宣布美国和伊朗达成协议后,油价下跌。美国总统特朗普表示,关键的霍尔木兹海峡水道将重新开放。

AGSI's Roebuck on US-Iran Deal

William Roebuck, Former US Ambassador to Bahrian and Executive Vice President at Arab Gulf States Institute, discusses his outlook for the agreement between US and Iran to halt the war. He says the conflict is "moving into a new stage" and that is towards a "diplomatic track". He speaks with Shery A

中文摘要 前美国驻巴林大使、阿拉伯海湾国家研究所执行副总裁威廉·鲁贝克讨论美伊协议前景,称冲突正进入新阶段,转向外交轨道。

Vanguard's Wang on Iran-US Deal Economic Impact

Qian Wang, Chief APAC Economist at Vanguard Group, discusses the US and Iran agreement to a deal that will halt war and says while there is still significant uncertainty around how sustainable the current deal is, it could boost the outlook for the global economy. She speaks with Shery Ahn and Haidi

中文摘要 先锋集团首席亚太经济学家钱·王讨论美伊协议的经济影响,称尽管协议可持续性存在不确定性,但可能提振全球经济前景。

Indonesia Awaits MSCI Verdict That Risks $13 Billion Outflows

Indonesia faces a crucial test this month when MSCI Inc. decides whether to follow through with a downgrade, with investors already questioning the resilience of the world’s worst-performing equity market.

中文摘要 印度尼西亚本月面临关键测试,MSCI Inc. 将决定是否降级,可能导致130亿美元资金外流,投资者已质疑其全球表现最差股市的韧性。

US and Iran Agree to Deal Halting War

The US and Iran reached an interim agreement to reopen the Strait of Hormuz, halting a war that killed thousands of people and setting the stage for negotiations on the fate of Iran's nuclear program. (Source: Bloomberg)

中文摘要 美国和伊朗达成临时协议,重新开放霍尔木兹海峡,停止导致数千人死亡的战争,并为伊朗核计划谈判奠定基础。

Yen Short Bets Jump to Nine-Year High as Carry Trade Revives

Speculators have boosted their bets against the yen to a nine-year high, signaling the revival of the yen carry trade despite intervention risks and a potential rate hike by the Bank of Japan on Tuesday.

中文摘要 投机者将日元空头头寸推至九年高位,日元套息交易复苏,尽管存在干预风险且日本央行可能在周二加息。

Anthropic scrambles after Trump administration freezes its top AI models

Export controls on Fable and Mythos raise doubts over how US will police the most powerful AI systems

中文摘要 Anthropic 公司在特朗普政府冻结其顶级AI模型后紧急应对。对 Fable 和 Mythos 的出口管制引发对美国如何监管最强大AI系统的质疑。

As more US business owners retire many are selling up to their staff

Some six million bosses of American firms will be entering retirement between now and 2035.

中文摘要 随着更多美国企业主退休,许多人将企业出售给员工。从现在到2035年,约600万美国公司老板将进入退休阶段。

Millions of people can get discounts on their bills - here's how

Lower social tariffs allow many people on benefits to get cheaper deals for water, broadband and phone.

中文摘要 数百万人可以通过较低的社会费率在水电、宽带和电话账单上获得折扣,特别是领取福利者。

Surge in scams as fraudsters use AI to target people

On average, nearly eight cases of fraud in which money is stolen are reported in the UK every minute.

中文摘要 诈骗案件激增,骗子使用AI技术针对个人。在英国,平均每分钟报告近8起涉及金钱盗窃的诈骗案。

Is the convertible heading into the sunset?

UK drivers have taken a shine to the SUV but could the fate of the convertible be reversed?

中文摘要 英国司机偏爱SUV,敞篷车的命运是否可能逆转?

Japan, Korean Stocks Gain on Strait of Hormuz Reopening Hopes

Japanese and Korean stocks gained after oil prices fell as Donald Trump said the US and Iran will sign a peace deal later this week to reopen the Strait of Hormuz.

中文摘要 日本和韩国股市上涨,因油价下跌,美国总统特朗普表示美国和伊朗将于本周晚些时候签署和平协议,重新开放霍尔木兹海峡。

纵有疾风起,嘿嘿

加入论坛也是很久了,认识了很多朋友,也开了很久/很多次公益。 从最早的无限Gemini 到后来的无限沉浸式翻译 再到现在的 君の公益 我绝大部分的决策其实不会因为诸位夸我或者骂我而有所改变。 AI 平权 七天内其实已经提供了超过一万亿的tokens了,还不错,接着奏乐接着舞 50 个帖子 - 48 位参与者 阅读完整话题

大学老师,用豆包编程?

本人计科专业学生,吐槽一下自己的Java老师,上java课的时候,跟老师闲谈的时候,老师一直让我们不要用ai,说我们要打基础什么的,问老师,ai来了不用ai以后能顺应时代吗,老师说,那你不会的用豆包,我自己就一直用豆包编程,当时我真的就觉得,这个课上不上无所谓了,大学目前依旧教的知识比我年龄大,好不容易这学期有个Java springboot,结果老师说不教这个(其实是他不会),上课的时候,我们班有个女同学电脑卡的要死,那个老师让他安装个360,现在我看他电脑上还是一堆广告弹窗,说实在的,大学这样真能教出能找到工作的学生吗,教着已经淘汰多少年的知识,还要求一直要上这没用的课,什么时候能改变一下

11天前,辞职做中转站的,要倒闭了。

发个牢骚。上班五年,但是确实上班觉得真的很累了。在一年前又在父母的每日叨叨匆匆结了婚,然后生了一个女儿。当然觉得自己挺没主见的,也并没有觉得自己能撑住这个家。另一方面只是想以这种方式说,你们得到了这个“结果”并不是适合我。经常会跟我父母调侃,都是因为你们让我结婚才会。自己也挺没品的。自己又丧又不自信又自大又自以为又没品,为什么好像什么又清楚又什么不改变。是不是还是自己没深刻认识到。脑子思维没办法控制,感觉我好像脑子病了。思维跳跃很大,思考的时候事物动物想象它是不正常的那种连接,而且没办法控制。 116 个帖子 - 48 位参与者 阅读完整话题

【九幺公益】700*10000$发放!

本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已

【HelloAI CN+】DeepSeek复活了

公益推广要求 (点击了解更多详细信息) 介绍 HelloAI CN+,一个专注国产模型的公益站。当前包含DeepSeek,Glm,Kimi,Qwen,SparkDesk等模型。 追求极致稳定,极致响应。只要模型广场里看得到的模型都可以调用。 所有时间开放注册,不搞复杂套路。所有模型按此计费。 hellofriend.eu.cc 额度补充(10LDC=100次)【HelloAI CN+】公益站官产模型100次兑换码 - LD士多 DeepSeek复活了 前几天DeepSeek被刷炸了(账号被封),于是空返回了。接着DS2API也炸了,于是就彻底没戏了… 不过,我又复活了HelloDeepSeek

【Krill·618】0.1倍率年中狂欢,全线无门槛直降+专属骨折价来袭!回复就送!

各位佬友的支持,始终是我们前进的最大动力。为了感谢佬友这段时间的大力支持,Krill现在开始全线骨折价回馈佬友:618 专属福利正式开启! 活动时间:北京时间2026年6月15日 00:00~2026年6月18日 23:59 Krill网址:www.krill-ai.com Telegram: View @krill_aichat 【限时福利一:线路全面优化,费率大幅下调】余额价格直享!1元=1刀。 海外线路:codex 倍率低至 0.15,claude倍率低至 0.65 三网 CDN 线路:codex 倍率低至 0.16,claude倍率低至 0.7 国内极速线路:codex 倍率低至 0.

模型测评:GLM-5.2 大战 Claude Opus 4.8

祖传 Bug 模型大比拼:GLM-5.2 Thinking vs Claude Opus 4.8 Max 最新实测 最新模型测试战报 GLM-5.2 Thinking (ZCode 3.0): 96分。1. 解决表面问题。2. 解决深层问题。3. 发现引用的库的bug,没有改动库,没改本地代码规避库的bug。第四个发现三层bug的老师。 Claude Opus 4.8 Max (Cursor Max Mode): 100分。同GPT 5.5 xhigh(1. 解决表面问题。2. 解决深层问题,改动代码量比5.4少,比gemini3.1多。3. 发现引用的库的bug,没有改动库,本地业务代码优雅

【模型大横评2.0】composer2.5胜者为王|GLM5.2| Kimi2.7 |DSv4|GPT5.5|Gemini3.5F

测试仓库 依旧使用本人的一个闭源的项目,以下是具体架构。 评测流程 开 work tree 跑两个测试题目,1 题和 2 题之间不会新开上下文窗口。 测试题目 上次我让 codex 出了一个比较具体的题,虽然没有那么细,但还是给得比较细。这次只给大体意图,让 AI 自己去做。 打分流程 单分支评测 = 每个模型自己打分,不排名 参赛选手 测试速度 Kimi2.6 + Claude Code(35分钟左右) 结果 因为上次我秉持着最强模型出题和最强模型打分,所以这次三个模型进行分别打分。 鸣谢 首先鸣谢我自己,然后鸣谢小青年 @GinWU 提供的ollama pro订阅~~~ 花絮 后续 com

送6000个Grok free-token

如题,送6000个Grokfreetoken,使用grok2api导入直接使用,不保证全活。 分发站链接b64:aHR0cHM6Ly9jZGsubGludXguZG8vcmVjZWl2ZS9mMDA1NDA3Ny05NWE0LTQyMmQtOWM0NS0wMzFjZTU2ZmI4NTQ= 后面有空还会发,取之于佬用之于佬 因为分发站不太方便分发给一个用户多个token,所以简单搓了个兑换页,从分发站领到CDK以后去这里兑换你的token。 UnsnowAPI CDK 兑换中心 不教怎么用!善用搜索。 不会用也不想学就别领,留给会用的人;) 32 个帖子 - 22 位参与者 阅读完整话题