每日简报

2026-06-14

← 历史归档

iptv-org/iptv

TypeScript · ★ 119,119 · 🍴 6,369 · 📈 530 stars today

Collection of publicly available IPTV channels from all over the world

中文介绍 该项目是一个全球公开可用的 IPTV 频道资源集合,整理了来自世界各地的网络电视源。它主要通过维护更新的 M3U 播放列表文件,让用户能够方便地使用 VLC、IINA 等媒体播放器观看直播电视,适合需要寻找免费电视直播源的用户。

addyosmani/agent-skills

Shell · ★ 58,382 · 🍴 6,304 · 📈 1,514 stars today

Production-grade engineering skills for AI coding agents.

中文介绍 该项目为 AI 编码代理(如 Copilot、Claude Code)提供一套可直接使用的、生产级别的工程技能模块。它包含经过验证的代码模式、最佳实践和自动化脚本,旨在提升 AI 代理生成代码的质量与可靠性,主要面向开发或集成 AI 编程助手的工程师。

chatwoot/chatwoot

Ruby · ★ 30,859 · 🍴 7,554 · 📈 83 stars today

Open-source live-chat, email support, omni-channel desk. An alternative to Intercom, Zendesk, Salesforce Service Cloud etc. 🔥💬

中文介绍 这是一个开源的全渠道客户支持平台,提供实时聊天、邮件工单和社媒集成等功能,旨在替代 Intercom、Zendesk 等商业服务。它基于 Ruby on Rails 构建,支持灵活的部署和二次开发,适合希望建立自主可控客服系统的企业和团队。

obra/superpowers

Shell · ★ 226,929 · 🍴 20,181 · 📈 924 stars today

An agentic skills framework & software development methodology that works.

中文介绍 该项目提供了一个构建有效 AI 代理技能的结构化框架及配套的软件开发方法论。它通过定义清晰的模式和工具,帮助开发者系统地设计、测试和部署代理技能,从而提升 AI 自动化任务的成功率,适用于 AI 工具链的构建者。

apple/container

Swift · ★ 36,310 · 🍴 1,033 · 📈 1,487 stars today

A tool for creating and running Linux containers using lightweight virtual machines on a Mac. It is written in Swift, and optimized for Apple silicon.

中文介绍 这是苹果官方推出的工具,用于在 Mac 上通过轻量级虚拟机创建和运行 Linux 容器。它使用 Swift 编写,并针对 Apple silicon 芯片进行了优化,为 Mac 开发者提供了一种更原生、高效的容器化解决方案,用于开发和测试 Linux 环境。

music-assistant/server

Python · ★ 2,003 · 🍴 429 · 📈 270 stars today

Music Assistant is a free, opensource Media library manager that connects to your streaming services and a wide range of connected speakers. The server is the beating heart, the core of Music Assistant and must run on an always-on device like a Raspberry Pi, a NAS or an Intel NUC or alike.

中文介绍 Music Assistant 是一个免费的开源媒体库管理器,其核心服务器组件可连接多个流媒体服务(如 Spotify、Tidal)和各类网络扬声器。它实现了跨服务、跨设备的音乐统一管理和播放控制,主要面向拥有多个音乐服务和智能音箱的家居音乐爱好者。

kenn-io/agentsview

Go · ★ 2,362 · 🍴 216 · 📈 190 stars today

Local-first session intelligence and analytics for coding agents, supporting Claude Code, Codex, and more than 20 other agents. Also: 100x faster replacement for ccusage!

中文介绍 该工具提供针对各类 AI 编码代理(如 Claude Code、Codex 等)的本地会话分析与智能监控功能。它能帮助开发者追踪代理的调用、用量和成本,号称速度比现有工具(如 ccusage)快 100 倍,适用于需要精细化管理 AI 编程助手使用情况的团队。

LMCache/LMCache

Python · ★ 8,892 · 🍴 1,305 · 📈 238 stars today

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

中文介绍 LMCache 是一个为大型语言模型(LLM)推理加速的缓存层,通过高效存储和复用推理过程中的 Key-Value 数据,显著减少重复计算。它旨在提升 LLM 服务的响应速度和吞吐量,主要面向部署和优化 LLM 推理服务的技术团队。

microsoft/PowerToys

C · ★ 134,673 · 🍴 8,078 · 📈 370 stars today

Microsoft PowerToys is a collection of utilities that supercharge productivity and customization on Windows

中文介绍 这是微软官方推出的一套 Windows 系统增强工具集,包含窗口管理、快捷键预览、批量重命名等数十种实用功能。它旨在提升高级用户的生产力和 Windows 系统定制化能力,是 Windows 用户扩展系统功能的常用首选工具。

andrewyng/aisuite

Python · ★ 14,113 · 🍴 1,487 · 📈 127 stars today

Simple, unified interface to multiple Generative AI providers

中文介绍 该项目提供了一个简洁统一的 Python 接口,用于访问多家主流的生成式 AI 服务提供商(如 OpenAI、Anthropic 等)。开发者只需编写一次代码,即可轻松切换不同的 AI 模型后端,方便进行对比测试和灵活选择,简化了多模型集成开发流程。

NVIDIA/SkillSpector

Python · ★ 4,454 · 🍴 335 · 📈 804 stars today

Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, and security risks.

bannedbook/fanqiang

Kotlin · ★ 47,485 · 🍴 8,082 · 📈 93 stars today

翻墙-科学上网

中文介绍 该仓库收集了关于网络代理、VPN 和域名解析等突破网络访问限制的技术方法、工具与信息。它主要为面临网络连接障碍的用户提供相关资源参考,属于信息汇编类项目。

swc-project/swc

Rust · ★ 33,637 · 🍴 1,398 · 📈 20 stars today

Rust-based platform for the Web

中文介绍 SWC 是一个基于 Rust 编写的超高速 Web 开发平台,其核心是一个 TypeScript/JavaScript 编译器和打包器。它通过 Rust 的高性能和并行处理能力,实现了远超传统工具(如 Babel、Webpack)的构建速度,适合追求极致构建性能的前端项目。

x1xhlol/system-prompts-and-models-of-ai-tools

★ 140,304 · 🍴 34,658 · 📈 109 stars today

FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sourced) System Prompts

中文介绍 该项目收集并整理了大量当前流行 AI 工具(如 Cursor、Claude Code、Devin AI、Manus 等)的系统提示词和所使用的模型信息。它为研究者、开发者和爱好者提供了一个集中的参考,用于了解和分析不同 AI 工具的内部工作逻辑与配置。

Codex-maxxing: treating Codex like an operating loop

@BradGroux · 5.9K 粉丝 · 714.6K 阅 · 1.0K 赞 · 638 转

Most people still use coding agents like fancy autocomplete or a one-shot chat box. That leaves a lot of value on the table. The better pattern is to treat Codex like a durable operating loop:

中文介绍 博主提出应将 Codex 等编程智能体视为“持久操作循环”而非一次性自动补全,以挖掘更多价值,分享了一种更高效的使用模式。

Fable 5 (Mythos) Prompting Masterclass by Anthropic

@aiedge_ · 69.5K 粉丝 · 700.1K 阅 · 506 赞 · 68 转

TLDR: Anthropic just published the official playbook for prompting the most powerful AI model on earth - I translated it. Most people won't read this guide (it's buried in the API docs), which is

中文介绍 作者将 Anthropic 官方为最强模型发布的提示指南翻译整理。该指南被埋没在 API 文档中,内容是关于提示工程的官方最佳实践。

First Steps Toward Automated AI Research

@Recursive_SI · 6.3K 粉丝 · 465.1K 阅 · 516 赞 · 71 转

Early results from Recursive’s automated AI research system on model training and GPU kernel benchmarks Today we are releasing early results from Recursive’s automated AI research system. Across three

中文介绍 Recursive 公司发布其自动化 AI 研究系统的早期结果。该系统已在模型训练和 GPU 内核基准测试等任务上进行了应用。

Build self-improving agent system with Fable 5 in 14 steps : loops, dynamic workflows, routines

@0xCodez · 6.4K 粉丝 · 371.8K 阅 · 515 赞 · 56 转

Most people are using Claude Fable 5 like Sonnet 4.6 with a bigger context window. They prompt it. It works for 5 minutes. They close the tab. 9 out of 10 users have never run an agent system that

中文介绍 批评多数用户未充分发挥 Claude Fable 5 潜力,分享了一个 14 步构建具备循环和动态工作流的自我改进代理系统的具体方法。

Anthropic is losing the mandate of heaven

@haridigresses · 12.5K 粉丝 · 281.7K 阅 · 513 赞 · 36 转

Four months ago, in early February, Anthropic was the darling. OpenAI was the dominant behemoth to root against. Over the last 1-2 years, we'd seen the Sam ouster / return drama, Ilya and Mira had

中文介绍 分析 Anthropic 公司近几个月来公众形象的变化,从四个月前的“宠儿”状态,探讨其当前面临的挑战与争议。

Building a Good Vertical Agent

@BrainsAndTennis · 10.5K 粉丝 · 187.4K 阅 · 539 赞 · 45 转

How do you build an agent that actually performs in a domain — one customers pick because it's better? The basics have been standardized over the past year: an agent is a while-loop around a model

中文介绍 探讨如何构建真正表现优异的垂直领域智能体。其基础架构已被标准化:智能体本质上是一个围绕模型的 while 循环。

Anthropic's War on Opensource AI

@TheAhmadOsman · 61.0K 粉丝 · 74.9K 阅 · 507 赞 · 98 转

Anthropic wants the public to see one thing: the careful lab, the safety lab, the grown-up in the room trying to keep frontier AI from running off a cliff. However, the pattern around Anthropic does

中文介绍 质疑 Anthropic 公众形象(谨慎的安全实验室)与其围绕开源 AI 的实际行为模式之间存在反差,进行批评性分析。

Coinbase for Agents: Your AI Agent Can Now Trade and Pay with Coinbase

@coinbase · 7.0M 粉丝 · 72.8K 阅 · 500 赞 · 62 转

TL;DR: Coinbase for Agents connects your AI agent directly to your Coinbase account so it can trade, pay, and execute workflows on your behalf, all within limits you control. Available today as an MCP

中文介绍 Coinbase 推出“Coinbase for Agents”服务,允许 AI 代理直接连接用户账户进行交易和支付,并提供用户控制的限制,该服务已作为 MCP 上线。

ORACLE: Official AI Agents Trade on Polymarket

@ORACLEAIFND · 31.9K 粉丝 · 63.6K 阅 · 1.5K 赞 · 563 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 宣传其 ORACLE 项目,指出自主 AI 代理已成为预测市场的有效策略,据称 Polymarket 超过 30% 的活动来自算法和 AI 钱包。

ORACLE: Official AI Agents Trade on Polymarket

@OracleTrdading · 39.7K 粉丝 · 61.5K 阅 · 1.5K 赞 · 577 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 宣传其 ORACLE 项目,指出自主 AI 代理已成为预测市场的有效策略,据称 Polymarket 超过 30% 的活动来自算法和 AI 钱包。

ORACLE: Official AI Agents Trade on Polymarket

@OracleTrdade · 41.7K 粉丝 · 60.6K 阅 · 1.5K 赞 · 560 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 宣传其 ORACLE 项目,指出自主 AI 代理已成为预测市场的有效策略,据称 Polymarket 超过 30% 的活动来自算法和 AI 钱包。

ORACLE: Official AI Agents Trade on Polymarket

@OracleMarkett · 45.0K 粉丝 · 60.1K 阅 · 1.5K 赞 · 568 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 宣传其 ORACLE 项目,指出自主 AI 代理已成为预测市场的有效策略,据称 Polymarket 超过 30% 的活动来自算法和 AI 钱包。

/goal + Loss Functions: How to Distill a Product in 30 Hours with One Prompt [Full Playbook]

@elvissun · 45.0K 粉丝 · 42.0K 阅 · 502 赞 · 45 转

99% people are using /goal and loops wrong. The hype they hear is "long-running loops prompting autonomous agent": point it at a task, walk away, come back to working code. But top agentic engineers

中文介绍 分享一个利用 `/goal` 命令和损失函数在 30 小时内用单个提示提炼产品原型的完整工作流,指出多数人用错了这些功能。

Building recursive agent systems

@leerob · 258.6K 粉丝 · 36.8K 阅 · 586 赞 · 40 转

At Cursor, we run thousands of agents to help us train the next version of Composer. We give them research tasks, and if they aren't succeeding or run into issues, they DM us on Slack or page us via

中文介绍 介绍 Cursor 内部运行数千个代理来辅助训练下一代 Composer 的实践。这些代理执行研究任务,遇到问题会主动通过 Slack 联系。

New OpenAI Academy courses for the next era of work

OpenAI introduces three Academy courses that help people build practical AI skills, create repeatable workflows, and apply agents in everyday work.

中文介绍 OpenAI推出三个学院课程,帮助人们构建实用AI技能、创建可重复工作流,并在日常工作中应用AI代理。

[AINews] Loopcraft: The Art of Stacking Loops

a quiet day lets us highlight a great concept from Peter Steinberger, Boris Cherny, and Andrej Karpathy

中文介绍 Loopcraft展示了循环堆叠的艺术,由Peter Steinberger、Boris Cherny和Andrej Karpathy提出。

How Preply combines AI and human tutors to personalize learning

Preply uses OpenAI to launch AI-generated lesson summaries, providing personalised feedback and language learning exercises.

中文介绍 Preply结合AI和人类导师,利用OpenAI技术生成课程摘要,提供个性化反馈和语言学习练习。

Google DeepMind is worried about what happens when millions of agents start to interact

Google DeepMind is funding research into the potential dangers of situations where millions of different AI agents interact with each other online. According to Rohin Shah, who directs the company’s AGI safety and alignment research, the mass-market arrival of agents that can carry out tasks without

中文介绍 Google DeepMind资助研究,探讨百万AI代理在线交互的潜在危险,由AGI安全研究主管Rohin Shah指出。

Supporting Europe’s work in ensuring a trustworthy AI ecosystem

OpenAI supports the EU Code of Practice on AI content transparency, advancing provenance standards and tools to help people understand AI-generated content.

中文介绍 OpenAI支持欧盟的AI内容透明度实践准则,推动来源标准和工具发展,以帮助公众理解AI生成内容。

BBVA puts AI at the core of banking with OpenAI

Learn how BBVA scaled ChatGPT Enterprise to 100,000 employees and partnered with OpenAI to accelerate AI-powered banking transformation worldwide.

中文介绍 BBVA将ChatGPT企业版部署给10万名员工,并与OpenAI合作,加速全球银行业务的AI转型。

OpenAI to acquire Ona

OpenAI plans to acquire Ona to expand Codex with secure, persistent cloud environments, enabling long-running AI agents across enterprise workflows.

中文介绍 OpenAI计划收购Ona,旨在扩展Codex功能,提供安全持久的云环境,使AI代理能在企业工作流中长期运行。

How an astrophysicist uses Codex to help simulate black holes

Discover how astrophysicist Chi-kwan Chan uses Codex to build black hole simulations, helping scientists study extreme physics and test Einstein’s theory of general relativity.

中文介绍 天体物理学家Chi-kwan Chan利用Codex构建黑洞模拟,助力科学家研究极端物理现象并测试广义相对论。

Access OpenAI models and Codex through your Oracle cloud commitment

Access OpenAI models and Codex through Oracle Cloud, using existing commitments to build and deploy AI with enterprise security and governance.

中文介绍 OpenAI模型和Codex可通过Oracle Cloud访问,允许用户利用现有云承诺来构建和部署具有企业安全和管理的AI应用。

每日论文 · arXiv cs.CR 最新公告批次

周末 arXiv 通常无新公告。当前展示最近一次可用公告批次。

Beyond the IT Checklist: Engineering a Reasonable Standard of Care for Cyber Safety

第一作者: Matthew E. Jablonski · 方向: 软件安全

Abstract:Current U.S. cyber policy, centered on security, often treats documentation of controls and incident reports as a proxy for safety in the built environment. This paper argues that such an approach is inadequate for cyber-physical systems, where digital failures can produce kinetic harm. We construct and code a corpus of critical infrastructure policy documents (N=292, 2000-2025) to examine how "reasonable care" is operationalized across the NIST SP 800-160 Vol.~2 resilience lifecycle. The resulting maps show that obligations are concentrated in the Anticipate phase and emphasize administrative compliance, while Withstand and Recover phases rely heavily on delegated references to IT-focused control catalogs that are poorly aligned with physics-based hazards. We identify three major disconnects: miscalibrated delegated standards, recovery defined as notification rather than...

论文介绍 本文研究了当前美国网络政策在应对网络物理系统时的不足,指出其过度依赖IT控制文档而忽视物理危害。作者通过构建关键基础设施政策文件的语料库,分析了NIST SP 800-160 Vol. 2中「合理关注」的实施方式,发现义务集中在预期阶段,而抵抗和恢复阶段的IT控制目录与物理风险对齐不佳。研究揭示了标准错配、恢复定义不清等脱节,为网络安全工程提供改进方向。

Differentially Private Hierarchical Heavy Hitters

第一作者: Ari Biswas · 方向: 安全研究

Abstract:The task of finding _Hierarchical_ Heavy Hitters (HHH) was introduced by Cormode et al. [VLDB 2003] as a generalisation of the heavy hitter problem. While finding HHH in data streams has been studied extensively, the question of releasing HHH when the underlying data is private remains unexplored. In this paper, we study differentially private HHH release in both the streaming and non-streaming setting. In the non-streaming setting, we show the surprising result that the relative error in estimating the residual count for any prefix is independent of the height of the hierarchy and the number of heavy hitters in the stream. Meanwhile, in the streaming setting, although the exact version of HHH has low global sensitivity (as counting queries are 1-sensitive), the approximation functions due to streaming have high global sensitivity, linear in the available space. Despite this...

论文介绍 本文探索了在差分隐私约束下发布层次化重击手(HHH)的问题,分别研究了流式和非流式设置。在非流式中,残余计数估计的相对误差独立于层次高度和重击手数量;在流式中,尽管精确HHH全局敏感性低,但近似函数的高敏感性导致挑战。研究为隐私保护的数据发布提供了新方法,可能应用于数据流分析和隐私计算。

Intent-Based Cryptographic API Design for Cryptographic Agility

第一作者: Navaneeth Rameshan · 方向: 密码学协议

Abstract:As organizations move toward post-quantum cryptography, they face the major challenge of updating cryptographic algorithms across large, complex software portfolios. However, most cryptographic APIs in use today were designed around specific algorithms. These APIs expect explicit use of specific algorithms, provide little or no support for policy-based algorithm selection, and offer no straightforward way to migrate existing keys to newer algorithms. This makes the transition to post-quantum cryptography challenging. The companion assessment framework identifies the barriers to cryptographic agility and explains why algorithm transition is largely a software engineering problem. To address the limitations of current cryptographic APIs, we identify the principles necessary to design a cryptographically agile API. The design principles are derived from five fundamental...

论文介绍 本文针对向后量子密码学过渡中密码API的局限性,提出了基于意图的密码API设计原则。当前API围绕特定算法设计,缺乏策略选择和密钥迁移支持,使算法过渡困难。作者通过评估框架识别障碍,并设计密码敏捷API,旨在简化算法更新和迁移,提升软件系统的适应性和安全性。

An Assessment Framework for Application-Level Cryptographic Agility

第一作者: Navaneeth Rameshan · 方向: 密码学协议

Abstract:The impending post-quantum transition to new cryptography will require complete replacement of algorithms within all software. The cryptographic APIs used today make this transition challenging because they were not designed with agility as a concern. There is no method for systematically assessing cryptographic agility as an overall ability. In addition to this, the term itself refers to multiple independent capabilities. Specifically, it includes replacing algorithms, selecting by policy, and substituting implementations. This lack of structured decomposition limits both the evaluation of systems and the development of cryptographically agile APIs. We introduce a component-based assessment framework that characterizes application-level cryptographic agility along seven orthogonal dimensions: three coupling dimensions that measure what the application code knows about...

论文介绍 本文提出一个应用级密码敏捷性的评估框架,以解决缺乏系统评估方法的问题。框架从七个正交维度分解密码敏捷性,包括耦合维度等,用于表征应用程序对密码细节的依赖。这有助于评估系统并指导密码敏捷API的开发,支持向后量子密码学的平稳过渡。

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

第一作者: Zihao Wang · 方向: AI 安全

Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences. This makes them vulnerable to prompt-injection attacks, in which seemingly benign content embeds adversarial instructions that manipulate agent behaviour. Existing security benchmarks adopt an \textit{attack-centric} perspective, focusing on the technical feasibility of injections while overlooking the nuanced distribution of resulting harms. In practice, however, prompt-injection risk is victim-dependent: a single exploit can produce asymmetric consequences for different stakeholders, and the same attack pattern may exhibit substantially different effectiveness depending on whom it targets. To capture these properties, we introduce \textbf{\sysname}, a \textit{stakeholder-centric} benchmark...

论文介绍 本文针对大型语言模型驱动的网络代理面临的提示注入攻击问题,提出了一个利益相关者中心的基准测试框架。现有基准忽视攻击对不同受害者的不对称影响,而新框架评估提示注入在真实网络环境中的风险分布,考虑多个利益相关者,以更全面地衡量安全威胁和防护措施。

The Invisible Ink of the Android Malware World: A Longitudinal Study on the Usage of Covert Communication Channels

第一作者: Zeya Umayya · 方向: 软件安全

Proxies, VPNs and Tor have long helped the privacy community and users in censored regions to fight censorship. However, the same tools can be maliciously exploited by malware and botnets to conceal their communication to external command and control servers. Despite being a critical concern fueled by the proliferation of malware based attacks, no longitudinal studies have analyzed how malware applications use covert channels (CC) to evade detection. We fill this gap by performing the first study of the usage of covert channels in the Android malware ecosystem. To that end, we develop a multistage pipeline that combines static and dynamic analysis to investigate both system and network-level features. We applied this pipeline on a corpus of 3.5M Android malware spanning 2009 to July 2025. Our carefully crafted static validation rules uncovered 288K APKs that used CCs spanning 511...

论文介绍 本文通过纵向研究分析了Android恶意软件中隐蔽通信通道的使用情况。作者开发了一个结合静态和动态分析的多阶段管道,应用于350万个恶意软件样本(2009-2025年),揭示了288K APKs使用隐蔽通道的广泛现象。研究为理解恶意软件的通信策略和改进检测方法提供了实证基础。

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

第一作者: Jiaqi Luo · 方向: AI 安全

Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross. Within this broader red-line scenario, autonomous penetration represents a core enabling capability and subtask: the ability of LLM-powered AI systems to independently conduct adversarial operations against a target server without human intervention, identify and exploit vulnerabilities, and obtain unauthorized access or control. A growing body of work has sought to assess the autonomous penetration capabilities of AI systems. However, existing evaluations often employ opaque methodologies, rely on unrealistic or overly simplified penetration-testing scenarios, or provide LLMs with excessive prior knowledge and task-specific guidance, and cannot accurately capture the extent to which modern AI...

论文介绍 本文探讨了大型语言模型驱动的AI系统自主渗透能力的评估问题。现有评估方法存在不透明、场景简化等局限,无法准确反映现代AI系统的实际能力。作者强调自主渗透作为AI安全红线的核心子任务,研究其独立识别漏洞和执行攻击的潜力,为安全评估和风险防范提供见解。

DIG: Oracle-Guided Directed Input Generation for One-Day Vulnerabilities

第一作者: Andrew Bao · 方向: 软件安全

One-day vulnerabilities pose significant risks due to delayed or incomplete patch adoption. Generating proof-of-concept (PoC) inputs is therefore essential for assessing real-world impact. The key challenge is identifying necessary constraints for triggering the vulnerability and solving them effectively. Existing directed fuzzing approaches prioritize inputs toward target locations, but neither explicitly identify necessary constraints nor solve them effectively, relying instead on target-distance feedback and random mutation. Agentic approaches show strong potential through code reasoning and structured input generation, but goal drift in long-horizon reasoning limits their effectiveness. DIG addresses this challenge by exploiting a key property of one-day vulnerabilities: patches often reveal necessary preconditions for triggering. DIG uses an LLM to analyze the patch and synthesize...

论文介绍 本文提出DIG系统,用于生成一日漏洞的概念验证输入。一日漏洞因补丁采用延迟而风险高,现有模糊测试方法在识别必要约束上不足。DIG利用LLM分析漏洞补丁,合成触发漏洞的前提条件,以指导定向输入生成,提升漏洞利用的准确性和效率。

SoK: The Constant Time Model

第一作者: Billy Bob Brumley · 方向: 软件安全

Abstract:Constant time programming patterns is the primary defense against timing attacks on cryptographic implementations, yet what "constant time" means varies across academia and industry. This work systematizes constant time models and their evolution, identifies a recurring gap between what models protect and what specifications assume, and distills an offensive methodology for discovering timing vulnerabilities that originate outside the cryptographic primitive boundary. Applying this methodology, we locate a specification-level vulnerability related to private key loading, and confirm the leak in both OpenSSL and BoringSSL. Counterintuitively, BoringSSL's per-observation signal is several orders of magnitude stronger than OpenSSL's, despite an explicitly stricter threat model.

论文介绍 本文系统化梳理了密码学实现中防御时序攻击的“常量时间”编程模式及其演变。研究指出,现有模型保护的目标与安全规范假设之间存在持续差距,并提炼出一种发现密码原语边界外时序漏洞的攻击方法。应用该方法,作者发现了一个与私钥加载相关的规范级漏洞,并在OpenSSL和BoringSSL中证实了其存在。

ViPER: Vision-based Packing-Aware Encoder for Robust Malware Detection

第一作者: Fatima Qaiser · 方向: 软件安全

Abstract:Visualization-based malware detection maps raw binary bytes to grayscale images and applies learned visual classifiers, providing an evasion-resistant and disassembly-free alternative to conventional analysis pipelines. However, executable packing remains a critical failure mode: packed binaries produce high-entropy images that obscure the structural patterns these models rely on. Because packing is also prevalent in benign software (e.g., for compression or copy protection), packing state alone is not a reliable indicator of maliciousness, and existing approaches do not address this challenge within a unified supervised framework. We present ViPER, a Vision-based Packing-Aware Encoder for Robust malware detection. ViPER builds on a LoRA-adapted ViT-B/14 backbone with a dual-head architecture that jointly learns malware classification and packing detection. A packing-aware...

论文介绍 基于可视化的恶意软件检测方法面临可执行文件加壳的挑战,加壳会产生高熵图像,掩盖模型依赖的结构特征。本文提出ViPER,一种基于视觉的、感知加壳的鲁棒编码器。它采用双头架构,在一个经LoRA适应的ViT-B/14骨干网络上,联合学习恶意软件分类与加壳检测任务,旨在构建一个统一框架以应对加壳带来的检测难题。

MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

第一作者: Chejian Xu · 方向: 软件安全

Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering. In these systems, safety and security are inherently distributed across role-specialized agents, significantly expanding the attack surface, particularly under coordinated adversarial behaviors such as privilege escalation and cross-agent collusion. Existing red-teaming approaches for MAS remain limited: they rely on heuristic selection of target agents and perturb isolated message streams, leaving critical questions unanswered as which agents are most responsible for system safety, and how compromised agents can coordinate to bypass defenses. We propose MAStrike, a closed-loop framework for collusive red-teaming in hierarchical MAS. We propose the first agent-level Shapley value analysis for MAS, quantifying each agent's marginal...

论文介绍 层级多智能体系统的安全风险分布式且攻击面大,现有红队方法在评估协作攻击方面存在不足。本文提出MAStrike,一个用于层级多智能体系统协作攻击的闭环红队框架。其核心是首次引入智能体级的Shapley值分析,以量化每个智能体对系统安全的边际贡献,从而识别关键风险点并指导攻击策略。

LNTest: A Testbed for Evaluating Bitcoin Lightning Network-Based Botnets

第一作者: Thomas Bakaysa · 方向: 密码学协议

Abstract:Bitcoin's Lightning Network (LN) can be exploited as a covert, low-cost command-and-control (C&C) channel for botnets, as demonstrated by the LNBot and D-LNBot designs. However, both remain proof-of-concept prototypes evaluated only through simulation, leaving key questions about real-world topology formation, propagation complexity, and resilience to takedowns unanswered. We present LNTest, the first reusable testbed for LN-based botnets, built from Core Lightning nodes containerized with Docker over a shared Bitcoin Core regtest chain. LNTest supports three overlay topology modes (a deterministic chain, autonomous peer discovery, and user-supplied graphs), enabling controlled experiments across different botnet structures. Using LNTest, we report three main findings. First, D-LNBot's autonomous formation protocol does not produce the uniform chain from its design; instead...

论文介绍 比特币闪电网络可被用作僵尸网络的隐蔽指挥与控制通道,但现有研究仅为模拟原型。本文提出了LNTest,首个用于评估基于闪电网络的僵尸网络的可复用测试平台。该平台基于容器化的Core Lightning节点构建,支持多种覆盖网络拓扑模式,使研究者能够进行受控实验,以探索真实环境中的网络形成与传播特性。

A Privacy-Preserving Framework Using Remote Data Science for Inter-Institutional Student Retention Prediction

第一作者: John Fields · 方向: AI 安全

This study explores privacy-preserving machine learning (PPML) techniques using the PySyft platform to enable collaborative prediction of student retention between institutions. We developed a remote data science (RDS) framework with a semi-air-gapped architecture consisting of high-side and low-side servers, allowing researchers from three universities to build predictive models on sensitive student data without direct data access. Using historical data from a small private university (N=720), we evaluated three synthetic data generation approaches and validated the framework through inter-institutional collaboration. The results demonstrate consistent classification performance across institutions (Macro F1: 0.690--0.695) while maintaining strict Family Educational Rights and Privacy Act (FERPA) compliance. We also propose Data-Type-Aware Templates, a novel synthetic data method that...

论文介绍 本研究探索利用隐私保护机器学习技术,在不直接访问敏感数据的前提下,实现机构间的学生留级率协同预测。文章构建了一个基于PySyft平台的远程数据科学框架,采用半气隙架构,并在三个大学间进行验证。结果表明,该框架能在遵守教育隐私法规的同时,保持跨机构的一致分类性能。

Semantic Identification of IoT Devices from Behavioral Primitives

第一作者: Samuel Witt · 方向: 密码学协议

Abstract:Accurate identification of IoT devices is important for security management and policy enforcement. Existing approaches typically learn device signatures from packets or flow records. These methods operate on low-level communication observations whose traffic patterns may vary across deployments, software versions, and user interactions. This paper studies device identification using Manufacturer Usage Description (MUD) profiles. MUD profiles describe device behavior using Access Control Entries (ACEs), where each ACE represents a behavioral primitive consisting of protocol, endpoint, direction, and port semantics derived from device communication policy. Our contributions are threefold. First, using 28 publicly available MUD profiles containing 1,023 ACE instances, we construct ACE-level semantic representations from compact behavioral text and analyze their geometric...

论文介绍 准确的物联网设备识别对安全管理至关重要。现有方法依赖易变的底层流量特征。本文研究利用制造商使用描述(MUD)配置文件进行设备识别,这些文件通过访问控制条目(ACE)描述设备行为原语。作者构建了ACE级的语义表示并分析其几何特性,旨在提供一种基于高层行为策略、更具部署稳定性的设备标识方法。

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

第一作者: Pengfei He · 方向: AI 安全

Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources. Existing defenses mainly focus on blocking malicious content at inference time, and current red-teaming methods primarily optimize attack success. As a result, developers have limited visibility into how latent prompt injections emerge and propagate through agents. We propose PI-Hunter, an automated agentic auditing framework for proactive vulnerability exposure in LLM agents. PI-Hunter constructs realistic source-aware test cases and iteratively evolves them through feedback-driven exploration to induce agents to retrieve and reveal latent malicious instructions embedded within external environments. Extensive experiments across multiple benchmarks, agent...

论文介绍 大语言模型代理通过外部环境交互引入了间接提示注入的新风险。现有防御多集中于推理时拦截,红队方法侧重于攻击成功率。本文提出PI-Hunter,一个自动化的代理审计框架,旨在主动暴露LLM代理中的潜在漏洞。它通过构建逼真的测试用例并进行反馈驱动的探索,诱导代理检索并揭示嵌入外部环境中的恶意指令。

SMSR: Certified Defence Against Runtime Memory Poisoning in Persistent LLM Agent Systems

第一作者: Tarun Sharma · 方向: 软件安全

Abstract:Retrieval-augmented generation (RAG) agents increasingly run with persistent memory that accumulates across user sessions. This creates a new attack surface: an adversary interacting only through normal channels can inject crafted memories that, once retrieved, steer the agent's responses for future users, without touching model weights or code. We call this Multi-Session Memory Poisoning (MSMP) and show that no existing defence certifies against it; static-corpus defences (RobustRAG, ReliabilityRAG) assume a fixed knowledge base, and heuristic filters are bypassed by fluent enterprise-style text. We present Signed Memory with Smoothed Retrieval (SMSR), the first defence with a certified robustness bound for this setting. Component 1 adds HMAC-SHA256 provenance at write time, blocking unsigned injection. Component 2 applies randomised memory ablation with verdict-based...

论文介绍 具备持久化内存的检索增强生成代理面临跨会话内存投毒攻击,攻击者可通过正常交互注入恶意记忆,影响未来响应。现有防御对此缺乏认证性保证。本文提出“签名内存与平滑检索”方案,通过为内存条目添加来源签名以阻止未授权注入,并采用随机化内存消融与基于判定的检索,首次为此类威胁提供了可证明的鲁棒性界限。

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents

第一作者: Siyu Shen · 方向: 隐私保护

Screenshot-based mobile GUI agents can operate ordinary smartphone apps through the same visual interface as a human user, but this capability also turns every screen observation into a privacy boundary. During normal task execution, screenshots may expose contacts, messages, photos, files, recommendations, health cues, and other sensitive context that is unrelated to the user's request. We call this problem incidental visual privacy exposure. It is difficult to address with existing defenses: text anonymization misses many visual and inferential cues, while generic privacy masking can remove the evidence and controls that a GUI agent needs to complete the task. This paper presents CAPED, a context-aware pre-upload exposure control layer for mobile GUI agents. CAPED is designed as a phone-side protection layer: before screenshots are released to a remote multimodal agent, it extracts...

论文介绍 本文关注基于截图的移动GUI代理在执行任务时引发的「附带视觉隐私暴露」问题。现有防御手段如文本匿名化或通用隐私遮盖存在不足,可能干扰任务完成。为此,论文提出了CAPED,一个上下文感知的预上传暴露控制层。该方案作为手机端的保护层,在截图发送给远程多模态代理之前,提取并识别其中与任务无关的敏感信息,旨在实现隐私保护与任务功能之间的平衡。

Amnesia: A Stealthy Replay Attack on Continual Learning Dreams

第一作者: Ahmed Sharshar · 方向: 安全研究

Abstract:Continual learning (CL) models often use experience replay to reduce catastrophic forgetting, but their robustness to replay sampling interference remains underexplored. Existing CL attacks alter inputs or training pipelines (poisoning/backdoors) and rarely include explicit auditable constraints, limiting realism. Here, auditability means a monitor can verify compliance from sampler-visible telemetry - e.g., logged replay index/label statistics - by checking that the realized replay class histogram stays close to a nominal baseline and that replay rate is unchanged per batch and/or over a rolling window. We study a limited-privilege insider who controls only replay index selection, not pixels, labels, or model parameters, while staying within auditable limits such as queue priorities. We introduce Amnesia, a replay composition attack that maximizes degradation under two...

论文介绍 持续学习模型常使用经验回放来减轻灾难性遗忘,但其对回放采样干扰的鲁棒性研究尚不充分。本文提出了一种隐蔽的「失忆」攻击,它由仅能控制回放索引选择、且在可审计限制内操作的有限特权内部人员发起。该攻击通过优化回放样本的组合来最大化模型退化,同时保持可审计指标不变。这项研究揭示了持续学习系统中一个现实且难以检测的新安全威胁。

Beyond Attack Success Rate: Examining Trigger Leakage in Vision-Language Agentic Systems

第一作者: Jiamin Chang · 方向: 系统安全

Vision-Language Agentic Systems (VLAS) connect visual perception to planning, tool use, and physical actions. This means backdoor-type triggers can propagate through both decision pipelines and their connected interfaces, thus making visual backdoors a system-level threat. Current evaluations on such backdoors focus on clean accuracy and attack success rate (ASR), metrics that capture whether a trigger works, but not whether an attack is actually "precise" -- i.e. whether it triggers hidden behaviors only when intended. In this work, we formalize the failure of trigger precision as "trigger leakage": inputs that are visually or semantically close to the intended trigger and therefore inadvertently activate the attacker-specified behavior. To quantify this leakage, we introduce Neighbor Leakage Rate (NLR). Our experiments show that at a 3% poisoning ratio, icon and text triggers remain...

论文介绍 针对视觉语言智能体系统,现有后门评估多关注攻击成功率,却忽略了触发器的精确性。本文将视觉或语义上接近恶意触发器的输入意外激活后门行为的现象定义为「触发泄漏」,并引入邻居泄漏率进行量化。实验表明,在较低投毒比例下,图标和文本触发器仍存在显著泄漏问题。这凸显了后门攻击作为系统级威胁的复杂性,强调了评估其精确性的重要性。

From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model Merging

第一作者: Zhenqian Zhu · 方向: 安全研究

Abstract:Model merging (MM) has gained significant attention as a cost-effective approach to integrate multiple task-specific models into a unified model. However, recent work reveals that MM is highly susceptible to backdoor attacks. Existing defenses based on task arithmetic often fail to eliminate backdoors without substantially degrading clean-task performance, owing to their reliance on direct parameter-space editing. To address this gap, we propose Linear Feature Path Minimization (LFPM), a backdoor mitigation framework for model merging, which introduces an anti-backdoor task vector into the backdoored merged model. Unlike prior approaches, LFPM formulates the backdoor robustness of the merged model from a unified feature-space perspective under the Cross-Task Linearity (CTL) framework, which leverages the approximate linearity of features across tasks. This perspective guides...

论文介绍 本文通过一个全面的因子实验(涵盖432种配置),系统研究了检索增强生成系统对投毒攻击的鲁棒性影响因素。研究分析了数据集、检索器类型、检索深度、数据库构成、分块策略和生成器模型等因素对检索级和生成级指标的影响。结果表明,检索器架构、数据集和检索深度是影响投毒暴露的最强因素,而生成器的选择和数据库构成对下游攻击效果有重大影响。

Influence Factors on RAG Poisoning

第一作者: Pedro Pereira · 方向: AI 安全

Abstract:Retrieval-Augmented Generation (RAG) systems enhance large language models by grounding responses in retrieved documents from external knowledge sources at inference time. However, this reliance on retrieved content introduces vulnerabilities to poisoning attacks, in which adversarial documents can manipulate both the retrieval process and the generated outputs. This paper investigates poisoning robustness in RAG through a full factorial experimental study covering 432 configurations. We analyze the impacts of dataset, retriever type, retrieval depth, database composition, chunking strategy, and generator model on retrieval-level and generation-level metrics. The results show that retriever architecture, dataset, and retrieval depth are the strongest factors affecting poisoning exposure, while generator choice and database composition have a major impact on downstream attack...

论文介绍 本文研究检索增强生成(RAG)系统中的中毒攻击鲁棒性问题。RAG系统通过外部文档增强大语言模型,但易受对抗文档的攻击。作者通过全因子实验分析了432种配置,考察数据集、检索器类型、检索深度、数据库组成、分块策略和生成器模型对检索和生成指标的影响。结果表明,检索器架构、数据集和检索深度是影响中毒暴露的主要因素,而生成器选择和数据库组成对攻击效果有关键作用。这项研究有助于理解RAG系统的安全弱点并指导改进设计。

Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks

第一作者: Achraf Hsain · 方向: AI 安全

Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata restricting an agent's actions. We argue this is the wrong product. The same automata-theoretic machinery -- specification compilation, product game construction, attractor computation, and winning-region extraction -- is better read as a design-time analytical instrument whose outputs are structural insights about a system rather than runtime constraints on a deployed agent. We instantiate this through a constrained two-player safety game for network defense. The two specifications are enforced asymmetrically: the defender specification defines the unsafe region of the game, whereas the attacker specification restricts the adversary's legal actions during attractor computation. Solving the game yields a defensibility verdict -- a formal...

论文介绍 本文重新审视了强化学习中的防护罩概念,提出将其视为一种设计时分析工具,而非仅是运行时安全机制。作者通过一个用于网络防御的约束两人安全博弈实例化了这一观点。求解该博弈会得出一个可防御性结论,即为防御者提供了一种形式化方法,用于分析对抗性网络中实现特定安全目标的可行性及其结构性洞察。

Split Tallies: A Discrete Certificate Calculus for Auditing Dynamic Ordered Sets in Constant Memory

第一作者: Faruk Alpay · 方向: 安全研究

Abstract:We study retrospective auditing for dynamic ordered sets maintained by an untrusted party. A passive auditor watches insert, delete, membership, predecessor, successor, min, and max operations, stores five machine words and a flag, and receives a constant-size public tally record per operation. At audit time the maintainer discloses the claimed live vacant intervals. The method represents order semantics by maximal gaps: gaps are born, cited, consumed, and timestamped, while two hidden field accumulators test equality of the birth and consumption ledgers. Honest executions are accepted with probability one. If any answer in a T-operation session is wrong, acceptance occurs with probability at most (4T+1)/p over one secret field element, against computationally unbounded maintainers. We prove that deterministic and visible-coin auditors require linear state, and that removing...

论文介绍 本文研究了如何对不可信方维护的动态有序集合进行回溯审计。一个被动审计者仅需存储常数大小的状态(五个机器字和一个标志),并在每次操作后接收一个常数大小的公开「计数记录」。该方法利用最大间隙表示顺序语义,并通过隐藏的域累加器测试账本的一致性。它能在诚实执行时以概率一接受,在存在错误时以高概率检测,为高效数据审计提供了新方案。

Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models

第一作者: Jianwei Fei · 方向: 安全研究

Abstract:Model fingerprinting, embedding user-specific identifiers (fingerprints) into generated outputs, has recently emerged as a popular solution to protect the intellectual property rights (IPR) of generative text-to-image (T2I) models and prevent unauthorized redistribution. In this work, we reveal a previously unexplored systematic vulnerability in existing generative model fingerprinting methods: they lack robustness against collusion attacks, where multiple attackers combine their models to remove or obscure the fingerprints. To address this issue, we take the first step towards a robust fingerprinting method for T2I models with anti-collusion capabilities. The proposed method encodes strings of bits, namely fingerprints, into the coefficients of a personalized normalization module (PNM) incorporated into T2I models, so that fingerprints can be reliably recovered from any...

论文介绍 现有生成式文本到图像模型的指纹技术缺乏对共谋攻击的鲁棒性。为解决此问题,本文首次提出一种具备抗共谋能力的扩散模型指纹方法。该方法将指纹(比特串)编码到个性化归一化模块的系数中,使得指纹能够从任意由该模型生成的图像中可靠恢复,即使多个攻击者合并其模型试图移除或混淆指纹。此工作增强了生成模型知识产权保护的完整性。

PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learning Agent

第一作者: Junfeng Guo Heng Huang · 方向: AI 安全

Abstract:While real-world applications of reinforcement learning (RL) are becoming increasingly popular, the security of RL systems deserve more attention and exploration. In particular, recent work has revealed that RL agents are vulnerable to backdoor attacks, where a victim agent behaves normally under standard conditions but executes malicious actions when a specific trigger is activated. Existing backdoor defenses for RL either require access to the agent's internal parameters, operate only at the model or trajectory level, or are limited to specific attack types. To ensure the security of RL agents, we propose \texttt{PolicyGuard}, a \textit{test-time step-level} backdoor defense which leverages Gaussian Process (GP) posterior variance and adapts pseudo trajectories to enable uncertainty computation for individual time step. Besides, we also provide theoretical foundations to...

论文介绍 该研究针对强化学习智能体面临的后门攻击威胁,提出了一种名为PolicyGuard的测试时步骤级防御框架。现有防御方法常需访问模型内部参数或仅能在模型/轨迹层面操作,而PolicyGuard利用高斯过程后验方差与伪轨迹,在独立时间步计算不确定性,从而实现对触发后恶意行为的实时检测与防御。该方法为保障RL系统安全提供了一种轻量级且通用的运行时防护方案。

Detecting Functional Memorization in Code Language Models

第一作者: Matthieu Meeus · 方向: 软件安全

Abstract:Large language models (LLMs) are increasingly used to generate code at scale. Meanwhile, prior work has investigated whether training data may be recoverable from model outputs, by auditing the textual overlap between training examples and model generations. Code, however, can be functionally equivalent while textually dissimilar. In this work, we study functional memorization: extraction of functional logic beyond what verbatim metrics detect. We construct a counterfactual setup for Olmo-3-32B, comparing a midtrained model (exposed to target code) against a pretrained reference (not exposed). We prompt both models with Python function signatures and measure both textual and functional similarity (i.e., LLM-as-a-judge, execution-based). Our results show clear evidence of functional memorization, highlighting the need for auditing metrics that go beyond textual overlap.

论文介绍 本文研究了大型语言模型中的功能性记忆现象,即模型学习并提取了训练数据中超越文本表面重叠的功能逻辑。通过为Olmo-3-32B构建反事实设置,对比了暴露与未暴露于目标代码的模型,并结合文本相似性、LLM判别及执行验证进行度量。实验明确发现了功能性记忆的存在,强调了开发能够审计功能逻辑而非仅限于文本重叠的指标的必要性。

Smarter Saboteurs, Better Fixers: Scaling & Security in Linear Multi-Agent Workflows

第一作者: Timothy McAllister · 方向: AI 安全

Abstract:As LLM-based multi-agent systems (MAS) are deployed in the wild, the resilience of their collaboration structures against adversarial compromise becomes a critical safety concern. Attackers may leverage prompt-injection or jailbreaking to sabotage individual agents within MAS workflows, but the interaction between model scaling and system-level resilience remains poorly understood. This paper investigates how model scale affects the security of linear multi-agent workflows. Our experiments across scales of two open-weight model families on the HumanEval benchmark reveal a compliance-correction symmetry: larger models are far more likely to faithfully execute malicious instructions, with the control-to-malicious performance drop reaching 53.7pp at 27B in uncorrected pipelines. However, appending a lightweight terminal Fixer stage collapses this to 0.6pp and restores statistical...

论文介绍 该研究探讨了大语言模型多智能体系统中,模型规模与线性工作流安全性之间的关系。实验发现了一种“服从-纠正对称性”:较大的模型更倾向于忠实执行恶意指令,导致性能大幅下降。然而,在流程末端附加一个轻量级的“修复”智能体阶段,能将安全风险急剧降低。研究揭示了模型规模化对系统韧性带来的双重影响。

Fed-FBD: Federated Functional Block Diversification for Isolation, Privacy, and Surgical Unlearning

第一作者: Weijie Chen · 方向: AI 安全

Abstract:Federated learning (FL) enables collaborative model training without sharing raw patient data, but standard approaches such as FedAvg treat each client as a black box and provide no mechanism for isolating an adversarial contributor, auditing per-client influence, or honoring a departed participant's right to be forgotten. We present Fed-FBD (Federated Functional Block Diversification), a modular federated architecture that decomposes a ResNet backbone into six functional blocks (the stem, four residual groups, and the classification head) and maintains a warehouse of N color variants, each assembled from independently tracked and contributor-stamped blocks. Fed-FBD provides three capabilities absent in FedAvg: (i) architecturally guaranteed block-level isolation, so that an adversarial or mislabelled client cannot contaminate the clean colous; (ii) privacy-by-design, where...

论文介绍 本文提出了Fed-FBD,一种模块化的联邦学习架构。它将ResNet骨干分解为功能块,并维护一个由独立追踪块组成的“仓库”,从而解决了传统联邦平均方法中客户端黑盒化的问题。Fed-FBD实现了三个关键能力:架构保证的块级隔离以防御恶意客户端、设计上的隐私保护,以及支持对离开客户端数据进行“外科手术式”的精确遗忘。

SAIGuard: Communication-State Simulation for Proactive Defense of LLM Multi-Agent Systems

第一作者: Ruxue Shi · 方向: AI 安全

Abstract:LLM-based multi-agent systems (MAS) solve complex tasks through inter-agent collaboration, but their communication-driven nature also allows security risks to spread across agents and trigger system-wide failures. Existing MAS defenses mainly follow a reactive paradigm after execution by detecting and isolating harmful agents, which may cause irreversible damage and degrade collaborative utility. To address this, we propose a proactive defense framework for MAS security, namely a Simulation-aware Interception Guard (SAIGuard). SAIGuard performs communication-state simulation over the MAS interaction graph, estimates the impact of incoming messages on local agent states and the global MAS state, and detects risky messages via reconstruction deviations from benign communication patterns. Instead of isolating agents, SAIGuard sanitizes or regenerates suspicious messages before it...

论文介绍 针对基于LLM的多智能体系统中风险通过通信传播的问题,本文提出了主动防御框架SAIGuard。与传统事后检测和隔离有害智能体的被动范式不同,SAIGuard在智能体交互图上执行通信状态模拟,估计传入消息对局部和全局状态的影响,并通过重建与良性通信模式的偏差来检测风险消息。该方法能在损害发生前对可疑消息进行清洗或重新生成。

Improving Robotic Generalist Policies via Flow Reversal Steering

第一作者: Andy Tang · 方向: 机器人操作 · 来源: cs.RO

Generalist policies can learn a wide range of skills from diverse robot datasets. In order to solve or improve on challenging news tasks, we need a way to infer and invoke the appropriate actions from the policy's rich behavioral prior, especially when directly commanding the policy fails. We focus on flow matching generalists and propose Flow Reversal Steering (FRS): a method that takes suboptimal but ``reasonable'' actions, finds their latent noises by passing them through the flow policy in reverse, and maps them to nearby generalist action modes. We evaluate FRS across many simulated and real-world manipulation settings. First, FRS can turn coarse semantic guidance from humans or vision-language models (VLMs) into corresponding good robot actions, improving zero-shot control. These gains can be distilled with behavioral cloning by training an auxiliary policy to output noises that...

论文介绍 本文提出了流逆转引导方法,以提升机器人通用策略在新任务上的表现。该方法对流匹配策略进行逆转,将次优但合理的动作映射回潜在噪声空间,再引导至通用策略的动作模式附近。实验表明,FRS能将来自人类或视觉语言模型的粗粒度语义指导转化为精细的机器人动作,提升零样本控制性能,并可通过行为克隆进行蒸馏学习。

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

第一作者: Daichi Azuma · 方向: 导航与运动 · 来源: cs.RO

Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change brings it closer to the goal. Navigation world models provide such visual foresight, but they remain prediction modules that require an external planner to convert predicted futures into closed-loop control. We propose Navigation World Action Model (NavWAM), a diffusion-transformer policy that turns navigation world-model prediction into executable action by representing future observations, goal-progress values, and action chunks in a shared latent sequence. By learning future prediction jointly with the action and value targets that determine closed-loop behavior, NavWAM makes visual foresight directly usable for robot control. We build NavWAM through simulation pretraining and real-robot adaptation, and...

论文介绍 为解决目标条件视觉导航问题,本文提出了导航世界动作模型NavWAM。它是一个扩散变换器策略,将未来观测、目标进展值和动作块编码在共享的潜在序列中,从而把导航世界模型的预测直接转化为可执行动作。通过联合学习未来预测与闭环行为所需的动作和价值目标,NavWAM实现了视觉预见与机器人控制的端到端集成。

See Selectively, Act Adaptively: Dual-Level Structural Decomposition for Bimanual Robot Manipulation

第一作者: Yoon-Ji Choi · 方向: VLA 通用模型 · 来源: cs.RO

In bimanual robotic manipulation, task-relevant visual information varies with the task stage and context, while the interaction of the two arms shifts between independent and coordinated modes, making policy learning challenging. However, existing monolithic Vision-Language-Action (VLA) policies process diverse visual inputs and interaction patterns through a single shared representation and action generation pathway, often failing to separately account for visual relevance and bimanual interaction structure. To address this issue, we propose a bimanual manipulation VLA framework based on Dual-Level Structural Decomposition. The View-Selective Visual Router dynamically adjusts wrist-view contributions to emphasize relevant visual cues, while the Interaction-Aware Action Mixture-of-Experts (MoE) decomposes action generation into coordinated and arm-wise pathways to adapt to varying...

论文介绍 在双臂机器人操作中,视觉信息的相关性与双臂交互模式随任务阶段动态变化。现有单体视觉-语言-动作策略难以分别处理这两方面。本文提出基于双级结构分解的VLA框架:视图选择视觉路由器动态调整腕部视图贡献以聚焦关键视觉线索,交互感知动作混合专家将动作生成分解为协调与独立路径,从而适应不同的交互需求。

EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation

第一作者: Kailin Wang · 方向: 机器人操作 · 来源: cs.RO

Pretrained-feature world models provide a useful substrate for robot imagination, but visual or latent prediction alone does not determine whether an imagined future satisfies task-relevant events. Long-horizon manipulation requires progress signals that are relational, predicate-level, and physically grounded: whether an object has moved, whether a drawer or contact state has changed, whether a placement predicate is satisfied, and whether a candidate future is reliable enough for execution. We introduce EA-WM, an event-aware world-model framework that augments frozen visual-feature dynamics with task-specification-grounded event prediction and verification. EA-WM rolls out candidate futures in pretrained visual-feature space, decodes them into structured event states, and scores them using task-progress, semantic-consistency, physical-feasibility, and uncertainty terms. The verifier...

论文介绍 针对长周期机器人操作中缺乏明确任务进展信号的问题,本文提出了事件感知世界模型EA-WM。该方法在预训练的视觉特征世界模型基础上,引入基于任务规范引导的事件预测与验证机制。它通过结构化事件状态解码与多维度评分(任务进度、语义一致性、物理可行性、不确定性),对候选未来轨迹进行筛选与排序,旨在为机器人操作提供可靠、可解释的想象与决策支持。

GeoCFNet: Geometry-Aware Confidence Field Network for Robot-Assisted Endoscopic Submucosal Dissection

第一作者: Rui Tang · 方向: 具身智能 · 来源: cs.CV

Advanced surgical robotics has made robot-assisted endoscopic submucosal dissection (ESD) a promising approach for the en-bloc resection of large lesions, with the potential to reduce recurrence and improve long-term outcomes. However, the technical complexity and risk of complications in ESD demand stable and precise visual guidance to maintain an accurate dissection corridor and a safe tissue margin. Dense confidence fields provide an effective representation for this purpose by describing both the preferred dissection region and its spatial transition to surrounding tissue. However, reliable confidence field estimation remains challenging in dynamic endoscopic scenes due to smoke, specular highlights, tissue deformation, weak texture, and the thin geometric structure of the target region. To address these challenges, we formulate dissection guidance as a geometry-aware confidence...

论文介绍 为解决机器人辅助内镜粘膜下剥离术中因烟雾、组织变形等因素导致的精确视觉引导难题,本文提出了几何感知置信场网络GeoCFNet。该方法将剥离引导建模为几何感知的置信场估计问题,旨在同时预测首选剥离区域及其向周围组织的空间过渡,以维持安全的剥离通道和组织边界,从而提升手术的稳定性和精确性。

Learning to Assist: Collaborative VLAs for Implicit Human-Robot Collaboration

第一作者: Leo Xu · 方向: VLA 通用模型 · 来源: cs.RO

Human-robot collaboration (HRC) combines the complementary strengths of humans and robots to improve task efficiency. However, many existing collaborative systems rely on hand-engineered pipelines, limiting their scalability and flexibility for new tasks. In this work, we show that models trained end-to-end with imitation learning, specifically vision-language-action (VLA) models, can support collaborative manipulation, and characterize the key factors affecting their real-world performance. We evaluate two state-of-the-art models and identify a failure mode of action-chunking policies in implicit HRC, where demonstration action leakage (i.e., action chunks crossing latent task transitions) can cause premature assistive behavior. We find that this issue increases with longer execution horizons and occurs in real-world collaborative VLA systems, such as when a robot attempts to hand...

论文介绍 本文研究了基于视觉-语言-动作模型(VLA)的端到端模仿学习在隐式人机协作中的应用。作者评估了当前先进模型,发现其动作分块策略在协作任务中存在“演示动作泄漏”导致过早协助的失败模式,并分析了该问题在长执行周期任务中的加剧。研究为理解与改进VLA模型在真实协作场景下的表现提供了关键洞察。

Mana: Dexterous Manipulation of Articulated Tools

第一作者: Zhao-Heng Yin · 方向: 机器人操作 · 来源: cs.RO

Abstract:Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal degrees of freedom and contact-rich interactions. While prior work has largely focused on rigid objects, articulated tool use remains underexplored because of its physical complexity and the difficulty of learning functional grasping and manipulation policies. We present Mana (Manipulation Animator), a general sim-to-real framework that reinterprets dexterous manipulation as an animation problem. Inspired by computer animation, Mana employs a coarse-to-fine pipeline that transforms procedurally-generated grasp keyframes into manipulation trajectories through motion planning and reinforcement learning. The data generation process is largely automatic, requiring only a few mouse clicks to specify functional affordances (<1 minute per tool). Across four articulated...

论文介绍 关节工具的操作因需协调内部自由度和复杂的接触交互而极具挑战。本文提出Mana框架,将灵巧操作重新诠释为一个动画问题。该框架采用粗到细的流水线,通过运动规划和强化学习,将程序化生成的抓取关键帧转换为操作轨迹,仅需少量用户交互即可为不同工具生成功能性抓取策略,并成功实现从仿真到现实的迁移。

$\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

第一作者: Arnav Kumar Jain · 方向: 机器人操作 · 来源: cs.RO

Abstract:The potential impacts of world models (WMs, i.e., learned simulators) on robotics are far-reaching -- policy evaluation, policy improvement, and test-time planning -- all with limited real-world interaction. To unlock these downstream capabilities, a WM needs to jointly satisfy three desiderata: $\textit{(i)}$ fidelity (i.e., producing simulated trajectories that correlate with reality), $\textit{(ii)}$ consistency (i.e., producing simulated trajectories that are coherent over long horizons), and $\textit{(iii)}$ efficiency (i.e., producing simulated trajectories quickly). We propose $\texttt{WEAVER}$ (World Estimation Across Views for Embodied Reasoning): a WM architecture that simultaneously achieves all three desiderata, providing state-of-the-art results on robotic manipulation tasks. $\texttt{WEAVER}$ is a multi-view WM trained to predict future latents and reward values...

论文介绍 为支持机器人的策略评估、改进和规划等下游任务,世界模型需同时满足高保真度、长时序一致性和高效生成三个目标。本文提出了WEAVER架构,这是一个多视角世界模型,旨在联合预测未来的隐状态和奖励值。该模型在机器人操作任务上达到了当前最优的性能,为构建有效可靠的机器人学习世界模型提供了新的解决方案。

MCR-Bionic Hand: Anatomical Structural Priors for Dexterous Manipulation

第一作者: Haosen Yang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dexterous robotic hands are usually formulated as high dimensional active control systems governed by degrees of freedom, actuation, and algorithms. Human hand dexterity, however, is partly encoded in the physical architecture of bones, ligaments, tendons, aponeuroses, and intrinsic muscles. This work describes that contribution as two linked forms of structural intelligence: structural prior generation, in which wrist to finger tenodesis, FDS/FDP routing, and the dorsal extensor hood transform low dimensional posture inputs into default grasp configurations and PIP to DIP coordination; and muscle mediated modulation, in which extrinsic muscles, lumbricals, and interossei regulate MCP posture, distal stability, fingertip force paths, and contact states around that default state. Based on this framework, MCR-Bionic Hand is developed as a 1:1 musculoskeletal biomimetic hand...

论文介绍 人类手的灵巧性部分源于其骨骼、韧带、肌腱等物理结构的“结构智能”。本文基于此解剖学先验,提出了MCR仿生手。该设计将结构智能分为“结构先验生成”和“肌肉介导调制”两个层面,前者通过肌腱路径等将低维姿态输入转化为默认抓取配置,后者通过外部肌肉等调节接触状态和力分布,实现了高度仿生的灵巧操作。

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale

第一作者: Nils Blank · 方向: 机器人操作 · 来源: cs.RO

Abstract:This work introduces Spatial Annotations from Robot Demonstrations with Reliability Calibration (SPARC), a risk-aware framework that automatically labels robot demonstrations with structured spatial annotations and assigns each annotation a reliability score. Structured spatial annotations, such as bounding boxes, object trajectories, and manipulation phase labels, benefit a broad range of robotics applications from training grounded robot policies and embodied foundation models to motion planning and hierarchical task composition. Existing automated pipelines generate such annotations at scale but provide no reliable quality signal: detector confidence is poorly calibrated for annotation correctness, forcing a choice between accepting noisy labels or discarding useful samples. In contrast to existing automated pipelines, SPARC leverages the spatio-temporal structure inherent...

论文介绍 为大规模机器人演示数据自动生成带可靠性评分的结构化空间标注(如边界框、轨迹),本文提出了SPARC框架。与现有自动化流水线不同,SPARC利用演示数据固有的时空结构,通过风险感知校准,为每个自动标注分配一个可靠的正确性评分,从而在过滤噪声标签和保留有用样本之间提供更优的平衡,服务于策略学习和任务规划。

GIVE: Grounding Human Gestures in Vision-Language-Action Models

第一作者: Pengfei Liu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Human communication is inherently multimodal, where language is often accompanied by non-verbal cues such as gestures to convey intentions. However, current Vision-Language-Action (VLA) models treat robotic manipulation as a pure text-driven task, overlooking the important role of gestures in Human-Robot Interaction (HRI). This often leads to inaccurate intent grounding and unreliable manipulation when language instructions are ambiguous or underspecified. To address this challenge, we propose GIVE (Gesture Intent via Visual-Semantic Enhancement), an effective approach that enhances pre-trained VLA models with human gesture understanding without architectural modifications. Specifically, GIVE incorporates gesture information through two complementary pathways: a visual pathway that overlays hand skeletons and fingertip rays onto robot observations for explicit object...

论文介绍 当前视觉-语言-动作模型多为纯文本驱动,忽略了手势等非语言线索在人机交互中的作用,可能导致意图理解不准确。本文提出GIVE方法,在不修改VLA模型架构的前提下,通过视觉和语义两条互补通路融入手势信息(如手部骨骼叠加和指尖射线),增强模型对人类手势意图的理解,从而提升在语言指令模糊场景下操作的可靠性。

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation

第一作者: Xiangyu Zhu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Whole-body mobile manipulation requires coordinating mobile base and manipulator under shifting viewpoints, posing challenges in geometric perception and action generation. Current policies either rely on 2D features or sparse 3D representations that lack dense spatial structure, and typically encode arm and base within one action vector that ignores their distinct control demands. Moreover, existing dense fusion strategies risk corrupting pretrained representations under noisy depth while incurring heavy computational overhead. We present GeoHAT, an end-to-end diffusion-based framework built on a simple principle: geometry should be injected only where reliable and attended to only where needed. GeoHAT employs a lightweight Fourier spatial encoder that maps dense per-pixel 3D coordinates into geometric tokens without an additional 3D vision backbone. These tokens are then...

论文介绍 针对全身移动操作中几何感知和动作生成的挑战,现有方法依赖2D特征或稀疏3D表示,且将手臂与基座控制编码为一体。本文提出GeoHAT,一个端到端扩散框架,采用轻量级傅里叶空间编码器将密集3D坐标映射为几何令牌,并实现几何信息的自适应注入与关注,旨在提升移动操作的性能与效率,适用于服务机器人等领域。

Real-Time Execution with Autoregressive Policies

第一作者: Sangkyu Lee · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Real-time execution, enabled by asynchronous inference that ensures both smooth action trajectories and fast reactivity, is critical for realistic deployments of large-scale Vision-Language-Action models. However, recent work on real-time execution primarily focuses on variants of diffusion policies, even though it is more critical for autoregressive policies given their slower rollout speed in synchronous inference. In contrast, we demonstrate that autoregressive policies can achieve real-time execution by adjusting the tokenization horizon and applying constrained decoding, thereby guaranteeing strict latency bounds that enable multi-trajectory decoding to maximize performance. Across simulated and real-world environments, we find that the autoregressive policy consistently outperforms its equivalent-level flow-matching policy counterpart while achieving significantly...

论文介绍 实时执行对于大规模视觉语言动作模型的部署至关重要,但自回归策略在同步推理中推出速度较慢。本文演示通过调整分词视界和应用约束解码,自回归策略可实现严格延迟界限的实时执行,并支持多轨迹解码以最大化性能,从而在模拟和真实环境中优于流匹配策略,推动机器人控制等应用。

Low cost, easily manufactured, highly flexible strain and touch sensitive fiber for robotics applications

第一作者: Christian Diaz Herrera · 方向: 机器人操作 · 来源: cs.RO

Abstract:Existing stretch and touch sensors for robots are generally expensive with respect to at least one of material costs, required manufacturing equipment, or manufacturing time. We present and experimentally characterize a conductive fiber made using only inexpensive commercial off-the-shelf parts (conductive thread at $0.07/ft, silicone tubing at $0.94/ft) and tools (loop-style needle threader at $2), which can be manufactured quickly (20 cm length in 2 minutes.) We demonstrate its use as a resistive strain sensor with three applications: Triggering a grasp in a pneumatically actuated assistive finger, sensing the pose of a pneumatically actuated robotic strap, and estimating the pose of a flexible solid. We also demonstrate that it can be used as a capacitive sensor with two applications: First, as a touch sensor which triggers a commercial robot arm to move, and second, as a...

论文介绍 现有机器人传感器通常成本高昂。本文提出并实验表征一种基于导电线的低成本、易制造柔性纤维传感器,可作为电阻应变传感器和电容传感器使用。该传感器在机器人抓取触发、姿态估计等应用中展示潜力,有助于降低机器人系统的传感成本并提升灵活性。

EMG-Based Adaptation of Anisotropic Virtual Fixtures for Robot-Assisted Surgical Resection and Dissection

第一作者: Dario Onfiani · 方向: 具身智能 · 来源: cs.RO

Abstract:In this paper, we address the development of an adaptive assistance system for robot-assisted laparoscopic surgery, specifically for delicate tasks such as Resection and Dissection. Even if Virtual Fixtures offer significant advantages for guiding a surgeon's movements, conventional Virtual Fixtures are often defined by fixed geometries, lacking the flexibility to adapt to the surgical workflow or the surgeon's immediate intent. To address these limitations, we propose a novel framework for an adaptive and anisotropic virtual fixture. In addition, we introduce an intuitive control interface that modulates the fixture's geometry in real-time based on the surgeon's intent, inferred from EMG signals. This approach allows the surgeon to dynamically expand or disengage the constraint by contracting their forearm muscles, enabling seamless transitions between precise guided motion...

论文介绍 针对机器人辅助手术中传统虚拟夹具固定几何、缺乏适应性的问题,本文提出一个自适应各向异性虚拟夹具框架。通过肌电信号推断外科医生意图,实时调整夹具几何,从而在切除和解剖任务中实现无缝的精准引导与动态约束,提升手术灵活性和精确度。

Humor Style Drives Laughter, Topic Shapes Acceptability: Evaluating Bilingual Personal and Political Robot-Delivered AI Jokes

第一作者: Anna-Maria Velentza · 方向: 具身智能 · 来源: cs.RO

Abstract:Humor plays a central role in human social relationships, and recent advances in computational humor create new opportunities for integrating humor into human-robot interaction (HRI). While large language models (LLMs) can generate diverse forms of humor, it remains unclear how humor style, joke content, and language preference shape perceptions of robot-delivered humor in group settings. In this exploratory study, we employed a mixed factorial design in which participants evaluated AI-generated jokes delivered by a robot in a university classroom. We examined the effects of humor type (Affiliative, Self-Enhancing, Aggressive, Self-Defeating) and joke content (person-related vs. political) on perceived funniness and appropriateness, as well as preferred language. Results show that humor type significantly influences funniness, with Aggressive and Affiliative humor rated...

论文介绍 幽默在人机交互中扮演重要角色,但机器人传递幽默的感知机制尚不清晰。本研究通过实验评估AI生成的笑话,分析幽默风格、笑话内容和语言偏好对机器人传递幽默的趣味性和可接受性的影响,结果表明幽默类型显著影响趣味性,为改善机器人社交能力提供依据。

WT-UMI: Tactile-based Whole-Body Manipulation via Force-Supervised Contact-Aware Planning

第一作者: Jaehwi Jang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Whole-body humanoid manipulation of bulky, deformable, and shared-load objects requires distributed contact sensing and explicit force regulation, yet most imitation policies treat contact force only implicitly. On the other hand, different demonstration sources provide complementary modalities with inherent trade-offs: human demonstrations capture natural contact forces but not robot-executable actions, while teleoperation directly records robot actions but with less natural force regulation. This paper presents \textbf{WT-UMI}, a wearable whole-body tactile interface worn by human operators or mounted on humanoids, providing accurate observations of tactile images, contact forces, and end-effector poses across both human demonstration and humanoid teleoperation modes. We introduce a force-conditioned target-pose correction module that converts measured human poses into...

论文介绍 全身类人机器人操作需要分布式接触传感和力调节。本文提出WT-UMI,一个可穿戴全身触觉接口,可提供触觉图像、接触力和末端执行器姿态观测,并引入力条件目标姿态校正模块,旨在整合人类演示和机器人遥操作,提升对大物体、可变形物体的操作性能。

Proprioceptive-visual correspondence enables self-other distinction in humanoid robots

第一作者: Yurun Chen · 方向: 具身智能 · 来源: cs.RO

Abstract:Distinguishing self from others is a prerequisite for social intelligence, yet humanoid robots that increasingly share workspaces with humans still lack this ability. Here we show that a humanoid robot can learn self-other distinction from proprioceptive-visual correspondence, without any identity labels or kinematic models. Once established, this distinction bootstraps a predictive self-model that maps joint configurations to three-dimensional body occupancy, capturing how the robot's body changes with action. In multi-agent scenes involving humans or morphologically identical robots, the system reliably identifies itself, learns a 3D self-model, and supports downstream tasks including target reaching, collision-aware motion planning, and human-to-robot motion retargeting. Together, these results outline a route toward bodily self-representation in robots that act and...

论文介绍 区分自我与他人是社会智能的前提,但类人机器人缺乏此能力。本文演示机器人从本体感觉-视觉对应中学习自我-他人区分,无需身份标签,进而建立预测性3D自我模型,支持目标到达、碰撞感知运动规划等任务,为机器人身体自表征开辟途径。

Embedding ISO 10218 Safety Compliance in Robots via Control Barrier Functions for Human-Robot Collaboration

第一作者: Federico Parma · 方向: 导航与运动 · 来源: cs.RO

Abstract:Human-Robot Collaboration (HRC) requires strict adherence to safety standards, such as ISO 10218, to prevent harmful interactions. Standard Speed and Separation Monitoring (SSM) filters calculate safe robotic speeds based on conservative assumptions, such as constant human velocity, which prevents accurate predictions of minimum separation distances and causes unnecessary operational halts. This paper proposes a Control Barrier Function (CBF) that explicitly incorporates human acceleration data to analytically forward-predict the minimum human-robot separation distance during a worst-case robotic stopping trajectory. To guarantee safety at the control level, this predictive CBF is integrated as an inequality constraint within a Sequential Quadratic Programming (SQP) framework. Specifically, two methods are proposed: Method I, a CBF-constrained PD safety filter; and Method II...

论文介绍 人机协作需严格遵守安全标准如ISO 10218。传统速度与分离监控基于保守假设,导致不必要的操作中断。本文提出控制屏障函数,整合人类加速度数据以预测最小人机分离距离,并集成到序列二次规划框架中作为约束,从而在控制层面保证安全,提升协作效率。

Multi-Modal Multi-Agent Robotic Cognitive Alignment enabled by Non-Invasive Consumer Brain Computer Interfaces: A Proof of Concept Exploration

第一作者: Nataliya Kosmyna · 方向: 具身智能 · 来源: cs.RO

Abstract:While non-verbal behaviors and expressive movements are essential for natural human-robot interaction, existing methods often overlook a crucial element: the human's internal cognitive state. Frequently, proactive multi-agent systems can interrupt humans at inopportune moments, leading to cognitive overload and decreased task performance. This paper introduces a framework for generating "cognitively aligned" multi-agent interactions, enhancing the ability of robotic systems to contextually defer communications to the user of an agent system during moments of high human mental workload and engagement. We present the design and implementation of a closed-loop architecture that explores the interplay between autonomous task execution and real-time neurophysiological focus. Using a consumer-grade Brain-Computer Interface (BCI), our approach continuously monitors...

论文介绍 该研究提出了一种利用消费级脑机接口实现「认知对齐」的多智能体机器人交互框架。其核心在于通过实时监测用户的神经生理信号,感知其认知负荷与专注度,从而引导机器人系统在人类心智高负荷时延迟或调整通信,避免造成干扰与认知过载。该闭环系统探索了自主任务执行与人类实时注意力之间的协同关系。

FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation

第一作者: Chengbo Yuan · 方向: 机器人操作 · 来源: cs.RO

Abstract:Despite the success of vision-based generalist robotic policies, existing tactile-based policies remain tied to fixed embodiments and sensor setups. This is because tactile signals are highly heterogeneous across hardware, making cross-sensor generalization difficult. We present FTP-1,the first generalist foundation tactile policy pretrained to acquire transferable tactile manipulation abilities across diverse sensors and embodiments. FTP-1 supports varied tactile inputs, including image-, array-, and state-based signals, by using heterogeneous encoders to project them into unified morphology-aware latent tokens that are jointly modeled by a shared tactile Transformer expert. Pretrained on around 3,000 hours of tactile manipulation data aggregated from 26 data sources, spanning human and robot demonstrations across 21 sensors, FTP-1 learns tactile skills that transfer beyond...

论文介绍 本文提出了首个通用的基础触觉策略FTP-1,旨在解决触觉信号因硬件高度异构而导致的跨传感器泛化难题。该策略通过异质编码器将图像、阵列和状态等多种触觉输入映射到统一的潜在空间,并利用共享的触觉Transformer进行联合建模。其在约3000小时、涵盖21种传感器的多源数据上预训练,学习到的技能可跨传感器和具身体验迁移。

Y-BotFrame: An Extensible Embodied Agent Framework for Quadruped Robot Assistants

第一作者: Luyao Zhang · 方向: 导航与运动 · 来源: cs.RO

Abstract:Quadruped robots are capable of traversing a wide range of complex terrains with high flexibility. As highly mobile ground-based intelligent platforms, they can be equipped with modules for navigation control, environmental perception, and intelligent interaction, thereby serving as real-world mobile deployment platforms for various algorithms. In this paper, we introduce Y-BotFrame, an extensible embodied platform that turns a robot into an intelligent ground assistant. Y-BotFrame integrates multimodal perception capabilities, including speech, vision, and LiDAR, and employs a large language model as the cognitive core for environmental understanding, contextual reasoning, and task planning. The system maps user natural-language instructions into executable embodied task units that can be carried out by the robot. Y-BotFrame supports natural interaction through voice commands...

论文介绍 本文介绍了Y-BotFrame,一个可扩展的具身智能体框架,旨在将四足机器人转变为智能地面助手。该框架集成了语音、视觉和激光雷达等多模态感知能力,并以大语言模型作为认知核心,用于环境理解、上下文推理和任务规划。系统能将用户的自然语言指令映射为机器人可执行的具身任务单元,支持通过语音命令进行自然交互。

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

第一作者: Dayu Xia · 方向: 机器人操作 · 来源: cs.RO

Abstract:Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models to judge not only final task success, but also how a manipulation execution is physically and temporally progressing. However, existing evaluations fail to test whether VLMs possess fine-grained process understanding. To address this gap, we present RoboProcessBench, a benchmark for process-aware understanding in vision-language robotic manipulation. RoboProcessBench decomposes such capability into two complementary dimensions, \emph{static monitoring} and \emph{dynamic reasoning}, instantiated as 12 diagnostic question families covering phase, contact, motion, coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions. Built from physically grounded execution...

论文介绍 现有研究缺乏对视觉语言模型在机器人操作中细粒度过程理解能力的评估。为此,本文提出RoboProcessBench基准,将过程理解能力分解为「静态监控」和「动态推理」两个维度。该基准由12个诊断性问题族构成,覆盖阶段、接触、运动、协调、时间顺序等多个方面,基于物理模拟的操作执行数据构建,用于系统评估VLMs对操作过程的进展判断能力。

GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training

第一作者: Zhihai Bi · 方向: 导航与运动 · 来源: cs.RO

Abstract:Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated...

论文介绍 人形机器人与物体的交互面临动态平衡与稳定交互耦合的挑战。现有方法通常需要耗时的策略训练或依赖于僵化的轨迹重放。本文提出GenHOI框架,使人形机器人能够通过直接模仿单个生成视频,在零样本条件下执行多样化的物体交互任务,无需任务特定的训练或物理演示数据。该框架首先在仿真中重建场景,然后根据语言指令合成交互视频并从中提取策略。

Trajectory-Level Redirection Attacks on Vision-Language-Action Models

第一作者: Gokul Puthumanaillam · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) policies bring natural language into closed-loop robot control, enabling robots to execute manipulation tasks directly from text instructions. The same interface gives text a recurring role in control because the prompt is reused at every replanning step, and each prompt-conditioned action changes the future observations on which the policy acts. Existing VLA attacks study adversarial prompts that elicit targeted low-level actions or make such actions persist across changing images. We identify a stronger trajectory-level failure mode: a prompt that still $\textit{appears}$ to specify the intended task but redirects the final physical outcome. We mathematically formalize this setting as $\textit{command-preserving trajectory redirection}$, a prompt-only threat model in which the attacker chooses one prompt before the episode, all policy and...

论文介绍 本文针对视觉语言动作模型提出了轨迹级重定向攻击的新威胁。攻击者可注入一个在语义上仍看似正确、但实际上能改变最终物理执行结果的对抗性提示。研究将此形式化为「命令保持的轨迹重定向」问题,即提示在文本上指定了预期任务,却将机器人的实际轨迹导向不同的结局。这是一种仅针对提示的攻击,具有潜在的安全风险。

EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment

第一作者: Shihefeng Wang · 方向: 模仿学习 · 来源: cs.RO

Abstract:Scalable robot imitation learning relies on large-scale heterogeneous data from diverse robots or body-free data, making Cartesian end-effector actions a key interface for embodiment-agnostic policy learning. However, end-effector-only abstraction leaves Cartesian policies unaware of the deployed robot body, making them brittle under robot-specific constraints such as whole-body collision avoidance. To overcome this limitation, we present EmbodiSteer, a training-free framework that steers embodiment-agnostic visuomotor policies toward zero-shot, embodiment-aware deployment. EmbodiSteer keeps policy learning in Cartesian space while efficiently lifting inference-time diffusion sampling into the target robot's joint space via forward kinematics and Jacobian-based updates. With whole-body collision-aware guidance over joint trajectories after each denoising step, the arm can be...

论文介绍 为解决具身无关的视觉运动策略在部署到具体机器人时因缺乏具身感知而脆弱的问题,本文提出了EmbodiSteer框架。该框架无需训练,它在保持策略于笛卡尔空间学习的同时,通过运动学和雅可比矩阵在推理时将采样过程提升到目标机器人的关节空间,并引入全身碰撞感知的引导,使策略在零样本条件下实现具身感知的部署。

SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation

第一作者: Sunghwan Kim · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Long-horizon robot mobile manipulation requires continual reasoning about localization, environment changes, and task progress, all of which are challenging to infer from image observations alone. In this paper, we show that conditioning a mobile manipulation policy on a spatiotemporal feature map improves reasoning over long horizons. The map represents the environment and the articulated robot body as neural points in a shared latent space and is updated online from egocentric observations and proprioceptive state. We update the environment neural points using object-level rigid tracking and the robot neural points using forward kinematics. We use our spatiotemporal environment and robot feature (SERF) map as a state input to a vision-language-action (VLA) model by extracting map tokens from multiple reference frames and spatial scales, providing the policy with both local...

论文介绍 长时程移动操作需要对定位、环境变化和任务进度进行持续推理。本文证明,基于时空特征图条件化移动操作策略能够改善长时程推理能力。所提SERF地图将环境和机器人身体在共享潜在空间中表示为神经点,并通过物体级刚体跟踪和运动学进行在线更新。该地图以多尺度、多参考帧的图元形式作为视觉语言动作模型的状态输入,为策略提供全局与局部信息。

Towards Reliable Sequential Object Picking in Clutter: The Runner-up Solution to RGMC 2025

第一作者: Wei Yu · 方向: 机器人操作 · 来源: cs.RO

Abstract:As a long-standing challenge in robotic manipulation, stable and efficient grasping in cluttered environments is of great importance in industrial settings. While recent studies have achieved relatively high success rates in grasping from clutter, there remain few mature solutions for more demanding tasks such as sequential object search and sorting. This work addresses sequential object picking in cluttered environments based on the Cluttered Environment Picking Benchmark (CEPB) and presents our solution to the Pick-in-Clutter track of the 10th Robotic Grasping and Manipulation Competition (RGMC) at ICRA 2025. The task poses several key challenges. First, it requires robust and collision-aware grasping with high success rates across a diverse set of objects, including both rigid and deformable ones. Second, it demands efficient search for target objects, which places...

论文介绍 该研究针对杂乱环境中的顺序物体拾取挑战,提出了一个基于CEPB基准和ICRA 2025 RGMC竞赛的解决方案。它需要系统具备高成功率、能处理刚性和可变形物体的碰撞感知抓取能力,并能高效搜索目标物体。这项工作为工业分拣等复杂操作任务提供了技术参考。

An Embodied Simulation Platform, Benchmark, and Data-Efficient Augmentation Framework for Wet-Lab Robotics

第一作者: Zhe Liu · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Wet-lab robots can improve the reproducibility, throughput, and safety of biomedical experiments, but scaling their learning requires customizable simulators for safe and reproducible task generation, open editable laboratory assets, and efficient pipelines that turn limited demonstrations into usable training data. We present Pipette, an embodied simulation platform, benchmark, and data-efficient augmentation framework for wet-lab robot learning. Pipette releases over 43 open-source and re-editable wet-lab assets, together with an extensible asset-building pipeline. A key component of Pipette is its simulation-based data augmentation pipeline, replaying human demonstrations in simulation, applies lighting, camera, speed, and action perturbations, and filters generated episodes with automatic task success checks, rapidly expanding usable training data from limited manual...

论文介绍 该论文介绍了Pipette平台,旨在解决湿实验室机器人学习中可定制仿真器、开放资产和高效数据增强流程的需求。它发布了可编辑的实验室资产,并提出了一个仿真数据增强流程,通过重播人类演示并施加扰动,从有限演示中快速生成大量可用的训练数据,以加速机器人学习。

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

第一作者: Allison Andreyev · 方向: 机器人操作 · 来源: cs.RO

Abstract:For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time. Although Vision-Language Models (VLMs) have enabled zero-shot generalization in robot task and motion planning (TAMP), current state-of-the-art approaches often remain computationally "heavyweight" or require extensive training on thousands of demonstrations. We present GRASP (Grounded Reasoning and Symbolic Planning), a framework designed as a step toward open-vocabulary tabletop manipulation. Our approach leverages a pretrained VLM to translate natural-language queries into neuro-symbolic goal states, grounded in the physical world via a bounding-box detection pipeline. Unlike methods that rely on fixed color lists or hard-coded coordinates, GRASP enables robots to interpret abstract spatial concepts such as "top shelf" and...

论文介绍 为使机器人能实时响应自然语言,本文提出GRASP框架。它利用预训练视觉语言模型将自然语言查询转化为基于边界框检测的神经符号目标状态,实现了开放词汇的桌面操作。与依赖固定列表的方法不同,GRASP能理解「顶层架子」等抽象空间概念,降低了对大量训练数据的依赖。

AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

第一作者: Jianli Sun · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-level Unmanned Aerial Vehicle (UAV) movement and end-effector-level arm manipulation differ substantially in action scale, dynamics, and control objectives. In this paper, we propose AIR-VLA+, a flow matching action generation architecture specifically designed for aerial manipulation, featuring cascaded dual-action decoders and an asymmetric feature-level Mixture of Experts (MoE). We construct cascaded manipulation and movement decoders, allowing the UAV to unidirectionally observe the manipulator's intent during movement to achieve workflow coordination, while isolating the impact of UAV movement information backpropagation on arm manipulation stability. Addressing the characteristic that UAV movement is highly dependent on high-level semantics and responsible for...

论文介绍 针对空中操作中平台运动与末端执行器操作存在表示耦合的问题,本文提出AIR-VLA+架构。它采用级联的双动作解码器和非对称混合专家模块,使无人机运动能单向观察机械臂意图以实现协调,并隔离运动信息对操作稳定性的反向传播影响。该架构专门为空中操作的高语义依赖特性而设计。

Sparse2Act: Learning Action-Aligned Sparse 3D Representations for Cross-Domain Robot Manipulation

第一作者: Yu Guo · 方向: 机器人操作 · 来源: cs.RO

Abstract:Explicit 3D representations are attractive for manipulation because they expose object shape, workspace geometry, and robot-object relations in metric coordinates. However, sparse 3D encoders are often learned through downstream task objectives, tying the representation to a particular data distribution, policy architecture, and action parameterization. We introduce Sparse2Act, an observation-action alignment framework for pretraining sparse point-cloud encoders. The key idea is to use task-space end-effector actions as geometric supervision: masked sparse 3D tokens are trained to organize scene features around the workspace motion paired with the observation. After pretraining, only the encoder initialization is reused by downstream policies, allowing them to retain their own architectures and action spaces, including joint-space commands. On the LIBERO-10 benchmark, our...

论文介绍 本文提出Sparse2Act框架,用于预训练与具体任务解绑的稀疏点云编码器。其核心思想是利用末端执行器动作作为几何监督,训练掩码的3D令牌围绕工作空间运动组织场景特征。预训练后仅复用编码器初始化,下游策略可保留各自架构和动作空间,这增强了跨域操作的通用性。

EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence

第一作者: Xin Zhou · 方向: 模仿学习 · 来源: cs.RO

Abstract:In this paper, we propose the Enhanced World Action Model (EWAM), a closed-loop online adaptation architecture built upon a pretrained and fully frozen Cosmos3 backbone network. Evaluated entirely under a zero-shot task protocol, EWAM is centrally focused on reducing the amount of additional deployment data required to adapt to new task layouts. Notably, no extra task-specific demonstration sets were introduced in any of the evaluations, and no fine-tuning was performed on the backbone network. Its performance gains stem entirely from an inference-time co-reasoning mechanism composed of four inserted lightweight neural layers: the Neural Experience Memory Layer located in the intermediate layers of the Diffusion Transformer (DiT) provides task-relevant execution context; the Neural Anomaly Detection Layer after the state prediction head monitors the divergence between...

论文介绍 本文提出增强型世界动作模型EWAM,这是一个基于冻结的Cosmos3骨干网络构建的闭环在线适应架构。在完全零样本任务协议下评估,其重点是减少适应新任务布局所需的部署数据量。性能提升完全来自推理时的共推理机制,包括提供执行上下文的神经经验记忆层和监测预测异常的检测层。

DARRMS -- An Efficient Algorithm for Dynamic Attention Radius in Resource-Constrained Multi-Agent Systems

第一作者: Benjamin Alcorn · 方向: 具身智能 · 来源: cs.RO

Abstract:Multi-agent systems are integral tools for various domains such as robotics, cybersecurity, and autonomous vehicle planning. These types of systems often have constraints on the computational resources, leading to a need for efficient lightweight algorithms. Traditional decision making frameworks often assume ideal conditions, such as full observability and unlimited computational capacity, which do not align with real-world challenges. In this paper, we introduce a new algorithm that allows for reduced demand on computational resources without a large cost of other performance metrics. Agents will limit their observability to some attention radius, which intentionally allows them to ignore parts of the environment that might be unnecessary for action planning. By optimizing both the attention radius and decision-making, our approach enhances coordination and scalability in...

论文介绍 针对资源受限多智能体系统中传统决策框架假设理想条件的问题,本文提出DARRMS算法。该算法允许智能体限制其观察范围至某个注意力半径,从而忽略环境中可能不必要的部分。通过优化注意力半径和决策过程,该方法在不过多牺牲其他性能指标的前提下,提升了协调性和可扩展性。

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

第一作者: Yangcen Liu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but directly using them for robot learning requires bridging two gaps: the visual gap between human and robot observations, and the action gap between human motion and robot-executable action. We propose EgoEngine, a scalable framework for transforming egocentric human manipulation videos into high-fidelity robot data. Given an egocentric RGB video, EgoEngine produces: (i) a high-fidelity robot observation video replacing human with robot while preserving scene context and temporal alignment, and (ii) a task-aligned, executable robot action trajectory under feasibility constraints. Experiments in simulation and on real robots show that EgoEngine enables scalable conversion of human videos into robot...

论文介绍 为降低收集大规模机器人灵巧操作演示的成本,本文提出EgoEngine框架,用于将第一人称人类视频转换为高保真机器人数据。该框架能生成保留场景上下文的机器人观测视频,并产生符合可行性约束的可执行机器人动作轨迹。实验表明,EgoEngine能够可扩展地将人类视频转化为可用于仿真和真实机器人学习的数据。

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation

第一作者: Honglin He · 方向: 导航与运动 · 来源: cs.RO

Abstract:Autonomous long-horizon sidewalk navigation is essential for micro-mobility applications such as robotic food delivery and assistive electronic wheelchairs. Unlike autonomous driving on the road, long-horizon sidewalk navigation requires precise maneuvering through unpredictable sidewalk terrains and pedestrians, with a lightweight perception stack as minimal as a single monocular RGB camera. While imitation learning (IL) from demonstrations offers a practical solution, the resulting autopilot policy often suffers from compounding errors, a lack of social compliance on sidewalks, and deficiencies in counterfactual reasoning to handle complex situations. To address these challenges, we introduce FlowPilot, a mapless navigation policy that achieves robust and efficient long-horizon navigation performance using only a monocular RGB camera. We first propose to use anchored flow...

论文介绍 针对微移动机器人长期人行道导航中的误差累积与社交合规性问题,本文提出了一个无地图导航策略FlowPilot。该系统仅使用单目RGB相机,通过提出锚定流估计方法来学习人类偏好策略,以实现鲁棒、高效的长距离导航。该研究为机器人送餐和辅助轮椅等应用提供了一种基于轻量感知的实用解决方案。

G-MAPP: GPU-accelerated Multi-Agent Planning and Perception for Reactive Motion Generation

第一作者: Tanmay Bishnoi · 方向: 导航与运动 · 来源: cs.RO

Abstract:Reactive motion generation in unstructured environments remains an open challenge in robotics. Due to the computational complexity of collision-free motion generation, existing methods either generate global trajectories for static scenarios, or employ models that make conservative assumptions about the environment. This paper identifies the primary bottleneck as the runtime performance demand of planning on high-fidelity environments, and the temporal integration between the perception and planning modules. Therefore, we propose a framework that does not compromise on runtime performance and world representations for perception and planning by accelerating world modeling and vector-field based planning using the GPU. This allows us to achieve faster parallel state exploration for quasi-global trajectory planning, and tighter coupling of the perception-action loop in real-time...

论文介绍 在无结构环境中实现反应式运动生成面临计算瓶颈,本文提出G-MAPP框架。该框架利用GPU加速世界建模和基于向量场的规划,实现了快速并行状态探索的准全局轨迹规划,并强化了感知与规划模块的实时耦合。这项工作提升了机器人在复杂动态环境中进行实时、高保真运动规划与感知的能力。

Action-Effect Memory Pretraining for Robot Manipulation

第一作者: Yijing Zhou · 方向: 机器人操作 · 来源: cs.RO

Abstract:We present AEM, an Action-Effect Memory pretraining framework for robot manipulation that learns compact temporal representations from vision-action history. Unlike prior robot representation pretraining methods that mainly focus on single-frame visual encoding, AEM targets the temporal nature of manipulation, where the current observation alone is often insufficient under partial observability. AEM models manipulation as an action-driven interaction process by interleaving visual and action features and applying masked modeling to recover missing content from incomplete histories, thereby learning action-conditioned state evolution. The Mamba-encoded output of the final vision token is used as a compact history representation, serving as the global context for decoding and downstream control. This design preserves a single-vector temporal bottleneck while keeping inference...

论文介绍 针对机器人操作中因部分可观测性导致当前观测信息不足的问题,本文提出了动作-效应记忆预训练框架AEM。该框架通过交错视觉与动作特征,并应用遮蔽建模来恢复不完整历史,从而学习动作条件下的状态演化。AEM输出的紧凑历史表示可作为全局上下文,用于下游解码与控制,以提升操作任务的性能。

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

第一作者: Baochang Ren · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach. AI can help read literature, generate hypotheses, and plan protocols, yet the execution of those protocols at the bench still requires a human operator. Vision-Language-Action (VLA) models provide one possible interface between written protocols and robot execution, but existing policies are trained mostly on household and tabletop demonstrations and rarely encounter the instruments, transparent liquids, or fixed protocol workflows found in scientific laboratories. Closing this gap requires both laboratory-specific supervision and a unified learning framework that can accommodate the diverse robot embodiments used to execute experimental protocols. We therefore identify data and embodiment as central bottlenecks...

论文介绍 本文探讨将视觉-语言-动作模型应用于科学实验室自动化。研究指出,现有VLA模型主要基于家居演示训练,缺乏实验室特定的仪器、液体和流程数据,且难以适配多样化的机器人构型。因此,本文致力于通过实验室特定数据和统一学习框架来填补这一缺口,使VLA模型能够执行实验操作规程。

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

第一作者: Hanyang Yu · 方向: 多模态具身 · 来源: cs.RO

Abstract:World Action Models (WAMs) present a promising paradigm for robotic control via video prediction. However, current WAMs suffer from fundamental spatial bottlenecks: standard text inputs introduce referential ambiguity in cluttered scenes, while unstructured RGB predictions lack semantic grounding and remain biased by task-irrelevant backgrounds. To overcome these limitations, we introduce MaskWAM, an object-centric world-action model. By jointly integrating masks as both explicit inputs and predictions via a unified Mixture of Transformers (MoT), MaskWAM unlocks robust policy generalization. This design provides two key benefits: (1) predicting future masks yields object-centric semantic supervision that suppresses visual noise, significantly enhancing even standard text-conditioned WAMs; and (2) coupling this predictive supervision with first-frame visual prompts, such as...

论文介绍 当前基于视频预测的世界-动作模型面临文本指代歧义和背景偏差等空间瓶颈。为解决此问题,本文提出MaskWAM,一种以物体为中心的世界-动作模型。它通过统一的混合Transformer架构,将掩码同时作为输入提示和预测目标,从而抑制视觉噪声,显著增强策略的泛化能力,为机器人控制提供了更鲁棒的解决方案。

$μ$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models

第一作者: Egor Cherepanov · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) models predict chunks of future actions from the current observation, an assumption that fails under partial observability, where decisions depend on information no longer visible. Existing memory-augmented VLAs simultaneously introduce recurrence, retrieval, compression modules, auxiliary objectives, hierarchical memory, or task-specific architectural changes, so the contribution of recurrence itself remains entangled with surrounding machinery. We present a controlled isolation study of recurrence in a strong pretrained VLA backbone. Our formulation augments the transformer with a small set of learnable memory tokens carried across timesteps and updated through self-attention, trained end to end with truncated backpropagation through time, with no auxiliary losses and no architectural changes. We instantiate this as $\mu$VLA, a family of...

论文介绍 在部分可观测环境中,VLA模型需要依赖历史信息。本文对预训练VLA骨干中的循环记忆进行了控制性隔离研究。方法是将一组可学习的记忆令牌在时间步之间传递,并通过自注意力更新,采用截断时间反向传播进行端到端训练。这种简洁设计(命名为μVLA)验证了循环记忆本身对处理部分可观测性问题的有效性。

VLADriveBench: Evaluating CoT-Action Relationship in VLA for Autonomous Driving

第一作者: Thach Nguyen · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Vision-language-action (VLA) models generate chain-of-thought (CoT) reasoning alongside driving trajectories, but existing benchmarks evaluate only trajectory quality and do not assess whether the CoT is relevant, consistent, or causally connected to the driving action. We introduce VLADriveBench, a framework that combines observational metrics (mentioning, hallucination, contradiction, action alignment) with a CoT intervention protocol to provide complementary views of the CoT-action relationship. Applying VLADriveBench to three models across two architectures, we find that the two analyses can diverge sharply: ORION scores highest on observational alignment yet its CoT is epiphenomenal, while Alpamayo v1.5 scores lower yet its CoT is strongly causal, with visual salience gating the extent of CoT influence.

论文介绍 现有自动驾驶中视觉-语言-动作模型生成思维链推理和驾驶轨迹,但基准评估仅关注轨迹质量,未检验思维链的相关性、一致性或因果连接。本文提出VLADriveBench框架,结合观察指标(如提及、幻觉、矛盾、动作对齐)和思维链干预协议,提供互补视图评估思维链-动作关系。应用于两个架构的三个模型,发现观察分析与因果分析可能显著分歧,例如ORION观察对齐高但思维链附带,而Alpamayo v1.5因果性强且视觉显著性影响思维链作用。

Rarity-Gated Context Conditioning for Offline Imitation Learning-Based Maritime Anomaly Detection

第一作者: Yongmin Kim · 方向: 模仿学习 · 来源: cs.LG

Abstract:Contextual anomaly detection aims to identify abnormal behavior conditional on context variables, but practical deployments often face highly imbalanced context distributions where rare regimes can be critical information. Under such frequency bias, context-conditioned models can produce unstable decisions and excessive false alarms in rare contexts. We propose Rarity-Gated Feature-wise Linear Modulation (RGFiLM), a rarity-aware conditioning module that combines feature-wise modulation (i.e., context-conditioned scaling and shifting of hidden features) with a gate controlled by a data-driven rarity score. The rarity score is estimated from the empirical distribution of context variables and regulates how strongly context modulates intermediate representations: the gate becomes more decisive under rare contexts while remaining conservative under frequent contexts. We evaluate...

论文介绍 本文研究基于离线模仿学习的上下文条件异常检测。针对数据中上下文分布高度不平衡、稀有情境可能关键但导致模型决策不稳的问题,提出了稀有性门控特征调制模块RGFiLM。该模块利用数据驱动的稀有性分数来调节上下文对隐藏特征的影响强度,从而在稀有上下文中做出更果断的判断,减少误报。

PersonaDrive: Human-Style Retrieval-Augmented VLA Agents for Closed-Loop Driving Simulation

第一作者: Mahmoud Srewa · 方向: VLA 通用模型 · 来源: cs.AI

Abstract:Closed-loop driving simulators typically populate their environments with non-ego traffic agents that behave largely the same way, produced either by rule-based traffic managers or by learned models trained toward a single behavioral mode. Recent work introduces style variation through post-hoc labels on observational data or LLM-inferred reward weights, but these signals act as proxies for what a style should reward rather than demonstrations of humans explicitly asked to drive in that style. We introduce PersonaDrive, a pipeline that conditions a vision-language-action (VLA) driving agent on retrieved demonstrations from a style-instructed human driving dataset, in which participants drive CARLA leaderboard routes under aggressive, neutral, and conservative instructions on a driver-in-the-loop rig. The pipeline has three stages: (i) offline triplet mining over per-style...

论文介绍 该研究针对闭环驾驶模拟器中非自我交通代理行为单一、缺乏人类风格多样性的问题,提出PersonaDrive方法。它通过检索增强的视觉-语言-动作(VLA)代理,基于风格指导的人类驾驶数据集中的示范,使代理能表现出激进、中性和保守等不同驾驶风格。该管道包括离线三元组挖掘等阶段,有望提升模拟环境的真实感和多样性,应用于自动驾驶系统的测试与开发。

市场总览

美股技术面呈现分化,SPY 和 QQQ 保持多头排列,趋势 bullish,价格接近52周高点,RSI 处于正常区域,但个股如 MSFT 和 META 转入空头排列,RSI 降至 37.1 和 34.7,显示板块内部分化。加密市场恐慌情绪浓厚,恐慌贪婪指数为 18,处于极度恐慌,总市值 2.28 万亿美元,BTC 主导率 56.6%,BTC 和 ETH 趋势 bearish,价格低于关键均线,尽管 BTC 出现 MACD 金叉,但 MACD 值为负,动量仍弱。中概股集体承压,BABA RSI 29.6 超卖,PDD 和 JD 空头排列,下行趋势明确。商品与外汇市场中,黄金和原油期货中性,美元指数 DXY 多头排列,RSI 58.7,接近52周高点,而人民币汇率 USDCNY=X 接近52周低,MACD 金叉但趋势 bearish,信号矛盾。整体市场技术面多空交织。

今日关注

BABA 阿里巴巴 (BABA)
偏下行

阿里巴巴当前价格 112.82 美元,RSI 14 读数为 29.6,处于超卖区域,趋势 bearish,MACD 值为 -4.9544,信号显示空头排列,表明下行技术压力显著,动量偏弱。

DX-Y.NYB 美元指数 DXY
偏上行

美元指数当前价格 99.81,RSI 14 为 58.7,接近52周高点,趋势 bullish,MACD 为正 0.3072,信号包括接近52周高和多头排列,显示上行动能持续,技术面支持偏上行倾向。

BTC-USD Bitcoin
中性

比特币价格 64619.7 美元,MACD 出现金叉但值为负 -3525.3567,趋势 bearish,RSI 14 为 37.4,信号有 MACD 金叉和空头排列,技术指标相互矛盾,呈现中性格局,动量不确定。

全部资产

^VIX

VIX 恐慌指数

$17.68 -9.05%
5 日
-17.81%
距 52w 高
-49.9%
RSI(14)
48.6
趋势
空头
SMA 20 / 50 / 200
17.53 / 18.25 / 18.54
MACD / 信号
0.309 / -0.098
死叉(SMA50↓SMA200) (2 天前)空头排列

^TNX

10Y 美债收益率 (%)

$4.49 +0.54%
5 日
-1.08%
距 52w 高
-10.2%
RSI(14)
50.8
趋势
多头
SMA 20 / 50 / 200
4.52 / 4.42 / 4.21
MACD / 信号
0.020 / 0.029
接近 52 周低多头排列

DX-Y.NYB

美元指数 DXY

$99.81 -0.05%
5 日
-0.26%
距 52w 高
-0.8%
RSI(14)
58.7
趋势
多头
SMA 20 / 50 / 200
99.42 / 98.90 / 98.64
MACD / 信号
0.307 / 0.261
接近 52 周高多头排列

SPY

S&P 500 ETF

$741.75 +0.54%
5 日
+0.57%
距 52w 高
-2.5%
RSI(14)
52.9
趋势
多头
SMA 20 / 50 / 200
745.07 / 722.80 / 686.30
MACD / 信号
3.772 / 7.393
接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$721.34 +0.59%
5 日
+2.31%
距 52w 高
-3.6%
RSI(14)
55.2
趋势
多头
SMA 20 / 50 / 200
721.50 / 681.81 / 625.38
MACD / 信号
8.315 / 13.604
多头排列

AAPL

Apple

$291.13 -1.52%
5 日
-5.27%
距 52w 高
-8.3%
RSI(14)
44.0
趋势
多头
SMA 20 / 50 / 200
303.88 / 285.49 / 266.87
MACD / 信号
2.261 / 5.760
多头排列

MSFT

Microsoft

$390.74 +0.10%
5 日
-6.22%
距 52w 高
-29.7%
RSI(14)
37.1
趋势
空头
SMA 20 / 50 / 200
419.75 / 411.80 / 453.73
MACD / 信号
-4.135 / 1.271
空头排列

NVDA

Nvidia

$205.19 +0.16%
5 日
+0.04%
距 52w 高
-13.3%
RSI(14)
45.2
趋势
中性
SMA 20 / 50 / 200
214.62 / 206.91 / 189.26
MACD / 信号
-1.101 / 1.257

GOOGL

Alphabet

$359.68 +0.53%
5 日
-2.40%
距 52w 高
-12.0%
RSI(14)
42.4
趋势
中性
SMA 20 / 50 / 200
376.42 / 362.26 / 307.94
MACD / 信号
-2.903 / 1.104

TSLA

Tesla

$406.43 +1.82%
5 日
+3.95%
距 52w 高
-18.5%
RSI(14)
48.9
趋势
中性
SMA 20 / 50 / 200
415.74 / 398.30 / 415.69
MACD / 信号
-2.173 / 2.196

META

Meta

$566.98 -0.26%
5 日
-4.39%
距 52w 高
-28.8%
RSI(14)
34.7
趋势
空头
SMA 20 / 50 / 200
604.21 / 621.83 / 658.09
MACD / 信号
-12.809 / -7.969
空头排列
加密恐慌贪婪
18
极度恐慌
加密总市值
$2.28 T
+0.86% / 24h
BTC 主导率
56.6%
ETH 8.9%
24h 成交量
$49.2 B
活跃币 17,463

BTC-USD

Bitcoin

$64,619.70 +1.69%
5 日
+2.42%
距 52w 高
-48.8%
RSI(14)
37.4
趋势
空头
SMA 20 / 50 / 200
67,523.53 / 74,162.01 / 77,770.88
MACD / 信号
-3,525.357 / -3,561.836
MACD 金叉 (今天)空头排列

ETH-USD

Ethereum

$1,684.30 +1.15%
5 日
-0.35%
距 52w 高
-66.0%
RSI(14)
33.2
趋势
空头
SMA 20 / 50 / 200
1,824.79 / 2,080.61 / 2,414.43
MACD / 信号
-131.828 / -130.750
空头排列

SOL-USD

Solana

$68.84 +3.13%
5 日
+3.06%
距 52w 高
-72.8%
RSI(14)
39.3
趋势
空头
SMA 20 / 50 / 200
73.23 / 81.75 / 100.04
MACD / 信号
-5.114 / -5.028
空头排列

BABA

阿里巴巴 (BABA)

$112.82 +0.12%
5 日
-6.81%
距 52w 高
-41.4%
RSI(14)
29.6
趋势
空头
SMA 20 / 50 / 200
125.81 / 130.33 / 149.57
MACD / 信号
-4.954 / -3.354
RSI 超卖空头排列

PDD

拼多多 (PDD)

$81.56 +0.32%
5 日
-4.13%
距 52w 高
-41.5%
RSI(14)
32.4
趋势
空头
SMA 20 / 50 / 200
88.52 / 95.37 / 111.40
MACD / 信号
-4.262 / -3.794
空头排列

JD

京东 (JD)

$28.56 +1.78%
5 日
-1.11%
距 52w 高
-22.5%
RSI(14)
41.6
趋势
空头
SMA 20 / 50 / 200
29.87 / 30.11 / 30.31
MACD / 信号
-0.573 / -0.378
空头排列

0700.HK

腾讯控股 (0700.HK)

HK$463.60 +1.40%
5 日
+2.29%
距 52w 高
-32.1%
RSI(14)
51.6
趋势
空头
SMA 20 / 50 / 200
450.45 / 472.52 / 568.94
MACD / 信号
-3.291 / -6.866
空头排列

GC=F

黄金期货

$4,238.80 +3.63%
5 日
-2.27%
距 52w 高
-24.1%
RSI(14)
35.4
趋势
中性
SMA 20 / 50 / 200
4,421.88 / 4,587.76 / 4,419.51
MACD / 信号
-118.340 / -87.662

CL=F

WTI 原油期货

$84.88 -3.23%
5 日
-6.25%
距 52w 高
-29.0%
RSI(14)
38.0
趋势
中性
SMA 20 / 50 / 200
93.98 / 96.74 / 73.41
MACD / 信号
-2.679 / -1.867

USDCNY=X

美元 / 人民币

¥6.76 -0.20%
5 日
-0.05%
距 52w 高
-6.2%
RSI(14)
35.4
趋势
空头
SMA 20 / 50 / 200
6.78 / 6.81 / 6.97
MACD / 信号
-0.012 / -0.014
MACD 金叉 (3 天前)接近 52 周低空头排列
风险提示

本报告基于公开行情数据计算的技术指标,过去走势不代表未来表现。所有分析仅供技术指标解读参考,不构成任何投资建议,读者应独立评估风险。

Australia news live: Jonno Duniam to retire from politics; father and daughter found dead in Sydney river

The Tasmanian Liberal senator has announced he will leave parliament before the end of the year. Follow the day’s news live Get our breaking news email, free app or daily news podcast The Senate will deliver its report from the NDIS inquiry on Tuesday. Butler doesn’t directly answer a question about

中文摘要 塔斯马尼亚自由党参议员乔诺·邓纳姆宣布将在年底前离开议会,这对陷入困境的联盟党是一个打击。同时,悉尼河中发现父子死亡事件。

Iran War Live Updates: Trump Says Peace Deal Will Be Signed Sunday, but Iran Disputes Timeline

An Iranian foreign ministry official sought to temper expectations, saying there were no plans for a Sunday signing and an agreement could be inked in the coming days.

中文摘要 特朗普表示伊朗和平协议将于周日签署,但伊朗外交部官员否认有周日签署计划,称协议可能在几天内签署,双方在时间线上存在分歧。

Liberal frontbencher Jonno Duniam to quit politics in blow to struggling Coalition

Tasmanian senator says it was ‘extremely difficult decision to make’ and he will leave ‘proud and grateful but exhausted’ Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Liberal frontbencher Jonno Duniam will quit politics before the

中文摘要 自由党前座议员、塔斯马尼亚参议员乔诺·邓纳姆宣布退出政坛,称这是极其艰难的决定,他对联盟党是一个打击。

Haiti fans in the streets, Scotland faithful in kilts

It’s expected to be a spicy encounter: Haiti supporters and Scotland’s Tartan Army celebrate World Cup matchup.

中文摘要 海地和苏格兰球迷在街头庆祝即将到来的世界杯小组赛,预期比赛将激烈,海地支持者和苏格兰“格子军”热情高涨。

Iran war live: Trump says deal to be signed today; Tehran urges caution

US and Iran appear close to signing the first stage of a peace deal, the two sides differ on when it will be signed.

中文摘要 特朗普称伊朗和平协议将于今天签署,但伊朗方面敦促谨慎,双方在签署时间上存在分歧,美国和伊朗接近签署第一阶段协议。

Brazil and Morocco kickoff amid electric atmosphere

Brazil and Morocco supporters sang, danced and chanted their way to one of the most anticipated group-stage matches.

中文摘要 巴西和摩洛哥支持者在热烈气氛中迎接世界杯小组赛,比赛是小组赛中最受期待的之一,球迷唱歌跳舞庆祝。

‘There was a lot of blood in the water’: paddleboarder rescues woman after ‘shocking’ Coogee shark attack

Charlie Verco managed to grab hold of the woman and bring her back to shore after the Sydney shark attack on Saturday Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Elite paddleboarder Charlie Verco has only seen one shark bigger th

中文摘要 在悉尼库吉的鲨鱼袭击事件中,桨板手查理·维科成功将一名受伤女子救回岸上,目击者称水中血迹斑斑,事件引发关注。

Top Haitian Security Official Kidnapped

A security expert who had recently become chief of staff to the new defense minister was abducted, the latest example of violence gripping the country.

中文摘要 海地一名高级安全官员被绑架,该官员最近刚成为新任国防部长的参谋长,这是该国暴力事件的最新例子。

Ticketmaster says Knicks fans won't be locked out of game after last-minute panic

An note on the Ticketmaster website for buying tickets to the game caused confusion and backlash among New Yorkers.

中文摘要 Ticketmaster网站上的购票说明引起纽约尼克斯球迷的困惑和反对,但公司澄清球迷不会被排除在比赛之外,回应了最后时刻的恐慌。

5 Children Are Killed After Van and S.U.V. Collide in Rural Ontario

The children — four girls and one boy — were among 10 people in the van at the time of the crash. An infant was seriously injured, the police said.

中文摘要 在加拿大安大略省农村,一辆面包车与SUV碰撞,导致5名儿童死亡,其中四女一男,一名婴儿重伤。

A Lebanon town's grief in the aftermath of a deadly Israeli airstrike

More than 3,700 people in Lebanon have died in the war between Israel and Hezbollah. In a village in southern Lebanon, one airstrike last month killed 14 people, including 10 women and children.

中文摘要 在以色列与真主党的战争中,黎巴嫩已有超过3700人死亡。南部一个村庄上月遭遇以色列空袭,造成14人死亡,其中包括10名妇女和儿童。

World’s Hottest Stock Market Turns Attention to MSCI Moment

After one of its most volatile weeks in years, South Korea’s stock market is approaching a milestone it has long been chasing: a potential path into MSCI Inc.’s developed-market status.

中文摘要 韩国股市在经历多年最波动的一周后,正接近长期追求的里程碑,即可能获得MSCI发达市场地位。

Fed and BOE Stay Guarded After 100 Days of Iran War

For several global central banks, the question of whether the Iran war poses more of an immediate danger to inflation or to growth is likely to remain open in the coming week.

中文摘要 在伊朗战争持续100天后,美联储和英国央行保持谨慎态度,对通胀与增长的风险评估在下周可能仍不明确。

CLO ETFs Boom on Higher Rates, Private Debt Woes

Wall Street has an answer for retail investors seeking to profit from elevated interest rates and dodge defaults in private credit: funds that buy collateralized loan obligations.

中文摘要 高利率和私人债务问题推动CLO ETFs繁荣,华尔街为散户投资者提供相关基金以获利并规避风险。

Bloomberg This Weekend 6/13/2026

The news doesn’t stop when markets close. Hosts David Gura, Christina Ruffini and Lisa Mateo bring clarity, context and a bit of humor to the weekend’s biggest headlines, LIVE from New York. Joined by Swarovski CEO Alexis Nasard, Johns Hopkins Center for Health Security Senior Scholar Dr. Amesh Adal

中文摘要 彭博周末节目由David Gura等主持人主持,邀请Swarovski CEO Alexis Nasard和Johns Hopkins健康安全中心专家,从纽约直播讨论周末头条新闻。

Caffeine Minimalists Rewrite Routines to Battle Coffee Jitters

While plenty of Americans are emphatically unready to give up caffeine, many are experimenting with a new range of options beyond the traditional cup of hot java, paying heed to caffeine’s impact on their sleep, mood and energy level. Consumers are becoming more cognizant of “energy management” in t

中文摘要 许多美国消费者尝试超越传统咖啡的新选择,以应对咖啡因对睡眠、情绪和能量水平的影响,成为“咖啡因极简主义者”。

Swarovski CEO's focus on "modern luxury" to return the 131-year-old company to profitability

Swarovski CEO Alexis Nasard joined Bloomberg This Weekend's Christina Ruffini for a tour and discussion about the significant brand transformation the company is undergoing as it aims to modernize its luxury image by partnering with celebrities like Ariana Grande and Venus Williams, as well as incor

中文摘要 Swarovski CEO Alexis Nasard致力于“现代奢华”战略,通过与Ariana Grande等名人合作进行品牌转型,以使这家131年历史的公司重返盈利。

US and Iran Move Closer to Deal Despite Hormuz Skirmishes

Pakistan said an interim US-Iran deal to reopen the Strait of Hormuz could be finalized within 24 hours, raising expectations that the sides may be nearing a broader agreement after recent skirmishes near the strategic waterway. Bloomberg News State Department Reporter Eric Martin and Jerusalem Repo

中文摘要 巴基斯坦表示,美伊就重开霍尔木兹海峡的临时协议可能在24小时内敲定,双方在近期冲突后可能接近达成更广泛协议。

As World Cup Begins, Health Officials Issue Warnings Amid Measles Outbreak

The FIFA World Cup is expected to create a "hospitable environment" for pathogens, with millions of people gathering in packed stadiums, and there are concerns about the spread of diseases such as Ebola, dengue, and measles. Johns Hopkins Center for Health Security Senior Scholar Dr. Amesh Adalja jo

中文摘要 FIFA世界杯期间,大量人群聚集可能助长病原体传播,卫生官员警告麻疹等疾病爆发风险。

SpaceX Shares Close 19% Higher After Historic $75 Billion IPO

SpaceX’s first day on the stock market transformed the startup into one of the world’s most-valuable public companies, handed buyers of the IPO a 19% return and turned its founder Elon Musk into the world’s first trillionaire. Bloomberg Tech Co-Host Ed Ludlow joined David Gura and Christina Ruffini

中文摘要 SpaceX完成750亿美元历史性IPO,上市首日股价上涨19%,使创始人埃隆·马斯克成为全球首位万亿富翁。

How Kalshi’s Lopes Lara Turned Kylie Jenner Gossip Into Billions

The firm’s co-founder talks about the celebrity buzz that sparked her interest in prediction markets, the legal battles her firm faces and Jay-Z’s advice.

中文摘要 Kalshi联合创始人Lopes Lara谈论如何将Kylie Jenner等名人八卦转化为预测市场商机,以及公司面临的法律挑战和Jay-Z的建议。

US’s Screwworm Fix Is Still a Year Away, Risking More Spread

The US’s best weapon against a deadly cattle parasite threatening the beef industry is more than a year away from showing meaningful results, raising concerns over how far the outbreak could spread before then.

中文摘要 美国针对威胁牛肉产业的螺旋虫病的解决方案还需一年以上才能见效,引发对疫情在生效前进一步蔓延的担忧。

Brazil’s World Cup Grilling Tradition Challenged by Beef Prices

Soaring prices for the world’s largest beef producer, Brazil, mean households will be buying less red meat when they gather to watch Brazil’s first game of the FIFA World Cup this Saturday.

中文摘要 巴西牛肉价格飙升,作为全球最大牛肉生产国,家庭在观看世界杯比赛时将减少购买红肉。

我的公益站被黑了,重新开始全员免费

我的ai图生视频公益站l0veyou昨天数据库被删了,源码也被删了,我真服了,不知道谁跟我这么大仇,我一直在修复,从现在开始免费24小时,如果积分还没补偿到位,那就延期免费。 55 个帖子 - 52 位参与者 阅读完整话题

深度求索研究员深夜发帖控诉:字节跳动工地凌晨施工扰民,投诉反遭封号

6 月 13 日,深度求索研究员陈德礼在社交平台 X 上发帖,强烈投诉北京字节跳动位于蓝景丽家的施工项目。他称,周末凌晨 2 点,挖掘机仍在工地轰鸣作业,距其住所仅 200 至 300 米,严重干扰居民休息。 他表示,此前已通过 12345 热线反映问题,官方回复称该项目为“市级重点基础设施工程”,法定施工时间为早 6 点至晚 10 点,含节假日。但如今工地凌晨违规施工,令他无法接受。他拍摄视频上传至国内社交平台小红书,随后账号被禁言,视频对他人不可见。他怀疑是北京方面看到后选择“封口”而非解决问题。目前其 12345 投诉仍未得到答复。 https://x.com/victor2077558

记一次对 GLM 5.2 的真实项目需求的横向评测

由于测试的模型越积越多了,表格会删除一些同厂商的旧模型,你可以在之前的评测帖子里找到它们的成绩。 项目 这是一个 Unity C# 项目,我进行测试的是一份皮肤系统需求案,我已经做了好预制体,而模型需要编写代码。 本轮与上两轮评测的项目和环境都完全一致: 第一轮 … 上一轮 模型来源 GLM 5.2: 官方 Coding Plan 速度 排名 模型 时间(分钟) 备注 1 Composer 2.5 3 2 Grok 4.20 0309 Reasoning 3 3 Step-3.5-Flash 6 4 Mimo V2 Omni 7 5 Doubao-Seed-2.0-Lite 7 6 Douba

星辰AI ccmax特惠分组降价至0.51

官网 https://ai.centos.hk 官方风控不大变动的话,这个价格会维持,如果风控加紧,会略微涨价 愿我们所有人 烟火向星辰 所愿皆成真~ 用前须知 https://doc.centos.hk 加入星辰AI售后群:QQ群 14 个帖子 - 12 位参与者 阅读完整话题

【CHY服务器公益站】开站

本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已

GLM 5.2测评:跻身第一梯队

老规矩私有bench 案例都很不错 第三个在这个案例中做热力学图的模型,前两个是Mythos和3.5Flash 正如知乎nao榜所言,日后通过中转贩子使用opus的人,都需要面对一个问题,你用的opus如果是glm5.2冒充的,而且难以分辨。 在实际bot agent体验上,如果不是对opus4.6特别熟悉的人基本无法分辨出两者。并且其追随上文的能力很强。如果上文用的opus,继续用glm5.2根本无法分别。 其缺点目前来看,上下文注意力可能不如4.6强(说实在的比4.6强的也几乎没有)。这次上1M上下文盲猜的DS V4的技术落地,DSA的注意力只能说目前来看中规中矩而已。 不过真的恭喜智谱啊

GLM-5.2能用了?

GLM5.2(思考:max)天气卡片测试: 以 iOS 18 的设计风格做一个带有动画效果的天气卡片,要求是使用 HTML、CSS 和基础 JavaScript,使用横板天气页面(拥有 4 个天气卡片 (晴天,大风,暴雨,暴雪))。应足够美观,实现一定的交互效果。 25 个帖子 - 22 位参与者 阅读完整话题

Fable5的封锁,或许是件好事

中美扯头发这么多年,对我们高科技的封锁早就不是新鲜事了,这些封锁无一不在帮助我们国产替代的发展,而且这一次有些不一样。 1.国模的coding能力比半年前已经好很多的,现在最多算不够好用 44 个帖子 - 37 位参与者 阅读完整话题