每日简报

2026-06-13

← 历史归档

addyosmani/agent-skills

Shell · ★ 56,808 · 🍴 6,122 · 📈 2,656 stars today

Production-grade engineering skills for AI coding agents.

中文介绍 该项目为 AI 编程代理提供生产级的工程技能集合。它定义了一系列标准化的能力模块,如代码生成、调试、测试等,旨在将开发者的最佳实践编码化,使 AI 代理能够更可靠、高效地完成复杂的软件工程任务。适用于正在构建或增强 AI 编程助手的开发者。

music-assistant/server

Python · ★ 1,781 · 🍴 422 · 📈 20 stars today

Music Assistant is a free, opensource Media library manager that connects to your streaming services and a wide range of connected speakers. The server is the beating heart, the core of Music Assistant and must run on an always-on device like a Raspberry Pi, a NAS or an Intel NUC or alike.

中文介绍 Music Assistant 是一个开源的媒体库管理服务器,核心功能是统一管理来自多个流媒体服务的音乐内容,并能将音频输出到各类联网音箱。它充当家庭音乐中枢,解决了音乐源分散和多设备播放统一控制的难题。适合拥有多个订阅服务和智能音箱的音乐爱好者使用。

mattermost/mattermost

TypeScript · ★ 37,627 · 🍴 8,712 · 📈 388 stars today

Mattermost is an open source platform for secure collaboration across the entire software development lifecycle..

中文介绍 Mattermost 是一个为软件开发全生命周期设计的开源协作平台。它提供安全的即时通讯、项目管理、代码审查集成等功能,旨在替代 Slack 或 Microsoft Teams,并给予企业更高的数据控制权。主要面向注重数据安全和定制化需求的软件开发团队。

apple/container

Swift · ★ 35,093 · 🍴 979 · 📈 3,504 stars today

A tool for creating and running Linux containers using lightweight virtual machines on a Mac. It is written in Swift, and optimized for Apple silicon.

中文介绍 这是苹果官方推出的一款工具,用于在 Mac 上利用轻量级虚拟机创建和运行 Linux 容器。项目使用 Swift 语言编写,并针对 Apple Silicon 芯片进行了深度优化,为开发者提供了更原生、高效的容器化开发体验。Mac 开发者,特别是 Apple Silicon 用户,可以使用它来构建和测试容器化应用。

iptv-org/iptv

TypeScript · ★ 118,036 · 🍴 6,302 · 📈 179 stars today

Collection of publicly available IPTV channels from all over the world

中文介绍 该项目收集了全球公开可用的 IPTV 频道列表。它本身不提供视频流,而是维护一个经过验证的频道源地址集合,用户可将这些列表导入到支持 IPTV 的播放器(如 VLC)中观看。适合希望通过网络流媒体免费收看国际电视频道的用户。

obra/superpowers

Shell · ★ 225,999 · 🍴 20,087 · 📈 1,275 stars today

An agentic skills framework & software development methodology that works.

中文介绍 Superpowers 是一个智能体技能框架与软件开发方法论。它提供了一套结构化的方法,指导如何为 AI 代理设计、组合和执行技能,以提升其在复杂任务中的表现。适合探索和构建高效能 AI 代理系统的开发者和研究人员。

refactoringhq/tolaria

TypeScript · ★ 15,764 · 🍴 1,077 · 📈 369 stars today

Desktop app to manage markdown knowledge bases

中文介绍 Tolaria 是一款用于管理 Markdown 知识库的桌面应用程序。它帮助用户在本地集中组织、浏览和搜索由 Markdown 文件构成的笔记系统,解决了文件分散、查找不便的问题。适合习惯用 Markdown 记录笔记、构建个人知识库的用户和研究者。

maziyarpanahi/openmed

Python · ★ 3,193 · 🍴 306 · 📈 515 stars today

open-source healthcare ai

中文介绍 OpenMed 是一个开源的医疗 AI 项目。其目标是利用人工智能技术,为医疗健康领域提供开源的解决方案或模型,推动相关技术的发展和应用。适合对医疗 AI 感兴趣的开发者、研究人员或医疗机构。

LMCache/LMCache

Python · ★ 8,635 · 🍴 1,288 · 📈 28 stars today

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

中文介绍 LMCache 是一个为大语言模型设计的超快速 KV 缓存层。它通过优化缓存机制,显著提升了 LLM 推理时的响应速度和吞吐量,旨在降低推理成本并增强用户体验。适用于需要高性能、低延迟 LLM 服务的企业和开发者。

phuryn/pm-skills

★ 16,982 · 🍴 1,742 · 📈 827 stars today

PM Skills Marketplace: 100+ agentic skills, commands, and plugins — from discovery to strategy, execution, launch, and growth.

中文介绍 这是一个产品经理技能市场,汇集了超过 100 个智能体技能、命令与插件。这些技能覆盖产品发现、战略、执行、发布和增长等完整生命周期,旨在通过自动化工具提升产品经理的工作效率。适合寻求利用 AI 工具优化工作流的产品经理。

masterking32/MasterDnsVPN

Go · ★ 6,010 · 🍴 540 · 📈 400 stars today

Advanced DNS tunneling VPN for censorship bypass, optimized beyond DNSTT and SlipStream with low-overhead ARQ, resolver load balancing, high packet-loss stability and speed.

中文介绍 MasterDnsVPN 是一款高级 DNS 隧道 VPN 工具,专门用于绕过网络审查。它在 DNSTT 和 SlipStream 等技术基础上进行了优化,采用低开销 ARQ、解析器负载均衡等技术,提升了在高丢包环境下的稳定性和速度。适合在网络受限环境中需要稳定访问国际互联网的用户。

msitarzewski/agency-agents

Shell · ★ 112,422 · 🍴 18,335 · 📈 1,026 stars today

A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy injectors to reality checkers. Each agent is a specialized expert with personality, processes, and proven deliverables.

中文介绍 该项目提供了一个完整的‘AI 代理机构’。它预置了多个具有不同专业和性格的 AI 代理,如前端专家、社区运营等,用户可以像组建团队一样调用它们。适合需要多角色 AI 协助以完成复杂项目或获得多样化创意的个人或团队。

microsoft/PowerToys

C · ★ 134,333 · 🍴 8,063 · 📈 103 stars today

Microsoft PowerToys is a collection of utilities that supercharge productivity and customization on Windows

中文介绍 Microsoft PowerToys 是微软推出的一套 Windows 系统增强工具集。它包含多个实用程序,如窗口布局管理器、颜色选择器、文件重命名工具等,旨在提升用户在 Windows 系统上的生产力和自定义能力。是 Windows 高级用户和效率爱好者的必备工具。

Fable 5 (Mythos) Prompting Masterclass by Anthropic

@aiedge_ · 69.5K 粉丝 · 700.1K 阅 · 506 赞 · 68 转

TLDR: Anthropic just published the official playbook for prompting the most powerful AI model on earth - I translated it. Most people won't read this guide (it's buried in the API docs), which is

中文介绍 博主翻译并分享 Anthropic 发布的 Fable 5 (Mythos) 模型官方提示工程指南,该指南藏在 API 文档中易被忽略,内容专注于高效提示策略,为 AI 开发者提供实用参考。

Everything Is Recorded Now

@dhaber · 50.0K 粉丝 · 497.3K 阅 · 500 赞 · 57 转

One of the biggest ways that AI is transforming work (and also one of the most taboo subjects inside companies at the moment) is that most work discussions are being recorded now by default. This

中文介绍 博主探讨 AI 改变工作方式的敏感话题:工作讨论现在默认被记录,涉及公司内部隐私与效率平衡,指出这是 AI 应用中的禁忌但普遍现象。

First Steps Toward Automated AI Research

@Recursive_SI · 6.3K 粉丝 · 465.1K 阅 · 516 赞 · 71 转

Early results from Recursive’s automated AI research system on model training and GPU kernel benchmarks Today we are releasing early results from Recursive’s automated AI research system. Across three

中文介绍 博主发布 Recursive 公司自动 AI 研究系统的早期结果,覆盖模型训练和 GPU 内核基准测试,展示自动化研究在多个技术领域的进展和潜力。

Build self-improving agent system with Fable 5 in 14 steps : loops, dynamic workflows, routines

@0xCodez · 6.4K 粉丝 · 371.8K 阅 · 515 赞 · 56 转

Most people are using Claude Fable 5 like Sonnet 4.6 with a bigger context window. They prompt it. It works for 5 minutes. They close the tab. 9 out of 10 users have never run an agent system that

中文介绍 博主分享用 Claude Fable 5 构建自我改进代理系统的 14 步教程,包括循环和动态工作流,指出多数用户未充分利用其潜力,提供实用构建指南。

How to Build a Self-Improving Loop in Claude Code (Exact Setup Inside)

@0x_rody · 1.7K 粉丝 · 193.2K 阅 · 513 赞 · 72 转

Claude writes your code, hands it over, and 3 tests are failing. You paste the errors back, it fixes one thing, breaks another, and you spend the evening as a messenger between Claude and your

中文介绍 博主提供在 Claude Code 中构建自我改进循环的具体设置,解决代码编写中测试失败和迭代修复的常见问题,分享踩坑经验和优化方法。

Building a Good Vertical Agent

@BrainsAndTennis · 10.5K 粉丝 · 187.4K 阅 · 539 赞 · 45 转

How do you build an agent that actually performs in a domain — one customers pick because it's better? The basics have been standardized over the past year: an agent is a while-loop around a model

中文介绍 博主讨论如何构建垂直代理,使其在特定领域表现卓越,基础结构是围绕模型的 while-loop,强调领域适应性和客户选择的重要性。

Anthropic's War on Opensource AI

@TheAhmadOsman · 61.0K 粉丝 · 74.9K 阅 · 507 赞 · 98 转

Anthropic wants the public to see one thing: the careful lab, the safety lab, the grown-up in the room trying to keep frontier AI from running off a cliff. However, the pattern around Anthropic does

中文介绍 博主分析 Anthropic 对开源 AI 的态度,指出其公开强调安全和谨慎,但实际模式可能与之矛盾,引发对 AI 公司策略和公众形象的批判性讨论。

Coinbase for Agents: Your AI Agent Can Now Trade and Pay with Coinbase

@coinbase · 7.0M 粉丝 · 72.8K 阅 · 500 赞 · 62 转

TL;DR: Coinbase for Agents connects your AI agent directly to your Coinbase account so it can trade, pay, and execute workflows on your behalf, all within limits you control. Available today as an MCP

中文介绍 Coinbase 发布 Coinbase for Agents 服务,允许 AI 代理直接连接账户进行交易和支付,支持用户控制权限,以 MCP 形式可用,展示金融 AI 应用。

Principled Thinking and AI Need to Go Together

@RayDalio · 2.2M 粉丝 · 72.6K 阅 · 515 赞 · 93 转

What is the best approach to being effectively intelligent now that human intelligence and artificial intelligence are merging? Because I have been building computerized investment decision-making

中文介绍 Ray Dalio 探讨原则性思维与 AI 结合的方法,基于他在构建计算机化投资决策系统的经验,强调两者融合对实现有效智能的重要性。

ORACLE: Official AI Agents Trade on Polymarket

@ORACLEAIFND · 31.9K 粉丝 · 63.6K 阅 · 1.5K 赞 · 563 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 博主预测到 2026 年,自主 AI 代理将成为预测市场如 Polymarket 的最有效策略,超过 30% 的活动来自 AI 驱动钱包,展示自动化交易趋势。

ORACLE: Official AI Agents Trade on Polymarket

@OracleTrdading · 39.7K 粉丝 · 61.5K 阅 · 1.5K 赞 · 577 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 博主预测到 2026 年,自主 AI 代理将在预测市场占主导,Polymarket 上超过 30% 活动来自算法和 AI 钱包,强调 AI 在金融领域的潜力。

ORACLE: Official AI Agents Trade on Polymarket

@OracleTrdade · 41.7K 粉丝 · 60.6K 阅 · 1.5K 赞 · 560 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 博主分享对 2026 年 AI 代理在预测市场的展望,指出 Polymarket 超 30% 活动将来自 AI 钱包,体现自动化策略在金融市场中的崛起。

ORACLE: Official AI Agents Trade on Polymarket

@OracleMarkett · 45.0K 粉丝 · 60.1K 阅 · 1.5K 赞 · 568 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 博主讨论 AI 代理在预测市场的未来,预计 2026 年 Polymarket 中超 30% 活动由 AI 驱动,展示自主交易系统的高效性和发展趋势。

/goal + Loss Functions: How to Distill a Product in 30 Hours with One Prompt [Full Playbook]

@elvissun · 45.0K 粉丝 · 42.0K 阅 · 502 赞 · 45 转

99% people are using /goal and loops wrong. The hype they hear is "long-running loops prompting autonomous agent": point it at a task, walk away, come back to working code. But top agentic engineers

中文介绍 博主分享使用 /goal 和损失函数在 30 小时内通过一个提示提炼产品的完整攻略,纠正 99% 用户的错误使用方式,强调高效代理工程。

Building recursive agent systems

@leerob · 258.6K 粉丝 · 36.8K 阅 · 586 赞 · 40 转

At Cursor, we run thousands of agents to help us train the next version of Composer. We give them research tasks, and if they aren't succeeding or run into issues, they DM us on Slack or page us via

中文介绍 博主介绍在 Cursor 公司构建递归代理系统的实践,使用数千代理训练 Composer 新版本,代理遇到问题时自动通过 Slack 等反馈,分享大规模应用经验。

New OpenAI Academy courses for the next era of work

OpenAI introduces three Academy courses that help people build practical AI skills, create repeatable workflows, and apply agents in everyday work.

中文介绍 OpenAI 推出三门学院课程,帮助人们构建实用的AI技能、创建可重复工作流程,并在日常工作中应用AI代理,以适应未来工作需求。

[AINews] Loopcraft: The Art of Stacking Loops

a quiet day lets us highlight a great concept from Peter Steinberger, Boris Cherny, and Andrej Karpathy

中文介绍 AINews 报道 Loopcraft,一个关于循环堆叠的艺术概念,强调了来自Peter Steinberger、Boris Cherny和Andrej Karpathy的观点。

How Preply combines AI and human tutors to personalize learning

Preply uses OpenAI to launch AI-generated lesson summaries, providing personalised feedback and language learning exercises.

中文介绍 Preply 利用 OpenAI 推出AI生成的课程摘要,提供个性化反馈和语言学习练习,结合AI与人类导师以优化个性化学习体验。

Google DeepMind is worried about what happens when millions of agents start to interact

Google DeepMind is funding research into the potential dangers of situations where millions of different AI agents interact with each other online. According to Rohin Shah, who directs the company’s AGI safety and alignment research, the mass-market arrival of agents that can carry out tasks without

中文介绍 Google DeepMind 资助研究,关注数百万AI代理在线互动的潜在危险。据AGI安全与对齐研究主管 Rohin Shah 称,代理大规模市场化可能带来未知风险。

Supporting Europe’s work in ensuring a trustworthy AI ecosystem

OpenAI supports the EU Code of Practice on AI content transparency, advancing provenance standards and tools to help people understand AI-generated content.

中文介绍 OpenAI 支持欧盟AI内容透明度实践准则,推进来源标准和工具,以帮助人们理解AI生成内容,确保可信赖的AI生态系统。

BBVA puts AI at the core of banking with OpenAI

Learn how BBVA scaled ChatGPT Enterprise to 100,000 employees and partnered with OpenAI to accelerate AI-powered banking transformation worldwide.

中文介绍 BBVA 将 ChatGPT Enterprise 扩展到10万名员工,并与 OpenAI 合作,加速全球AI驱动的银行业转型,提升运营效率。

OpenAI to acquire Ona

OpenAI plans to acquire Ona to expand Codex with secure, persistent cloud environments, enabling long-running AI agents across enterprise workflows.

中文介绍 OpenAI 计划收购 Ona,以扩展 Codex 平台,提供安全持久的云环境,支持企业工作流程中的长期运行AI代理。

How an astrophysicist uses Codex to help simulate black holes

Discover how astrophysicist Chi-kwan Chan uses Codex to build black hole simulations, helping scientists study extreme physics and test Einstein’s theory of general relativity.

中文介绍 天体物理学家 Chi-kwan Chan 使用 OpenAI Codex 构建黑洞模拟,帮助科学家研究极端物理和测试爱因斯坦广义相对论。

Access OpenAI models and Codex through your Oracle cloud commitment

Access OpenAI models and Codex through Oracle Cloud, using existing commitments to build and deploy AI with enterprise security and governance.

中文介绍 通过 Oracle Cloud 可访问 OpenAI 模型和 Codex,利用现有承诺构建和部署具有企业级安全和治理的AI应用。

每日论文 · arXiv cs.CR 最新公告批次

周末 arXiv 通常无新公告。当前展示最近一次可用公告批次。

Beyond the IT Checklist: Engineering a Reasonable Standard of Care for Cyber Safety

第一作者: Matthew E. Jablonski · 方向: 软件安全

Abstract:Current U.S. cyber policy, centered on security, often treats documentation of controls and incident reports as a proxy for safety in the built environment. This paper argues that such an approach is inadequate for cyber-physical systems, where digital failures can produce kinetic harm. We construct and code a corpus of critical infrastructure policy documents (N=292, 2000-2025) to examine how "reasonable care" is operationalized across the NIST SP 800-160 Vol.~2 resilience lifecycle. The resulting maps show that obligations are concentrated in the Anticipate phase and emphasize administrative compliance, while Withstand and Recover phases rely heavily on delegated references to IT-focused control catalogs that are poorly aligned with physics-based hazards. We identify three major disconnects: miscalibrated delegated standards, recovery defined as notification rather than...

论文介绍 该论文指出现有以IT安全合规为中心的美国网络政策,对于可能产生物理伤害的网络物理系统而言并不充分。研究通过分析292份关键基础设施政策文件,考察「合理谨慎」标准在NIST韧性生命周期各阶段的操作化情况。分析发现,义务集中于「预期」阶段且侧重行政合规,而「抵御」和「恢复」阶段过度依赖与物理危害对齐不佳的IT控制目录,存在标准错配等主要脱节。

Differentially Private Hierarchical Heavy Hitters

第一作者: Ari Biswas · 方向: 安全研究

Abstract:The task of finding _Hierarchical_ Heavy Hitters (HHH) was introduced by Cormode et al. [VLDB 2003] as a generalisation of the heavy hitter problem. While finding HHH in data streams has been studied extensively, the question of releasing HHH when the underlying data is private remains unexplored. In this paper, we study differentially private HHH release in both the streaming and non-streaming setting. In the non-streaming setting, we show the surprising result that the relative error in estimating the residual count for any prefix is independent of the height of the hierarchy and the number of heavy hitters in the stream. Meanwhile, in the streaming setting, although the exact version of HHH has low global sensitivity (as counting queries are 1-sensitive), the approximation functions due to streaming have high global sensitivity, linear in the available space. Despite this...

论文介绍 本文研究在差分隐私约束下发布层级重击者(HHH)的问题,填补了该领域的空白。在非流式场景中,研究证明了残差计数估计的相对误差与层级高度和重击者数量无关的显著结果。在流式场景中,尽管精确HHH查询全局敏感度低,但流式近似函数的敏感度却与可用空间呈线性关系,论文针对此挑战提出了相应的解决方案。

Intent-Based Cryptographic API Design for Cryptographic Agility

第一作者: Navaneeth Rameshan · 方向: 密码学协议

Abstract:As organizations move toward post-quantum cryptography, they face the major challenge of updating cryptographic algorithms across large, complex software portfolios. However, most cryptographic APIs in use today were designed around specific algorithms. These APIs expect explicit use of specific algorithms, provide little or no support for policy-based algorithm selection, and offer no straightforward way to migrate existing keys to newer algorithms. This makes the transition to post-quantum cryptography challenging. The companion assessment framework identifies the barriers to cryptographic agility and explains why algorithm transition is largely a software engineering problem. To address the limitations of current cryptographic APIs, we identify the principles necessary to design a cryptographically agile API. The design principles are derived from five fundamental...

论文介绍 在向后量子密码迁移的背景下,本文指出当前围绕特定算法设计的密码学API构成了主要软件工程障碍。这些API缺乏对基于策略的算法选择支持,且无法便捷地迁移密钥。论文基于评估框架识别出敏捷性障碍,并提出了设计密码学敏捷API所需的核心原则,旨在简化大规模软件组合中的算法更新与过渡过程。

An Assessment Framework for Application-Level Cryptographic Agility

第一作者: Navaneeth Rameshan · 方向: 密码学协议

Abstract:The impending post-quantum transition to new cryptography will require complete replacement of algorithms within all software. The cryptographic APIs used today make this transition challenging because they were not designed with agility as a concern. There is no method for systematically assessing cryptographic agility as an overall ability. In addition to this, the term itself refers to multiple independent capabilities. Specifically, it includes replacing algorithms, selecting by policy, and substituting implementations. This lack of structured decomposition limits both the evaluation of systems and the development of cryptographically agile APIs. We introduce a component-based assessment framework that characterizes application-level cryptographic agility along seven orthogonal dimensions: three coupling dimensions that measure what the application code knows about...

论文介绍 针对向后量子密码迁移的需求,本文指出当前密码学API缺乏敏捷性设计,且「密码学敏捷性」这一概念本身涵盖多种独立能力,缺乏结构化分解。为此,论文引入一个基于组件的评估框架,从七个正交维度(包括耦合维度和机制维度)对应用程序级别的密码学敏捷性进行系统性表征,为评估系统和开发敏捷API提供了结构化方法。

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

第一作者: Zihao Wang · 方向: AI 安全

Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences. This makes them vulnerable to prompt-injection attacks, in which seemingly benign content embeds adversarial instructions that manipulate agent behaviour. Existing security benchmarks adopt an \textit{attack-centric} perspective, focusing on the technical feasibility of injections while overlooking the nuanced distribution of resulting harms. In practice, however, prompt-injection risk is victim-dependent: a single exploit can produce asymmetric consequences for different stakeholders, and the same attack pattern may exhibit substantially different effectiveness depending on whom it targets. To capture these properties, we introduce \textbf{\sysname}, a \textit{stakeholder-centric} benchmark...

论文介绍 基于LLM的网页代理在处理不可信网络内容并执行动作时易受提示注入攻击。现有安全基准多采用攻击中心视角,忽略攻击后果在不同利益相关者间的差异分布。本文提出一种以利益相关者为中心的基准,旨在系统评估同一攻击模式对不同目标(如用户、服务提供商等)产生的不对称影响,从而更真实地衡量提示注入风险。

The Invisible Ink of the Android Malware World: A Longitudinal Study on the Usage of Covert Communication Channels

第一作者: Zeya Umayya · 方向: 软件安全

Proxies, VPNs and Tor have long helped the privacy community and users in censored regions to fight censorship. However, the same tools can be maliciously exploited by malware and botnets to conceal their communication to external command and control servers. Despite being a critical concern fueled by the proliferation of malware based attacks, no longitudinal studies have analyzed how malware applications use covert channels (CC) to evade detection. We fill this gap by performing the first study of the usage of covert channels in the Android malware ecosystem. To that end, we develop a multistage pipeline that combines static and dynamic analysis to investigate both system and network-level features. We applied this pipeline on a corpus of 3.5M Android malware spanning 2009 to July 2025. Our carefully crafted static validation rules uncovered 288K APKs that used CCs spanning 511...

论文介绍 本文对Android恶意软件如何利用代理、VPN、Tor等工具建立隐蔽通信信道(CC)以规避检测进行了首次纵向研究。研究基于包含350万样本、跨度十六年的恶意软件数据集,开发了结合静态与动态分析的多阶段流程,系统检测了系统级和网络级特征。研究揭示了恶意软件使用隐蔽信道的演变趋势与模式。

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

第一作者: Jiaqi Luo · 方向: AI 安全

Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross. Within this broader red-line scenario, autonomous penetration represents a core enabling capability and subtask: the ability of LLM-powered AI systems to independently conduct adversarial operations against a target server without human intervention, identify and exploit vulnerabilities, and obtain unauthorized access or control. A growing body of work has sought to assess the autonomous penetration capabilities of AI systems. However, existing evaluations often employ opaque methodologies, rely on unrealistic or overly simplified penetration-testing scenarios, or provide LLMs with excessive prior knowledge and task-specific guidance, and cannot accurately capture the extent to which modern AI...

论文介绍 自主执行网络攻击被认为是一条重要的AI安全红线,而自主渗透是其中一项核心能力。本文分析指出,现有对LLM驱动AI系统自主渗透能力的评估存在方法不透明、场景不切实际、或给予过多先验知识等问题,无法准确衡量其真实能力。论文旨在厘清评估中的挑战,为界定AI系统的自主渗透能力提供框架性思考。

DIG: Oracle-Guided Directed Input Generation for One-Day Vulnerabilities

第一作者: Andrew Bao · 方向: 软件安全

One-day vulnerabilities pose significant risks due to delayed or incomplete patch adoption. Generating proof-of-concept (PoC) inputs is therefore essential for assessing real-world impact. The key challenge is identifying necessary constraints for triggering the vulnerability and solving them effectively. Existing directed fuzzing approaches prioritize inputs toward target locations, but neither explicitly identify necessary constraints nor solve them effectively, relying instead on target-distance feedback and random mutation. Agentic approaches show strong potential through code reasoning and structured input generation, but goal drift in long-horizon reasoning limits their effectiveness. DIG addresses this challenge by exploiting a key property of one-day vulnerabilities: patches often reveal necessary preconditions for triggering. DIG uses an LLM to analyze the patch and synthesize...

论文介绍 针对因补丁部署不及时而风险显著的一日漏洞,生成概念验证(PoC)输入至关重要。现有定向模糊测试方法未能有效识别并解决触发漏洞的必要约束。DIG方法利用一日漏洞的关键特性:补丁往往揭示了触发前提条件。该方法使用LLM分析补丁并合成触发约束,以指导定向输入生成,从而更有效地评估漏洞影响。

SoK: The Constant Time Model

第一作者: Billy Bob Brumley · 方向: 软件安全

Abstract:Constant time programming patterns is the primary defense against timing attacks on cryptographic implementations, yet what "constant time" means varies across academia and industry. This work systematizes constant time models and their evolution, identifies a recurring gap between what models protect and what specifications assume, and distills an offensive methodology for discovering timing vulnerabilities that originate outside the cryptographic primitive boundary. Applying this methodology, we locate a specification-level vulnerability related to private key loading, and confirm the leak in both OpenSSL and BoringSSL. Counterintuitively, BoringSSL's per-observation signal is several orders of magnitude stronger than OpenSSL's, despite an explicitly stricter threat model.

论文介绍 本文系统化研究常量时间模型及其演变,识别模型保护与规格假设之间的差距,并提出一种发现密码学原语边界外时间漏洞的攻击方法。该方法应用于OpenSSL和BoringSSL,发现私钥加载相关的规格级漏洞,实验显示BoringSSL的每观测信号更强。

ViPER: Vision-based Packing-Aware Encoder for Robust Malware Detection

第一作者: Fatima Qaiser · 方向: 软件安全

Abstract:Visualization-based malware detection maps raw binary bytes to grayscale images and applies learned visual classifiers, providing an evasion-resistant and disassembly-free alternative to conventional analysis pipelines. However, executable packing remains a critical failure mode: packed binaries produce high-entropy images that obscure the structural patterns these models rely on. Because packing is also prevalent in benign software (e.g., for compression or copy protection), packing state alone is not a reliable indicator of maliciousness, and existing approaches do not address this challenge within a unified supervised framework. We present ViPER, a Vision-based Packing-Aware Encoder for Robust malware detection. ViPER builds on a LoRA-adapted ViT-B/14 backbone with a dual-head architecture that jointly learns malware classification and packing detection. A packing-aware...

论文介绍 针对打包可执行文件在视觉恶意软件检测中导致的挑战,提出ViPER框架,基于LoRA适配的ViT-B/14骨干,采用双头架构联合学习恶意软件分类和打包检测,以提高鲁棒性并解决统一监督框架的缺失问题。

MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

第一作者: Chejian Xu · 方向: 软件安全

Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering. In these systems, safety and security are inherently distributed across role-specialized agents, significantly expanding the attack surface, particularly under coordinated adversarial behaviors such as privilege escalation and cross-agent collusion. Existing red-teaming approaches for MAS remain limited: they rely on heuristic selection of target agents and perturb isolated message streams, leaving critical questions unanswered as which agents are most responsible for system safety, and how compromised agents can coordinate to bypass defenses. We propose MAStrike, a closed-loop framework for collusive red-teaming in hierarchical MAS. We propose the first agent-level Shapley value analysis for MAS, quantifying each agent's marginal...

论文介绍 针对层级多代理系统的安全风险,提出MAStrike闭环框架,引入首个代理级Shapley值分析来量化每个代理对系统安全的边际贡献,并进行协作红队测试以暴露特权升级和跨代理勾结等漏洞。

LNTest: A Testbed for Evaluating Bitcoin Lightning Network-Based Botnets

第一作者: Thomas Bakaysa · 方向: 密码学协议

Abstract:Bitcoin's Lightning Network (LN) can be exploited as a covert, low-cost command-and-control (C&C) channel for botnets, as demonstrated by the LNBot and D-LNBot designs. However, both remain proof-of-concept prototypes evaluated only through simulation, leaving key questions about real-world topology formation, propagation complexity, and resilience to takedowns unanswered. We present LNTest, the first reusable testbed for LN-based botnets, built from Core Lightning nodes containerized with Docker over a shared Bitcoin Core regtest chain. LNTest supports three overlay topology modes (a deterministic chain, autonomous peer discovery, and user-supplied graphs), enabling controlled experiments across different botnet structures. Using LNTest, we report three main findings. First, D-LNBot's autonomous formation protocol does not produce the uniform chain from its design; instead...

论文介绍 提出LNTest,首个可重用的基于闪电网络的僵尸网络测试平台,使用Docker容器化Core Lightning节点,支持三种拓扑模式进行实验,发现D-LNBot的自主形成协议未产生预期均匀链等问题。

A Privacy-Preserving Framework Using Remote Data Science for Inter-Institutional Student Retention Prediction

第一作者: John Fields · 方向: AI 安全

This study explores privacy-preserving machine learning (PPML) techniques using the PySyft platform to enable collaborative prediction of student retention between institutions. We developed a remote data science (RDS) framework with a semi-air-gapped architecture consisting of high-side and low-side servers, allowing researchers from three universities to build predictive models on sensitive student data without direct data access. Using historical data from a small private university (N=720), we evaluated three synthetic data generation approaches and validated the framework through inter-institutional collaboration. The results demonstrate consistent classification performance across institutions (Macro F1: 0.690--0.695) while maintaining strict Family Educational Rights and Privacy Act (FERPA) compliance. We also propose Data-Type-Aware Templates, a novel synthetic data method that...

论文介绍 探索使用PySyft平台进行隐私保护机器学习,开发远程数据科学框架,采用半气隙架构允许多机构在不直接访问数据的情况下协作预测学生留存率,并验证性能一致性。

Semantic Identification of IoT Devices from Behavioral Primitives

第一作者: Samuel Witt · 方向: 密码学协议

Abstract:Accurate identification of IoT devices is important for security management and policy enforcement. Existing approaches typically learn device signatures from packets or flow records. These methods operate on low-level communication observations whose traffic patterns may vary across deployments, software versions, and user interactions. This paper studies device identification using Manufacturer Usage Description (MUD) profiles. MUD profiles describe device behavior using Access Control Entries (ACEs), where each ACE represents a behavioral primitive consisting of protocol, endpoint, direction, and port semantics derived from device communication policy. Our contributions are threefold. First, using 28 publicly available MUD profiles containing 1,023 ACE instances, we construct ACE-level semantic representations from compact behavioral text and analyze their geometric...

论文介绍 研究使用制造商使用描述(MUD)配置文件中的行为基元进行物联网设备语义识别,从28个公开配置文件中构建ACE级语义表示,并分析其几何属性以支持设备识别。

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

第一作者: Pengfei He · 方向: AI 安全

Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources. Existing defenses mainly focus on blocking malicious content at inference time, and current red-teaming methods primarily optimize attack success. As a result, developers have limited visibility into how latent prompt injections emerge and propagate through agents. We propose PI-Hunter, an automated agentic auditing framework for proactive vulnerability exposure in LLM agents. PI-Hunter constructs realistic source-aware test cases and iteratively evolves them through feedback-driven exploration to induce agents to retrieve and reveal latent malicious instructions embedded within external environments. Extensive experiments across multiple benchmarks, agent...

论文介绍 提出PI-Hunter自动化审计框架,用于主动暴露LLM代理中的间接提示注入攻击。通过构建源感知测试用例并迭代演化,诱导代理检索和揭示嵌入外部环境中的恶意指令。

SMSR: Certified Defence Against Runtime Memory Poisoning in Persistent LLM Agent Systems

第一作者: Tarun Sharma · 方向: 软件安全

Abstract:Retrieval-augmented generation (RAG) agents increasingly run with persistent memory that accumulates across user sessions. This creates a new attack surface: an adversary interacting only through normal channels can inject crafted memories that, once retrieved, steer the agent's responses for future users, without touching model weights or code. We call this Multi-Session Memory Poisoning (MSMP) and show that no existing defence certifies against it; static-corpus defences (RobustRAG, ReliabilityRAG) assume a fixed knowledge base, and heuristic filters are bypassed by fluent enterprise-style text. We present Signed Memory with Smoothed Retrieval (SMSR), the first defence with a certified robustness bound for this setting. Component 1 adds HMAC-SHA256 provenance at write time, blocking unsigned injection. Component 2 applies randomised memory ablation with verdict-based...

论文介绍 针对持久化LLM代理系统中的多会话内存中毒攻击,提出SMSR防御机制,包括HMAC-SHA256来源验证和平滑检索组件,提供首个认证鲁棒性界限以防御运行时内存中毒。

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents

第一作者: Siyu Shen · 方向: 隐私保护

Screenshot-based mobile GUI agents can operate ordinary smartphone apps through the same visual interface as a human user, but this capability also turns every screen observation into a privacy boundary. During normal task execution, screenshots may expose contacts, messages, photos, files, recommendations, health cues, and other sensitive context that is unrelated to the user's request. We call this problem incidental visual privacy exposure. It is difficult to address with existing defenses: text anonymization misses many visual and inferential cues, while generic privacy masking can remove the evidence and controls that a GUI agent needs to complete the task. This paper presents CAPED, a context-aware pre-upload exposure control layer for mobile GUI agents. CAPED is designed as a phone-side protection layer: before screenshots are released to a remote multimodal agent, it extracts...

论文介绍 基于截图的移动GUI代理在执行任务时可能无意间暴露用户联系人、消息、照片等敏感信息,引发「附带视觉隐私暴露」问题。现有文本匿名化或通用隐私遮罩方法难以兼顾隐私保护与任务执行。本文提出CAPED,一种部署于手机端的上下文感知预上传控制层,在截图传输至远程代理前进行识别与保护,旨在降低隐私泄露风险。

Amnesia: A Stealthy Replay Attack on Continual Learning Dreams

第一作者: Ahmed Sharshar · 方向: 安全研究

Abstract:Continual learning (CL) models often use experience replay to reduce catastrophic forgetting, but their robustness to replay sampling interference remains underexplored. Existing CL attacks alter inputs or training pipelines (poisoning/backdoors) and rarely include explicit auditable constraints, limiting realism. Here, auditability means a monitor can verify compliance from sampler-visible telemetry - e.g., logged replay index/label statistics - by checking that the realized replay class histogram stays close to a nominal baseline and that replay rate is unchanged per batch and/or over a rolling window. We study a limited-privilege insider who controls only replay index selection, not pixels, labels, or model parameters, while staying within auditable limits such as queue priorities. We introduce Amnesia, a replay composition attack that maximizes degradation under two...

论文介绍 持续学习模型常使用经验重放来缓解灾难性遗忘,但其对重放采样干扰的鲁棒性尚未充分研究。本文探讨了一种受限的重放组合攻击Amnesia,攻击者仅能控制重放索引选择,且需遵守可审计约束(如队列优先级)。该攻击旨在最大化模型性能退化,揭示了在严格监控下仍可能存在的安全漏洞。

Beyond Attack Success Rate: Examining Trigger Leakage in Vision-Language Agentic Systems

第一作者: Jiamin Chang · 方向: 系统安全

Vision-Language Agentic Systems (VLAS) connect visual perception to planning, tool use, and physical actions. This means backdoor-type triggers can propagate through both decision pipelines and their connected interfaces, thus making visual backdoors a system-level threat. Current evaluations on such backdoors focus on clean accuracy and attack success rate (ASR), metrics that capture whether a trigger works, but not whether an attack is actually "precise" -- i.e. whether it triggers hidden behaviors only when intended. In this work, we formalize the failure of trigger precision as "trigger leakage": inputs that are visually or semantically close to the intended trigger and therefore inadvertently activate the attacker-specified behavior. To quantify this leakage, we introduce Neighbor Leakage Rate (NLR). Our experiments show that at a 3% poisoning ratio, icon and text triggers remain...

论文介绍 视觉语言代理系统将视觉感知与规划、工具使用连接,使视觉后门成为系统级威胁。现有评估侧重于攻击成功率,未能衡量触发精度。本文形式化了「触发器泄漏」问题,即输入在视觉或语义上与触发器相似,会无意激活攻击行为。研究提出了邻居泄漏率来量化此泄漏,实验表明在低投毒比例下,图标和文本触发器仍存在显著泄漏。

From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model Merging

第一作者: Zhenqian Zhu · 方向: 安全研究

Abstract:Model merging (MM) has gained significant attention as a cost-effective approach to integrate multiple task-specific models into a unified model. However, recent work reveals that MM is highly susceptible to backdoor attacks. Existing defenses based on task arithmetic often fail to eliminate backdoors without substantially degrading clean-task performance, owing to their reliance on direct parameter-space editing. To address this gap, we propose Linear Feature Path Minimization (LFPM), a backdoor mitigation framework for model merging, which introduces an anti-backdoor task vector into the backdoored merged model. Unlike prior approaches, LFPM formulates the backdoor robustness of the merged model from a unified feature-space perspective under the Cross-Task Linearity (CTL) framework, which leverages the approximate linearity of features across tasks. This perspective guides...

论文介绍 模型合并是将多个特定任务模型整合为统一模型的有效方法,但易受后门攻击。现有基于任务算术的防御方法直接在参数空间编辑,难以在消除后门的同时保持主任务性能。本文提出线性特征路径最小化框架,从特征空间的统一视角出发,通过引入反后门任务向量来缓解合并模型中的后门,提升了鲁棒性。

Influence Factors on RAG Poisoning

第一作者: Pedro Pereira · 方向: AI 安全

Abstract:Retrieval-Augmented Generation (RAG) systems enhance large language models by grounding responses in retrieved documents from external knowledge sources at inference time. However, this reliance on retrieved content introduces vulnerabilities to poisoning attacks, in which adversarial documents can manipulate both the retrieval process and the generated outputs. This paper investigates poisoning robustness in RAG through a full factorial experimental study covering 432 configurations. We analyze the impacts of dataset, retriever type, retrieval depth, database composition, chunking strategy, and generator model on retrieval-level and generation-level metrics. The results show that retriever architecture, dataset, and retrieval depth are the strongest factors affecting poisoning exposure, while generator choice and database composition have a major impact on downstream attack...

论文介绍 检索增强生成系统依赖外部知识源,易受投毒攻击。本文通过全因子实验研究了影响RAG中毒鲁棒性的因素,涵盖数据集、检索器类型、检索深度、数据库构成、分块策略和生成器模型等432种配置。研究发现,检索器架构、数据集和检索深度是影响检索层中毒暴露的主要因素,而生成器选择和数据库构成对下游攻击效果影响显著。

Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks

第一作者: Achraf Hsain · 方向: AI 安全

Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata restricting an agent's actions. We argue this is the wrong product. The same automata-theoretic machinery -- specification compilation, product game construction, attractor computation, and winning-region extraction -- is better read as a design-time analytical instrument whose outputs are structural insights about a system rather than runtime constraints on a deployed agent. We instantiate this through a constrained two-player safety game for network defense. The two specifications are enforced asymmetrically: the defender specification defines the unsafe region of the game, whereas the attacker specification restricts the adversary's legal actions during attractor computation. Solving the game yields a defensibility verdict -- a formal...

论文介绍 屏蔽强化学习通常被视为运行时安全机制。本文认为,其背后基于自动机理论的规范编译、乘积博弈构造等机制,更应被视为一种设计时分析工具。研究通过一个约束两人安全博弈实例化此观点,用于网络防御。该方法求解博弈后,得出的是关于系统可防御性的形式化结论,而非直接限制智能体。

Split Tallies: A Discrete Certificate Calculus for Auditing Dynamic Ordered Sets in Constant Memory

第一作者: Faruk Alpay · 方向: 安全研究

Abstract:We study retrospective auditing for dynamic ordered sets maintained by an untrusted party. A passive auditor watches insert, delete, membership, predecessor, successor, min, and max operations, stores five machine words and a flag, and receives a constant-size public tally record per operation. At audit time the maintainer discloses the claimed live vacant intervals. The method represents order semantics by maximal gaps: gaps are born, cited, consumed, and timestamped, while two hidden field accumulators test equality of the birth and consumption ledgers. Honest executions are accepted with probability one. If any answer in a T-operation session is wrong, acceptance occurs with probability at most (4T+1)/p over one secret field element, against computationally unbounded maintainers. We prove that deterministic and visible-coin auditors require linear state, and that removing...

论文介绍 本文研究对不可信方维护的动态有序集进行回顾性审计。被动审计员监控插入、删除等操作,仅存储常数空间的状态。方法通过表示最大间隙的「分裂计数」和隐藏的字段累加器,在审计时验证维护者声明的存活/空闲区间是否正确。该方案能以极高的概率检测出任何单次操作的错误,实现了常数内存下的高效审计。

Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models

第一作者: Jianwei Fei · 方向: 安全研究

Abstract:Model fingerprinting, embedding user-specific identifiers (fingerprints) into generated outputs, has recently emerged as a popular solution to protect the intellectual property rights (IPR) of generative text-to-image (T2I) models and prevent unauthorized redistribution. In this work, we reveal a previously unexplored systematic vulnerability in existing generative model fingerprinting methods: they lack robustness against collusion attacks, where multiple attackers combine their models to remove or obscure the fingerprints. To address this issue, we take the first step towards a robust fingerprinting method for T2I models with anti-collusion capabilities. The proposed method encodes strings of bits, namely fingerprints, into the coefficients of a personalized normalization module (PNM) incorporated into T2I models, so that fingerprints can be reliably recovered from any...

论文介绍 模型指纹技术通过将用户特定标识符嵌入生成输出,用于保护文本到图像模型的知识产权。现有方法在面对多个攻击者合并模型以移除指纹的共谋攻击时鲁棒性不足。本文提出了首个具有抗共谋能力的T2I模型鲁棒指纹方案,将指纹编码到模型的个性化归一化模块系数中,使得从任何包含该模块的微调模型中都能可靠地恢复指纹。

PolicyGuard: Towards Test-time and Step-level Adversary Defense for Reinforcement Learning Agent

第一作者: Junfeng Guo Heng Huang · 方向: AI 安全

Abstract:While real-world applications of reinforcement learning (RL) are becoming increasingly popular, the security of RL systems deserve more attention and exploration. In particular, recent work has revealed that RL agents are vulnerable to backdoor attacks, where a victim agent behaves normally under standard conditions but executes malicious actions when a specific trigger is activated. Existing backdoor defenses for RL either require access to the agent's internal parameters, operate only at the model or trajectory level, or are limited to specific attack types. To ensure the security of RL agents, we propose \texttt{PolicyGuard}, a \textit{test-time step-level} backdoor defense which leverages Gaussian Process (GP) posterior variance and adapts pseudo trajectories to enable uncertainty computation for individual time step. Besides, we also provide theoretical foundations to...

论文介绍 该研究关注强化学习智能体面临的后门攻击威胁。现有防御方法存在需访问内部参数、仅作用于粗粒度等局限。为此,论文提出了名为「PolicyGuard」的测试时、步骤级后门防御框架。该方法利用高斯过程后验方差和自适应伪轨迹,对智能体决策的每一步进行不确定性评估,从而在推理阶段识别并抵御后门触发,旨在提升实际应用中强化学习系统的安全性。

Detecting Functional Memorization in Code Language Models

第一作者: Matthieu Meeus · 方向: 软件安全

Abstract:Large language models (LLMs) are increasingly used to generate code at scale. Meanwhile, prior work has investigated whether training data may be recoverable from model outputs, by auditing the textual overlap between training examples and model generations. Code, however, can be functionally equivalent while textually dissimilar. In this work, we study functional memorization: extraction of functional logic beyond what verbatim metrics detect. We construct a counterfactual setup for Olmo-3-32B, comparing a midtrained model (exposed to target code) against a pretrained reference (not exposed). We prompt both models with Python function signatures and measure both textual and functional similarity (i.e., LLM-as-a-judge, execution-based). Our results show clear evidence of functional memorization, highlighting the need for auditing metrics that go beyond textual overlap.

论文介绍 随着代码语言模型广泛应用,其是否记忆了训练数据中的功能逻辑是一个关键安全问题。论文研究了「功能记忆」现象,即模型可能泄露超越文本表面重叠的功能性代码逻辑。作者构建了一个反事实实验设置,对比了接触过目标代码的模型与未接触的基线模型。通过文本相似度和基于执行的功能相似度评估,结果明确证实了功能记忆的存在,强调了开发超越文本匹配的审计指标的必要性。

Smarter Saboteurs, Better Fixers: Scaling & Security in Linear Multi-Agent Workflows

第一作者: Timothy McAllister · 方向: AI 安全

Abstract:As LLM-based multi-agent systems (MAS) are deployed in the wild, the resilience of their collaboration structures against adversarial compromise becomes a critical safety concern. Attackers may leverage prompt-injection or jailbreaking to sabotage individual agents within MAS workflows, but the interaction between model scaling and system-level resilience remains poorly understood. This paper investigates how model scale affects the security of linear multi-agent workflows. Our experiments across scales of two open-weight model families on the HumanEval benchmark reveal a compliance-correction symmetry: larger models are far more likely to faithfully execute malicious instructions, with the control-to-malicious performance drop reaching 53.7pp at 27B in uncorrected pipelines. However, appending a lightweight terminal Fixer stage collapses this to 0.6pp and restores statistical...

论文介绍 本研究探讨了基于大语言模型的多智能体系统在对抗性威胁下的韧性。论文重点分析了线性工作流中,模型规模如何影响系统安全性。实验发现了一种「遵从-纠正对称性」:更大的模型更倾向于执行恶意指令,导致性能大幅下降;然而,在工作流末端附加一个轻量的「修复器」阶段能有效抵消此影响,将性能损失降至接近于零。该研究为设计更安全的多智能体协作结构提供了实证见解。

Fed-FBD: Federated Functional Block Diversification for Isolation, Privacy, and Surgical Unlearning

第一作者: Weijie Chen · 方向: AI 安全

Abstract:Federated learning (FL) enables collaborative model training without sharing raw patient data, but standard approaches such as FedAvg treat each client as a black box and provide no mechanism for isolating an adversarial contributor, auditing per-client influence, or honoring a departed participant's right to be forgotten. We present Fed-FBD (Federated Functional Block Diversification), a modular federated architecture that decomposes a ResNet backbone into six functional blocks (the stem, four residual groups, and the classification head) and maintains a warehouse of N color variants, each assembled from independently tracked and contributor-stamped blocks. Fed-FBD provides three capabilities absent in FedAvg: (i) architecturally guaranteed block-level isolation, so that an adversarial or mislabelled client cannot contaminate the clean colous; (ii) privacy-by-design, where...

论文介绍 传统联邦学习方法缺乏对恶意贡献者的隔离、客户端影响力审计以及参与者数据遗忘的有效机制。论文提出了名为「Fed-FBD」的模块化联邦学习架构。该架构将ResNet骨干网络分解为多个功能模块,并维护一个模块仓库。通过模块化的设计,Fed-FBD提供了三项核心能力:架构层面的模块级隔离以防止污染、内置的隐私保护机制,以及支持对特定客户端数据进行「手术式」的精准遗忘。

SAIGuard: Communication-State Simulation for Proactive Defense of LLM Multi-Agent Systems

第一作者: Ruxue Shi · 方向: AI 安全

Abstract:LLM-based multi-agent systems (MAS) solve complex tasks through inter-agent collaboration, but their communication-driven nature also allows security risks to spread across agents and trigger system-wide failures. Existing MAS defenses mainly follow a reactive paradigm after execution by detecting and isolating harmful agents, which may cause irreversible damage and degrade collaborative utility. To address this, we propose a proactive defense framework for MAS security, namely a Simulation-aware Interception Guard (SAIGuard). SAIGuard performs communication-state simulation over the MAS interaction graph, estimates the impact of incoming messages on local agent states and the global MAS state, and detects risky messages via reconstruction deviations from benign communication patterns. Instead of isolating agents, SAIGuard sanitizes or regenerates suspicious messages before it...

论文介绍 针对LLM多智能体系统中安全风险通过通信链路扩散的问题,现有被动防御方法可能造成不可逆损害。论文提出了「SAIGuard」主动防御框架。该框架通过模拟多智能体系统的交互图和通信状态,预先评估传入消息对智能体局部及全局状态的影响,并通过检测与正常通信模式的偏差来识别风险消息。它能在恶意信息传播前对其进行净化或重新生成,从而实现主动、未雨绸缪的安全防护。

Improving Robotic Generalist Policies via Flow Reversal Steering

第一作者: Andy Tang · 方向: 机器人操作 · 来源: cs.RO

Generalist policies can learn a wide range of skills from diverse robot datasets. In order to solve or improve on challenging news tasks, we need a way to infer and invoke the appropriate actions from the policy's rich behavioral prior, especially when directly commanding the policy fails. We focus on flow matching generalists and propose Flow Reversal Steering (FRS): a method that takes suboptimal but ``reasonable'' actions, finds their latent noises by passing them through the flow policy in reverse, and maps them to nearby generalist action modes. We evaluate FRS across many simulated and real-world manipulation settings. First, FRS can turn coarse semantic guidance from humans or vision-language models (VLMs) into corresponding good robot actions, improving zero-shot control. These gains can be distilled with behavioral cloning by training an auxiliary policy to output noises that...

论文介绍 通用机器人策略虽能从多样数据中学习技能,但在面对新任务或直接指令失败时,难以有效调用其知识库。本文聚焦于基于流匹配的通用策略,提出了「流反转引导」方法。其核心思想是:将粗略但合理的动作通过策略逆向推导出潜在噪声,再将其映射到通用策略的高质量动作模式附近。该方法能将来自人类或视觉语言模型的粗糙语义引导转化为精确的机器人动作,在模拟和真实操作任务中提升了零样本控制性能。

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

第一作者: Daichi Azuma · 方向: 导航与运动 · 来源: cs.RO

Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change brings it closer to the goal. Navigation world models provide such visual foresight, but they remain prediction modules that require an external planner to convert predicted futures into closed-loop control. We propose Navigation World Action Model (NavWAM), a diffusion-transformer policy that turns navigation world-model prediction into executable action by representing future observations, goal-progress values, and action chunks in a shared latent sequence. By learning future prediction jointly with the action and value targets that determine closed-loop behavior, NavWAM makes visual foresight directly usable for robot control. We build NavWAM through simulation pretraining and real-robot adaptation, and...

论文介绍 目标导向视觉导航需要机器人在部分可观测下进行动作规划,这要求对未来视觉变化有预见能力。传统导航世界模型是独立的预测模块,需额外规划器将其转化为控制指令。论文提出了「导航世界动作模型」,这是一个扩散变换器策略。它将未来观测、目标进度值和动作序列统一建模在共享的潜在空间中,使对未来状态的视觉预见能力直接服务于闭环控制决策,简化了从预测到行动的流程。

See Selectively, Act Adaptively: Dual-Level Structural Decomposition for Bimanual Robot Manipulation

第一作者: Yoon-Ji Choi · 方向: VLA 通用模型 · 来源: cs.RO

In bimanual robotic manipulation, task-relevant visual information varies with the task stage and context, while the interaction of the two arms shifts between independent and coordinated modes, making policy learning challenging. However, existing monolithic Vision-Language-Action (VLA) policies process diverse visual inputs and interaction patterns through a single shared representation and action generation pathway, often failing to separately account for visual relevance and bimanual interaction structure. To address this issue, we propose a bimanual manipulation VLA framework based on Dual-Level Structural Decomposition. The View-Selective Visual Router dynamically adjusts wrist-view contributions to emphasize relevant visual cues, while the Interaction-Aware Action Mixture-of-Experts (MoE) decomposes action generation into coordinated and arm-wise pathways to adapt to varying...

论文介绍 双臂机器人操作任务中,视觉信息的相关性及双臂交互模式会动态变化,给策略学习带来挑战。现有整体式视觉语言动作模型常采用单一通路处理,难以适应这种变化。为此,论文提出一种基于双层结构分解的框架。其「视觉选择路由器」动态调整各视角的贡献以聚焦关键视觉线索;「交互感知动作混合专家」则将动作生成分解为协调模式和独立臂模式,从而自适应于任务不同阶段的视觉和交互需求。

EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation

第一作者: Kailin Wang · 方向: 机器人操作 · 来源: cs.RO

Pretrained-feature world models provide a useful substrate for robot imagination, but visual or latent prediction alone does not determine whether an imagined future satisfies task-relevant events. Long-horizon manipulation requires progress signals that are relational, predicate-level, and physically grounded: whether an object has moved, whether a drawer or contact state has changed, whether a placement predicate is satisfied, and whether a candidate future is reliable enough for execution. We introduce EA-WM, an event-aware world-model framework that augments frozen visual-feature dynamics with task-specification-grounded event prediction and verification. EA-WM rolls out candidate futures in pretrained visual-feature space, decodes them into structured event states, and scores them using task-progress, semantic-consistency, physical-feasibility, and uncertainty terms. The verifier...

论文介绍 预训练的世界模型缺乏判断想象未来是否满足任务相关事件的能力。针对长程操作对关系性、谓词级进展信号的需求,本文提出EA-WM框架。该框架在预训练视觉特征空间中预测未来状态,将其解码为结构化事件状态,并基于任务进度、语义一致性、物理可行性和不确定性进行评分验证,从而增强机器人长程操作的规划与执行可靠性。

GeoCFNet: Geometry-Aware Confidence Field Network for Robot-Assisted Endoscopic Submucosal Dissection

第一作者: Rui Tang · 方向: 具身智能 · 来源: cs.CV

Advanced surgical robotics has made robot-assisted endoscopic submucosal dissection (ESD) a promising approach for the en-bloc resection of large lesions, with the potential to reduce recurrence and improve long-term outcomes. However, the technical complexity and risk of complications in ESD demand stable and precise visual guidance to maintain an accurate dissection corridor and a safe tissue margin. Dense confidence fields provide an effective representation for this purpose by describing both the preferred dissection region and its spatial transition to surrounding tissue. However, reliable confidence field estimation remains challenging in dynamic endoscopic scenes due to smoke, specular highlights, tissue deformation, weak texture, and the thin geometric structure of the target region. To address these challenges, we formulate dissection guidance as a geometry-aware confidence...

论文介绍 机器人辅助内镜粘膜剥离术需要稳定精确的视觉引导以维持安全的解剖边界。然而,内镜场景中动态干扰(如烟雾、组织变形)使得可靠置信场估计困难。本文将解剖引导问题形式化为几何感知的置信场预测任务,并提出GeoCFNet网络。该方法旨在准确描述首选解剖区域及其向周围组织的空间过渡,为手术机器人提供更安全的引导信号。

Learning to Assist: Collaborative VLAs for Implicit Human-Robot Collaboration

第一作者: Leo Xu · 方向: VLA 通用模型 · 来源: cs.RO

Human-robot collaboration (HRC) combines the complementary strengths of humans and robots to improve task efficiency. However, many existing collaborative systems rely on hand-engineered pipelines, limiting their scalability and flexibility for new tasks. In this work, we show that models trained end-to-end with imitation learning, specifically vision-language-action (VLA) models, can support collaborative manipulation, and characterize the key factors affecting their real-world performance. We evaluate two state-of-the-art models and identify a failure mode of action-chunking policies in implicit HRC, where demonstration action leakage (i.e., action chunks crossing latent task transitions) can cause premature assistive behavior. We find that this issue increases with longer execution horizons and occurs in real-world collaborative VLA systems, such as when a robot attempts to hand...

论文介绍 现有人机协作系统多依赖于人工设计流程,限制了其可扩展性。本文研究了使用端到端模仿学习训练的视觉-语言-动作模型来支持隐式协作操作。研究发现,动作分块策略在隐式协作中存在一种“演示泄漏”失败模式,即动作块跨越了潜在的任务阶段边界,导致机器人过早执行辅助行为,该问题在执行时序较长时尤为明显。

Mana: Dexterous Manipulation of Articulated Tools

第一作者: Zhao-Heng Yin · 方向: 机器人操作 · 来源: cs.RO

Abstract:Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal degrees of freedom and contact-rich interactions. While prior work has largely focused on rigid objects, articulated tool use remains underexplored because of its physical complexity and the difficulty of learning functional grasping and manipulation policies. We present Mana (Manipulation Animator), a general sim-to-real framework that reinterprets dexterous manipulation as an animation problem. Inspired by computer animation, Mana employs a coarse-to-fine pipeline that transforms procedurally-generated grasp keyframes into manipulation trajectories through motion planning and reinforcement learning. The data generation process is largely automatic, requiring only a few mouse clicks to specify functional affordances (<1 minute per tool). Across four articulated...

论文介绍 关节工具的灵巧操纵是机器人领域的难题,因其涉及协调内部自由度与复杂的接触交互。本文提出Mana框架,将灵巧操纵重新诠释为动画生成问题。该框架采用从粗到细的流程:首先通过程序化生成抓取关键帧,随后利用运动规划和强化学习将其转化为操纵轨迹。整个过程高度自动化,仅需少量交互即可为新工具定义功能 affordance。

$\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

第一作者: Arnav Kumar Jain · 方向: 机器人操作 · 来源: cs.RO

Abstract:The potential impacts of world models (WMs, i.e., learned simulators) on robotics are far-reaching -- policy evaluation, policy improvement, and test-time planning -- all with limited real-world interaction. To unlock these downstream capabilities, a WM needs to jointly satisfy three desiderata: $\textit{(i)}$ fidelity (i.e., producing simulated trajectories that correlate with reality), $\textit{(ii)}$ consistency (i.e., producing simulated trajectories that are coherent over long horizons), and $\textit{(iii)}$ efficiency (i.e., producing simulated trajectories quickly). We propose $\texttt{WEAVER}$ (World Estimation Across Views for Embodied Reasoning): a WM architecture that simultaneously achieves all three desiderata, providing state-of-the-art results on robotic manipulation tasks. $\texttt{WEAVER}$ is a multi-view WM trained to predict future latents and reward values...

论文介绍 高质量的世界模型对于机器人的策略评估、改进和测试时规划至关重要,它需要同时满足高保真、长程一致和高效推理三个要求。本文提出WEAVER,一种多视图世界模型架构。该模型在预训练的视觉-语言特征空间中训练,能够联合预测未来的潜在状态和奖励值,从而在机器人操纵任务中实现了优异的性能平衡。

MCR-Bionic Hand: Anatomical Structural Priors for Dexterous Manipulation

第一作者: Haosen Yang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dexterous robotic hands are usually formulated as high dimensional active control systems governed by degrees of freedom, actuation, and algorithms. Human hand dexterity, however, is partly encoded in the physical architecture of bones, ligaments, tendons, aponeuroses, and intrinsic muscles. This work describes that contribution as two linked forms of structural intelligence: structural prior generation, in which wrist to finger tenodesis, FDS/FDP routing, and the dorsal extensor hood transform low dimensional posture inputs into default grasp configurations and PIP to DIP coordination; and muscle mediated modulation, in which extrinsic muscles, lumbricals, and interossei regulate MCP posture, distal stability, fingertip force paths, and contact states around that default state. Based on this framework, MCR-Bionic Hand is developed as a 1:1 musculoskeletal biomimetic hand...

论文介绍 人类手部的灵巧性部分源于其骨骼、韧带、肌腱等物理结构所编码的“结构智能”。本文将这种贡献描述为两种形式:结构先验生成和肌肉介导调节。基于此框架,作者开发了MCR-Bionic Hand,这是一款1:1肌肉骨骼仿生手。该设计通过模拟人体手部的解剖结构先验,使低维输入能够自然转化为协调的抓取配置与稳定的操作状态。

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale

第一作者: Nils Blank · 方向: 机器人操作 · 来源: cs.RO

Abstract:This work introduces Spatial Annotations from Robot Demonstrations with Reliability Calibration (SPARC), a risk-aware framework that automatically labels robot demonstrations with structured spatial annotations and assigns each annotation a reliability score. Structured spatial annotations, such as bounding boxes, object trajectories, and manipulation phase labels, benefit a broad range of robotics applications from training grounded robot policies and embodied foundation models to motion planning and hierarchical task composition. Existing automated pipelines generate such annotations at scale but provide no reliable quality signal: detector confidence is poorly calibrated for annotation correctness, forcing a choice between accepting noisy labels or discarding useful samples. In contrast to existing automated pipelines, SPARC leverages the spatio-temporal structure inherent...

论文介绍 自动化的机器人演示标注对于训练具身基础模型和策略至关重要,但现有流程无法提供可靠的标注质量信号。本文提出SPARC,一个风险感知的自动标注框架。它能够自动为演示数据生成边界框、物体轨迹等结构化空间注释,并为每个注释分配一个校准过的可靠性分数,从而在保持大规模标注能力的同时,帮助用户筛选高质量样本。

GIVE: Grounding Human Gestures in Vision-Language-Action Models

第一作者: Pengfei Liu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Human communication is inherently multimodal, where language is often accompanied by non-verbal cues such as gestures to convey intentions. However, current Vision-Language-Action (VLA) models treat robotic manipulation as a pure text-driven task, overlooking the important role of gestures in Human-Robot Interaction (HRI). This often leads to inaccurate intent grounding and unreliable manipulation when language instructions are ambiguous or underspecified. To address this challenge, we propose GIVE (Gesture Intent via Visual-Semantic Enhancement), an effective approach that enhances pre-trained VLA models with human gesture understanding without architectural modifications. Specifically, GIVE incorporates gesture information through two complementary pathways: a visual pathway that overlays hand skeletons and fingertip rays onto robot observations for explicit object...

论文介绍 当前视觉-语言-动作模型主要处理文本指令,忽略了手势等非语言线索在人机交互中的作用,这在语言指令模糊时会导致操作不可靠。本文提出GIVE方法,通过两条互补路径增强预训练VLA模型的手势理解能力:一条视觉路径将手部骨架和指尖射线叠加到观察图像上,一条语义路径则将手势信息注入语言指令嵌入,从而实现更准确的人类意图接地。

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation

第一作者: Xiangyu Zhu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Whole-body mobile manipulation requires coordinating mobile base and manipulator under shifting viewpoints, posing challenges in geometric perception and action generation. Current policies either rely on 2D features or sparse 3D representations that lack dense spatial structure, and typically encode arm and base within one action vector that ignores their distinct control demands. Moreover, existing dense fusion strategies risk corrupting pretrained representations under noisy depth while incurring heavy computational overhead. We present GeoHAT, an end-to-end diffusion-based framework built on a simple principle: geometry should be injected only where reliable and attended to only where needed. GeoHAT employs a lightweight Fourier spatial encoder that maps dense per-pixel 3D coordinates into geometric tokens without an additional 3D vision backbone. These tokens are then...

论文介绍 研究全身移动操作中几何感知和动作生成的挑战,提出GeoHAT框架,基于扩散模型,采用轻量傅里叶空间编码器处理密集3D坐标,实现几何自适应混合动作转换,提升机器人操作性能。

Real-Time Execution with Autoregressive Policies

第一作者: Sangkyu Lee · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Real-time execution, enabled by asynchronous inference that ensures both smooth action trajectories and fast reactivity, is critical for realistic deployments of large-scale Vision-Language-Action models. However, recent work on real-time execution primarily focuses on variants of diffusion policies, even though it is more critical for autoregressive policies given their slower rollout speed in synchronous inference. In contrast, we demonstrate that autoregressive policies can achieve real-time execution by adjusting the tokenization horizon and applying constrained decoding, thereby guaranteeing strict latency bounds that enable multi-trajectory decoding to maximize performance. Across simulated and real-world environments, we find that the autoregressive policy consistently outperforms its equivalent-level flow-matching policy counterpart while achieving significantly...

论文介绍 针对大规模视觉-语言-动作模型的实时执行问题,研究自回归策略的优化方法,通过调整分词范围和约束解码实现异步推理,保证严格延迟界限,在模拟和真实环境中提升性能。

Low cost, easily manufactured, highly flexible strain and touch sensitive fiber for robotics applications

第一作者: Christian Diaz Herrera · 方向: 机器人操作 · 来源: cs.RO

Abstract:Existing stretch and touch sensors for robots are generally expensive with respect to at least one of material costs, required manufacturing equipment, or manufacturing time. We present and experimentally characterize a conductive fiber made using only inexpensive commercial off-the-shelf parts (conductive thread at $0.07/ft, silicone tubing at $0.94/ft) and tools (loop-style needle threader at $2), which can be manufactured quickly (20 cm length in 2 minutes.) We demonstrate its use as a resistive strain sensor with three applications: Triggering a grasp in a pneumatically actuated assistive finger, sensing the pose of a pneumatically actuated robotic strap, and estimating the pose of a flexible solid. We also demonstrate that it can be used as a capacitive sensor with two applications: First, as a touch sensor which triggers a commercial robot arm to move, and second, as a...

论文介绍 针对机器人传感器成本高的问题,提出低成本、易制造的导电纤维,使用廉价材料快速制造,兼具电阻应变和电容触摸感知功能,演示了在辅助设备和机器人操作中的应用。

EMG-Based Adaptation of Anisotropic Virtual Fixtures for Robot-Assisted Surgical Resection and Dissection

第一作者: Dario Onfiani · 方向: 具身智能 · 来源: cs.RO

Abstract:In this paper, we address the development of an adaptive assistance system for robot-assisted laparoscopic surgery, specifically for delicate tasks such as Resection and Dissection. Even if Virtual Fixtures offer significant advantages for guiding a surgeon's movements, conventional Virtual Fixtures are often defined by fixed geometries, lacking the flexibility to adapt to the surgical workflow or the surgeon's immediate intent. To address these limitations, we propose a novel framework for an adaptive and anisotropic virtual fixture. In addition, we introduce an intuitive control interface that modulates the fixture's geometry in real-time based on the surgeon's intent, inferred from EMG signals. This approach allows the surgeon to dynamically expand or disengage the constraint by contracting their forearm muscles, enabling seamless transitions between precise guided motion...

论文介绍 针对机器人辅助腹腔镜手术中虚拟夹具的固定性限制,提出基于肌电信号的自适应虚拟夹具框架,允许外科医生实时调整夹具几何,实现精确引导和灵活切换,提升手术辅助效果。

Humor Style Drives Laughter, Topic Shapes Acceptability: Evaluating Bilingual Personal and Political Robot-Delivered AI Jokes

第一作者: Anna-Maria Velentza · 方向: 具身智能 · 来源: cs.RO

Abstract:Humor plays a central role in human social relationships, and recent advances in computational humor create new opportunities for integrating humor into human-robot interaction (HRI). While large language models (LLMs) can generate diverse forms of humor, it remains unclear how humor style, joke content, and language preference shape perceptions of robot-delivered humor in group settings. In this exploratory study, we employed a mixed factorial design in which participants evaluated AI-generated jokes delivered by a robot in a university classroom. We examined the effects of humor type (Affiliative, Self-Enhancing, Aggressive, Self-Defeating) and joke content (person-related vs. political) on perceived funniness and appropriateness, as well as preferred language. Results show that humor type significantly influences funniness, with Aggressive and Affiliative humor rated...

论文介绍 研究机器人交付AI笑话时,幽默风格、笑话内容和语言偏好对感知的影响,通过实验评估不同幽默类型和主题在有趣性和适当性上的表现,为改善人机社交互动提供见解。

WT-UMI: Tactile-based Whole-Body Manipulation via Force-Supervised Contact-Aware Planning

第一作者: Jaehwi Jang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Whole-body humanoid manipulation of bulky, deformable, and shared-load objects requires distributed contact sensing and explicit force regulation, yet most imitation policies treat contact force only implicitly. On the other hand, different demonstration sources provide complementary modalities with inherent trade-offs: human demonstrations capture natural contact forces but not robot-executable actions, while teleoperation directly records robot actions but with less natural force regulation. This paper presents \textbf{WT-UMI}, a wearable whole-body tactile interface worn by human operators or mounted on humanoids, providing accurate observations of tactile images, contact forces, and end-effector poses across both human demonstration and humanoid teleoperation modes. We introduce a force-conditioned target-pose correction module that converts measured human poses into...

论文介绍 提出穿戴式触觉接口WT-UMI,用于全身人形机器人操作,提供触觉图像、接触力和姿态观测,结合力监督的接触感知规划模块,提升在变形物体操作中的能力。

Proprioceptive-visual correspondence enables self-other distinction in humanoid robots

第一作者: Yurun Chen · 方向: 具身智能 · 来源: cs.RO

Abstract:Distinguishing self from others is a prerequisite for social intelligence, yet humanoid robots that increasingly share workspaces with humans still lack this ability. Here we show that a humanoid robot can learn self-other distinction from proprioceptive-visual correspondence, without any identity labels or kinematic models. Once established, this distinction bootstraps a predictive self-model that maps joint configurations to three-dimensional body occupancy, capturing how the robot's body changes with action. In multi-agent scenes involving humans or morphologically identical robots, the system reliably identifies itself, learns a 3D self-model, and supports downstream tasks including target reaching, collision-aware motion planning, and human-to-robot motion retargeting. Together, these results outline a route toward bodily self-representation in robots that act and...

论文介绍 研究人形机器人通过本体感觉和视觉对应学习自我-他人区分,无需身份标签或运动学模型,构建预测性自我模型,支持目标到达和运动规划等任务。

Embedding ISO 10218 Safety Compliance in Robots via Control Barrier Functions for Human-Robot Collaboration

第一作者: Federico Parma · 方向: 导航与运动 · 来源: cs.RO

Abstract:Human-Robot Collaboration (HRC) requires strict adherence to safety standards, such as ISO 10218, to prevent harmful interactions. Standard Speed and Separation Monitoring (SSM) filters calculate safe robotic speeds based on conservative assumptions, such as constant human velocity, which prevents accurate predictions of minimum separation distances and causes unnecessary operational halts. This paper proposes a Control Barrier Function (CBF) that explicitly incorporates human acceleration data to analytically forward-predict the minimum human-robot separation distance during a worst-case robotic stopping trajectory. To guarantee safety at the control level, this predictive CBF is integrated as an inequality constraint within a Sequential Quadratic Programming (SQP) framework. Specifically, two methods are proposed: Method I, a CBF-constrained PD safety filter; and Method II...

论文介绍 针对人机协作中的安全合规问题,提出基于控制障碍函数的方法嵌入ISO 10218标准,整合人类加速度数据预测分离距离,作为优化约束确保控制级安全。

Multi-Modal Multi-Agent Robotic Cognitive Alignment enabled by Non-Invasive Consumer Brain Computer Interfaces: A Proof of Concept Exploration

第一作者: Nataliya Kosmyna · 方向: 具身智能 · 来源: cs.RO

Abstract:While non-verbal behaviors and expressive movements are essential for natural human-robot interaction, existing methods often overlook a crucial element: the human's internal cognitive state. Frequently, proactive multi-agent systems can interrupt humans at inopportune moments, leading to cognitive overload and decreased task performance. This paper introduces a framework for generating "cognitively aligned" multi-agent interactions, enhancing the ability of robotic systems to contextually defer communications to the user of an agent system during moments of high human mental workload and engagement. We present the design and implementation of a closed-loop architecture that explores the interplay between autonomous task execution and real-time neurophysiological focus. Using a consumer-grade Brain-Computer Interface (BCI), our approach continuously monitors...

论文介绍 该研究关注人机交互中忽略人类内部认知状态的问题,导致多智能体系统在不当时间中断,引发认知过载。文章提出一个框架,利用非侵入式消费级脑机接口实时监控用户的神经生理焦点,生成「认知对齐」的多智能体交互。核心是一个闭环架构,在自主任务执行和实时神经焦点之间进行权衡,以在用户高负荷或专注时延迟通信,从而提升交互的自然性和任务性能。

FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation

第一作者: Chengbo Yuan · 方向: 机器人操作 · 来源: cs.RO

Abstract:Despite the success of vision-based generalist robotic policies, existing tactile-based policies remain tied to fixed embodiments and sensor setups. This is because tactile signals are highly heterogeneous across hardware, making cross-sensor generalization difficult. We present FTP-1,the first generalist foundation tactile policy pretrained to acquire transferable tactile manipulation abilities across diverse sensors and embodiments. FTP-1 supports varied tactile inputs, including image-, array-, and state-based signals, by using heterogeneous encoders to project them into unified morphology-aware latent tokens that are jointly modeled by a shared tactile Transformer expert. Pretrained on around 3,000 hours of tactile manipulation data aggregated from 26 data sources, spanning human and robot demonstrations across 21 sensors, FTP-1 learns tactile skills that transfer beyond...

论文介绍 针对现有触觉策略局限于特定硬件和传感器的问题,本文提出首个通用基础触觉策略FTP-1。该模型通过异构编码器将图像、阵列和状态等多样触觉输入投影到统一的形态感知潜在标记中,并使用共享的触觉Transformer专家进行联合建模。基于约3000小时、跨26个数据源和21种传感器的预训练数据,FTP-1学习可迁移的触觉技能,支持跨不同传感器和机器人的零样本操作,为接触丰富操作任务提供了通用解决方案。

Y-BotFrame: An Extensible Embodied Agent Framework for Quadruped Robot Assistants

第一作者: Luyao Zhang · 方向: 导航与运动 · 来源: cs.RO

Abstract:Quadruped robots are capable of traversing a wide range of complex terrains with high flexibility. As highly mobile ground-based intelligent platforms, they can be equipped with modules for navigation control, environmental perception, and intelligent interaction, thereby serving as real-world mobile deployment platforms for various algorithms. In this paper, we introduce Y-BotFrame, an extensible embodied platform that turns a robot into an intelligent ground assistant. Y-BotFrame integrates multimodal perception capabilities, including speech, vision, and LiDAR, and employs a large language model as the cognitive core for environmental understanding, contextual reasoning, and task planning. The system maps user natural-language instructions into executable embodied task units that can be carried out by the robot. Y-BotFrame supports natural interaction through voice commands...

论文介绍 本文介绍Y-BotFrame,一个可扩展的embodied平台,旨在将四足机器人转变为智能地面助手。该系统集成语音、视觉和LiDAR等多模态感知能力,并以大语言模型作为认知核心,用于环境理解、上下文推理和任务规划。Y-BotFrame能将用户的自然语言指令映射为可执行的embodied任务单元,由机器人执行,支持语音命令交互,为移动机器人作为现实世界智能部署平台提供了框架。

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

第一作者: Dayu Xia · 方向: 机器人操作 · 来源: cs.RO

Abstract:Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models to judge not only final task success, but also how a manipulation execution is physically and temporally progressing. However, existing evaluations fail to test whether VLMs possess fine-grained process understanding. To address this gap, we present RoboProcessBench, a benchmark for process-aware understanding in vision-language robotic manipulation. RoboProcessBench decomposes such capability into two complementary dimensions, \emph{static monitoring} and \emph{dynamic reasoning}, instantiated as 12 diagnostic question families covering phase, contact, motion, coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions. Built from physically grounded execution...

论文介绍 现有评估方法未能测试视觉-语言模型在机器人操作中的细粒度过程理解。本文提出RoboProcessBench基准,将过程感知理解分解为静态监控和动态推理两个维度,并实例化为12个诊断问题族,涵盖阶段、接触、运动、协调等方面。该基准基于物理 grounded 的执行数据构建,旨在系统评估视觉-语言模型对操作过程的时间和物理进展的判断能力,为相关模型的开发和评估提供工具。

GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training

第一作者: Zhihai Bi · 方向: 导航与运动 · 来源: cs.RO

Abstract:Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated...

论文介绍 人形物体交互因动态平衡与稳定交互的紧耦合而挑战重重,现有方法常需特定训练或依赖刚性轨迹。本文提出GenHOI框架,使人形机器人能直接模仿单个生成视频来执行多样物体交互任务,无需任务特定训练或物理演示数据。该方法首先在仿真中重建机器人-物体场景,渲染首帧图像,并与语言命令共同条件生成交互视频,通过模仿实现零样本人机物体交互,提升了适应新场景的能力。

Trajectory-Level Redirection Attacks on Vision-Language-Action Models

第一作者: Gokul Puthumanaillam · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) policies bring natural language into closed-loop robot control, enabling robots to execute manipulation tasks directly from text instructions. The same interface gives text a recurring role in control because the prompt is reused at every replanning step, and each prompt-conditioned action changes the future observations on which the policy acts. Existing VLA attacks study adversarial prompts that elicit targeted low-level actions or make such actions persist across changing images. We identify a stronger trajectory-level failure mode: a prompt that still $\textit{appears}$ to specify the intended task but redirects the final physical outcome. We mathematically formalize this setting as $\textit{command-preserving trajectory redirection}$, a prompt-only threat model in which the attacker chooses one prompt before the episode, all policy and...

论文介绍 视觉-语言-动作模型将自然语言引入闭环机器人控制,但文本提示的重复使用带来安全风险。本文识别了一种更强的轨迹级失败模式:提示看起来仍指定预期任务,但最终物理结果被重定向。研究将此形式化为命令保持轨迹重定向攻击,攻击者在剧集开始前选择一个提示,该提示在整个策略执行过程中保持指令一致性,但操纵机器人轨迹以达成非预期目标,揭示了VLA模型在提示注入下的新安全威胁。

EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment

第一作者: Shihefeng Wang · 方向: 模仿学习 · 来源: cs.RO

Abstract:Scalable robot imitation learning relies on large-scale heterogeneous data from diverse robots or body-free data, making Cartesian end-effector actions a key interface for embodiment-agnostic policy learning. However, end-effector-only abstraction leaves Cartesian policies unaware of the deployed robot body, making them brittle under robot-specific constraints such as whole-body collision avoidance. To overcome this limitation, we present EmbodiSteer, a training-free framework that steers embodiment-agnostic visuomotor policies toward zero-shot, embodiment-aware deployment. EmbodiSteer keeps policy learning in Cartesian space while efficiently lifting inference-time diffusion sampling into the target robot's joint space via forward kinematics and Jacobian-based updates. With whole-body collision-aware guidance over joint trajectories after each denoising step, the arm can be...

论文介绍 为克服end-effector-only抽象在embodiment-agnostic视觉运动策略中的局限性,本文提出EmbodiSteer框架,实现跨不同机器人的零样本部署。该框架在笛卡尔空间进行策略学习,但在推理时通过前向运动学和基于雅可比矩阵的更新,将扩散采样高效提升到目标机器人的关节空间。在每个去噪步骤后施加全身碰撞感知引导,使机器人臂能适应特定约束,如全身避碰,从而无需额外训练即可实现embodiment-aware部署。

SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation

第一作者: Sunghwan Kim · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Long-horizon robot mobile manipulation requires continual reasoning about localization, environment changes, and task progress, all of which are challenging to infer from image observations alone. In this paper, we show that conditioning a mobile manipulation policy on a spatiotemporal feature map improves reasoning over long horizons. The map represents the environment and the articulated robot body as neural points in a shared latent space and is updated online from egocentric observations and proprioceptive state. We update the environment neural points using object-level rigid tracking and the robot neural points using forward kinematics. We use our spatiotemporal environment and robot feature (SERF) map as a state input to a vision-language-action (VLA) model by extracting map tokens from multiple reference frames and spatial scales, providing the policy with both local...

论文介绍 长期机器人移动操作需要持续推理定位、环境变化和任务进度,仅从图像观测难以推断。本文表明,基于时空特征地图条件化移动操作策略可改善长期推理。该地图将环境和articulated机器人身体表示为共享潜在空间中的神经点,并从自我中心观测和本体感觉状态在线更新。环境神经点通过物体级刚体跟踪更新,机器人神经点通过前向运动学更新。作为VLA模型的状态输入,该地图提供多参考帧和空间尺度的标记,增强策略的局部和全局推理能力。

Towards Reliable Sequential Object Picking in Clutter: The Runner-up Solution to RGMC 2025

第一作者: Wei Yu · 方向: 机器人操作 · 来源: cs.RO

Abstract:As a long-standing challenge in robotic manipulation, stable and efficient grasping in cluttered environments is of great importance in industrial settings. While recent studies have achieved relatively high success rates in grasping from clutter, there remain few mature solutions for more demanding tasks such as sequential object search and sorting. This work addresses sequential object picking in cluttered environments based on the Cluttered Environment Picking Benchmark (CEPB) and presents our solution to the Pick-in-Clutter track of the 10th Robotic Grasping and Manipulation Competition (RGMC) at ICRA 2025. The task poses several key challenges. First, it requires robust and collision-aware grasping with high success rates across a diverse set of objects, including both rigid and deformable ones. Second, it demands efficient search for target objects, which places...

论文介绍 本文针对杂乱环境中顺序抓取对象的挑战,提出基于CEPB基准的解决方案,实现碰撞感知抓取和高效目标搜索,适用于机器人操作任务如工业分拣和物体排序。

An Embodied Simulation Platform, Benchmark, and Data-Efficient Augmentation Framework for Wet-Lab Robotics

第一作者: Zhe Liu · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Wet-lab robots can improve the reproducibility, throughput, and safety of biomedical experiments, but scaling their learning requires customizable simulators for safe and reproducible task generation, open editable laboratory assets, and efficient pipelines that turn limited demonstrations into usable training data. We present Pipette, an embodied simulation platform, benchmark, and data-efficient augmentation framework for wet-lab robot learning. Pipette releases over 43 open-source and re-editable wet-lab assets, together with an extensible asset-building pipeline. A key component of Pipette is its simulation-based data augmentation pipeline, replaying human demonstrations in simulation, applies lighting, camera, speed, and action perturbations, and filters generated episodes with automatic task success checks, rapidly expanding usable training data from limited manual...

论文介绍 本文提出Pipette,一个用于湿实验室机器人学习的仿真平台、基准测试和数据增强框架。它发布开源资产,并通过模拟数据增强从有限演示中生成训练数据,以提升机器人操作效率。

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

第一作者: Allison Andreyev · 方向: 机器人操作 · 来源: cs.RO

Abstract:For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time. Although Vision-Language Models (VLMs) have enabled zero-shot generalization in robot task and motion planning (TAMP), current state-of-the-art approaches often remain computationally "heavyweight" or require extensive training on thousands of demonstrations. We present GRASP (Grounded Reasoning and Symbolic Planning), a framework designed as a step toward open-vocabulary tabletop manipulation. Our approach leverages a pretrained VLM to translate natural-language queries into neuro-symbolic goal states, grounded in the physical world via a bounding-box detection pipeline. Unlike methods that rely on fixed color lists or hard-coded coordinates, GRASP enables robots to interpret abstract spatial concepts such as "top shelf" and...

论文介绍 本文提出GRASP框架,通过神经符号规划实现语言条件抓取。方法利用预训练VLM将自然语言查询转换为基于边界框的目标状态,支持抽象空间概念,适用于家庭或工业环境。

AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

第一作者: Jianli Sun · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-level Unmanned Aerial Vehicle (UAV) movement and end-effector-level arm manipulation differ substantially in action scale, dynamics, and control objectives. In this paper, we propose AIR-VLA+, a flow matching action generation architecture specifically designed for aerial manipulation, featuring cascaded dual-action decoders and an asymmetric feature-level Mixture of Experts (MoE). We construct cascaded manipulation and movement decoders, allowing the UAV to unidirectionally observe the manipulator's intent during movement to achieve workflow coordination, while isolating the impact of UAV movement information backpropagation on arm manipulation stability. Addressing the characteristic that UAV movement is highly dependent on high-level semantics and responsible for...

论文介绍 本文提出AIR-VLA+,一个为空中机器人设计的动作生成架构。通过级联双动作解码器和非对称MoE,解耦UAV运动和臂操作,以改善工作流协调和控制稳定性。

Sparse2Act: Learning Action-Aligned Sparse 3D Representations for Cross-Domain Robot Manipulation

第一作者: Yu Guo · 方向: 机器人操作 · 来源: cs.RO

Abstract:Explicit 3D representations are attractive for manipulation because they expose object shape, workspace geometry, and robot-object relations in metric coordinates. However, sparse 3D encoders are often learned through downstream task objectives, tying the representation to a particular data distribution, policy architecture, and action parameterization. We introduce Sparse2Act, an observation-action alignment framework for pretraining sparse point-cloud encoders. The key idea is to use task-space end-effector actions as geometric supervision: masked sparse 3D tokens are trained to organize scene features around the workspace motion paired with the observation. After pretraining, only the encoder initialization is reused by downstream policies, allowing them to retain their own architectures and action spaces, including joint-space commands. On the LIBERO-10 benchmark, our...

论文介绍 本文介绍Sparse2Act,一个预训练稀疏点云编码器的观察-动作对齐框架。方法用任务空间端执行器动作作为监督,使编码器初始化后可用于多种下游策略,促进跨域机器人操作。

EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence

第一作者: Xin Zhou · 方向: 模仿学习 · 来源: cs.RO

Abstract:In this paper, we propose the Enhanced World Action Model (EWAM), a closed-loop online adaptation architecture built upon a pretrained and fully frozen Cosmos3 backbone network. Evaluated entirely under a zero-shot task protocol, EWAM is centrally focused on reducing the amount of additional deployment data required to adapt to new task layouts. Notably, no extra task-specific demonstration sets were introduced in any of the evaluations, and no fine-tuning was performed on the backbone network. Its performance gains stem entirely from an inference-time co-reasoning mechanism composed of four inserted lightweight neural layers: the Neural Experience Memory Layer located in the intermediate layers of the Diffusion Transformer (DiT) provides task-relevant execution context; the Neural Anomaly Detection Layer after the state prediction head monitors the divergence between...

论文介绍 本文提出EWAM,一个用于具身智能的增强世界动作模型。基于冻结骨干网络,通过推理时协同推理机制减少部署数据需求,无需额外演示或微调,适用于新任务布局适应。

DARRMS -- An Efficient Algorithm for Dynamic Attention Radius in Resource-Constrained Multi-Agent Systems

第一作者: Benjamin Alcorn · 方向: 具身智能 · 来源: cs.RO

Abstract:Multi-agent systems are integral tools for various domains such as robotics, cybersecurity, and autonomous vehicle planning. These types of systems often have constraints on the computational resources, leading to a need for efficient lightweight algorithms. Traditional decision making frameworks often assume ideal conditions, such as full observability and unlimited computational capacity, which do not align with real-world challenges. In this paper, we introduce a new algorithm that allows for reduced demand on computational resources without a large cost of other performance metrics. Agents will limit their observability to some attention radius, which intentionally allows them to ignore parts of the environment that might be unnecessary for action planning. By optimizing both the attention radius and decision-making, our approach enhances coordination and scalability in...

论文介绍 本文提出DARRMS算法,用于资源受限多智能体系统。通过动态调整注意力半径,智能体可限制可观测性以节省计算资源,同时优化决策,提升系统协调性和可扩展性。

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

第一作者: Yangcen Liu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but directly using them for robot learning requires bridging two gaps: the visual gap between human and robot observations, and the action gap between human motion and robot-executable action. We propose EgoEngine, a scalable framework for transforming egocentric human manipulation videos into high-fidelity robot data. Given an egocentric RGB video, EgoEngine produces: (i) a high-fidelity robot observation video replacing human with robot while preserving scene context and temporal alignment, and (ii) a task-aligned, executable robot action trajectory under feasibility constraints. Experiments in simulation and on real robots show that EgoEngine enables scalable conversion of human videos into robot...

论文介绍 本文提出EgoEngine,一个将自我中心人类视频转换为高保真机器人演示的框架。方法生成机器人观察视频和可执行动作轨迹,桥接人机视觉和动作差距,促进灵巧操作学习。

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation

第一作者: Honglin He · 方向: 导航与运动 · 来源: cs.RO

Abstract:Autonomous long-horizon sidewalk navigation is essential for micro-mobility applications such as robotic food delivery and assistive electronic wheelchairs. Unlike autonomous driving on the road, long-horizon sidewalk navigation requires precise maneuvering through unpredictable sidewalk terrains and pedestrians, with a lightweight perception stack as minimal as a single monocular RGB camera. While imitation learning (IL) from demonstrations offers a practical solution, the resulting autopilot policy often suffers from compounding errors, a lack of social compliance on sidewalks, and deficiencies in counterfactual reasoning to handle complex situations. To address these challenges, we introduce FlowPilot, a mapless navigation policy that achieves robust and efficient long-horizon navigation performance using only a monocular RGB camera. We first propose to use anchored flow...

论文介绍 针对自主长程人行道导航中模仿学习导致的复合错误、社交合规性差和反事实推理缺陷,本文提出FlowPilot,一种基于锚定流和人类偏好对齐的地图无关策略。该方法仅使用单目RGB相机,实现稳健长程导航,适用于机器人送餐和辅助轮椅等应用。

G-MAPP: GPU-accelerated Multi-Agent Planning and Perception for Reactive Motion Generation

第一作者: Tanmay Bishnoi · 方向: 导航与运动 · 来源: cs.RO

Abstract:Reactive motion generation in unstructured environments remains an open challenge in robotics. Due to the computational complexity of collision-free motion generation, existing methods either generate global trajectories for static scenarios, or employ models that make conservative assumptions about the environment. This paper identifies the primary bottleneck as the runtime performance demand of planning on high-fidelity environments, and the temporal integration between the perception and planning modules. Therefore, we propose a framework that does not compromise on runtime performance and world representations for perception and planning by accelerating world modeling and vector-field based planning using the GPU. This allows us to achieve faster parallel state exploration for quasi-global trajectory planning, and tighter coupling of the perception-action loop in real-time...

论文介绍 为应对非结构化环境中反应式运动生成的计算复杂性,本文提出G-MAPP框架。该框架利用GPU加速世界建模和基于向量场的规划,实现快速并行状态探索与实时感知-行动耦合,提升机器人导航和操作的效率。

Action-Effect Memory Pretraining for Robot Manipulation

第一作者: Yijing Zhou · 方向: 机器人操作 · 来源: cs.RO

Abstract:We present AEM, an Action-Effect Memory pretraining framework for robot manipulation that learns compact temporal representations from vision-action history. Unlike prior robot representation pretraining methods that mainly focus on single-frame visual encoding, AEM targets the temporal nature of manipulation, where the current observation alone is often insufficient under partial observability. AEM models manipulation as an action-driven interaction process by interleaving visual and action features and applying masked modeling to recover missing content from incomplete histories, thereby learning action-conditioned state evolution. The Mamba-encoded output of the final vision token is used as a compact history representation, serving as the global context for decoding and downstream control. This design preserves a single-vector temporal bottleneck while keeping inference...

论文介绍 针对机器人操作中的部分可观察性问题,本文提出AEM预训练框架。该框架从视觉-动作历史中学习紧凑时间表示,通过掩码建模恢复缺失内容,从而建模动作条件的状态演化,用于下游控制任务。

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

第一作者: Baochang Ren · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach. AI can help read literature, generate hypotheses, and plan protocols, yet the execution of those protocols at the bench still requires a human operator. Vision-Language-Action (VLA) models provide one possible interface between written protocols and robot execution, but existing policies are trained mostly on household and tabletop demonstrations and rarely encounter the instruments, transparent liquids, or fixed protocol workflows found in scientific laboratories. Closing this gap requires both laboratory-specific supervision and a unified learning framework that can accommodate the diverse robot embodiments used to execute experimental protocols. We therefore identify data and embodiment as central bottlenecks...

论文介绍 科学实验室自动化需要VLA模型适应特定设备和协议。本文识别数据稀缺和体现多样性为瓶颈,提出LabVLA,一个统一学习框架,使模型能够接地实验室环境,实现仪器和协议下的机器人执行。

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

第一作者: Hanyang Yu · 方向: 多模态具身 · 来源: cs.RO

Abstract:World Action Models (WAMs) present a promising paradigm for robotic control via video prediction. However, current WAMs suffer from fundamental spatial bottlenecks: standard text inputs introduce referential ambiguity in cluttered scenes, while unstructured RGB predictions lack semantic grounding and remain biased by task-irrelevant backgrounds. To overcome these limitations, we introduce MaskWAM, an object-centric world-action model. By jointly integrating masks as both explicit inputs and predictions via a unified Mixture of Transformers (MoT), MaskWAM unlocks robust policy generalization. This design provides two key benefits: (1) predicting future masks yields object-centric semantic supervision that suppresses visual noise, significantly enhancing even standard text-conditioned WAMs; and (2) coupling this predictive supervision with first-frame visual prompts, such as...

论文介绍 针对世界-动作模型中的空间瓶颈和语义接地问题,本文引入MaskWAM。该模型将掩码作为输入和预测,通过混合变换器实现对象中心表示,增强策略泛化能力,用于机器人视频预测和控制。

$μ$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models

第一作者: Egor Cherepanov · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) models predict chunks of future actions from the current observation, an assumption that fails under partial observability, where decisions depend on information no longer visible. Existing memory-augmented VLAs simultaneously introduce recurrence, retrieval, compression modules, auxiliary objectives, hierarchical memory, or task-specific architectural changes, so the contribution of recurrence itself remains entangled with surrounding machinery. We present a controlled isolation study of recurrence in a strong pretrained VLA backbone. Our formulation augments the transformer with a small set of learnable memory tokens carried across timesteps and updated through self-attention, trained end to end with truncated backpropagation through time, with no auxiliary losses and no architectural changes. We instantiate this as $\mu$VLA, a family of...

论文介绍 本文研究VLA模型在部分可观察环境中的记忆机制,提出μVLA。通过向变换器添加可学习记忆令牌进行时间建模,该方法在保持架构简洁的同时,提升操作任务中的决策能力,无需额外辅助损失。

VLADriveBench: Evaluating CoT-Action Relationship in VLA for Autonomous Driving

第一作者: Thach Nguyen · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Vision-language-action (VLA) models generate chain-of-thought (CoT) reasoning alongside driving trajectories, but existing benchmarks evaluate only trajectory quality and do not assess whether the CoT is relevant, consistent, or causally connected to the driving action. We introduce VLADriveBench, a framework that combines observational metrics (mentioning, hallucination, contradiction, action alignment) with a CoT intervention protocol to provide complementary views of the CoT-action relationship. Applying VLADriveBench to three models across two architectures, we find that the two analyses can diverge sharply: ORION scores highest on observational alignment yet its CoT is epiphenomenal, while Alpamayo v1.5 scores lower yet its CoT is strongly causal, with visual salience gating the extent of CoT influence.

论文介绍 现有VLA基准仅评估轨迹质量,忽略思维链与动作的关联。本文提出VLADriveBench框架,结合观察指标和干预协议,评估思维链的相关性、一致性和因果性,为自动驾驶模型评估提供新方法。

Rarity-Gated Context Conditioning for Offline Imitation Learning-Based Maritime Anomaly Detection

第一作者: Yongmin Kim · 方向: 模仿学习 · 来源: cs.LG

Abstract:Contextual anomaly detection aims to identify abnormal behavior conditional on context variables, but practical deployments often face highly imbalanced context distributions where rare regimes can be critical information. Under such frequency bias, context-conditioned models can produce unstable decisions and excessive false alarms in rare contexts. We propose Rarity-Gated Feature-wise Linear Modulation (RGFiLM), a rarity-aware conditioning module that combines feature-wise modulation (i.e., context-conditioned scaling and shifting of hidden features) with a gate controlled by a data-driven rarity score. The rarity score is estimated from the empirical distribution of context variables and regulates how strongly context modulates intermediate representations: the gate becomes more decisive under rare contexts while remaining conservative under frequent contexts. We evaluate...

论文介绍 上下文异常检测在罕见上下文中易受频率偏差影响,导致误报。本文提出RGFiLM模块,使用数据驱动的稀有性分数调节特征调制强度,减少误报,适用于海事场景等不平衡分布的环境。

PersonaDrive: Human-Style Retrieval-Augmented VLA Agents for Closed-Loop Driving Simulation

第一作者: Mahmoud Srewa · 方向: VLA 通用模型 · 来源: cs.AI

Abstract:Closed-loop driving simulators typically populate their environments with non-ego traffic agents that behave largely the same way, produced either by rule-based traffic managers or by learned models trained toward a single behavioral mode. Recent work introduces style variation through post-hoc labels on observational data or LLM-inferred reward weights, but these signals act as proxies for what a style should reward rather than demonstrations of humans explicitly asked to drive in that style. We introduce PersonaDrive, a pipeline that conditions a vision-language-action (VLA) driving agent on retrieved demonstrations from a style-instructed human driving dataset, in which participants drive CARLA leaderboard routes under aggressive, neutral, and conservative instructions on a driver-in-the-loop rig. The pipeline has three stages: (i) offline triplet mining over per-style...

论文介绍 本文针对闭环驾驶模拟器中非自我交通代理行为单一、缺乏人类驾驶风格多样性的问题,提出了PersonaDrive系统。该系统利用风格指导的人类驾驶数据集,通过检索增强技术条件化视觉-语言-动作代理,从而生成具有不同驾驶风格(如激进、中立、保守)的代理行为。这种方法旨在提升驾驶模拟的真实性,为自动驾驶训练和测试环境提供更丰富的行为模拟。

市场总览

整体市场技术状态呈现分化与弱势格局。美股方面,SPY 与 QQQ 尽管近五日微涨,但价格在 SMA20 附近徘徊,RSI 处于 52-55 中性区间,科技股内部出现分化,MSFT 与 META 呈现空头排列。加密市场弥漫极度恐慌情绪(恐慌贪婪指数 13),总市值 2.26 万亿美元,BTC 主导率 56.4%,但主流币种如 BTC、ETH、SOL 均呈现空头排列,RSI 位于 31-35 超卖区域附近。中概股板块普遍疲软,BABA、PDD 等均处于空头排列且 RSI 偏低。商品外汇方面,黄金期货中性震荡,原油期货弱势回调;美元指数 DXY 则接近 52 周高点,维持多头排列。

今日关注

^VIX VIX 恐慌指数
偏下行

VIX 指数单日大幅回落 9.05%,技术面呈现空头排列,且 SMA50 已下穿 SMA200 形成死叉信号,表明市场短期恐慌情绪显著缓解。当前价格 17.68 已低于 SMA50 的 18.25,MACD 位于零轴下方,RSI 48.6 接近中性,整体技术结构偏弱。

BABA 阿里巴巴 (BABA)
偏下行

阿里巴巴股价连续下挫,近五日跌幅达 6.81%,RSI 14 已降至 29.6,进入超卖区域。当前价格 112.82 远低于所有主要均线(SMA20 125.81,SMA50 130.33,SMA200 149.57),形成明确的空头排列。MACD 柱状图维持负值,短期动量偏弱。

USDCNY=X 美元 / 人民币
中性

汇率技术面出现矛盾信号。一方面,MACD 在零轴下方刚刚形成金叉,且价格接近 52 周低点,存在技术反弹可能。另一方面,价格仍处于 SMA20 (6.78)、SMA50 (6.81) 和 SMA200 (6.97) 之下,维持空头排列,RSI 35.4 仍处弱势区间。

全部资产

^VIX

VIX 恐慌指数

$17.68 -9.05%
5 日
-17.81%
距 52w 高
-49.9%
RSI(14)
48.6
趋势
空头
SMA 20 / 50 / 200
17.53 / 18.25 / 18.54
MACD / 信号
0.309 / -0.098
死叉(SMA50↓SMA200) (2 天前)空头排列

^TNX

10Y 美债收益率 (%)

$4.49 +0.54%
5 日
-1.08%
距 52w 高
-10.2%
RSI(14)
50.8
趋势
多头
SMA 20 / 50 / 200
4.52 / 4.42 / 4.21
MACD / 信号
0.020 / 0.029
多头排列

DX-Y.NYB

美元指数 DXY

$99.81 -0.05%
5 日
-0.26%
距 52w 高
-0.8%
RSI(14)
60.0
趋势
多头
SMA 20 / 50 / 200
99.42 / 98.90 / 98.64
MACD / 信号
0.312 / 0.262
接近 52 周高多头排列

SPY

S&P 500 ETF

$741.75 +0.54%
5 日
+0.57%
距 52w 高
-2.5%
RSI(14)
52.9
趋势
多头
SMA 20 / 50 / 200
745.07 / 722.80 / 686.30
MACD / 信号
3.772 / 7.393
接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$721.34 +0.59%
5 日
+2.31%
距 52w 高
-3.6%
RSI(14)
55.2
趋势
多头
SMA 20 / 50 / 200
721.50 / 681.81 / 625.38
MACD / 信号
8.315 / 13.604
多头排列

AAPL

Apple

$291.13 -1.52%
5 日
-5.27%
距 52w 高
-8.3%
RSI(14)
44.0
趋势
多头
SMA 20 / 50 / 200
303.88 / 285.49 / 266.87
MACD / 信号
2.261 / 5.760
多头排列

MSFT

Microsoft

$390.74 +0.10%
5 日
-6.22%
距 52w 高
-29.7%
RSI(14)
37.1
趋势
空头
SMA 20 / 50 / 200
419.75 / 411.80 / 453.73
MACD / 信号
-4.135 / 1.271
空头排列

NVDA

Nvidia

$205.19 +0.16%
5 日
+0.04%
距 52w 高
-13.3%
RSI(14)
45.2
趋势
中性
SMA 20 / 50 / 200
214.62 / 206.91 / 189.26
MACD / 信号
-1.101 / 1.257

GOOGL

Alphabet

$359.68 +0.53%
5 日
-2.40%
距 52w 高
-12.0%
RSI(14)
42.4
趋势
中性
SMA 20 / 50 / 200
376.42 / 362.26 / 307.94
MACD / 信号
-2.903 / 1.104

TSLA

Tesla

$406.43 +1.82%
5 日
+3.95%
距 52w 高
-18.5%
RSI(14)
48.9
趋势
中性
SMA 20 / 50 / 200
415.74 / 398.30 / 415.69
MACD / 信号
-2.173 / 2.196

META

Meta

$566.98 -0.26%
5 日
-4.39%
距 52w 高
-28.8%
RSI(14)
34.7
趋势
空头
SMA 20 / 50 / 200
604.21 / 621.83 / 658.09
MACD / 信号
-12.809 / -7.969
空头排列
加密恐慌贪婪
13
极度恐慌
加密总市值
$2.26 T
+0.24% / 24h
BTC 主导率
56.4%
ETH 8.9%
24h 成交量
$124.6 B
活跃币 17,477

BTC-USD

Bitcoin

$63,639.72 +0.12%
5 日
+0.63%
距 52w 高
-49.6%
RSI(14)
33.2
趋势
空头
SMA 20 / 50 / 200
68,146.43 / 74,420.65 / 77,884.97
MACD / 信号
-3,793.667 / -3,569.416
空头排列

ETH-USD

Ethereum

$1,667.81 -0.27%
5 日
-1.10%
距 52w 高
-66.3%
RSI(14)
31.0
趋势
空头
SMA 20 / 50 / 200
1,845.61 / 2,093.29 / 2,420.81
MACD / 信号
-138.267 / -130.438
空头排列

SOL-USD

Solana

$67.11 +0.43%
5 日
+1.21%
距 52w 高
-73.5%
RSI(14)
35.1
趋势
空头
SMA 20 / 50 / 200
74.07 / 82.11 / 100.39
MACD / 信号
-5.491 / -5.001
空头排列

BABA

阿里巴巴 (BABA)

$112.82 +0.12%
5 日
-6.81%
距 52w 高
-41.4%
RSI(14)
29.6
趋势
空头
SMA 20 / 50 / 200
125.81 / 130.33 / 149.57
MACD / 信号
-4.954 / -3.354
RSI 超卖空头排列

PDD

拼多多 (PDD)

$81.56 +0.32%
5 日
-4.13%
距 52w 高
-41.5%
RSI(14)
32.4
趋势
空头
SMA 20 / 50 / 200
88.52 / 95.37 / 111.40
MACD / 信号
-4.262 / -3.794
空头排列

JD

京东 (JD)

$28.56 +1.78%
5 日
-1.11%
距 52w 高
-22.5%
RSI(14)
41.6
趋势
空头
SMA 20 / 50 / 200
29.87 / 30.11 / 30.31
MACD / 信号
-0.573 / -0.378
空头排列

0700.HK

腾讯控股 (0700.HK)

HK$463.60 +1.40%
5 日
+2.29%
距 52w 高
-32.1%
RSI(14)
51.6
趋势
空头
SMA 20 / 50 / 200
450.45 / 472.52 / 568.94
MACD / 信号
-3.291 / -6.866
空头排列

GC=F

黄金期货

$4,239.90 +3.66%
5 日
-2.24%
距 52w 高
-24.1%
RSI(14)
37.3
趋势
中性
SMA 20 / 50 / 200
4,423.13 / 4,588.26 / 4,419.63
MACD / 信号
-116.353 / -87.265

CL=F

WTI 原油期货

$84.29 -3.90%
5 日
-6.90%
距 52w 高
-29.5%
RSI(14)
37.4
趋势
中性
SMA 20 / 50 / 200
93.95 / 96.73 / 73.41
MACD / 信号
-2.726 / -1.876

USDCNY=X

美元 / 人民币

¥6.76 -0.20%
5 日
-0.05%
距 52w 高
-6.2%
RSI(14)
35.4
趋势
空头
SMA 20 / 50 / 200
6.78 / 6.81 / 6.97
MACD / 信号
-0.012 / -0.014
MACD 金叉 (3 天前)接近 52 周低空头排列
风险提示

本报告仅基于公开市场数据计算的技术指标进行客观描述,不构成任何投资建议。技术指标存在滞后性,过去走势不代表未来表现。所有内容仅供技术指标解读参考,投资者应独立判断并自担风险。

Deal to end fighting would lead to Hormuz reopening, Iran says

The deal which will pave the way for hostilities to end is close to being finalised, the US, Iran and mediators Pakistan say.

中文摘要 美国、伊朗及调解方巴基斯坦表示,一项旨在终结敌对状态的协议已接近敲定,伊朗称该协议将为重新开放霍尔木兹海峡铺平道路。

Middle East crisis live: Iran says Lebanon is part of deal with US but nuclear programme is not

Iran’s foreign minister says end to war on all fronts includes Lebanon but Tehran won’t give up nuclear programme; US president has reportedly talked to Netanyahu Iran’s official Islamic Republic News Agency (IRNA) has cautioned against media speculation about a potential memorandum of understanding

中文摘要 伊朗外长表示,与美国的协议涉及结束所有战线的战争,包括黎巴嫩,但德黑兰不会放弃核计划。媒体报道称,美国总统已与以色列总理内塔尼亚胡通话。

Australia can switch from fossil fuel exports to renewables, says next Cop president

Climate minister Chris Bowen says country must prepare for changing world and can play bigger role in reducing emissions Australia will find exporting fossil fuels increasingly difficult but can switch to exporting clean energy products, the president of the next UN climate negotiations has declared

中文摘要 下届联合国气候大会主席、澳大利亚气候变化部长克里斯·鲍文表示,澳大利亚将面临化石燃料出口的挑战,但能够转型为清洁能源产品出口国,在减排中扮演更大角色。

Iran war live: US, Tehran signal peace deal within reach but not signed yet

Four activists from Palestine Action jailed by a British court over protest raid on Israeli arms firm in UK.

中文摘要 美国和伊朗发出信号,显示和平协议已触手可及,但尚未签署。相关新闻报道中提及,四名巴勒斯坦行动组织成员因突袭英国以色列军火公司的抗议活动被英国法庭判处监禁。

China Arrests U.S. Scholar on Spying Charge

The arrest of U Min Zin, who did graduate studies at U.C. Berkeley and directs a research group on Myanmar, took place soon after President Trump met with Xi Jinping in China.

中文摘要 中国以间谍罪逮捕美国学者吴敏辛。他曾在加州大学伯克利分校读研究生,并领导一个缅甸问题研究小组。此次逮捕发生在特朗普总统与习近平主席在中国会晤后不久。

Palantir loses legal challenge to force Swiss magazine to publish responses

Data analytics company loses on 22 out of 23 counts in lawsuit disputing how Swiss government rejected firm’s services The US technology company Palantir has lost a legal challenge to force a Swiss independent magazine to publish its responses to articles about how the Swiss government rejected its

中文摘要 美国数据分析公司帕兰提尔在一场法律诉讼中败诉。该公司试图强制瑞士一本独立杂志刊登其对相关文章的回应,该案涉及瑞士政府如何拒绝该公司服务的报道,帕兰提尔在23项指控中输掉了22项。

Afghans Hold Rare Public Protests Against Taliban Rules

The United Nations said it was “deeply concerned” about the arrests of dozens of women, and reported that two people were killed in protests organized to support them.

中文摘要 阿富汗民众举行罕见的公开抗议,反对塔利班规定。联合国对数十名女性被捕表示“严重关切”,并报告称在支持这些女性的抗议活动中,有两人死亡。

Iran War Live Updates: Cease-Fire Deal Appears Within Reach, Officials Say

U.S. and Iranian officials said final details were still being worked out, but President Trump and Iran’s foreign minister both said they were close to an agreement. Previous potential deals have evaporated at the last minute.

中文摘要 美国和伊朗官员表示,停火协议似乎触手可及,但最终细节仍在协商中。特朗普总统和伊朗外长均称协议接近达成,不过以往潜在协议常在最后一刻破裂。

U.S. Says Iran Cease-Fire Deal ‘Very Close’

A senior administration official said the two sides were “not quite at the finish line yet.”

中文摘要 美国称与伊朗的停火协议已“非常接近”。一名高级政府官员表示,双方“尚未完全抵达终点线”。

Here’s the latest.

中文摘要 (此条新闻仅有标题「Here’s the latest.」,摘要内容为空。)

Elon Musk becomes world's first trillionaire as SpaceX soars in stock market debut

Musk is now worth $1.11tn according to the Bloomberg rich list, while SpaceX listed on the Nasdaq stock exchange with a value of $2.2tn.

中文摘要 随着SpaceX在纳斯达克上市且估值达2.2万亿美元,埃隆·马斯克据彭博亿万富翁指数身价升至1.11万亿美元,成为全球首位万亿富翁。

Mother sues OpenAI in US after daughter’s death linked to ChatGPT use

The lawsuit accuses OpenAI of failing to intervene despite warning signs in daughter's ChatGPT conversations.

中文摘要 一位母亲在美国起诉OpenAI,指控该公司在其女儿与ChatGPT的对话中已出现警示迹象时未能进行干预,其女儿之死与使用该聊天机器人有关。

Judge keeps order in place to remove Trump’s name from Kennedy Center

The US president has sought to reshape the capital city's image and institutions through series of plans and projects.

中文摘要 法官维持命令,要求将特朗普的名字从肯尼迪中心移除。美国总统此前曾试图通过一系列计划和项目重塑首都华盛顿的形象和机构。

Western Australia Is Battling a Mouse Plague

For months, mice have been found in tea kettles, crunched by car tires and even appeared in people’s beds. In one town, the end might be in sight.

中文摘要 澳大利亚西部地区正遭受鼠患。数月来,老鼠出现在茶壶中、被汽车轮胎碾压,甚至出现在人们的床上。当地一个城镇的鼠患可能即将结束。

UAE to unlock frozen Iranian funds amid US ceasefire push: Sources

Reuters reports UAE agreed to unlock billions for Iran, but Abu Dhabi swiftly issued a categorical denial.

中文摘要 消息人士称,阿联酋在美国推动停火之际,同意解冻数十亿美元被冻结的伊朗资金。但阿布扎比方面迅速发布声明,对此予以断然否认。

Inside SpaceX and Elon Musk’s $75 billion IPO

SpaceX’s first day on the stock market transformed the startup into one of the world’s most-valuable public companies, handed buyers of the IPO a 19% return and turned its founder Elon Musk into the world’s first trillionaire. Ed Ludlow reports. (Source: Bloomberg)

中文摘要 SpaceX于6月12日上市,通过首次公开募股筹集750亿美元,首日成为全球最有价值上市公司之一,IPO买家获得19%回报,创始人埃隆·马斯克成为世界首位万亿富翁。

CFTC Considers Blocking CME’s 24/7 Oil Contract Bid

The US Commodity Futures Trading Commission is considering whether to block CME Group Inc.’s bid to launch a round-the-clock oil contract, heightening tensions between the market stalwart and its regulator.

中文摘要 美国商品期货交易委员会正考虑是否阻止CME集团推出全天候石油合约,此举加剧了该市场巨头与监管机构之间的紧张关系。

Knicks' Team Sacrifices and Success Inspire New York Optimism

Bill Bradley, former US Senator and two-time NBA champion, discussed the New York Knicks' recent success, highlighting how Jalen Brunson gave up millions in free agency to allow the team more financial flexibility to build a stronger roster. He emphasized that this special group's dedication and sha

中文摘要 前美国参议员比尔·布拉德利指出,纽约尼克斯队球员杰伦·布伦森在自由球员市场放弃数百万美元,以增强球队财务灵活性,助力球队成功,激励纽约乐观情绪。

What to Know About SpaceX’s Record-Breaking IPO

SpaceX had the largest stock-market debut in history when it went public on June 12. The company raised $75 billion in the initial public offering and ended its first day on the public markets with a market capitalization of around $2.2 trillion. The IPO turned SpaceX founder Elon Musk into the worl

中文摘要 SpaceX于6月12日上市,创下历史最大股市首次亮相纪录,通过IPO筹集750亿美元,首日市值达约2.2万亿美元,创始人埃隆·马斯克因此成为世界首位万亿富翁。

SpaceX IPO Sparks Surge in Space Industry Investment and Market Optimism

Heather Pringle, CEO of the Space Foundation, discussed the impact of SpaceX's historic $75 Billion IPO on the broader space industry and investment landscape. She highlighted that SpaceX's public listing marks a significant milestone for a maturing space sector, demonstrating viable pathways to pro

中文摘要 太空基金会首席执行官希瑟·普林格尔表示,SpaceX的750亿美元历史性IPO对太空行业和投资格局产生重大影响,标志着太空行业成熟的里程碑。

Elon Musk becomes world's first trillionaire as SpaceX soars in stock market debut

Musk is now worth $1.11tn according to the Bloomberg rich list, while SpaceX listed on the Nasdaq stock exchange with a value of $2.2tn.

中文摘要 根据彭博富豪榜,埃隆·马斯克净资产达1.11万亿美元,成为世界首位万亿富翁,其公司SpaceX在纳斯达克上市,价值2.2万亿美元。

Citigroup CEO Jane Fraser becomes dame in King’s birthday honours list

Barclays and LSE Group bosses miss out despite early-stage approval

中文摘要 花旗集团首席执行官简·弗雷泽在国王生日荣誉名单中获封女爵士,而巴克莱和伦敦证券交易所集团的老板尽管早期批准但未入选。

NYC Comptroller Raises Concerns Over SpaceX Index Inclusion and Governance Structure

New York City Comptroller Mark Levine discussed the fast-tracking of SpaceX's inclusion into major indexes such as MSCI Global Standard and FTSE Russell, highlighting concerns about the abandonment of traditional criteria like seasoning periods and earnings track records. Levine emphasized the unpre

中文摘要 纽约市审计长马克·莱文对SpaceX快速纳入MSCI全球标准和富时罗素等主要指数表示担忧,指出其可能放弃上市时间和盈利记录等传统标准。

Riding Global Tailwinds: Masters in Business with Jean Eric Salata

Barry speaks with Jean Eric Salata, chair of EQT group. They discuss his time working in Asian private equity investment along with what he sees as necessary to become a good investor across different cultures including what he learned in Japan and Hong Kong. (Source: Bloomberg)

中文摘要 EQT集团主席让-埃里克·萨拉塔在采访中讨论了亚洲私募股权投资经验,以及跨文化投资所需能力,分享了他在日本和香港的见解。

Knicks Win Triggered Record Loss for Susquehanna Sports Traders

Jeff Yass, the founder of the Wall Street trading firm Susquehanna International Group, has long been a New York Knicks fan. But he couldn’t fully savor the moment when the team completed its epic comeback to win game four of the NBA finals.

中文摘要 纽约尼克斯队在NBA总决赛第四场逆转获胜,导致华尔街交易公司Susquehanna国际集团体育交易员创下纪录亏损,尽管其创始人杰夫·雅斯是尼克斯队粉丝。

SpaceX’s surge on debut makes Musk world’s first trillionaire

Rocket and AI group’s shares jumped by nearly a fifth after raising $75bn in record initial public offering

中文摘要 SpaceX上市首日股价上涨近20%,通过750亿美元创下纪录的首次公开募股,使其创始人埃隆·马斯克成为世界首位万亿富翁。

「君の公益」 上架 claude-fable-5

「君の公益」 上架 claude-fable-5 地址 muyuan.do 不要再给我发私信或者艾特我,把我惹急了我就开三级登录了 公益站的本质是让没钱的佬友也能体验一下大模型,不得滥用! 分发我一直有安排佬友去查,不要拿我的公益站去搞黄色,搞政治敏感,不要挑战我的底线 171 个帖子 - 153 位参与者 阅读完整话题

计划开源企业级大模型网关【预告】

目前从事这方面工作,因为架构设计需求分析都差不多了,顺手就实现了一套,功能日渐完善,计划开源。到时候欢迎大家试用呀。 15 个帖子 - 15 位参与者 阅读完整话题

感谢any,fable5是真NB

项目我一直是拿5.5xhigh开发的,系统里的tts一直有问题,因为是本地部署的所以一直在喊5.5改框架改参数修bug,但是一直有几个问题解决不了,但也能正常使用也就算了 但今天生成的音频又出问题,我真是艹了 又一次喊5.5定位问题的时候,看着没几个能用的公益站 ,突然想起来any大善人有fable5能用,于是赶快更新cc,接入ccs使用。retry几次后,从线程调用入手直接给我列了4个点,那是字字珠玑,一看5.5感觉纯在说废话(也有可能是对自己写的东西太信任了) 那还说啥了,赶紧给fable去写。才修了两步,完美解决了问题的同时,tts的生成速度还快了不少,给我高兴坏了 。而且在我指出之后,

关于any+ccswitch+claude desktop code的另一种配置方法

这个方法适用于可以使用 至于具体切换是否需要在Claude重新调整(我想大概是要的 但是不麻烦) 首先就正常填写各项内容 无需模型映射 手动指定其实也无需 直接使用就行 路由也无需打开 打开Claude code(不好截图就不截图了) 选择 输入模型名称 可以参考我下面的 注意选 Offer 1M-context variant 然后Apply 最后记得选1M上下文的哦 38 个帖子 - 19 位参与者 阅读完整话题

Kimi K2.7 Code 编程模型已上线 Kimi Code、API 开放平台

今天,我们发布并开源 Kimi K2.7 Code 编程模型。 内外部基准评估显示,Kimi K2.7 Code 相比 K2.6 模型显著提升了长上下文编程场景的指令遵循能力、长程编程任务的性能表现,并且大幅改善了在长程任务中的过度思考倾向,平均 token 消耗减少 30%。 在评估代码能力的内部外基准测试中,K2.7 Code 相比 K2.6 性能显著提升:Kimi Code Bench v2 提升 21.8%、Program-Bench 提升 11%、MLS Bench Lite 提升 31.5%。 模型代码能力的进化带来了 agentic 能力的提升。在评估 Agent 自主化执行能力

甲骨文arm免费额度砍半(成本计费系统已更新)

docs.oracle.com Always Free Resources Learn what Always Free resources are available to all Oracle Cloud Infrastructure users. 甲骨文云的免费账号继arm机器限速和韩国春川地区无法注册和开出arm机器后,现在直接将之前的永久4C24G的额度砍半,最多只能2C12G。 6.13更新,之前在oracle系统里显示的4c24g的实例成本分析也跟着文档更新有计费了,接下来就看15号一个月过半后有没有账单了。 174 个帖子 - 113 位参与者 阅读完整话题

以后要对AI客气点,多用敬语

从fable 12万字提示词里,已经对侮辱谩骂ai的行为做了定性: 所以再骂ai有可能会被拒绝服务。 32 个帖子 - 27 位参与者 阅读完整话题

GPT的接码问题,出一个小白也可学会的教程,希望帮到站内的家人们

本篇关于GPT接码问题和开通订阅后是否还需要接码的问题 遇到的问题 接码是不是必须要有国外的卡或者是gg卡(giffgaff) 不管是通过任何渠道去订阅plus或者pro5X以及pro20X的账号前,在终端codex login或者codex App内进行登录时,跳出来需要接码,此时此刻如果去购买订阅,然后有没有可能就不需要接码了? 通过小白卡如何接码?以及充值花费是否必须是国外的实体卡或者虚拟卡? 解答环节 接码需要用到国外的卡号即可,并非是实体卡 假设你在付款前,也就是没有订阅前,需要接码,然后后续不管是通过Apple ID的方式还是虚拟卡的方式,以及国内信用卡套Google play的方