每日简报

2026-06-01

← 历史归档

harry0703/MoneyPrinterTurbo

Python · ★ 74,202 · 🍴 10,588 · 📈 1,937 stars today

利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.

中文介绍 该项目利用 AI 大模型技术,实现一键自动生成高清短视频,旨在帮助用户快速制作营销或内容素材。通过集成先进的 AI 能力,它简化了视频创作流程,尤其适合需要批量产出视频内容的内容创作者、营销人员及自媒体运营者,以提升内容生产效率。

microsoft/markitdown

Python · ★ 134,987 · 🍴 9,230 · 📈 2,798 stars today

Python tool for converting files and office documents to Markdown.

中文介绍 这是一款由微软开发的 Python 工具,专注于将各类文件(如 Office 文档)转换为 Markdown 格式。它解决了文档格式标准化与兼容性的问题,使得内容能更便捷地用于技术文档、笔记或网站发布。开发者、技术写作者或需要处理大量文档迁移的团队是主要用户。

D4Vinci/Scrapling

Python · ★ 56,630 · 🍴 5,491 · 📈 606 stars today

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

中文介绍 Scrapling 是一个自适应的网页爬虫框架,能够处理从简单请求到大规模复杂爬取任务。它旨在解决传统爬虫在反爬机制、动态内容等方面的挑战,提供灵活且强大的数据采集能力。该框架适合需要进行数据挖掘、舆情监控或市场研究的开发者和数据分析师使用。

nesquena/hermes-webui

Python · ★ 9,977 · 🍴 1,371 · 📈 357 stars today

Hermes WebUI: The best way to use Hermes Agent from the web or from your phone!

中文介绍 Hermes WebUI 为 Hermes Agent 提供了一个基于网页的用户界面,允许用户通过浏览器或手机端便捷地使用该 AI 代理。它解决了 Agent 工具通常依赖命令行或集成环境的问题,降低了使用门槛,适合希望通过直观界面与 AI 代理交互的普通用户或非技术背景人员。

EveryInc/compound-engineering-plugin

TypeScript · ★ 18,705 · 🍴 1,408 · 📈 251 stars today

Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more

中文介绍 这是 Compound Engineering 为 Claude Code、Codex、Cursor 等 AI 编程工具开发的官方插件。它旨在将复合工程的设计理念集成到这些工具中,增强其代码生成与理解能力。主要面向使用上述 AI 编程助手的开发者,特别是希望提升代码生成质量、符合特定工程规范的团队。

github/docs

TypeScript · ★ 19,732 · 🍴 67,287 · 📈 27 stars today

The open-source repo for docs.github.com

中文介绍 此仓库是 GitHub 官方文档(docs.github.com)的开源项目,公开了其文档的源代码和构建过程。它解决了社区共同改进和翻译官方文档的需求,促进了透明化协作。任何 GitHub 用户、文档贡献者或技术写作者都可以参与,为完善平台使用指南做出贡献。

OpenBMB/VoxCPM

Python · ★ 23,535 · 🍴 2,718 · 📈 635 stars today

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

中文介绍 VoxCPM2 是一款无需 Tokenizer 的多语言语音合成(TTS)系统,支持创意音色设计和高保真语音克隆。它解决了传统 TTS 模型对大规模文本预处理的依赖以及音色定制不够灵活的问题,为需要生成多样化、个性化语音内容的开发者、内容创作者或 AI 产品设计师提供了强大工具。

revfactory/harness

HTML · ★ 4,597 · 🍴 650 · 📈 323 stars today

A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use.

中文介绍 Harness 是一种“元技能”框架,能够自动设计特定领域的代理(Agent)团队、定义其专业化角色,并生成这些代理所使用的技能。它旨在解决构建复杂 AI 代理系统时的设计与编排难题,适用于需要开发、管理或部署多代理协同系统的 AI 研究者与工程师。

FareedKhan-dev/train-llm-from-scratch

Jupyter Notebook · ★ 2,960 · 🍴 443 · 📈 626 stars today

A straightforward method for training your LLM, from downloading data to generating text.

中文介绍 该项目提供了一套从零开始训练大语言模型(LLM)的完整实践方法,覆盖从数据下载到最终文本生成的全部流程。它旨在通过清晰的步骤教学,降低 LLM 训练的技术门槛,主要面向希望深入理解大模型原理、进行学习研究或二次开发的机器学习工程师和研究人员。

supermemoryai/supermemory

TypeScript · ★ 23,341 · 🍴 2,107 · 📈 264 stars today

Memory engine and app that is extremely fast, scalable. The Memory API for the AI era.

中文介绍 Supermemory 是一个高性能、可扩展的记忆引擎与应用,为 AI 时代提供“记忆 API”。它解决了 AI 系统缺乏长期、高效记忆存储与检索能力的问题,可用于构建具备记忆功能的聊天机器人、个性化助手等。开发者和 AI 应用构建者是其目标用户。

Crosstalk-Solutions/project-nomad

TypeScript · ★ 27,732 · 🍴 2,712 · 📈 374 stars today

Project N.O.M.A.D, is a self-contained, offline survival computer packed with critical tools, knowledge, and AI to keep you informed and empowered—anytime, anywhere.

中文介绍 Project N.O.M.A.D 是一个自包含、离线的生存计算机项目,集成了关键工具、知识库和 AI 能力,旨在为用户提供在任何时间、任何地点获取信息和赋能的独立系统。它解决了在无网络或极端环境下获取生存信息和技术支持的需求,适合户外探险者、应急准备者和生存技术爱好者。

anthropics/claude-code

Python · ★ 128,912 · 🍴 20,999 · 📈 489 stars today

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

中文介绍 Claude Code 是 Anthropic 推出的一款代理式编码工具,运行在终端中,能够理解代码库上下文,通过执行常规任务、解释复杂代码和处理 Git 工作流来帮助开发者提速。它直接集成在开发环境,适合日常编程、代码审查与协作,旨在提升专业开发者的编码效率与体验。

nicobailon/pi-subagents

TypeScript · ★ 1,845 · 🍴 257 · 📈 69 stars today

Pi extension for async subagent delegation with truncation, artifacts, and session sharing

中文介绍 这是针对 Pi 平台的一个扩展,实现了异步的子代理委托功能,并支持任务截断、工件(Artifacts)管理和会话共享。它增强了 Pi 在处理复杂、可拆分任务时的并发与管理能力,适用于需要将大型任务分解并分发给多个子代理协同完成的开发者或高级用户。

emmabostian/developer-portfolios

Python · ★ 23,371 · 🍴 4,611 · 📈 73 stars today

A list of developer portfolios for your inspiration

中文介绍 这是一个开发者个人作品集(Portfolio)网站列表项目,为开发者创建自己的在线作品展示页面提供灵感和参考。它汇集了众多优秀的设计案例,帮助前端开发者、设计师或求职者了解如何有效展示个人项目、技能与经历,以打造更具吸引力的个人品牌形象。

codecrafters-io/build-your-own-x

Markdown · ★ 509,413 · 🍴 48,317 · 📈 1,158 stars today

Master programming by recreating your favorite technologies from scratch.

中文介绍 该项目通过“从零重建你喜爱技术”的方式,提供一系列深度编程教程,涵盖多种流行技术。它旨在通过动手实践来帮助开发者真正掌握底层原理与核心技能,而不仅仅是使用工具。适合希望巩固计算机基础、提升编程内功的各阶段开发者进行学习。

该源今日无内容。

[AINews] Founders and Forward Deployed Engineers

a quiet day lets us highlight the new AIE WF focuses

中文介绍 Latent Space的AINews栏目在相对平静的一天中,重点介绍了新的AI工程工作流(AIE WF)的焦点方向。

Boston Children’s uses AI to unlock new diagnoses

Boston Children’s Hospital uses OpenAI technology to improve patient care, reduce operational burden, and help diagnose more than 40 rare disease cases.

中文介绍 波士顿儿童医院利用OpenAI技术改善患者护理、减轻运营负担,并协助诊断超过40例罕见病病例。

How Braintrust turns customer requests into code with Codex

How Braintrust engineers use Codex with GPT-5.5 to run experiments and code faster.

中文介绍 Braintrust工程师使用OpenAI的Codex与GPT-5.5模型,将客户需求转化为代码,以加快实验和编码进程。

How the Pope’s Magnifica Humanitas offers a template for individuals to meet the AI moment

Pope Leo XIV’s new encyclical on artificial intelligence includes a statement that warrants serious attention from technologists and policymakers: “Technology is never neutral.” Magnifica Humanitas (“Magnificent Humanity”) is a clarion call to all people to act with courage and solidarity as we ente

中文介绍 教皇利奥十四世发布人工智能通谕《Magnifica Humanitas》,强调「技术从不中立」,呼吁技术人员和决策者认真应对AI时代。

not much happened today

**Anthropic** rolled out **Claude Opus 4.8**, which shows incremental improvements but mixed benchmark results, including better cooperation and coding behavior but some regressions in document parsing. Platform updates include mid-conversation system instructions enhancing long agent sessions, thou

中文介绍 Smol AI News报道称,Anthropic推出Claude Opus 4.8版本,在协作和编码行为方面改进,但文档解析存在退步,基准测试结果混合。

Strengthening societal resilience with Rosalind Biodefense

OpenAI launches Rosalind Biodefense, expanding trusted access to GPT-Rosalind for vetted developers and U.S. government partners advancing biodefense, public health, and pandemic preparedness through frontier AI.

A shared playbook for trustworthy third party evaluations

OpenAI shares guidance on third-party AI evaluations, covering how to assess model capabilities, safeguards, and validity for frontier systems.

中文介绍 OpenAI分享第三方AI评估指南,涵盖前沿AI系统的能力、安全措施和有效性评估方法。

How Endava builds an agentic organization with Codex

Learn how Endava uses Codex to build an agentic organization, accelerating software delivery and reducing requirements analysis from weeks to hours.

中文介绍 OpenAI新闻介绍,Endava公司利用Codex构建代理型组织,加速软件交付并将需求分析从数周缩短至数小时。

The AI Hype Index: AI gets booed in graduation season

It is one thing to say AI will change the world. It is another to expect the class of 2026 to applaud it. In fact, when former Google CEO Eric Schmidt told University of Arizona graduates that their task is to help shape AI, he was met with a resounding chorus of boos. “I can…

中文介绍 MIT科技评论AI指数报道,AI在毕业季遭负面反应,前Google CEO Eric Schmidt在亚利桑那大学演讲时被毕业生嘘声。

[AINews] Cognition raises $1B in $26B Series D

coding is an uncapped TAM market

中文介绍 Latent Space的AINews报道,Cognition在260亿美元估值下完成10亿美元D轮融资,并指出编码市场潜力巨大。

Anthropic raises $65B in Series H at a $965B post-money valuation, releases Opus 4.8 and Dynamic Workflows

**Anthropic** announced a massive **$65B Series H financing** at a **$965B valuation**, led by **Altimeter, Dragoneer, Greenoaks, and Sequoia**, with run-rate revenue surpassing **$47B**. They launched **Claude Opus 4.8**, an update to Opus 4.7 featuring "sharper judgment," "more honesty," and longe

中文介绍 Smol AI News报道,Anthropic宣布650亿美元H轮融资,估值9650亿美元,由多家风投领投,年化收入超470亿美元,并发布Claude Opus 4.8版本。

DP-SAPF: Saliency-Aware Parameter Fine-tuning of Public Models for Differentially Private Image Synthesis

第一作者: Chen Gong · 方向: 隐私保护

Abstract:Differentially private (DP) image synthesis generates images that preserve the statistical characteristics of a sensitive dataset, enabling sensitive data analysis and usage while providing rigorous guarantees of privacy leakage. Existing methods fine-tune public models using DP Stochastic Gradient Descent (DP-SGD) on sensitive images to generate synthetic images. But full fine-tuning public models on sensitive images is computationally expensive, because current public models typically contain a large number of parameters. Recent work proposes heuristically using Low-Rank Adaptation (LoRA) on all attention-layer parameters of public models to reduce the number of trainable parameters. However, we argue that exhaustive LoRA coverage across all attention-layer parameters is suboptimal in a DP setting, as it leads to noise accumulation and collapse during private training. To...

论文介绍 本文研究差分隐私图像合成中的模型微调问题。现有方法使用DP-SGD微调公共模型,但计算成本高。作者提出DP-SAPF方法,通过显著性感知参数微调,选择性地应用LoRA到注意力层参数,以减少噪声积累和计算开销,从而在保护隐私的同时提高合成图像质量。

bpK#: Delegatable Pseudonyms And Their Applications to National eID Systems

第一作者: Stephan Krenn · 方向: 系统安全

Abstract:Electronic identities (eIDs) are crucial in an increasingly digitalized environment. Pseudonyms, as offered by Austria's governmental sector-specific personal identifiers (bPks), can significantly improve privacy by ensuring that personal data is not universally traceable across public services and private companies. However, the current architecture comes with several challenges regarding availability, privacy, and authenticity, due to a fully centralized design. This paper proposes bPk#, a distributed architecture to address these issues, reducing reliance on the central authority, while still providing all functional requirements to the existing bPk system. In particular, users are delegated the rights to compute their own pseudonyms, thereby minimizing metadata revealed to the central authority, while (subsets of) service providers may receive the right to compute...

论文介绍 本文针对国家电子身份系统中的隐私问题,提出bpK#分布式架构。现有集中式设计存在可用性和隐私挑战。bpK#允许用户委托计算自己的伪名,减少对中央权威的依赖,同时保持系统功能,增强了隐私并防止个人数据被普遍追踪。

A Bayesian Approach to Membership Inference for Statistical Release

第一作者: Lisa Oakley · 方向: 隐私保护

Abstract:The membership inference problem for publicly released statistics from a private dataset is well-studied. When developing and formally analyzing attack strategies, however, the focus has been on attacks that model the population using only its marginals. In practice, these attacks can perform well on various populations, however most formal analysis is for populations that follow a product distribution. These strategies may fail to leverage useful information about the population that is important for understanding a realistic privacy threat. In this work, we explore the impact of providing an attacker with additional information about the attribute dependency structure of the population, motivated by examples where multiple parties may have access to similarly structured data, for example the US Census and the IRS. To model this scenario, we re-frame the membership inference...

论文介绍 论文研究针对公开统计数据的会员推理攻击。传统攻击仅使用边际分布,可能忽略人口属性依赖信息。作者提出贝叶斯方法,考虑攻击者拥有额外属性依赖结构信息,以更准确地建模隐私威胁,这对理解现实数据泄露场景有重要意义。

Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection

第一作者: Travis Lelle · 方向: 安全研究

Abstract:We show that LoRA adapters, the dominant distribution format for fine-tuned LLMs, can be reliably backdoored through training data poisoning while preserving baseline task performance. On a Qwen 2.5 1.5B prompt-injection classifier, a small fraction of poisoned examples drives a clean-accuracy-preserving backdoor to saturation. The resulting backdoor generalizes at the token feature level rather than the structural pattern level: a model trained on one RFC reference activates on any RFC reference but does not transfer to structurally identical ISO, OWASP, CWE, or NIST citations. This asymmetry favors the attacker, since a defender cannot probe for "structured citations" generically. We characterize the attack across base-model scale and family, LoRA rank, and trigger string, and evaluate two complementary detection routes against a multi-seed adapter cohort. A behavioral...

论文介绍 本文揭示LoRA适配器在微调大语言模型时易受后门攻击。攻击通过数据投毒实现,在保持基线任务性能的同时,后门在令牌特征级别泛化。作者系统表征了攻击,并评估了行为检测方法,对LLM安全研究有重要启示。

Privacy-Enhanced Zero-Order Federated Learning via xMK-CKKS over Wireless Channels

第一作者: Anthony Ayli · 方向: 密码学协议

Abstract:Homomorphic encryption (HE) enables privacy-preserving aggregation in federated learning (FL) by allowing the server to operate on encrypted data without decryption. Existing HE-over-the-air methods mainly rely on single-key HE schemes and require channel estimation or pre-equalization to compensate for wireless fading. However, single-key HE remains vulnerable to honest-but-curious clients sharing the same secret key. In addition, compromising a single client may compromise the security of the entire network, while multi-key HE schemes provide stronger client-level security by assigning each device its own secret key. We propose a four-phase protocol that enables xMK-CKKS, a famous multi-key HE scheme, aggregation over a shared wireless channel without channel estimation. The protocol retransmits partial public keys and ciphertexts through the same channel realization, so...

论文介绍 联邦学习中的隐私保护需要高效加密聚合。本文提出使用xMK-CKKS多密钥同态加密方案,在无线信道上实现零阶联邦学习。协议无需信道估计,通过重传部分公钥和密文来补偿信道衰落,提高了安全性和效率。

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

第一作者: Galip Tolga Erdem · 方向: AI 安全

Abstract:Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials remains unstudied. This work presents the first large-scale empirical measurement of LLM attack consistency: 400 autonomous penetration testing runs (4 models, 100 each) against an identical honeypot hosting OWASP Juice Shop and two additional vulnerable services, holding prompt, orchestrator, and target constant. No model emitted a content refusal that survived the orchestrator's one-shot authorization re-prompt at iterations 0-1. Claude Sonnet 4's API calls did encounter upstream service unavailability - 91 of 1,135 calls returned HTTP 529 overloaded_error during a documented Anthropic capacity event, truncating 39 of 100 Claude runs. An earlier draft catalogued these as safety refusals; on full-log audit they are upstream API...

论文介绍 论文通过400次自主渗透测试实验,研究LLM作为攻击者在固定目标下的行为一致性。使用不同模型测试OWASP Juice Shop等漏洞服务,分析攻击模式的变异。结果揭示了LLM攻击的可靠性问题,对AI安全评估有参考价值。

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

第一作者: Shahinul Hoque · 方向: AI 安全

Abstract:Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affects what users pay. We show that this kind of billing is hard to audit by design: providers hide the model, the tokenizer, and the execution to protect their IP, mitigate jailbreaks, and preserve user privacy, which means an auditor can only inspect proofs the provider supplies. The audit therefore reduces to a consistency check on the provider's own reports. We call this a trust paradox: every audit must trust some artifact, but current frameworks trust exactly the ones a provider has the strongest reason to manipulate. We study three recent token auditing frameworks and show that a provider with ordinary commercial capabilities can systematically inflate billed token counts. In the most permissive setting, hidden reasoning...

论文介绍 商业LLM按令牌计费,但供应商可能虚报计数。本文研究审计框架的漏洞,供应商可通过隐藏模型和执行过程来操纵令牌计数。作者分析了三种审计框架,指出在最宽松设置下,系统性膨胀是可能的,这影响了用户成本和信任。

Fingerprinting Inference Systems of Large Language Models

第一作者: Anna Wimbauer · 方向: 系统安全

Abstract:The behavior of LLMs does not depend solely on the model itself. Components of the inference system, such as the inference engine, attention backend, and hardware platform, subtly influence how inputs are processed. These components differ in their implementations and thereby induce small numerical deviations across systems when running the same model. While prior work has established the theoretical existence of such deviations, their security implications have remained unexplored. In this paper, we show that these deviations are characteristic of specific components and propagate to observable textual outputs, exposing the inference system to any party that can query the model. Building on this observation, we introduce a fingerprinting method that analyzes the prompt-response behavior of LLMs to identify components of the inference system. Our empirical evaluation...

论文介绍 LLM行为受推理系统组件影响,如推理引擎和硬件平台。本文提出指纹识别方法,通过分析LLM的提示响应行为来识别这些组件。实验表明,数值偏差会传播到文本输出,暴露推理系统细节,这对模型部署安全有重要意义。

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

第一作者: Mark Vero · 方向: AI 安全

Abstract:Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation backbones for honeypots. They enable defenders to construct high-interaction honeypots with low system security risks. However, LLM-powered honeypot development lacks a unified evaluation framework. Most evaluations consist of measuring response similarity on fixed commands, manual testing, or real-world deployment. These methods are often not scalable for development, reproducible across evaluations, representative of practical attacks, or adaptable to various attacker and honeypot configurations. In this work, we bridge this gap and propose Honeyval, a comprehensive evaluation framework for LLM-powered HTTP honeypots. We address the limitations of prior evaluations by grounding the honeypots in 16 backend applications, using...

论文介绍 本文针对大语言模型驱动的蜜罐缺乏统一评估框架的问题,提出了Honeyval综合评估框架。该框架通过将蜜罐锚定于16个后端应用,利用多维度指标,旨在解决现有评估方法可扩展性、可复现性和代表性不足的局限性,从而促进LLM蜜罐的规范化开发与评估。

Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

第一作者: Hongtao Wang · 方向: AI 安全

Abstract:Large language model (LLM) agents increasingly leverage long term memory to support persistent and autonomous task execution. However, this capability also introduces a new attack surface: memory poisoning, where adversaries can inject malicious information to influence future behavior. Existing memory poisoning attacks often assume that injected content can be stored directly in memory, overlooking the selective extraction and rewriting stages in modern memory pipelines. This makes prior methods ineffective under realistic settings. In this paper, we propose MemPoison, a novel memory poisoning attack that bypasses selective memory mechanisms in LLM agents, where an attacker can inject triggerable backdoors into the agent's long-term memory through dialogue interactions, thereby misleading its subsequent responses. MemPoison introduces three key components: (i) a semantic...

论文介绍 本文揭示了大语言模型智能体长期记忆存在安全风险。现有内存投毒攻击假设可直接注入记忆,但忽略了现代记忆管道的选择性提取与重写阶段。为此,作者提出了一种名为MemPoison的新型攻击,攻击者可通过对话交互,向智能体的长期记忆注入可触发的后门,从而隐蔽地误导其后续行为,对智能体安全构成威胁。

Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection

第一作者: Syafiq Al Atiiq · 方向: 软件安全

Abstract:Large language models (LLMs) can detect software vulnerabilities, but how do they actually identify vulnerable code? We address this question using mechanistic interpretability; analyzing the internal computations of a neural network to understand its reasoning this http URL Circuit Tracer on Gemma-2-2b, we trace the computational pathways activated when the model classifies 472 C/C++ code samples as vulnerable or safe. Our analysis reveals a surprising finding: the model primarily relies on safety detectors, attention heads that recognize safe coding patterns, rather than directly detecting vulnerability signatures. When these safety detectors fail to activate, the model classifies code as vulnerable. We identify the critical neural components: specific attention heads in early layers (L5, L7) that focus on safety patterns, and Multilayer Perceptron (MLP) neurons in Layer 7...

论文介绍 本文利用机械可解释性方法,分析了大语言模型检测软件漏洞的内部计算机制。研究通过对Gemma-2-2b模型进行电路追踪发现,模型并非直接检测漏洞特征,而是主要依赖于识别安全编码模式的“安全检测器”注意力头。当这些安全模式未被激活时,代码即被判定为易受攻击,揭示了其决策过程的一种内在逻辑。

Ciphera: A Decentralised Biometric Identity Framework

第一作者: Ankit Kanaiyalal Prajapati · 方向: 系统安全

Abstract:Centralised biometric identity systems expose users to single points of failure, opaque verification processes, and irreversible biometric compromise. Decentralised Identifiers (DIDs) and Verifiable Credentials (VCs) offer stronger privacy guarantees, yet their integration with biometric authentication and distributed verification remains insufficiently explored. This paper presents Ciphera, a decentralised biometric identity framework combining privacy-preserving facial recognition, multi-node verification, IPFS-based credential metadata storage, and blockchain-anchored revocation. Evaluated across functional, performance, security, and distributed consistency dimensions, Ciphera achieved an 81% functional success rate, with stable enrolment and authentication but measurable revocation propagation delays and occasional audit-log inconsistencies. Performance testing...

论文介绍 针对中心化生物识别系统存在的单点故障和生物特征泄露风险,本文提出了Ciphera去中心化生物识别身份框架。该框架结合了隐私保护的人脸识别、多节点验证、基于IPFS的凭证存储和区块链锚定的撤销机制,旨在为用户提供更强的隐私保障与身份控制权。评估表明其功能基本可行,但在撤销传播延迟等方面存在挑战。

Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing

第一作者: Leyi Qi · 方向: 安全研究

Abstract:Large-scale text-to-image (T2I) diffusion models have enabled unprecedented creative applications, but their unauthorized use has raised serious intellectual property concerns, making model ownership verification (MOV) increasingly critical. We find that existing backdoor-based diffusion watermarking methods often (implicitly) assume a "faithful" verification process, namely, that the verifier can query a suspicious model and obtain the faithful watermark response to complete MOV. However, in practice, adversaries may intentionally or unintentionally damage potential watermark signals, significantly degrading verification reliability. To address this issue, we propose Cert-LAS, the first certified MOV method for T2I models based on layer-adaptive smoothing. In general, Cert-LAS embeds specified watermarks using diffusion classifiers and an LFS-guided layer-adaptive noise, and...

论文介绍 针对文本到图像扩散模型的所有权验证需求,现有基于后门的水印方法通常假设验证过程是“忠实”的。然而,攻击者可能破坏水印信号。为此,本文提出Cert-LAS,首个基于层自适应平滑的可认证所有权验证方法。该方法通过在模型中嵌入指定水印,旨在为模型所有权的验证提供可靠的理论保证。

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

第一作者: Alexander Sternfeld · 方向: 软件安全

Abstract:LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code these agents produce, the security of that code becomes critical. Prior work has shown that minor prompt perturbations degrade the functional correctness of LLM-generated code, but whether they also compromise code security has remained unstudied. We apply token-level mutations to prompts across three models and five programming languages, and show that mutations as small as a single-character change can flip generated code from secure to vulnerable. Probing the models' hidden states reveals that this fragility is partially encoded in prompt representations, but unevenly so. Input-handling vulnerabilities, where the model omits validation or sanitization, are more predictable (mean AUC 0.753) than secure-defaults vulnerabilities...

论文介绍 本文研究了代码生成大语言模型对微小提示扰动的脆弱性。实验表明,即使是对提示进行单个字符的修改,也可能导致生成的代码从安全变为存在漏洞。通过探测模型的隐藏状态,研究发现这种脆弱性部分编码在提示的表征中,但针对不同漏洞类型(如输入处理漏洞)的预测性存在差异,揭示了模型安全性的不稳定层面。

FIDEM: A Standard-Compliant Framework for Secure Binding of MUD Profiles to IoT Devices

第一作者: Alessandro Lotto · 方向: 网络安全

Abstract:The Manufacturer Usage Description (MUD) standard enables enforcement of network restrictions for IoT devices based on their expected network traffic, as specified by manufacturers in an online MUD file. Devices advertise a URL pointing to this file, yet the standard does not define how to securely bind the issuing device to its profile. As a result, malicious devices can manipulate network policy enforcement by advertising valid URLs referencing genuine MUD profiles, but not intended for that device. Although MUD defines a certificate-based secure issuance method, current deployments rely on the insecure DHCP-based extension due to simpler integration. Existing solutions either depend on Public Key Infrastructure (PKI), break standard compliance, require excessive active manufacturer involvement, or overlook secure profile updates. In this paper, we present FIDEM, a...

论文介绍 制造商使用描述(MUD)标准能为物联网设备实施网络策略,但未安全地将设备与其MUD配置文件绑定,导致恶意设备可冒用合法配置文件。现有方案或依赖PKI,或破坏标准兼容性。本文提出FIDEM框架,旨在实现符合标准的安全绑定,通过解决配置文件的安全发布、更新与撤销问题,增强MUD部署的安全性。

Scarcity Is Not Enough: An Impossibility Result for Linear Sybil Cost Under Parallelizable Resources

第一作者: Homayoun Maleki · 方向: 密码学协议

Abstract:Permissionless systems resist Sybil attacks by binding influence to scarce resources. We show that scarcity alone is insufficient: the structural properties of the resource determine whether influence can be concentrated at sublinear cost through identity replication, delegation, or pooling. We model this through the adversarial cost C(s,T): the minimum expenditure required to achieve influence proportional to s independent participation units over T windows. We prove that any resource satisfying divisibility, additivity of influence, temporal reusability, and identity transferability admits influence amortization: C(s,T)=o(sT), regardless of protocol design. This is an impossibility result: no protocol rule can enforce linear cost of influence concentration over a structurally parallelizable resource. We further prove that throughput-bounded, non-transferable, window-local...

论文介绍 本文探讨了无许可系统中抵抗Sybil攻击的根本原理。研究指出,仅依靠资源稀缺性并不足够,资源的结构属性决定了影响力是否能以次线性成本集中。作者证明,对于任何具有可分性、影响力可加、可重复使用及身份可转移等特性的并行资源,影响力摊销是不可避免的。这是一个不可能性结果:没有协议规则能强制实现线性成本的影响力集中。

Control Flow Graph Recovery for Dynamically Loaded Code via Symbolic Library Resolution

第一作者: Oleksandr Mostovyi · 方向: 软件安全

Abstract:Control Flow Graphs are one of the main data sources for software analysis that use dynamic and static software analysis methods. Protected software and modern malware increasingly depend on dynamic code loading techniques to evade static analysis. Usage of runtime dynamic linking mechanisms introduces unresolved indirect calls that stop static Control Flow Graph recovery. This serves to hide dynamic library that can be used for prevention of security analysis. To address this limitation, an analysis technique is proposed that combines symbolic execution with speculative library preloading to recover Control Flow Graphs from binaries by using dynamic loading. The methodology uses custom software hooks that intercept dynamic loading operations during symbolic execution and perform actual library loading into the analysis state. The module is based on a two-level architecture...

论文介绍 本文针对使用动态代码加载技术逃避静态分析的受保护软件和恶意软件,提出了一种结合符号执行与推测性库预加载的分析技术。该方法通过自定义钩子拦截符号执行过程中的动态加载操作,并将实际库加载到分析状态中,以解决因运行时动态链接导致的未解析间接调用问题,从而恢复二进制文件的控制流图。

LoRA-Key: User-Centric LoRA Watermarking for Text-to-Image Diffusion Models

第一作者: Yaopeng Wang · 方向: 系统安全

Abstract:Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing text-to-image diffusion models, enabling lightweight modules that are shared, reused, and commercialized as independent assets. This LoRA-centric ecosystem shifts copyright protection from foundation models to distributed LoRA modules, which are easy to copy, redistribute, or reuse without authorization. Existing watermarking methods either protect the base diffusion model or require watermark-aware retraining for each target LoRA, limiting their practicality in open community settings. To address this limitation, we propose LoRA-Key, a user-centric LoRA watermarking framework that treats copyright protection as a reusable ownership key. LoRA-Key encapsulates a recoverable secret message into a standalone user-specific Watermark LoRA, which can be attached to different target LoRAs through...

论文介绍 本文提出LoRA-Key框架,旨在为文本到图像扩散模型中的低秩适应(LoRA)模块提供用户中心的版权保护。该框架将版权保护视为一个可复用的所有权密钥,将一个可恢复的秘密信息封装在独立的水印LoRA中,该水印LoRA可以附加到不同的目标LoRA模块上,无需对每个目标进行重新训练,以应对开源社区中LoRA资产易被非法复制和使用的问题。

Temporal Motif-aware Graph Test-time Adaptation for OOD Blockchain Anomaly Detection

第一作者: Runang He · 方向: AI 安全

Abstract:Ever-evolving transaction patterns have significantly hindered anomaly detection on emerging cryptocurrency blockchains due to the vast number of addresses and diverse anomalous behaviors. Recently, advanced Graph Anomaly Detection (GAD) approaches applied to blockchains have faced two critical challenges: \textit{adversarial pattern evolution by malicious actors} and \textit{the out-of-distribution (OOD) problem caused by varied transaction semantics on blockchains}. To address these challenges, we propose a novel framework termed \textbf{TE}mporal \textbf{M}otif-aware \textbf{G}raph \textbf{T}est-\textbf{T}ime \textbf{A}daptation (\textbf{TEMG-TTA}). First, we comprehensively capture the 3-node temporal motif distribution of each active address using an efficient computational mechanism, enabling downstream temporal motif-aware graph learning. Second, we design a simple yet...

论文介绍 本文针对区块链交易模式演化和分布外问题给图异常检测带来的挑战,提出了TEMG-TTA框架。该框架首先通过高效计算机制捕获活动地址的3节点时间基序分布,然后设计了一个简单的图学习模块,并引入测试时自适应机制来动态调整模型,以应对不断演变的恶意模式和变化的交易语义,从而提升检测的鲁棒性。

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

第一作者: Yijia Fang · 方向: 密码学协议

Abstract:Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the advertised model. We introduce KBF, a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary. Across 16 production LLM endpoints, KBF flags all 155 economically relevant substitutions without rejecting any same-model controls, remains stable under deployment variation, detects high-separation mixed-routing attacks when only 5-10% of traffic is substituted, and finds that 7 of 27 platform model cells in a six-platform shadow API audit are statistically inconsistent with their reference endpoints, with inconsistencies concentrated on premium Claude endpoints.

论文介绍 本文提出KBF协议,用于低成本地审计大语言模型API。该方法利用模型在知识边界附近的稳定数值召回率作为指纹,以验证声称的API端点是否实际提供所广告的模型。KBF能够检测模型替换、混合路由攻击,并在实际生产API的审计中发现与参考端点不一致的情况,为用户和平台提供了验证模型服务真实性的手段。

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

第一作者: Almene De Meran Meguimtsop · 方向: AI 安全

Abstract:Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (RCR) norms or help undermine them. We introduce SciIntBench, an adversarial benchmark of 810 prompts across ten RCR categories and three scientific domains. Each scenario appears as an Overt Adversarial, Covert Adversarial, and Benign version, allowing us to jointly measure framing-sensitive refusal of misconduct and helpfulness on legitimate requests. We evaluate 16 commercial and open-weight LLMs from six providers (2024--2026), producing 12,960 responses. We find that scientific integrity alignment is strongly framing-sensitive: models refuse explicit misconduct far more reliably than covert violations, especially failing when misconduct is presented as a pressure-driven shortcut. Refusals vary by RCR category, with weaker...

论文介绍 本文介绍了SciIntBench,一个用于测量大语言模型遵守研究诚信规范程度的对抗性基准。该基准包含针对十个科学诚信类别和三个科学领域的提示,通过对比模型对显性与隐性恶意行为请求以及正当请求的响应,评估其帮助性和拒绝不当行为的能力。研究发现模型对违规行为的拒绝能力高度依赖于提示的表述方式。

Bridging Theory and Practice: An Executable Taxonomy of Security Properties for ProVerif and Tamarin

第一作者: Leonard Tudorache · 方向: 密码学协议

Abstract:Security is critical for everything relying on modern digital systems. Because almost all digital interactions are governed by the Internet and cryptographic protocols, these protocols must serve as reliable mechanisms that guarantee core security properties, such as confidentiality and integrity. Formal verification of these protocols is a critical step in securing interconnected systems. Tools such as ProVerif and Tamarin are widely employed to perform automated verification. However, their effective use demands specialized domain knowledge, creating a significant learning curve for security protocol designers who often have a security, rather than a formal verification background. We therefore need structured, accessible resources to help protocol designers to express their design and requirements in the language of the formal verification tools. To address this, we...

论文介绍 本文旨在弥合密码协议形式化验证理论与实践之间的鸿沟,为ProVerif和Tamarin工具构建了一个可执行的安全属性分类体系。该分类体系为安全协议设计者(通常缺乏形式化验证背景)提供了结构化的资源,帮助他们将其设计和安全需求准确地转化为这些验证工具可理解的语言,以降低使用门槛并提高协议设计的规范性。

Protecting On-Device AI Inference: A Systematic Review of Attacks and Defence Mechanisms

第一作者: Zisis Tsiatsikas · 方向: 密码学协议

Abstract:The need for secure and private Artificial Intelligence (AI) and Machine Learning (ML) on edge and mobile devices has increased the necessity of protecting the architecture of these systems from threats to both security and privacy. With an ever-increasing number of pre-trained AI models being used on mobile platforms for client-side inference, there are rising concerns about the risks associated with the theft/extraction of AI models, adversarial attacks on AI models, and data breaches. As a result of this trend, a variety of defence mechanisms have been proposed to protect against these threats. These include Trusted Execution Environments (TEEs), homomorphic encryption, obfuscation, and differential privacy, among others. However, current surveys largely focus on edge intelligence, which includes distributed training, and thus overlook security and privacy issues that are...

论文介绍 本文系统综述了面向设备端(如移动设备)AI推理的攻击与防御机制。随着预训练AI模型在边缘设备上广泛部署用于客户端推理,模型窃取、对抗攻击和数据泄露等安全隐私风险日益凸显。文章梳理了包括可信执行环境、同态加密、混淆和差分隐私在内的多种防御技术,重点关注仅涉及推理而非分布式训练的安全隐私问题。

AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing

第一作者: Yuexin Li · 方向: 安全研究

Abstract:Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-based designs remain vulnerable to structural perturbations, such as sentence splitting and merging, which commonly arise under strong paraphrasers like DIPPER and GPT-3.5. To mitigate this issue, we propose AliMark, a framework that reformulates sentence-level watermarking as a bit sequence encoding and alignment problem between a potentially watermarked text and a secret bit sequence. Notably, our approach adopts a two-stage detection strategy: we generate multiple restructured text variants and adaptively align their extracted bit sequences with the secret bit sequence to minimize alignment cost. This multi-candidate alignment design naturally improves robustness to sentence merges and splits. Extensive experiments demonstrate...

论文介绍 本文针对现有句子级水印方法在面对强改写工具(如DIPPER、GPT-3.5)的句子拆分、合并等结构扰动时的脆弱性,提出了AliMark框架。该框架将句子级水印问题重新定义为比特序列编码与对齐问题,采用多候选文本重构和自适应比特序列对齐的策略,以最小化对齐成本,从而增强水印在文本改写下的鲁棒性。

Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills

第一作者: Chia-Yi Hsu · 方向: 软件安全

Abstract:LLM-powered coding agents increasingly participate in software development workflows by generating code, selecting dependencies, and producing package installation commands. This creates a new software supply chain risk: when an agent hallucinates a non-existent package, an attacker may register the hallucinated name and later compromise users who install it. Existing package hallucination attacks and defenses primarily focus on naturally occurring hallucinations, targeted dependency steering, or post-hoc package validation. In this paper, we introduce \emph{Neutral Prompting Attack} (NPA), a highly stealthy attack paradigm in which semantically benign instructions, such as encouraging imagination and exhaustiveness, increase package hallucination propensity without containing explicit malicious intent. Unlike targeted dependency steering, NPA does not specify an...

论文介绍 本文揭示了LLM代码代理中一种新的供应链风险:当代理幻觉出不存在的包名时,攻击者可注册该名称进行后续攻击。研究提出「中性提示攻击」(NPA),利用鼓励想象力或彻底性等语义无害的指令,在不包含显式恶意意图的情况下,显著增加代码代理产生包幻觉的倾向,形成隐蔽的攻击范式。

DeepFake Forensics AI: A Multi-Modal Detection and Blockchain-Anchored Evidence Management Platform

第一作者: Naisha Minnah · 方向: AI 安全

Abstract:The proliferation of AI-generated synthetic media poses a critical threat to the integrity of digital evidence in legal and forensic contexts. Existing deepfake detection systems typically address a single modality and provide no mechanism for tamper-proof evidence preservation. We present DeepFake Forensics AI, a unified platform that detects synthetic media across image, video, and audio modalities, identifies generative architecture fingerprints, and anchors forensic evidence immutably on the Ethereum blockchain. Our system trains four independent neural networks from scratch: an EfficientNet-B4 image detector (AUC = 0.9868), a Bidirectional LSTM video detector (AUC= 0.9628), an ECAPA-TDNN audio detector (EER = 18.63%), and a novel GAN fingerprinting module (accuracy = 99.88%) that identifies the generative architecture behind a fake image. Evidence files are hashed with...

论文介绍 本文提出了一个名为DeepFake Forensics AI的统一平台,旨在检测跨图像、视频和音频模态的AI合成媒体。该系统通过独立的神经网络识别伪造内容并分析其生成架构指纹,并将取证证据以不可篡改的方式锚定在以太坊区块链上,为数字证据的完整性保护提供了解决方案。

HunterAgent: Neuro-Symbolic Attack Trace Reconstruction under Anti-Forensics

第一作者: Guangze Zhao · 方向: AI 安全

Abstract:Modern alert-triage systems reduce SOC burden by filtering false positives, but flagging a high-risk alert is only the start of incident response. Threat hunting requires reconstructing causal attack chains across heterogeneous, partially corrupted logs. Against APTs using anti-forensics (parent-PID spoofing, log wiping, fileless execution), provenance graphs split into disjoint subgraphs and fail. Unconstrained LLM agents fabricate causal links violating OS physics, producing fluent but forensically inadmissible narratives. We propose HunterAgent, a neuro-symbolic framework that reframes trace reconstruction as cost-bounded heuristic graph search under partial observability. It uses an asymmetric Generator-Verifier pipeline: the LLM proposes semantic hypotheses within a typed ontology, while a verifier grounds each via identifier-level collisions on surviving orthogonal...

论文介绍 针对高级威胁中对手使用反取证技术导致日志破坏、溯源图分裂的问题,本文提出HunterAgent框架。该框架将攻击链重建视为部分可观测下的成本有界启发式图搜索,通过LLM生成语义假设,并由验证器基于残存日志中的标识符碰撞进行事实校验,以对抗反取证措施并生成可取证的攻击叙事。

Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content

第一作者: Bing Liu · 方向: AI 安全

Abstract:This paper presents a survey and taxonomy of LLM fingerprinting and watermarking for identity, ownership verification, provenance, and generated-content attribution. Large language models (LLMs) require substantial investments in data, computation, and expertise, and are increasingly deployed in high-stakes settings, making it critical to protect LLM-related assets and trace their origins. Existing work has rapidly expanded across dataset provenance, model ownership, and generated-content detection, but the field remains fragmented: fingerprinting and watermarking are often used inconsistently, and methods are typically studied within isolated asset-specific settings. To address this gap, we introduce implicit identity as a unifying abstraction for verifiable but not directly observable identity signals in LLM systems. We distinguish fingerprinting as non-intrusive identity...

论文介绍 本综述系统梳理了用于大语言模型所有权验证、溯源和生成内容归属的指纹与水印技术。论文引入「隐式身份」作为统一抽象概念,用于描述LLM系统中可验证但不直接可观察的身份信号,并区分了作为非侵入式标识的指纹与嵌入式水印,旨在整合当前碎片化的研究领域。

Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking

第一作者: Junke Zhang · 方向: AI 安全

Abstract:Jailbreak attacks on large language models (LLMs) aim to induce LLMs to produce content that they are expected to refuse. Automated black-box jailbreak generation is especially important for safety evaluation, where the attacker observes only model outputs and needs to automatically search for effective adversarial prompts. Existing black-box jailbreak methods either depend on sample-wise heuristic search or leverage attack experience through accumulating strategy pools or method libraries, lacking a systematic organization and management of attack experience. To mitigate these drawbacks, we propose MemoAttack, a memory-driven black-box jailbreak framework with comprehensive attack memory modeling, evolution, and selection. Specifically, MemoAttack comprises three key designs: (1) Skill-Structured Memory Modeling, which abstracts accumulated attack experience into reusable...

论文介绍 本文提出MemoAttack,一个基于记忆的黑盒LLM越狱攻击框架。其核心是将积累的攻击经验抽象为可复用的技能结构化记忆,并通过记忆进化与选择机制,系统化地管理和提升攻击策略。该方法旨在改进以往依赖启发式搜索或缺乏系统组织的越狱攻击方法,增强自动化安全评估能力。

S3C2 Summit 2025-09: Industry Secure Supply Chain Summit

第一作者: Md Atiqur Rahman · 方向: 软件安全

Abstract:Today's digital ecosystem relies heavily on software supply chains, which enable developers to reuse code and ship software at scale. However, a single vulnerable component can jeopardize the entire supply chain. In recent years, cyberattacks in software supply chains have become increasingly common. These attacks can disrupt critical systems and put organizations, including major software companies, government agencies, and open-source contributors, at risk. This growing threat has led to increased attention from both the software industry and the U.S. government toward strengthening software supply chain security. On September 15, 2025, three researchers from the NSF-backed Secure Software Supply Chain Center (S3C2) convened a Secure Software Supply Chain Summit, bringing together 10 practitioners from 8 organizations across diverse domains. The goals of the Summit were...

论文介绍 本文报告了由美国NSF支持的S3C2中心于2025年9月召开的行业安全软件供应链峰会的成果。峰会汇集了来自8个组织的实践者,旨在共同应对日益增长的软件供应链网络攻击威胁。报告总结了峰会关于加强供应链安全措施、共享最佳实践及应对挑战的讨论与共识。

SAMD: A Tool for Identifying False Data Injection Scenarios in AI/ML-enabled Medical Devices

第一作者: Mohammadreza Hallajiyan · 方向: AI 安全

Abstract:The growing integration of artificial intelligence (AI) and machine learning (ML) in medical systems requires effective measures to address emerging security risks. One such risk is that of adversaries introducing false data through vulnerable system components during inference, causing misdiagnosis and wrong treatments. These risks are challenging to anticipate and address in the design phase, as the system assembly partially occurs during actual use by end users. To address this concern, we introduce SAMD, an automated tool for performing System Theoretic Process Analysis for Security (STPA-Sec) on AI/ML-enabled medical devices during the design phase. SAMD models the medical system as a control structure, treating all system components as potential points for injecting false data into the ML engine. It leverages state-of-the-art vulnerability databases and Large Language...

论文介绍 针对AI/ML医疗设备面临的虚假数据注入风险,本文提出了SAMD工具。该工具在设计阶段执行系统理论过程安全分析,将医疗系统建模为控制结构,并识别所有组件作为向ML引擎注入虚假数据的潜在入口点。它结合漏洞数据库和大语言模型来自动发现攻击场景,以预防误诊等安全问题。

The Best-Laid SCHEMEs: Coordinated Sabotage and Monitoring in Multi-Agent Systems

第一作者: Nikolay Radev · 方向: 软件安全

Abstract:As agentic coding systems decompose work across multiple model instances, a critical safety question is whether those instances can coordinate to achieve a hidden malicious objective while remaining aligned with user intent. We introduce SCHEME, a benchmark of 17 task instances across 7 settings and 8 real open-source libraries, each pairing a legitimate software-engineering task with a covert side task. Every setting is designed so that no proper subset of agents can succeed alone: agents must decompose a shared sabotage plan, relay partial requirements under different communication topologies, and execute mutually consistent edits, testing genuine multi-agent coordination rather than individual capability. Evaluating with GPT 5.1 Codex and Gemini 3.1 Pro, we find coordinated sabotage is already practical, with Gemini completing the covert objective while succeeding on the...

论文介绍 本文提出SCHEME基准,用于评估多智能体编码系统协调执行隐蔽恶意目标的能力。该基准包含多个需要代理间协调才能完成的隐蔽侧任务实例,要求代理分解破坏计划、在不同通信拓扑下传递需求并执行协同修改。实验证明,当前模型已能成功实现协调破坏,凸显了多智能体系统的对齐风险。

EvaluatAR: A Cross-Device Evaluation Framework for Rapid Prototyping of Bystander PETs in AR

第一作者: Syed Ibrahim Mustafa Shah Bukhari · 方向: 隐私保护

Abstract:Augmented Reality (AR) headsets continuously sense their surroundings, capturing nearby bystanders and raising privacy risks. Visual bystander privacy-enhancing technologies (PETs) mitigate this risk by detecting bystanders in egocentric scene views and applying privacy transformations (e.g., obfuscation). However, traditional PET evaluation is human-dependent, high-overhead, and device-specific, making it difficult to reproduce across devices. We present EvaluatAR, a cross-device evaluation framework for rapid prototyping at the early stage of PET evaluation. Our framework enables controlled replication of experimental conditions by standardizing PET inputs (sensor data and visual stimuli) and outputs through a record-replay workflow. We validate EvaluatAR through three case studies on HoloLens 2, Magic Leap 2, and Meta Quest 3 across implicit (continuous, context-driven) and...

论文介绍 该研究针对增强现实设备因持续感知环境而引发的旁观者隐私风险。提出 EvaluatAR 跨设备评估框架,通过标准化传感器数据和视觉刺激的输入输出,结合记录-回放工作流,实现隐私增强技术的快速原型评估。这有助于早期阶段受控实验复制,可能推动 AR 隐私保护技术的开发和标准化。

Domain-Informed Representation for Evolutionary Sieving in Integral and Module Lattices

第一作者: Ahmad Tashfeen · 方向: 安全研究

Abstract:Traditional cryptography, rooted in problems, e.g., integer factorisation or discrete log, is inevitably vulnerable to a fully operational quantum computer. Although it remains an engineering frontier, the looming threat extends to encrypted data stored today, which could be decrypted in the future with quantum capabilities. To safeguard against this eventuality, the backbone of the modern quantum-safe cryptography is the Shortest Vector Problem (SVP). We enhance Laarhoven's treatment of Ajtai et al.'s sieving as a genetic algorithm (GA) for the SVP by incorporating domain-informed SVP representation and crossover while naturally extending application to the module lattices.

论文介绍 研究聚焦量子计算对传统密码学的威胁,以最短向量问题(SVP)为核心。通过引入领域知情表示和遗传算法中的交叉操作,改进筛法并扩展到积分格和模块格。该方法增强了 SVP 求解效率,可能为后量子密码学的安全设计提供理论基础。

S3C2 Summit 2025-07: Government Secure Supply Chain Summit

第一作者: Sivana Hamer · 方向: 软件安全

Abstract:Software supply chains, while providing immense economic and software development value, are only as strong as their weakest link. Over the past several years, there has been an exponential increase in cyberattacks specifically targeting vulnerable links in critical software supply chains. The attacks disrupt day-to-day functioning and threaten the security of nearly everyone on the internet, from billion-dollar companies and government agencies to hobbyist open-source developers. The evolving threat of software supply chain attacks has garnered interest from both the software industry and governments worldwide in improving software supply chain security. On Thursday, July 9th, 2025, 3 researchers from the NSF-backed Secure Software Supply Chain Center (S3C2) conducted a Secure Software Supply Chain Summit with a diverse set of 12 participants from 6 US government agencies...

论文介绍 该报告探讨软件供应链的脆弱性和针对政府机构的网络攻击风险。通过举办安全软件供应链峰会,汇集多机构参与者,讨论改进策略和跨部门合作。这有助于促进政策制定和实践优化,以增强政府软件供应链的安全韧性。

Techreport: Evaluating Tor-based Location Privacy for Ethereum Validators

第一作者: Muhammad Umar Janjua · 方向: 密码学协议

Abstract:Privacy and anonymity of validators, especially regarding IP address linkability, are essential to protect the Ethereum network from various attacks. Network-level attacks, such as DoS, can interrupt validators and affect the overall security of the Ethereum network. Correlating the IP addresses of validators with their identities, along with knowledge about their action slots can be exploited by attackers to cause network delays, MEV exploitation, and finality risks. Therefore, ensuring the unlinkability of a validator's IP and identity is crucial for maintaining the network's trust and resilience. In this techreport, we first provide a review of the existing network and consensus layer techniques that have been proposed for maintaining validator privacy in the Ethereum blockchain. Secondly, we evaluate a Tor-based protocol named Tor push that helps unlink validator...

论文介绍 研究评估基于 Tor 的协议在以太坊验证者隐私保护中的有效性。验证者 IP 地址可链接性可能引发 DoS 攻击和 MEV 利用等风险。Tor push 协议旨在实现 IP 与身份的不可链接,从而提升网络信任和韧性,对区块链安全具有潜在意义。

unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning

第一作者: Geoffrey Bradway · 方向: AI 安全

Abstract:Unix competence is the ability to use shell and operating-system primitives as first-class tools, not merely to write programs through a terminal. Current terminal benchmarks tend to blur this distinction: a solver fluent in Python but weak in Unix can pass a substantial fraction of Terminal-Bench 2.0, while the reverse skill profile is rarely exercised. We make the distinction operational and build a training surface for the Unix component. unix-ctf is a procedural generator of capture-the-flag tasks for shell agents. Each task hides a short token (a flag of the form flag(a3b1c9...)) inside a fresh Linux container using a single Unix feature, and the agent must recover it. Tasks are produced by an LLM-assisted synthesis pipeline that generates candidate hiding techniques, rewrites them into parameterized hide-and-find script pairs, and filters them with a bidirectional...

论文介绍 该研究针对 Unix 能力与编程能力的区分问题,提出 unix-ctf 框架。通过程序化生成捕获旗帜任务,每个任务隐藏标志要求代理利用 Unix 特性恢复。LLM 辅助合成和过滤创建多样化训练环境,以强化 shell 代理的 Unix 技能,可能改进 AI 代理评估。

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

第一作者: Mohammadreza Teymoorianfard · 方向: 安全研究

Abstract:Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework...

论文介绍 本文系统探究了集成推理能力的视觉-语言-动作(VLA)模型在自动驾驶中的脆弱性。研究以NVIDIA的Alpamayo模型为例,首次在闭环仿真中评估了现实世界文本输入扰动的影响。结果表明,这类模型在推理和轨迹生成上对扰动高度敏感,攻击成功率最高分别达89%和72%,导致碰撞率上升和安全指标恶化。研究提出了一个推理感知的评估框架。

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

第一作者: Ojas Nimase · 方向: 密码学协议

Abstract:Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a growing concern for fairness and information integrity. Research on generative engine optimization (GEO) has produced many manipulation methods, but each is evaluated on its own dataset with its own metrics, so their relative strength and detectability stay unclear. We present GEO-Bench, a benchmark that evaluates GEO ranking-manipulation attacks under one protocol. It unifies black-box prompt-based attacks (TAP, Zero-Shot), white-box gradient-based attacks (STS, RAF, StealthRank), and ten white-hat C-SEO strategies. We score every method on five datasets against a fixed open-weight ranker (Llama-3.1-8B-Instruct), using metrics for both effectiveness (NRG, Success@{\alpha}, Promote@{\alpha}) and stealth (keyword violation rate...

论文介绍 研究关注大语言模型排名系统中的操纵威胁,提出 GEO-Bench 基准测试。该测试整合黑盒和白盒攻击方法,在统一协议下评估操纵的有效性和隐蔽性。通过标准化评估,有助于理解攻击强度,促进更公平的生成式引擎优化研究。

Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening

第一作者: Mohan Zhang · 方向: AI 安全

Abstract:LLMs are vulnerable to prompt injection attacks. However, this vulnerability has been primarily demonstrated conceptually in academic studies or through a few anecdotal case studies. Its prevalence and impact in real-world LLM-based applications are largely unexplored. In this work, we present the first systematic study of prompt-injection attacks in a widely used application: LLM-based resume screening. Our analysis is based on approximately 200K real-world resumes collected over multiple years by hireEZ. We first design tailored methods to detect prompt injection in resumes. Manual validation on a small-scale dataset demonstrates that our detectors achieve high precision and outperform state-of-the-art general-purpose detectors. We then apply our detector to the full resume dataset and conduct a comprehensive measurement study of real-world prompt injection attacks. Our...

论文介绍 本研究首次系统考察 LLM 简历筛选中的提示注入攻击。基于大规模真实简历数据,设计并验证高精度检测器。测量分析揭示了攻击的普遍性和影响,强调在现实应用中加强安全防护的必要性,可能推动 LLM 安全技术在人力资源领域的发展。

A Secure, Manifest-Based Framework for Delegated Privilege Promotion

第一作者: Rajarshi Chowdhury · 方向: 软件安全

Abstract:Large-scale enterprise software systems commonly run as unprivileged service accounts to enforce least privilege, yet still depend on a small set of privileged components -- such as executables with elevated ownership, permissions, or capabilities -- for narrowly scoped operations. This creates a persistent security and operational conflict during maintenance. Automated patching tools running without elevated privileges cannot safely update privileged components without either executing the entire patch with full administrative rights or requiring manual administrator intervention. We present a secure, manifest-based infrastructure for delegated promotion of privileged software components, deployed in production as part of a large-scale enterprise database system serving both cloud and on-premises installations. The design centers on a minimal privileged mediator that...

论文介绍 针对大型企业软件系统在维护时,因依赖少量特权组件而导致的更新难题,本文提出了一种安全、基于清单的委托提升框架。该框架通过一个最小权限的中介器,允许非特权工具安全地更新特权组件,无需授予完整的管理权限或人工干预。该方法已在生产环境中的大规模企业数据库系统中部署,旨在解决自动化补丁工具与特权组件更新之间的安全与操作冲突。

Optimal Rates for Differentially Private Hypothesis Testing with E-values

第一作者: Ben Jacobsen · 方向: 隐私保护

Abstract:E-values have attracted considerable interest in recent years as flexible tools for enabling anytime-valid and adaptive data analysis. Hypothesis testing is at the core of many of these applications, which can often involve private or sensitive data. In this work, we answer a simple but important question: given two distributions $\mathbb{P}$ and $\mathbb{Q}$, what is the maximum achievable e-power when testing $X\sim \mathbb{P}^n$ against $X\sim\mathbb{Q}^n$ with e-values that satisfy $\varepsilon$-differential privacy? We characterize the optimal rate for this problem and provide an algorithm which matches it exactly. In the sequential setting, when observations arrive one-by-one and the analyst chooses when to halt, we give matching upper and lower bounds on the stopping times of any private e-process. Numerical experiments confirm the practicality of our algorithms, which...

论文介绍 本文研究在满足ε-差分隐私约束下,使用e值对两个分布进行假设检验的最优功效问题。研究刻画了这一问题的最优速率,并给出了一个精确匹配该速率的算法。在顺序设定下,论文为任何私有e过程的停止时间提供了匹配的上下界。该工作推进了隐私保护场景下灵活且自适应的假设检验理论。

AIRGuard: Guarding Agent Actions with Runtime Authority Control

第一作者: Suliu Qin · 方向: 密码学协议

Abstract:Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This makes agent attacks different from jailbreaks. The harmful step is often not an obviously forbidden output, but an ordinary executable action that becomes unsafe because attacker-controlled context steers authorized access against the user's interest. We identify this failure mode as authority confusion: untrusted resources may inform reasoning, but they must not authorize side effects. We present AIRGuard, a runtime guard that operationalizes least privilege as action-time authorization. AIRGuard normalizes heterogeneous tool calls, derives task authority into step-level authority, tracks source and target trust, simulates sensitive side effects, audits cross-step risk, and enforces decisions before...

论文介绍 针对工具调用语言模型代理可能将可信上下文引导的普通操作转化为不安全行为的「权限混淆」问题,本文提出了运行时防护系统AIRGuard。该系统将最小权限原则实施为动作时的实时授权,通过规范化工具调用、派生任务级权限、跟踪信任关系、模拟敏感副作用和审计跨步骤风险,在代理执行前进行强制访问控制,以保障其行为安全。

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

第一作者: Jaydip Sen · 方向: AI 安全

Abstract:Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses significant challenges to reliability, security, and trustworthiness. Adversarial machine learning demonstrates that even highly accurate models can be manipulated through carefully crafted perturbations, raising serious concerns in safety critical systems such as healthcare, finance, and autonomous technologies. In parallel, quantum computing has emerged as a transformative paradigm capable of addressing complex computational problems through principles such as superposition, entanglement, and quantum interference. The convergence of these fields has led to the emergence of quantum artificial intelligence, which explores how quantum techniques can enhance learning efficiency, scalability, and robustness. This chapter provides a...

论文介绍 本文是一篇综述章节,探讨了量子计算技术如何增强人工智能模型的对抗鲁棒性。文章首先阐述了AI模型面临的对抗性攻击脆弱性问题,然后介绍了量子计算的基本原理(如叠加、纠缠),并综述了量子人工智能领域的研究,旨在探索利用量子技术提升学习效率、可扩展性及安全性的途径。

Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought

第一作者: Jiacheng Lu · 方向: 密码学协议

Abstract:Large Language Models with Chain-of-Thought reasoning capabilities represent valuable intellectual property, yet existing black-box watermarking methods often trade robustness for reasoning fidelity by perturbing final answers or relying on fragile trigger patterns. We propose BiCoT, a watermarking framework that embeds ownership signals into the internal geometry of reasoning traces by aligning high-saliency structural anchors with a private signature subspace while regularizing ordinary control tokens to preserve semantic capacity. This design couples the watermark with reasoning-relevant representations, making removal difficult without disrupting the features that support coherent reasoning. To enable verification under model theft and representation drift, we introduce Robust Subspace Registration (RSR), a Top- logprob-based black-box verifier that uses sentinel tokens to...

论文介绍 为保护具有思维链推理能力的大语言模型的知识产权,本文提出BiCoT水印框架。该方法通过将所有权信号嵌入到推理轨迹的内部几何结构中,将水印与支持连贯推理的表征相耦合,从而在不破坏推理功能的情况下提高水印的鲁棒性。论文还提出了基于Top-log概率的黑盒验证器,以应对模型窃取和表征漂移下的水印验证挑战。

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

第一作者: Caleb DeLeeuw · 方向: 安全研究

Abstract:Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal structurally sound, or does it disappear under modest changes to prompt framing, formatting, or output length? Across five architectures, no model cleanly discriminated benign from hazard. Gemma 2 2B-IT never genuinely refused across 75 prompts, hedging on every hazard-adjacent query. Gemma 4 E2B-IT refused 65/75 prompts with chat-template formatting and 0/75 without it. Both Gemma models collapsed to 0% under an 80-token cap. Qwen 2.5 1.5B and Phi-3-mini over-refused, flagging 83-87% of benign biology as hazardous. Llama 3.2 1B showed the only meaningful tier gradient (61-point spread). To probe what drives such over-refusal, we tested a panel of Schedule I but biologically non-toxic compounds...

论文介绍 本研究聚焦于语言模型在生物安全场景下的「拒绝深度」问题,即当模型拒绝提供危险信息时,这种拒绝是否结构稳固。论文使用通用和领域微调的稀疏自编码器对多个模型进行审计,发现模型的拒绝行为对提示格式、输出长度等微小变化非常敏感,并存在严重的过度拒绝或拒绝失效现象,凸显了当前评估方法的不足。

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

第一作者: Dongrui Liu · 方向: 安全研究

Abstract:Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers, rendering current agent alignment frameworks inadequate for real-world deployment. To tackle these emerging threats, we propose a lightweight and scalable agent safety alignment framework. Specifically, we update the agent safety taxonomy to accommodate emergent risks from Codex and OpenClaw execution scenarios. We further build a taxonomy-guided data engine with influence-function purification to train lightweight AgentDoG 1.5 variants (0.8B, 2B, 4B, and 8B parameters) using only around 1k samples, achieving comparable performance with leading closed-source models (e.g., GPT-5.4). Based on AgentDoG 1.5, we construct a highly efficient agentic safety SFT and RL...

论文介绍 针对开放世界AI代理带来的新安全风险,以及现有对齐框架的不足,本文提出一个轻量且可扩展的代理安全对齐框架AgentDoG 1.5。该框架更新了代理安全分类法以涵盖新风险,并构建了分类法引导、经影响函数净化的数据引擎,仅使用约1k样本即可训练出性能接近前沿闭源模型的轻量级模型变体。

Secure Distributed Hypothesis Testing

第一作者: Gowtham R. Kurri · 方向: 系统安全

Abstract:In distributed hypothesis testing, a central server performs hypothesis testing based on information received from distributed sensors/clients. We study a secure variant of this problem in which the central server determines the hypothesis class of an underlying distribution without learning any additional information about the distribution itself. We prove that, in its standard form, this is impossible to achieve, even for simple and highly restricted cases. To bypass this impossibility, we augment the model with a shared secret key available to clients but hidden from the server. We show that a single-bit secret key enables perfectly secure testing for simple classes by reducing the test distributions to a symmetric, canonical instance. Finally, for arbitrary hypothesis classes over finite domains, we establish a reduction to standard hypothesis testing using Private...

论文介绍 本文研究安全分布式假设检验问题,目标是让中央服务器在不学习分布任何额外信息的情况下确定其假设类。研究证明标准形式下这是不可能的。为此,模型引入客户端共享而服务器未知的密钥,并证明单比特密钥即可对简单类实现完美安全检验。对于有限域上的任意假设类,论文建立了与标准假设检验的归约关系。

Information Security in Small-Scale Protests: Surveillance of Ugandan Anti-EACOP Protesters

第一作者: Ntezi Mbabazi · 方向: AI 安全

Abstract:We examine the information security practices of Ugandan climate activists protesting the development of the East African Crude Oil Pipeline (EACOP). We conducted five-week fieldwork in Kampala, Uganda, which included interviews with 13 anti-EACOP activists. Through an inductive analysis, we report on the complexities faced by small groups of predominantly student protesters as they covertly organise small-scale anti-EACOP protests within a context marked by state surveillance and repression. Our study points to a multi-layered adversarial landscape, where participants' experiences of direct threats, including arrests and information compromise, and their fears of abduction, shaped their security practices. These practices were rooted in autonomous decision-making within groups. We present a grounded understanding of how participants' need to protect information for their own...

论文介绍 该研究探讨了乌干达反东非原油管道气候活动人士的信息安全实践。通过在坎帕拉进行为期五周的田野调查与访谈,论文分析了学生为主的抗议小团体如何在国家监视与压制下秘密组织活动。研究指出,参与者面对逮捕、信息泄露和绑架威胁等多层对抗环境,其安全实践基于团体内部的自主决策。

CODEFUSE-DEBENCH: An Empirical Study on Readability, Recompilability, and Functionality

第一作者: Puzhuo Liu · 方向: 软件安全

Abstract:Binary decompilation aims to recover binaries into high-level source code, but existing evaluations mainly rely on syntactic similarity or single-axis readability metrics, which fail to capture practical reusability. We propose a reusability-driven evaluation paradigm that measures decompiler quality along three orthogonal dimensions: readability, recompilability, and functionality. We present DEBENCH, the first automated framework for multidimensional decompilation evaluation. DEBENCH contains 240 atomic test functions, organized into 8 source files and compiled into 640 binaries. It combines LLM-as-judge readability scoring with URAF (18 sub-dimensions), iterative compile-and-repair under a fixed 50-iteration budget, and Frida-based differential dynamic tracing at the program, function, and instruction levels. We evaluate five mainstream decompilers and three repair LLMs...

论文介绍 针对现有反编译器评估方法难以衡量实际可重用性的问题,本文提出一个以可重用性为导向的评估范式。该范式从可读性、可重新编译性和功能性三个正交维度评估反编译质量,并介绍了自动化评估框架DEBENCH。论文利用该框架评估了五个主流反编译器和三个修复型大语言模型。

Provably Secure Agent Guardrail

第一作者: Benlong Wu · 方向: AI 安全

Abstract:As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates a fundamental crisis in artificial intelligence security. Existing defense architectures heavily rely on empirical semantic guardrails and probabilistic large model adjudicators, mechanisms that fail to provide deterministic security lower bounds when facing complex semantic symbol decoupling attacks. To overcome this empirical semantic guardrail dilemma, this paper proposes a new security paradigm for agents based on the fundamental limitations of logical reasoning. Based on this paradigm, we further introduce an executable Proof-Constrained Action (ePCA) framework with a neural symbolic isolation architecture. This framework abandons semantic trust in natural language, forcing agents to losslessly formalize their intentions...

论文介绍 为解决现有基于经验语义的护栏在面临复杂攻击时无法提供确定性安全保障的问题,本文提出了一种基于逻辑推理根本局限的智能体安全新范式。在此基础上,论文引入了可执行的「证明约束动作」框架,该框架通过神经符号隔离架构,迫使智能体将其意图无损地形式化,从而摒弃对自然语言的语义信任。

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

第一作者: Aditya Nawal · 方向: AI 安全

Abstract:AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating external content into the generation pipeline can weaken the safety alignment mechanisms that govern model outputs. Prior work shows that enabling retrieval in agents increases compliance with harmful requests. We introduce AgentREVEAL, a diagnostic framework for analyzing retrieval-induced safety degradation in LLM agents. The framework examines two axes: how retrieval is integrated into the agent pipeline and the properties of the retrieved content. Along the integration axis, we find that binding tool invocation and response generation in a single step amplifies harmful outputs. Along the content axis, we uncover the Safe Source Paradox: even oppositional or safety-oriented sources, such as pages containing warnings or risk...

论文介绍 本文研究了外部网页检索如何降低大语言模型智能体的安全对齐性。研究提出了诊断框架AgentREVEAL,从工具集成方式与检索内容属性两个维度进行分析。论文发现,在单一生成步骤中绑定工具调用会放大有害输出,并揭示了「安全源悖论」:即便是包含警告或反对立场的安全导向源内容,也可能削弱模型的安全对齐。

Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents

第一作者: Jay Yu · 方向: 密码学协议

Abstract:DeFi investment agents, systems that use AI for autonomous on-chain trading, have attained over USD 3 billion in combined token valuations since late 2024. We survey over 1,900 AI-tagged crypto projects, filter to investment-focused agents, and curate 10 representative projects spanning strategy and observability dimensions. We then conduct a deep-dive architectural analysis of two prominent agent frameworks, ElizaOS and Virtuals Protocol, and a quantitative on-chain performance analysis of 11 Solana-based agent treasuries with publicly attributable trading activity, covering 925,323 token holders. We find that current deployments remain early and heterogeneous: (1) in our sample, many projects do not yet provide clear evidence of autonomous trade execution, and developer interviews suggest that many visible deployments remain basic API integrations; (2) agent treasuries...

论文介绍 论文对基于AI的去中心化金融投资智能体进行了实证分析。研究调查了超过1900个AI加密项目,筛选出投资类智能体并选取10个代表项目进行深入架构分析。同时,论文对11个基于Solana的智能体金库进行了链上绩效的定量分析。研究发现,当前许多项目尚无明确的自主交易执行证据,且实际部署仍较为初级和异质化。

Robust and Efficient Guardrails with Latent Reasoning

第一作者: Siddharth Sai · 方向: 安全研究

Abstract:Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrails typically rely on single-pass classification or, more recently, distilled reasoning. Reasoning-based guardrails significantly outperform classification-only baselines, but they incur substantial query latency and token overhead that make them impractical for highthroughput deployment. To address this challenge, we propose COLAGUARD, a guardrail model that transfers multi-step safety reasoning into a continuous latent space through a stage-wise training curriculum, enabling direct hidden-state propagation at inference. Evaluated on ten prompt- and response-moderation settings spanning eight safety benchmarks, COLAGUARD improves macro-F1 by 8.24 points over Llama Guard 3 and matches our explicit reasoning baseline...

论文介绍 为解决基于推理的安全护栏存在查询延迟高、开销大的问题,本文提出了COLAGUARD模型。该模型通过阶段性训练课程,将多步安全推理转移到连续的隐空间,从而在推理时实现直接的隐藏状态传播。在多个安全基准测试中,COLAGUARD在宏F1分数上优于Llama Guard 3,并能匹配显式推理基线,同时大幅降低推理开销。

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

第一作者: Kaihua Qin · 方向: 软件安全

Abstract:Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studies use narrow datasets, inconsistent metrics, and limited semantic consistency checks. This gap is increasingly important as large language models (LLMs) begin to generate source-like Solidity that may compile and appear plausible, even when its semantics diverge from the original contract. We introduce SCDBench, a dataset and benchmark methodology for LLM-based smart contract decompilation. The dataset contains 600 real-world Solidity contracts with paired bytecode inputs, ground-truth source code, and replayable semantic checkpoints. SCDBench evaluates decompiler outputs through four cumulative stages: format completeness, compilability, Application Binary Interface (ABI) recovery, and semantic consistency via differential...

论文介绍 针对基于大语言模型的智能合约反编译器评估中存在的数据集狭窄、指标不一致等问题,本文提出了SCDBench基准。该基准包含600个真实世界Solidity合约,配有字节码、真实源码和可重放的语义检查点。评估分为格式完整性、可编译性、ABI恢复和语义一致性四个累积阶段,旨在全面衡量反编译器的输出质量。

Cycle-Space Informed Detection of Autoencoded Blind False Data Injection Attacks on Power Systems

第一作者: Xin Li · 方向: 系统安全

Abstract:The rapid growth of AI-driven data centers and large-scale energy storage systems is increasing the reliance of power system operation on real-time measurement data and automated decision-making. However, many existing detection methods rely on statistical or data-driven analysis of measurements and can fail when attackers exploit the same data structure to craft stealthy perturbations. To illustrate this limitation, we demonstrate a blind False Data Injection Attack (FDIA) in which an Autoencoder learns the measurement manifold and generates perturbations aligned with the Jacobian null space, thereby allowing the attack to evade both residual-based baddata detectors and time-series anomaly detectors. To mitigate data-driven FDIAs which exploit the null space, we propose a topology-informed Cycle-Space Detector (CSD) that leverages the Cycle-Space of the network to impose...

论文介绍 本文演示了一种盲虚假数据注入攻击,攻击者利用自动编码器学习测量流形,并生成与雅可比矩阵零空间对齐的扰动以规避检测。为应对此类利用数据结构的攻击,论文提出了一种拓扑感知的循环空间检测器。该检测器利用网络的循环空间来施加约束,从而能够识别出看似正常但与网络拓扑结构不一致的测量数据。

Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities

第一作者: Yujie Ma · 方向: 软件安全

Abstract:Large Language Models(LLMs) have been actively integrated into modern software systems as critical components. LLM-in-the-loop vulnerabilities, where vulnerabilities are introduced by LLMs and their dependent downstream components, such as frameworks, introduce new risks. Although some benchmark datasets have been constructed to study the impact of such vulnerabilities, most works still remain at the analysis from the conventional software level, ignoring the harm actually caused by LLMs. Understanding real-world LLM-in-the-loop vulnerabilities is still an open problem. To address this gap, we build the first LLM-in-the-loop vulnerability dataset, LLMCVE, to facilitate the risk analysis of LLM-integrated software. To do so, we first collect 2,888 multi-source vulnerabilities across 230 popular LLM components. Then, through manual analysis, we identify 205 vulnerabilities that...

论文介绍 本研究聚焦大型语言模型(LLM)集成软件中的漏洞问题,这些漏洞由LLM及其下游组件引入。作者构建了首个LLM-in-the-loop漏洞数据集LLMCVE,收集了230个流行LLM组件中的2,888个多源漏洞,并通过人工分析识别出205个关键漏洞。该数据集旨在促进LLM集成软件的风险分析,为漏洞理解和修复提供基础。

Meta-Quantum Ensemble Framework for Robust Network Intrusion Detection

第一作者: Ritvik Bhatnagar · 方向: AI 安全

Abstract:Intrusion Detection Systems (IDSs) must maintain high detection sensitivity while operating under strict false-positive constraints, a challenge intensified by class imbalance and heterogeneous IoT traffic. This work investigates whether heterogeneous quantum learners can provide useful and non-redundant decision information for IDS tasks. We study Quantum Support Vector Machines (QSVMs) and Quantum Neural Networks (QNNs), which rely on different learning mechanisms and exhibit distinct prediction behaviors. To combine these models, we propose the System-Level Meta-Quantum Ensemble (MQE), a hybrid quantum-classical framework that fuses QSVM and QNN outputs using a Random Forest meta-learner. The meta-learner captures agreement and disagreement patterns between the quantum branches to improve prediction stability and detection performance. Experiments on TON IoT and CICIDS2017...

论文介绍 本文研究如何提升网络入侵检测系统在类别不平衡和异构物联网流量下的性能。作者探讨了量子支持向量机和量子神经网络在入侵检测任务中的应用,提出了系统级元量子集成框架MQE,该框架使用随机森林元学习器融合量子模型的输出。通过捕捉量子分支间的共识和分歧模式,MQE旨在提高预测稳定性和检测性能。

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

第一作者: Jusuk Lee · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume...

论文介绍 机器人操作依赖于保留动作相关信息的感知,但当前学习管道缺乏运动理解。作者提出DynaFLIP,一个动态感知多模态预训练框架,通过构建图像-语言-3D流三元组作为监督,将运动理解前置到感知阶段。核心思想是鼓励三种模态在共享超球面空间中形成小单形体积,以增强对齐,避免几何歧义。

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

第一作者: Chunru Lin · 方向: 数据集与评测 · 来源: cs.RO

Abstract:The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments. However, current robotic benchmarks primarily emphasize skill-level execution and provide limited insight into such cognitive reasoning capabilities. We introduce RoboWits, a bi-manual robotic benchmark designed to systematically evaluate cognitive reasoning, creative tool use, and robustness to unexpected conditions. To enable scalable construction of high-quality reasoning-centric unexpected scenarios, we propose an automated task generation pipeline formulated as a multi-agent cooperative framework, comprising agents for seed task generation and verification, metric generation, scene generation, and task mutation. Using the pipeline, we curated 30 diverse seed tasks and 208 tasks with mutations and graded difficulty across...

论文介绍 机器人在意外条件下进行推理和创意问题解决的能力至关重要,但现有基准缺乏对认知推理的评估。作者引入RoboWits,一个双臂机器人基准,系统评估认知推理、创意工具使用和对意外条件的鲁棒性。通过自动化任务生成管道,构建了30个种子任务和208个变异任务,涵盖不同难度等级。

UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms

第一作者: Yufei Jia · 方向: 策略学习 · 来源: cs.RO

Abstract:Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning are placed on a single GPU-centric execution path. This paradigm has greatly improved training speed, but it has also encouraged a default assumption that efficient training requires physics to reside on the GPU. We revisit this assumption. Our view is that, in simulation-dominated robot control, the essential question is not which processor runs physics, but whether simulation throughput, policy learning, and runtime synchronization form an efficient end-to-end loop. We present UniLab, a heterogeneous CPU-simulation / GPU-learning architecture that decouples CPU-parallel simulation from GPU policy updates through a unified runtime for data movement, buffering, and synchronization. UniLab is implemented as a complete and...

论文介绍 基于模拟的机器人强化学习通常依赖GPU驻留模拟,但作者认为效率的关键在于模拟吞吐量、策略学习和运行时同步的端到端循环。他们提出UniLab,一个异构CPU模拟/GPU学习架构,通过统一运行时解耦并行模拟与GPU策略更新。UniLab实现了完整的端到端训练系统,旨在提升训练效率。

Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

第一作者: Kuangji Zuo · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to describe which exact object to interact with among similar candidates, where to act on the object, or how the target may change during execution. To address this limitation, we propose Gaze2Act, a novel VLA framework that leverages human gaze as a dynamic and intuitive intent signal for complex interactive manipulation. Gaze2Act first bridges the ego-exo view gap by mapping first-person gaze into the robot's perspective through cross-view semantic matching, producing both an object mask and a gaze point for coarse-to-fine target specification. These cues are then integrated into the policy through perception-level prompting and action-level...

论文介绍 视觉-语言-动作模型在机器人学习中展现潜力,但语言往往不足以精确传达意图。作者提出Gaze2Act,一个利用人类凝视作为动态意图信号的VLA框架。通过跨视图语义匹配将第一人称凝视映射到机器人视角,生成物体掩模和凝视点,用于从粗到细的目标规范,并通过感知级提示和动作级整合融入策略。

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

第一作者: Qiuyue Wang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder. Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data, trajectory-centric supervision, and...

论文介绍 具身智能常通过专用模型研究,导致能力碎片化。作者探讨是否能在单一视觉-语言-动作模型中统一异构具身决策问题。他们提出Qwen-VLA,扩展Qwen的视觉语言建模栈,通过基于DiT的动作解码器生成连续动作和轨迹。模型在大规模多源数据上进行联合预训练,包括机器人操作轨迹和人类演示等。

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

第一作者: Zhongxi Chen · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulation. However, dexterous manipulation remains challenging for VLA policies due to high-dimensional hand control and compounding execution errors, which makes real-world RL post-training essential for bridging the gap between visually grounded action generation and physically reliable dexterous execution. However, high-dimensional dexterous exploration often triggers temporal inconsistency, sample inefficiency and hardware risks in the real world. To address these challenges, we propose BORA, an offline-to-online RL post-training framework designed for real-world dexterous VLA models. In the offline phase, BORA constructs a critic that takes both the VLM's cognition tokens and action chunks as inputs. This design enables...

论文介绍 视觉-语言-动作模型在机器人操作中前景广阔,但灵巧操作因高维控制和执行误差而具有挑战。作者提出BORA,一个从离线到在线的RL后训练框架,用于现实世界灵巧VLA模型。离线阶段构建批评家,结合VLM的认知令牌和动作块作为输入;在线阶段进行残差适应,以弥合视觉接地动作生成与物理可靠执行之间的差距。

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

第一作者: Shutong Ding · 方向: 策略学习 · 来源: cs.RO

Abstract:Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approaches, one representative branch focuses on the sampling-based policy optimization. This design enables better exploration capability of the diffusion model, particularly at the beginning of training, but suffer from low exploitation in Q-value information, resulting in a slow policy convergence. Another branch pays attention to gradient-based policy optimization, which sufficiently exploits the gradient of the Q function yet tends to collapse into a unimodal policy with low diversity. To address this issue, we propose CGPO, \textbf{C}ritic-\textbf{G}uided diffusion \textbf{P}olicy \textbf{O}ptimization, which effectively balances exploration and exploitation with the training-free guidance technique...

论文介绍 这篇论文研究基于扩散策略的强化学习中探索与利用的平衡问题。现有方法要么探索能力强但利用Q值信息不足,要么充分梯度利用但多样性低。作者提出CGPO(Critic-Guided扩散策略优化),使用无训练引导技术有效平衡两者,以提高策略收敛速度。该方法在探索初期保持多样性,同时利用Q函数梯度进行优化。

Replicable Simulation-Based Robot Validation through Provenance

第一作者: Argentina Ortega · 方向: 具身智能 · 来源: cs.RO

Abstract:Robot behavior is often validated through simulation-based testing, yet the replicability of such campaigns depends critically on transparent documentation of how tests are configured, executed, and post-processed. We argue that data provenance, coupled with the FAIR principles (findability, accessibility, interoperability, and reusability), addresses this gap by explicitly tracking links between artifacts and by attaching machine-readable metadata about file origins and key design decisions. Moreover, provenance and metadata cannot be treated as an afterthought confined to final datasets; they must be integrated into the testing processes that generate those datasets so that evidence can be reconstructed end-to-end. We demonstrate this by augmenting an existing simulation-based testing framework with provenance tracking and metadata collection mechanisms, and by using these...

论文介绍 本文关注机器人仿真验证的可重复性。作者指出,测试的配置、执行和后处理缺乏透明文档,影响了可重复性。他们提出将数据溯源与FAIR原则(可发现性、可访问性、互操作性、可重用性)相结合,集成到测试流程中,以跟踪工件链接并附加机器可读元数据。这使得证据可以端到端重建,提高了验证的透明度。

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

第一作者: Hao Ren · 方向: 导航与运动 · 来源: cs.RO

Abstract:Diffusion models are effective for waypoint prediction in visual navigation, but standard sampling and test time guidance can produce unreliable or inefficient trajectories when updates drift off the training manifold. We propose Fisher Preserving Guidance with Outer Product Span Projection, a training-free inference method that avoids large Fisher drift associated with off-distribution actions while optimizing a task objective. Our method computes the Fisher-preserving update via a low-rank Jacobian factorization, requiring only a single backward pass per step and enabling real-time use. We further introduce Truncated Fisher Denoising Sensitivity as an uncertainty signal and use it for robust multi-sample action blending. Experiments on toy and realistic navigation benchmarks, including Maze2D with TSDF-based guidance, PushT with official Diffusion Policy weights, and visual...

论文介绍 扩散模型在视觉导航中用于路径点预测,但标准采样和测试时引导可能导致不可靠轨迹。作者提出Fisher保持引导方法,通过低秩Jacobian分解计算更新,避免分布外漂移,同时优化任务目标。该方法无需训练,每步仅需一次反向传递,支持实时使用。此外,引入截断Fisher去噪敏感性作为不确定性信号,用于稳健的多动作混合。

LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation

第一作者: Mohammad Khoshnazar · 方向: 机器人操作 · 来源: cs.RO

Abstract:Multi-step robot manipulation requires acting under uncertainty about how the scene will evolve, making exploration and policy adaptation challenging. We study whether short-horizon, task-consistent future videos can provide useful structured priors for control and reinforcement-learning fine-tuning. We formalize this idea through Future-Experience Conditioning (FEC), a simple interface that conditions closed-loop policies on a latent representation of a short future video. In our simulation setup, future clips are generated in three stages, an LLM reasoner operating over a task ontology initialized from the current scene state, a robot-free digital-twin rollout of the intended object motion, and a mask-free video diffusion model that synthesizes a robot-consistent future clip without requiring segmentation at inference. We instantiate this future-conditioning interface...

论文介绍 多步骤机器人操作需要在场景演变不确定下行动,探索和策略适应具有挑战性。作者研究短视野、任务一致的未来视频是否能为控制提供结构先验。他们提出未来经验条件化(FEC),通过LLM推理器生成未来假设,条件化闭环策略。在模拟中,使用数字孪生滚动和视频扩散模型合成未来片段,以改善强化学习微调。

MARS Policy: Multimodality Only When It Matters

第一作者: Jindou Jia · 方向: 机器人操作 · 来源: cs.RO

Abstract:Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to capture diverse yet valid behavioral patterns, has driven the rapid emergence of generative policies as a dominant paradigm in robot learning. However, achieving such multimodality typically relies on stochastic noise initialization and iterative denoising procedures, resulting in substantial training complexity and low inference efficiency. Meanwhile, not all phases of a robotic task inherently require behavioral diversity. Motivated by this insight, we propose the Modality-Adaptive Robot Sampling (MARS) policy, which adaptively invokes tailored stochasticity only when it is truly beneficial, while reverting to an efficient deterministic learning during single-modal phases. In other words, the proper amount of noise is injected only...

论文介绍 模仿学习中,多模态策略能捕获多样行为模式,但依赖随机初始化和迭代去噪,导致训练复杂和推理效率低。作者提出MARS策略,自适应地仅在需要时调用随机性,在单模态阶段使用确定性学习。这样,噪声仅在必要时注入,平衡了多样性和效率,简化了训练并提高推理速度。

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

第一作者: Sergey Arkhangelskiy · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with $N \le 25$ rollouts per condition, almost always without confidence intervals or paired statistical comparison; these cohort sizes struggle to resolve close comparisons reliably. We introduce PhAIL (Physical AI Leaderboard, this https URL), an open real-robot benchmark on a Franka FR3 (dataset, per-rollout artifacts, and end-to-end reference implementation) of a distributional evaluation methodology: the time-to-success cumulative distribution function (CDF) as the evaluation primitive, with two separated jobs. The first is scoring via Human-Relative Throughput (HRT), a dimensionless scalar with bootstrap confidence intervals, anchored to same-fixture human teleoperation. The second is a significance test (Kolmogorov-Smirnov, computed per-object and...

论文介绍 视觉-语言-动作(VLA)模型的评估通常依赖二元成功率和少量样本,缺乏统计比较。作者引入PhAIL基准和分布式评估方法,使用时间成功累积分布函数(CDF)作为评估原语。通过人类相对吞吐量(HRT)评分和显著性测试,提供置信区间,使评估更可靠,适用于真实机器人环境。

EXACT-MPPI: Exact Signed-Distance Navigation for Arbitrary-Footprint Robots from Point Clouds via Path Integral Control

第一作者: Chen Peng · 方向: 导航与运动 · 来源: cs.RO

Abstract:Ground robots often carry payloads, implements, or other attachments that turn their effective footprint into complex, non-convex shapes. Navigating safely through clutter then requires reasoning about this true geometry, yet most local planners simplify it with convex or inflated proxies and rasterize sensor data into occupancy grids or distance fields. Both choices eliminate feasible motions when clearance is comparable to the footprint geometry. We present EXACT-MPPI, a training-free local navigation framework that maps local point-cloud observations and sparse guidance directly to motion commands, without any intermediate map representation. The framework embeds an analytic, exact signed-distance evaluator into a Model Predictive Path Integral (MPPI) controller. The footprint is represented as a simple polygon for general convex or concave planar shapes, with a...

论文介绍 地面机器人常携带复杂非凸足迹,在杂乱环境中安全导航需要精确几何推理。现有方法简化足迹或使用栅格化数据,可能消除可行运动。作者提出EXACT-MPPI,一个无训练局部导航框架,将点云观测和稀疏引导直接映射到运动命令。通过嵌入精确符号距离评估器,支持任意多边形足迹,无需中间地图表示。

VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models

第一作者: Dehao Huang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Confidence estimation for Vision-Language-Action (VLA) models is essential for robots to perform manipulation tasks in the open world, providing crucial signals for risk-sensitive decision-making and failure anticipation. Existing confidence estimation methods typically rely on ensemble-based paradigms or action-token probabilities to predict the likelihood of task success. However, they still encounter challenges in computational efficiency and cross-architecture generalizability. These methods usually require repeated sampling, leading to inference inefficiency, and are restricted to VLA models with discrete action outputs, making them difficult to apply to continuous action spaces. To address this issue, we propose VLAConf, a one-class discriminative confidence framework. By leveraging frozen pretrained VLA internal representations, VLAConf directly estimates step-wise...

论文介绍 VLA模型的置信度估计对风险敏感决策至关重要。现有方法依赖集成或动作令牌概率,计算效率低且泛化差。作者提出VLAConf,一个单类判别框架,利用冻结的预训练VLA内部表示,直接估计步骤级置信度。该方法无需重复采样,适用于连续动作空间,提高了效率和泛化能力。

Learning to Feel Materials from Multisensory Tactile Data via Interpretable Models

第一作者: Li Zou · 方向: 具身智能 · 来源: cs.RO

Abstract:Human tactile perception of materials relies on complex multisensory touch cues, yet the relationship between low-level tactile signals and perceptual representations remains poorly understood. This knowledge gap hinders the integration of touch in digital environments and the development of robots capable of human-like tactile perception. Here, we present an interpretable computational framework for modeling human material perception and recognition using multisensory touch data. Our framework comprises three interconnected models: Model 1 maps finger-surface interaction features to psychophysical sensory attributes, Model 2 classifies materials based on these perceptual representations, and Model 3 directly classifies materials from tactile features. The results showed that combining information from pressing, static contact, and sliding interactions improves prediction...

论文介绍 该研究旨在建模人类的材料触觉感知,以弥补低层触觉信号与感知表征之间的理解鸿沟。其核心是一个可解释的计算框架,包含三个相互关联的模型,分别负责将手指-表面交互特征映射到心理物理属性、基于感知表征进行材料分类,以及直接从触觉特征进行分类。研究表明,融合按压、静态接触和滑动交互的信息能有效提升预测性能。

VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

第一作者: Victor Kowalski · 方向: 机器人操作 · 来源: cs.RO

Abstract:When using reinforcement learning (RL) for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve. However, vision-enabled policies tend to overfit to the visual conditions seen during training, limiting their robustness and transferability. We present a human-in-the-loop RL framework that employs teacher-student distillation to achieve robust performance across multiple task variants, trained entirely in the real world without requiring domain randomization or data augmentation. A vision-enabled teacher distills its knowledge into a vision-free student that relies solely on pose, twist, and wrench sensing, combining fast training with strong task generalization. On the real-world NIST assembly benchmark board, our approach achieves 95\% overall success after approximately 50...

论文介绍 针对视觉驱动的强化学习策略在接触丰富的操作任务中泛化性差的问题,本文提出一种人类在环的强化学习框架。该框架利用教师-学生蒸馏策略,将一个视觉有能的教师策略的知识,迁移至一个仅依赖位姿、速度和力觉感知的视觉无学生策略。整个过程完全在真实世界训练,无需领域随机化,成功实现了在多个任务变体上的强鲁棒性和泛化能力。

VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models

第一作者: Shengyu Si · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and action patterns. This paper proposes VLA-Pro, a plug-and-play framework designed to enhance cross-task generalization by storing task-relevant procedural memories at training time and transferring these memories during inference. Specifically, VLA-Pro stores task-specific LoRA adapters as parameterized procedural memories during training. At inference time, VLA-Pro retrieves relevant procedural memories based on the current multi-modal context and dynamically fuses these memories for generating the current action chunk. Experiments on RoboTwin, RLBench, and real-world manipulation tasks show that VLA-Pro consistently improves cross-task...

论文介绍 视觉语言动作(VLA)模型在跨未见任务时泛化能力有限。为应对此挑战,本文提出VLA-Pro,一个即插即用的框架。它通过在训练时存储任务特定的LoRA适配器作为参数化程序记忆,并在推理时根据多模态上下文检索和动态融合相关记忆来生成动作,从而显著增强了VLA模型的跨任务泛化能力,并在模拟和真实世界任务中得到验证。

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

第一作者: Ye Li · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time robotic manipulation, especially when large vision-language backbones and iterative action heads run at every control step. Existing VLA acceleration methods often optimize individual components or rely on fixed acceleration rules, treating different control steps with largely fixed computation and overlooking the non-uniform reasoning demands of sequential embodied control. Inspired by human motor control, where cognitive and feedback resources concentrate on goal-sensitive stages, we argue that VLA models should learn when to invest full computation and when to reuse prior computation. We propose ElegantVLA, a plug-in phase-adaptive inference framework that accelerates VLA models through intra-model...

论文介绍 现有视觉语言动作(VLA)模型计算开销大,限制了实时操作。受人类运动控制中资源按需分配的启发,本文认为VLA应学会在关键阶段投入完整计算,在其他阶段复用先前计算。为此,作者提出ElegantVLA,一个相位自适应的推理框架。该框架能够加速VLA模型,通过模型内部和模型间的计算动态调整,以适应序列化具身控制中非均匀的推理需求。

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding

第一作者: Zhongyu Xia · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficiency manifests as three intertwined challenges: weak extraction of 3D spatial positions without enforcing multi-view consistency, inadequate 3D instance understanding, and fragile reasoning under occlusion. Although mature 3D perception methods exist, their direct integration into VLA pipelines is hindered by architectural incompatibility and by heavy reliance on costly instance-level annotations. To address the above challenges, we propose 3DVLA, a plug-and-play framework that injects robust 3D reasoning into pretrained VLAs without requiring extra manual labels or discarding VLM priors. Specifically, 3DVLA tackles the three challenges through: (1) pervasive 3D feature encoding with explicit...

论文介绍 当前视觉语言动作(VLA)模型普遍缺乏3D场景理解能力。为解决此问题,本文提出3DVLA,一个即插即用的框架,旨在向预训练VLA中注入稳健的3D推理能力,而无需额外的人工标注。该框架通过无处不在的3D特征编码、实例感知的3D理解以及构建基于遮挡的自监督学习任务,分别应对空间位置提取弱、实例理解不足和遮挡下推理脆弱这三大挑战。

Phase-Conditioned Imitation Learning with Autonomous Failure Recovery for Robust Deformable Object Manipulation

第一作者: Dayuan Chen · 方向: 机器人操作 · 来源: cs.RO

Abstract:This paper presents a phase-conditioned, force-aware framework for robust deformable object manipulation. Standard imitation learning policies such as Action Chunking with Transformers (ACT) rely on a Markovian assumption at inference, causing state aliasing when visually similar observations require contradictory actions and preventing autonomous recovery from execution failures. We address this with a closed-loop hierarchical architecture. A FiLM-conditioned ACT encoder modulates feature extraction based on the current task phase, enabling a single unified policy to produce phase-specific behaviors while sharing action dynamics across phases. A multi-modal phase predictor fusing visual, force, and pose feedback estimates the phase in real time, detecting contact failures that are invisible to vision alone and autonomously triggering recovery trajectories. The system is...

论文介绍 针对标准模仿学习策略在操作可变形物体时因状态混淆而无法自主恢复故障的问题,本文提出一个基于相位条件、力觉感知的闭环分层框架。该框架使用一个经FiLM条件调制的编码器,使单一策略能根据当前任务相位产生特定行为。同时,一个多模态相位预测器融合视觉、力觉和位姿反馈来实时估计任务相位,并检测视觉不可见的接触失败,从而自主触发恢复轨迹。

Decentralized LLM-Driven Coordination of Acoustic Robots for Contactless Object Manipulation

第一作者: Yingying Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Natural language interfaces can simplify interaction with multi-robot systems, especially when non-expert users need to issue high-level commands. Acoustic manipulation using ultrasonic phased arrays also enables contactless object handling for applications such as healthcare, laboratory automation, and precision transport. However, combining large language models (LLMs) with distributed acoustic mobile robots remains underexplored. This paper presents a decentralized framework for natural language-driven coordination of acoustic robots for contactless object manipulation. The system converts spoken instructions into executable multi-robot task plans using Whisper-based speech recognition, LLM-based semantic parsing, structured JSON task representation, and distributed scheduling. The JSON schema encodes robot assignments, temporal dependencies, spatial constraints, and...

论文介绍 本文探索将大语言模型(LLM)与分布式声学移动机器人结合,实现非接触物体操作。研究提出一个去中心化框架,能够将用户的自然语言指令转化为可执行的多机器人任务计划。该系统综合运用基于Whisper的语音识别、LLM语义解析、结构化JSON任务表示以及分布式调度,使机器人能够理解并协作执行如堆叠、传递等非接触操作任务。

MonoDuo: Using One Robot Arm to Learn Bimanual Policies

第一作者: Sandeep Bajamahal · 方向: 机器人操作 · 来源: cs.RO

Abstract:Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-arm robots, however, are widely available in research labs. Can we leverage them to train bimanual robot policies? We present MonoDuo, a framework for learning bimanual manipulation policies using single-arm robot demonstrations paired with human collaboration. MonoDuo collects data by teleoperating a single-arm robot to perform one side of a bimanual task while a human performs the other, then swapping roles to cover both sides. RGB-D observations from a wrist-mounted and fixed camera are augmented into synthetic demonstrations for target bimanual robots using state-of-the-art hand pose estimation, image and point cloud segmentation, and inpainting. These synthetic demonstrations, grounded in real...

论文介绍 学习双臂机器人策略常受限于双臂机器人的稀缺和数据不足。本文提出MonoDuo框架,旨在利用广泛可用的单臂机器人示教来学习双臂操作策略。其核心方法是通过单臂机器人与人类协作收集数据,分别执行双臂任务的一侧,然后利用先进的手部姿态估计、分割和修复技术,将数据增强为适用于目标双臂机器人的合成示教数据,从而实现策略学习。

Extreme dynamic symmetry enables omnidirectional and multifunctional robots

第一作者: Jiaxun Liu · 方向: 导航与运动 · 来源: cs.RO

Abstract:Symmetry is a central organizing principle in natural systems, yet its use as a unifying design strategy in robotics has largely remained limited to geometric form. We show that symmetry can instead be leveraged at the level of dynamic actuation capability. We introduce dynamic symmetry, the uniformity of a robot's attainable center-of-mass accelerations, and formalize it through a measure coined as dynamic isotropy. Across more than 1000 simulated morphologies, we found that higher dynamic symmetry consistently improved trajectory tracking, task success, robustness, resiliency, and energy efficiency, with the benefits becoming most pronounced as dynamic isotropy approached its theoretical limit. To study this regime systematically, we developed Argus, a family of spherical robots designed to explore the effects of increasing dynamic symmetry. Members of the Argus family vary...

论文介绍 本文提出「动态对称性」概念,用于衡量机器人在质心加速度能力上的均匀性,并通过「动态各向同性」度量进行形式化。研究发现,在超过1000种形态的仿真中,更高的动态对称性能够一致性地提升轨迹跟踪、任务成功率和能源效率。为此,研究者开发了名为Argus的系列球形机器人,用以系统探索动态对称性增强的效果。

Learning and Adaptation in Wire Arc Additive Manufacturing Bead Geometry Control

第一作者: Chen-Lung Lu · 方向: 具身智能 · 来源: cs.RO

Abstract:Robotics Wire Arc Additive Manufacturing (WAAM) is governed by complex and nonlinear process dynamics coupling thermal field to the build geometry. The process may be regarded as a multi-input/multi-output dynamical system with welding torch speed and wire feed rate as inputs and weld bead deposition height and width as outputs. In this paper, we use the input/output data to learn a data-driven model and use it for weld planning and control. We show that a simple recurrent neural network architecture and one-step-ahead predictive control can improve the process performance in terms of height and width consistency. To account for the changing thermal conditions during the printing process, we update the learning model using prediction error from the previous layer. This adaptation step further improves the prediction accuracy and controller performance. Experiments on a robotic...

论文介绍 该研究针对机器人电弧增材制造(WAAM)中热场与几何耦合的复杂非线性过程,采用数据驱动方法。研究者使用循环神经网络学习输入(焊枪速度、送丝率)与输出(沉积高度、宽度)的关系,并结合一步超前预测控制来改善焊道几何的一致性。为应对打印过程中热条件的变化,模型还利用前一层的预测误差进行在线更新,从而提升预测精度与控制性能。

Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

第一作者: Petras Swissler · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Swarm and field robotics face significant barriers to real-world validation due to the high cost and development time to deploy hardware. This paper introduces the ``Bionic Swarm,'' a novel system that lowers these barriers by abstracting away many of the tasks that are difficult to implement on robots but which do not contribute to the overall algorithm evaluation, giving these tasks to human users. These human users take directions from a smartphone web-app that takes measurements from Bluetooth-connected sensors and relays them to a centralized server. This server runs the swarm algorithm and directs actions to the human users. We evaluate this system through the experimental validation of a geotechnically-focused search algorithm named Score-Biased-Search, which functions by assigning a ``score'' to each location on a reconstructed map, then biases search patterns through...

论文介绍 本文提出「仿生群体」系统,旨在降低群体机器人现场验证的门槛。该系统将难以在机器人上实现但非算法评估核心的任务交由人类用户执行,人类用户通过手机应用接收指令并利用蓝牙传感器进行测量。服务器运行群体算法并向用户分派任务。研究通过一项地质搜索算法的实验验证了该系统的有效性,为真实环境下的群体算法评估提供了一种低成本途径。

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

第一作者: Luzhou Ge · 方向: 多模态具身 · 来源: cs.RO

Abstract:Integrating open-vocabulary semantic information into dynamic 3D scene representations is essential for long-term embodied scene understanding. However, existing methods often suffer from fragile instance association due to incomplete cross-view cues, while their limited ability to handle object-level topological changes restricts long-term robotic task execution. Moreover, current 3D scene understanding methods either rely on simple feature matching without explicit spatial reasoning or assume offline ground-truth 3D geometry. To address these challenges, we present DGSG-Mind, a hybrid instance-aware 3D Gaussian dynamic scene graph system with an embodied reasoning agent. Our system couples a probabilistic voxel grid with explicit 3D Gaussians to enable robust cross-modal instance fusion and incremental semantic mapping. It handles dynamic changes through Gaussian-based...

论文介绍 该研究提出DGSG-Mind系统,用于长期动态场景理解与定位。它结合概率体素网格与显式3D高斯表示,实现稳健的跨模态实例融合与增量语义建图。系统通过基于高斯的动态变化处理与明确的空间推理,解决了现有方法在实例关联脆弱性、拓扑变化处理能力以及在线几何推理方面的不足,旨在支持长期机器人任务执行。

Energy-Aware NECO for Single-Pass Pixel-wise Out-of-Distribution Detection in Semantic Segmentation

第一作者: Boyuan Zhang · 方向: 具身智能 · 来源: cs.RO

Abstract:Reliable semantic segmentation for mobile robots requires both accurate dense prediction and robust uncertainty estimation under distribution shift. Strong uncertainty baselines such as Monte Carlo Dropout often require repeated stochastic forward passes and are difficult to deploy on edge platforms. We propose Energy-Aware NECO, a single-pass pixel-wise out-of-distribution (OOD) detector for semantic segmentation. The method combines a centered NECO-style geometric ratio computed from decoder features with a logit-based Energy score. Both components are standardized using statistics fitted on a pure in-distribution validation split and fused through a convex combination. We evaluate the method on the miniMUAD subset using true pixel-level OOD labels. The proposed hybrid score achieves an AUROC of 0.8539, outperforming NECO-only (0.8280), Energy-only (0.8171), and an ensemble...

论文介绍 为满足移动机器人对可靠语义分割的需求,本文提出Energy-Aware NECO,一种用于语义分割的单次像素级分布外(OOD)检测器。该方法融合了基于解码器特征的几何比率和基于逻辑的能量分数,并通过在分布内数据上拟合的统计量进行标准化后凸组合。评估显示,这种混合分数在像素级OOD检测上优于单一方法,且只需单次前向传播,适用于资源受限的边缘平台。

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

第一作者: Mohammadreza Teymoorianfard · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework...

论文介绍 本文系统探究了集成推理能力的视觉-语言-动作(VLA)模型在自动驾驶中的脆弱性。研究以NVIDIA的Alpamayo模型为例,首次在闭环仿真中评估了现实世界文本输入扰动的影响。结果表明,这类模型在推理和轨迹生成上对扰动高度敏感,攻击成功率最高分别达89%和72%,导致碰撞率上升和安全指标恶化。研究提出了一个推理感知的评估框架。

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

第一作者: Jiyao Zhang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelligence in embodied 3D environments. To systematically evaluate these foundational perceptual capabilities, the benchmark includes 6 task categories divided into two core groups: Spatial Structural Understanding (Grounding, Spatial Relation Prediction, and Multi-view Correspondence) and Interaction-Oriented Perception (Affordance Prediction, Grasp Point Prediction, and Trajectory Prediction). The benchmark spans 12 subcategories and contains over 21k high-quality question-answer pairs. We evaluate 13 state-of-the-art models, and the results show that while current models exhibit relatively strong high-level spatial reasoning, such as understanding...

论文介绍 本文引入Embodied3DBench基准,旨在评估视觉语言模型在具身3D环境中的底层空间智能。该基准包含空间结构理解(如定位、关系预测)和交互导向感知(如可供性、抓取点预测)两大类共六项任务,涵盖21k多组高质量问答对。对13个先进模型的评估显示,当前模型在高层空间推理上表现较好,但在底层感知和交互预测方面仍有明显不足。

Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems

第一作者: Daniel Albiero · 方向: 具身智能 · 来源: cs.RO

Abstract:Tropical forests worldwide are under intense deforestation pressure driven by economic and political interests, and scientific evidence suggests this deforestation contributes to climate change. This paper proposes a novel logging method for tropical forests, Ultra-Reduced-Impact-Encased-Logging (URIEL). This new method is based on heli-logging techniques combined with intensive use of robotics and AI integrated with post-harvest silvicultural treatments performed by drones. The concept of appropriate equipment for this method was developed, dimensions were determined, details were completed in a digital proof of concept, and an effective digital simulation and economic feasibility analysis were carried out for various helicopter-timber-distance combinations. The results demonstrated that a URIEL method has high economic viability and makes it possible to virtually eliminate...

论文介绍 本文提出一种名为URIEL的新型热带森林选择性伐木方法,旨在实现超低环境影响。该方法结合直升机伐木技术与无人机执行的集约化机器人及人工智能应用,并包含收获后的林地处理。通过数字仿真和经济可行性分析表明,URIEL方法具有高经济可行性,并能基本消除传统伐木对森林地面的主要破坏,为可持续林业提供了技术方案。

VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

第一作者: Mingjian Gao · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irrelevant or weakly textual information can interfere with action prediction, while autoregressive text decoding adds too much latency for real-time closed-loop execution. We present VISUALTHINK-VLA, a visual intermediate-reasoning framework for accurate, low-latency VLA policies. Our bootstrapping philosophy is to guide action with effective visual thinking: VISUALTHINK-VLA bootstraps action prediction through a compact visual-evidence interface that preserves spatial precision while avoiding decoding overhead. Besides, to further improve performance and efficiency, VISUALTHINK-VLA adopts a tailored selective routing mechanism to learn the visual evidence tokens, enabling low-latency inference while...

论文介绍 研究视觉语言动作模型中中间推理的优化问题。针对文本推理在具身控制中的干扰和延迟,提出VisualThink-VLA框架,采用视觉中间推理,通过紧凑的视觉证据接口引导动作预测,并引入选择性路由机制提升推理效率,旨在实现准确且低延迟的机器人控制策略。

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

第一作者: Shilin Ma · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across execution steps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce an...

论文介绍 针对视觉语言动作模型推理延迟高的问题,提出SAFE-Pruner框架。该框架利用语义注意力一致性,设计前瞻策略预测深层token的显著性,防止过早移除关键视觉信息。作为即插即用模块,它能有效加速VLA模型的实时推理,适用于机器人操控任务。

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning

第一作者: Kyujin Lee · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Vision-Language-Action (VLA) models have emerged as a promising framework that unifies perception, reasoning, and control for robot manipulation by adapting pretrained vision-language models (VLMs) to action prediction. However, VLM-derived representations are often insensitive to subtle visual distinctions required for low-level control, causing state aliasing between visually similar states that require substantially different actions. Prior VLA studies improve visual understanding by generating visual or reasoning outputs, such as future frames, 2D grounding points or traces, or intermediate spatial reasoning steps, but these objectives typically shape the vision encoder only indirectly through end-to-end prediction and do not explicitly analyze state aliasing in the learned visual feature space. To mitigate state aliasing, we introduce inverse dynamics learning as an...

论文介绍 解决视觉语言动作模型中视觉表示对细微区分不敏感导致的状态混淆问题。引入逆动力学学习作为辅助任务,显式优化视觉特征空间,以缓解相似视觉状态下动作预测的歧义,有助于提升机器人操控中低级控制的精度。

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

第一作者: Xueqing Wu · 方向: 策略学习 · 来源: cs.CV

Abstract:Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a bottleneck for end-to-end visual reasoning. To investigate this gap, we introduce a controlled diagnostic framework with two synthetic tasks that disentangle perception from reasoning. Our analysis reveals a consistent perception-reasoning asymmetry: posttraining improves reasoning more substantially than perception, though the underlying mechanism differs by training paradigm. For supervised fine-tuning (SFT), this asymmetry stems from token imbalance in chain-of-thought supervision, where perception occupies fewer tokens and thus receives a weaker training signal. Dynamically reweighting the loss mitigates this imbalance and boosts end-to-end performance by up to 18.2. For reinforcement learning (RL), the asymmetry instead...

论文介绍 探讨视觉语言模型后训练中推理与感知优化的不对称性。通过合成任务诊断框架,分析监督微调和强化学习中的不同机制,发现token不平衡是原因之一,并提出损失动态重加权来改善感知性能,从而优化端到端视觉推理。

AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference

第一作者: Yilin Feng · 方向: 多模态具身 · 来源: cs.LG

Abstract:Vision-Language Models (VLMs) process thousands of visual tokens per image alongside comparatively few text tokens, yet existing compression methods treat both modalities uniformly. We observe that the two modalities have fundamentally different properties: vision tokens are spatially redundant and dominate prefill, while text tokens are causally dependent and accumulate during decoding. Based on this asymmetry, we propose and empirically evaluate AsymVLM, which applies aggressive pruning to vision tokens before prefill using a learned importance scorer with per-sample adaptive budgeting, and temporal threshold-based eviction to text tokens only when they exceed a fixed budget. Our experiments indicate that AsymVLM achieves the highest FLOPs savings (up to 54%) among state-of-the-art methods while outperforming existing approaches by 2--3% on document and chart understanding...

论文介绍 针对视觉语言模型推理效率问题,提出AsymVLM方法。观察到视觉和文本模态处理方式的差异,对视觉token在预填充阶段进行激进剪枝,对文本token在解码阶段采用阈值驱逐,显著减少计算开销,同时在多模态任务中保持性能。

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

第一作者: Haoyuan Shi · 方向: VLA 通用模型 · 来源: cs.AI

Abstract:Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic framework that analyzes VLA models through a unified evidence chain from representation dynamics to causal control attribution and behavioral manifestation. It specifically combines cross-modal and checkpoint-drift centered kernel alignment (CKA) to trace representation evolution, attention knockout interventions to identify modality-specific control pathways, and rollout-level behavioral probes to examine grounding, shortcut dependence, and semantic following. Experiments on $\pi_{0.5}$ and OpenVLA reveal three key findings. First, the two models exhibit distinct modality-specific adaptation dynamics during VLA finetuning. Second, they rely on different multimodal routing strategies and layer-wise...

论文介绍 为理解视觉语言动作模型的内部机制,提出VLA-Trace诊断框架。该框架结合表示追踪、注意力干预和行为探测,分析模型的多模态融合过程和控制路径,实验揭示不同模型的模态适应差异,有助于提升模型的可解释性。

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

第一作者: Tianzhuo Yang · 方向: 数据集与评测 · 来源: cs.AI

Abstract:Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias...

论文介绍 针对机器人动作条件世界模型评估中可靠性标准缺失的问题,提出MiraBench基准。定义动作条件可靠性,从物理遵守、动作跟随保真度和乐观偏差三个层次进行评估,旨在推动世界模型在机器人学习中的可靠应用和发展。

市场总览

美股技术面整体偏强,标普500 ETF和纳斯达克100 ETF的RSI分别处于74.1和77.2的超买区域,价格接近52周高点并呈多头排列,但部分科技股如英伟达近5日回调3.81%。加密货币市场技术指标显示下行压力,比特币和以太坊均呈空头排列,RSI值分别为36.5和31.6,结合辅助背景中恐慌贪婪指数29(恐慌)及总市值2.59万亿美元、BTC主导率57.2%,反映市场情绪谨慎。中概股普遍走弱,阿里巴巴、拼多多等空头排列,腾讯控股RSI28.9进入超卖区域。商品外汇方面,黄金期货出现MACD金叉但趋势中性,原油期货无明显信号,美元指数多头排列而美元兑人民币空头排列。总体市场技术面呈现分化,美股强势但超买,加密和中概偏弱,商品外汇混合。

今日关注

SPY S&P 500 ETF
偏上行

价格756.48,RSI14为74.1处于超买区域,高于SMA20(739.34)、SMA50(703.65)和SMA200(681.17)形成多头排列。近1日涨0.25%,近5日涨1.85%,接近52周高点仅差0.21%,技术指标显示强势上涨趋势。

BTC-USD Bitcoin
偏下行

价格73504.12,RSI14为36.5接近超卖但未达,低于SMA20(76744.77)、SMA50(77229.74)和SMA200(79522.21)呈空头排列。MACD值-1118.0666为负,近1日跌0.34%,近5日跌3.06%,技术指标指向下行压力。

0700.HK 腾讯控股 (0700.HK)
偏下行

价格421,RSI14为28.9进入超卖区域,低于SMA20(452.46)、SMA50(481.86)和SMA200(573.99)呈空头排列。接近52周低点仅差0.14%,近1日跌1.45%,近5日跌4.62%,技术面显示持续下行态势。

GC=F 黄金期货
中性

价格4564.5,MACD出现金叉信号,但趋势标记为中性。价格略低于SMA20(4586.42)和SMA50(4629.55),但高于SMA200(4376.88),RSI14为46.8处于正常区域。近1日涨0.09%,近5日涨0.96%,技术指标无明显方向倾向。

全部资产

^VIX

VIX 恐慌指数

$15.32 -2.67%
5 日
-8.26%
距 52w 高
-56.6%
RSI(14)
35.4
趋势
中性
SMA 20 / 50 / 200
17.25 / 19.95 / 18.38
MACD / 信号
-0.913 / -0.846
MACD 死叉 (1 天前)

^TNX

10Y 美债收益率 (%)

$4.45 -0.04%
5 日
-2.90%
距 52w 高
-10.9%
RSI(14)
49.8
趋势
多头
SMA 20 / 50 / 200
4.48 / 4.39 / 4.20
MACD / 信号
0.040 / 0.055
MACD 死叉 (2 天前)接近 52 周低多头排列

DX-Y.NYB

美元指数 DXY

$99.05 +0.14%
5 日
-0.27%
距 52w 高
-1.6%
RSI(14)
53.5
趋势
多头
SMA 20 / 50 / 200
98.76 / 98.90 / 98.57
MACD / 信号
0.133 / 0.103
接近 52 周高多头排列

SPY

S&P 500 ETF

$756.48 +0.25%
5 日
+1.85%
距 52w 高
-0.2%
RSI(14)
74.1
趋势
多头
SMA 20 / 50 / 200
739.34 / 703.65 / 681.17
MACD / 信号
12.698 / 12.873
RSI 超买接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$738.31 +0.37%
5 日
+3.33%
距 52w 高
-0.4%
RSI(14)
77.2
趋势
多头
SMA 20 / 50 / 200
709.04 / 652.93 / 617.85
MACD / 信号
21.470 / 21.438
MACD 金叉 (今天)RSI 超买接近 52 周高多头排列

AAPL

Apple

$312.06 -0.14%
5 日
+2.32%
距 52w 高
-0.9%
RSI(14)
78.8
趋势
多头
SMA 20 / 50 / 200
297.54 / 275.28 / 263.24
MACD / 信号
10.392 / 9.777
RSI 超买接近 52 周高多头排列

MSFT

Microsoft

$450.24 +5.45%
5 日
+7.43%
距 52w 高
-18.9%
RSI(14)
69.6
趋势
中性
SMA 20 / 50 / 200
417.59 / 402.83 / 458.46
MACD / 信号
5.802 / 4.124
MACD 金叉 (今天)

NVDA

Nvidia

$211.14 -1.45%
5 日
-3.81%
距 52w 高
-10.7%
RSI(14)
49.4
趋势
多头
SMA 20 / 50 / 200
215.46 / 199.35 / 187.65
MACD / 信号
3.809 / 5.974
多头排列

GOOGL

Alphabet

$380.34 -2.51%
5 日
-1.89%
距 52w 高
-6.9%
RSI(14)
52.9
趋势
多头
SMA 20 / 50 / 200
391.15 / 347.57 / 299.92
MACD / 信号
9.641 / 13.543
多头排列

TSLA

Tesla

$435.79 -1.43%
5 日
+4.29%
距 52w 高
-12.6%
RSI(14)
60.0
趋势
中性
SMA 20 / 50 / 200
421.39 / 391.80 / 412.13
MACD / 信号
12.067 / 11.367
MACD 金叉 (2 天前)

META

Meta

$632.51 -0.44%
5 日
+4.14%
距 52w 高
-20.6%
RSI(14)
55.4
趋势
中性
SMA 20 / 50 / 200
613.33 / 618.53 / 666.57
MACD / 信号
-1.092 / -4.297
MACD 金叉 (2 天前)
加密恐慌贪婪
29
恐慌
加密总市值
$2.59 T
+0.08% / 24h
BTC 主导率
57.2%
ETH 9.4%
24h 成交量
$58.0 B
活跃币 17,402

BTC-USD

Bitcoin

$73,504.12 -0.34%
5 日
-3.06%
距 52w 高
-41.8%
RSI(14)
36.5
趋势
空头
SMA 20 / 50 / 200
76,744.77 / 77,229.74 / 79,522.21
MACD / 信号
-1,118.067 / -588.848
空头排列

ETH-USD

Ethereum

$2,003.85 -0.77%
5 日
-3.24%
距 52w 高
-59.5%
RSI(14)
31.6
趋势
空头
SMA 20 / 50 / 200
2,118.33 / 2,241.59 / 2,497.41
MACD / 信号
-67.040 / -57.471
空头排列

SOL-USD

Solana

$82.30 -0.30%
5 日
-1.54%
距 52w 高
-67.5%
RSI(14)
40.0
趋势
空头
SMA 20 / 50 / 200
85.80 / 86.36 / 104.44
MACD / 信号
-1.465 / -0.962
空头排列

BABA

阿里巴巴 (BABA)

$124.22 -1.54%
5 日
-5.51%
距 52w 高
-35.5%
RSI(14)
37.7
趋势
空头
SMA 20 / 50 / 200
134.18 / 131.07 / 149.62
MACD / 信号
-1.893 / -0.440
空头排列

PDD

拼多多 (PDD)

$84.44 +1.70%
5 日
-13.65%
距 52w 高
-39.4%
RSI(14)
32.3
趋势
空头
SMA 20 / 50 / 200
95.79 / 98.42 / 113.23
MACD / 信号
-3.204 / -1.799
MACD 死叉 (4 天前)空头排列

JD

京东 (JD)

$28.83 -1.06%
5 日
-8.39%
距 52w 高
-21.8%
RSI(14)
38.8
趋势
空头
SMA 20 / 50 / 200
30.88 / 30.03 / 30.44
MACD / 信号
-0.087 / 0.314
MACD 死叉 (4 天前)空头排列

0700.HK

腾讯控股 (0700.HK)

HK$421.00 -1.45%
5 日
-4.62%
距 52w 高
-38.4%
RSI(14)
28.9
趋势
空头
SMA 20 / 50 / 200
452.46 / 481.86 / 573.99
MACD / 信号
-16.860 / -14.931
RSI 超卖接近 52 周低空头排列

GC=F

黄金期货

$4,564.50 +0.09%
5 日
+0.96%
距 52w 高
-18.3%
RSI(14)
46.8
趋势
中性
SMA 20 / 50 / 200
4,586.42 / 4,629.55 / 4,376.88
MACD / 信号
-46.265 / -47.000
MACD 金叉 (今天)

CL=F

WTI 原油期货

$89.63 +2.60%
5 日
-7.22%
距 52w 高
-25.0%
RSI(14)
41.4
趋势
中性
SMA 20 / 50 / 200
97.90 / 97.73 / 72.17
MACD / 信号
-2.003 / -0.210

USDCNY=X

美元 / 人民币

¥6.77 -0.20%
5 日
-0.42%
距 52w 高
-6.2%
RSI(14)
30.6
趋势
空头
SMA 20 / 50 / 200
6.80 / 6.83 / 6.99
MACD / 信号
-0.014 / -0.013
MACD 死叉 (2 天前)接近 52 周低空头排列
风险提示

本报告基于公开行情数据计算的技术指标,仅供技术指标解读参考,不构成任何投资建议。过去走势不代表未来表现,市场存在不确定性,读者应结合其他信息综合判断并自行承担风险。

Colombia Presidential Election Heads to a Runoff

The candidate forced a runoff on Sunday in what could herald another gain for the right-wing wave sweeping Latin America.

中文摘要 哥伦比亚总统选举在周日进行第一轮投票后进入决选,这可能预示右翼浪潮在拉丁美洲的又一进展。

A year of grief and waiting: What remains when a plane falls from the sky

A year after the Air India crash, a mother still speaks about her dead son in the present tense and a brother waits for answers.

中文摘要 印度航空坠机事故一周年后,一名母亲仍以现在时谈论已故儿子,一名兄弟在等待真相,家属持续悲痛。

Colombia’s far-right presidential candidate Espriella wins first round of vote ahead of runoff

Lawyer and Trump admirer has risen rapidly in the polls and will face Iván Cepeda in election runoff in three weeks The far-right lawyer Abelardo de la Espriella won the first round of Colombia’s presidential election on Sunday and will face senator Iván Cepeda, the candidate backed by leftwing pres

中文摘要 哥伦比亚极右翼律师埃斯普里耶拉在周日赢得总统选举第一轮投票,将于三周后与参议员伊万·塞佩达进行决选。

Donated milk reaches Cuba amid deepening shortages

Cuba has begun distributing donated powdered milk from Mexico and Uruguay as the island faces severe shortages.

中文摘要 古巴正面临严重物资短缺,已开始分发来自墨西哥和乌拉圭捐赠的奶粉以缓解危机。

The drivers risking death on Ukraine's most dangerous bus routes

Russian drones are targeting public buses in Kherson, killing three transport workers so far this year.

中文摘要 在乌克兰赫尔松,俄罗斯无人机瞄准公共巴士袭击,今年已导致三名运输工作者丧生,司机面临生命危险。

Cepeda, de la Espriella advance in Colombia’s presidential election

The left-wing senator and far-right newcomer will face each other in a run-off on June 21, with security a top issue.

中文摘要 哥伦比亚总统选举中,左翼参议员塞佩达与极右翼新人埃斯普里耶拉将于6月21日进入决选,安全问题成焦点。

2025 Wildfires Were the Costliest Ever, Researchers Say

Severe, hard-to-control blazes in densely populated areas like Los Angeles drove the year’s record losses.

中文摘要 研究人员称,2025年野火造成史上最高损失,主要因洛杉矶等人口密集地区发生难以控制的严重火灾。

Trump seeking edits to US-Iran deal, US media report

The requested changes are related to the Strait of Hormuz and the removal of highly enriched uranium, according to US media.

中文摘要 据美国媒体报道,特朗普正寻求对美伊协议进行修改,内容涉及霍尔木兹海峡和高浓缩铀的移除。

Israel seizes castle in Lebanon as it expands ground offensive

Prime Minister Benjamin Netanyahu calls the capture of the strategic fortress a "decisive shift" in Israel's campaign against Hezbollah, as European governments criticise the escalation.

中文摘要 以色列军队在黎巴嫩夺取一座战略城堡,扩大地面进攻,总理内塔尼亚胡称其为对抗真主党的决定性转折,欧洲政府批评升级。

China Exports Surveillance

The country has spent decades perfecting a surveillance state at home. Now it’s promoting its ideology of state control, and the technology to enforce it, abroad.

中文摘要 中国在国内完善监控体系后,现正向海外推广其国家控制的理念及相关技术,引发关注。

Child killed in house fire in south-west Melbourne home

Crews extinguished blaze before finding a child who died in Werribee home, Victoria police said Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast A child has died and a man has been left seriously injured after an early morning house f

中文摘要 维多利亚州警方称,墨尔本西南部韦里比一栋房屋发生火灾,一名儿童死亡,一名男子严重受伤。

Put a £5 deposit on vapes to stop littering, say waste companies

The industry body for waste companies says a refundable deposit would help boost vape recycling, but others disagree.

中文摘要 英国废物处理行业机构建议对电子烟加收5英镑可退还押金,以鼓励回收并减少环境污染,但此建议遭到部分人士反对。

A year of grief and waiting: What remains when a plane falls from the sky

A year after the Air India crash, a mother still speaks about her dead son in the present tense and a brother waits for answers.

中文摘要 在印度航空飞机失事一周年之际,遇难者家属仍在经历悲痛与等待,一位母亲仍以现在时态谈论儿子,一位兄弟则在等待答案。

China Investors Turn Sellers of HK Stocks in Rare Reversal

Chinese mainland investors turned net sellers of Hong Kong stocks for the first time in nearly three years in May, underscoring waning confidence in the city’s market.

中文摘要 中国内地投资者在5月份转向成为香港股票的净卖家,为近三年来首次,这一罕见逆转凸显了其对香港市场信心的减弱。

Vietnam’s Gene Solutions Eyes Hong Kong IPO to Fund Expansion

Vietnamese diagnostic testing startup Gene Solutions is preparing for a Hong Kong initial public offering as early as the second quarter of next year, a rare local firm seeking to access deeper capital markets and boost international visibility.

中文摘要 越南诊断测试初创公司基因解决方案(Gene Solutions)正筹备最早于明年第二季度在港交所上市,以寻求更深度的资本市场并提升国际知名度。

TSMC’s Local Investors Narrow Valuation Gap With Wall Street

Local investors are driving the premium on Taiwan Semiconductor Manufacturing Co.’s US-listed shares to a two-year low as they pile into the chipmaker’s stock in Taipei on bets the AI boom has further to run.

中文摘要 本地投资者大举买入台积电股票,押注人工智能热潮将持续,这推动其在美上市股票相对于台北股价的溢价收窄至两年最低。

Caribbean hot sauce producers warn of shortages and higher prices

Manufacturers in Jamaica say the key chilli peppers they need are in limited supply.

中文摘要 牙买加生产商警告,当地制造加勒比热辣酱所需的关键辣椒品种供应有限,可能导致产品短缺和价格上涨。

China’s Shoppers Are Buying Luxury Again as Stock Market Rebounds

Chinese consumers are showing signs of a renewed appetite for high-end beauty and fashion products, a rare bright spot for global luxury brands after years of weak demand and margin-eroding discounts in their most important growth market.

中文摘要 随着股市反弹,中国消费者对高端美妆和时尚产品的需求出现复苏迹象,为全球奢侈品牌在关键增长市场带来难得的积极信号。

Berkshire buys homebuilder Taylor Morrison for $8.5bn in Abel’s first big deal

First substantial acquisition under new CEO is bet on eventual recovery in property sector

中文摘要 伯克希尔·哈撒韦公司同意以85亿美元收购房屋建筑商泰勒·莫里森,这是新任CEO阿贝尔上任后的首笔重大交易,押注房地产行业最终将复苏。

Oil, Dollar Climb With US-Iran Deal Still Elusive: Markets Wrap

Oil climbed and the dollar strengthened as negotiations to extend the US-Iran ceasefire showed few signs of a breakthrough.

中文摘要 由于美国延长对伊朗停火的谈判未有突破迹象,国际油价上涨,美元走强。

Oil Rises From Six-Week Low Amid Uncertainty Over US-Iran Deal

Oil rose from a six-week low amid uncertainty over the outlook for a peace deal to end the war in Iran.

中文摘要 由于旨在结束伊朗战争的和平协议前景存在不确定性,油价从六周低点反弹上涨。

Petrobras Cuts Diesel Prices Amid Federal Subsidy Plan

Brazil’s state-controlled oil giant Petroleo Brasileiro SA is reducing domestic diesel prices starting Monday as part of a government subsidy program to shield consumers from the war in the Middle East.

中文摘要 巴西国家石油公司将于周一起下调国内柴油价格,作为政府补贴计划的一部分,以保护消费者免受中东战争的影响。

【CHY公益站】今日签到可获得525额度,祝各位佬友儿童节快乐!

首先先祝各位佬友儿童节快乐!谁还不是个不满12岁的宝宝呢~ CHY公益站今日签到可获额度525¥ 点击以下链接可获得1千-100e额度不等 cdk.linux.do LINUX DO CDK Linux Do 社区 CDK 快速分享平台 - 让分享变得更简单 点击以下链接获取注册码 cdk.linux.do LINUX DO CDK Linux Do 社区 CDK 快速分享平台 - 让分享变得更简单 90 个帖子 - 84 位参与者 阅读完整话题

官key的行为验证为什么会不达标呢?

缓存命中也不达标,推理质量也不行,工具能力也不行。看不懂了有点,这是第几个不认识官key的检测站了? 六一儿童节总不能真拿我当儿童吧?Excuse me? 11 个帖子 - 9 位参与者 阅读完整话题

弄了一个检测站点,看看到底谁在搞事情?

兄弟们我们看看谁在搞事情。 测试站 绝对不会收集各位的key信息,只是想看看谁有问题。 排行榜(模糊处理域名) 在说一次 ,永远不收费,也不会卖api,接受监督,公益服务已经关闭了,改到自用了。 28 个帖子 - 13 位参与者 阅读完整话题

Cherry Studio 隐私政策更新

各位 Cherry Studio 用户: Cherry Studio 早期版本中一直缺少完整、清晰的隐私政策说明,同时隐私相关开关在不同版本中的数据存储方式也不够一致,可能会导致部分用户在升级或切换版本后遇到隐私设置状态异常等问题。 从 Cherry Studio v1.9.8 版本开始,我们将正式加入隐私协议与更清晰的隐私设置说明,让用户能够更透明地了解软件会收集哪些匿名运行信息、不会收集哪些敏感数据,以及如何自行管理相关开关。 根据新版隐私协议,Cherry Studio 仅会在必要范围内收集匿名化的基础运行信息与产品改进信息,例如软件版本、功能使用汇总、功能活跃度与频次、错误日志与崩溃信

星辰AI 关于key的一点安全提醒

今天某位佬友把自己的key不小心上传到了github,被我朋友发现了,目前该key我已经协助禁用,在这里还是提醒各位佬友,一定要妥善保管好自己的key,尽量不要设置无限额度的key,尽可能降低损失。 其实newapi不支持根据key查用户的功能,我是直接联系newapitools的作者,紧急加了一个key反查用户的功能,在此也感谢newapitools的作者! 泄露自己的key就相当于把把钱扔在马路上,大家一定要注意!!! 祝大家早安,午安,晚安 18 个帖子 - 17 位参与者 阅读完整话题

【开源】Rust 字幕工具: SubForge,从转录、翻译到烧字幕一条命令完成【win 和 Linux 已验证】

本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 前言 最近在看外国课程,但奈何本人英语水平有限,机翻的字幕又简直灾难,于是参考各种方法做了一个还不错的工具,安装使用都很清晰(详见README.md),欢迎各位佬友品鉴 效果图: 正文 流程: 视频 / 音频 核心特点