每日简报

2026-06-10

← 历史归档

mvanhorn/last30days-skill

Python · ★ 37,353 · 🍴 3,030 · 📈 3,191 stars today

AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary

中文介绍 这个AI代理技能能自动研究任何主题,从Reddit、X、YouTube、HN、Polymarket和网络中提取信息,并合成基于事实的摘要。适用于需要快速获取跨平台信息的研究人员或内容创作者。

RyanCodrai/turbovec

Python · ★ 10,179 · 🍴 869 · 📈 1,801 stars today

A vector index built on TurboQuant, written in Rust with Python bindings

中文介绍 Turbovec是一个基于TurboQuant构建的向量索引,用Rust实现以提供高性能,并支持Python绑定。适合机器学习开发者用于高效相似性搜索和向量数据库应用。

roboflow/supervision

Python · ★ 42,988 · 🍴 3,836 · 📈 733 stars today

We write your reusable computer vision tools. 💜

中文介绍 Supervision项目提供一套可重用的计算机视觉工具,帮助开发者简化CV应用的开发流程。基于常见CV技术,适用于图像处理、目标检测和视频分析等场景。

opencv/opencv

C++ · ★ 88,635 · 🍴 56,622 · 📈 102 stars today

Open Source Computer Vision Library

中文介绍 OpenCV是一个开源的计算机视觉库,提供丰富的图像和视频处理算法。广泛应用于学术研究、工业界、机器人技术和嵌入式系统开发。

refactoringhq/tolaria

TypeScript · ★ 14,329 · 🍴 995 · 📈 829 stars today

Desktop app to manage markdown knowledge bases

中文介绍 Tolaria是一个桌面应用,专注于管理Markdown格式的知识库。帮助用户高效组织和检索笔记,适用于个人知识管理、团队协作和文档整理。

aaif-goose/goose

Rust · ★ 48,495 · 🍴 5,094 · 📈 489 stars today

an open source, extensible AI agent that goes beyond code suggestions - install, execute, edit, and test with any LLM

中文介绍 Goose是一个开源可扩展的AI代理,不仅能提供代码建议,还能安装、执行、编辑和测试代码。支持任何大语言模型,适用于开发者自动化软件开发流程。

Andyyyy64/whichllm

Python · ★ 4,095 · 🍴 227 · 📈 633 stars today

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

中文介绍 WhichLLM工具帮助用户找到在本地硬件上运行和性能最佳的大语言模型。基于真实及时的基准测试进行排名,而非仅看参数数量。一条命令即可测试运行,适合本地AI开发者。

TapXWorld/ChinaTextbook

Roff · ★ 73,470 · 🍴 16,447 · 📈 519 stars today

所有小初高、大学PDF教材。

中文介绍 ChinaTextbook项目汇集了中国从小学到大学的PDF教材资源。为学生和教师提供便捷的教育材料获取途径,用于学习和教学参考。

x1xhlol/system-prompts-and-models-of-ai-tools

★ 139,155 · 🍴 34,559 · 📈 79 stars today

FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sourced) System Prompts

中文介绍 这个项目收集了多种AI工具的系统提示和模型信息,如Augment Code、Claude Code等。帮助开发者理解AI工具的底层机制,适用于AI研究和开发。

yikart/AiToEarn

TypeScript · ★ 19,937 · 🍴 3,019 · 📈 402 stars today

Let's use AI to Earn!

中文介绍 AiToEarn项目旨在利用AI技术帮助用户实现盈利。可能涉及AI驱动的商业应用或自动化工具,适合创业者和AI爱好者探索赚钱机会。

phuryn/pm-skills

★ 13,432 · 🍴 1,541 · 📈 806 stars today

PM Skills Marketplace: 100+ agentic skills, commands, and plugins — from discovery to strategy, execution, launch, and growth.

中文介绍 PM Skills Marketplace提供超过100种代理技能、命令和插件,涵盖产品管理的全流程。从需求发现到策略制定、执行和增长,帮助产品经理高效工作。

santifer/career-ops

JavaScript · ★ 51,668 · 🍴 10,424 · 📈 1,110 stars today

AI-powered job search system built on Claude Code. 14 skill modes, Go dashboard, PDF generation, batch processing.

中文介绍 Career Ops是一个AI驱动的求职系统,基于Claude Code开发。提供14种技能模式、Go语言仪表板、PDF生成和批量处理功能,帮助求职者高效管理职位申请。

openai/plugins

JavaScript · ★ 2,608 · 🍴 305 · 📈 284 stars today

OpenAI Plugins

中文介绍 OpenAI Plugins项目是OpenAI官方的插件系统,允许开发者扩展和集成OpenAI模型的功能。适用于构建自定义AI应用和服务,增强AI交互能力。

maziyarpanahi/openmed

Python · ★ 1,869 · 🍴 215 · 📈 191 stars today

open-source healthcare ai

中文介绍 OpenMed是一个开源的医疗AI项目,旨在利用人工智能技术改善医疗健康领域。适用于医疗研究、诊断辅助、患者管理等场景。

francescopace/espectre

Python · ★ 8,211 · 🍴 633 · 📈 134 stars today

🛜 ESPectre 👻 - Motion detection system based on Wi-Fi spectre analysis (CSI), with Home Assistant integration.

中文介绍 ESPectre是一个基于Wi-Fi频谱分析(CSI)的运动检测系统,并集成Home Assistant。通过分析Wi-Fi信号实现无接触运动检测,适用于智能家居自动化。

addyosmani/agent-skills

Shell · ★ 49,828 · 🍴 5,567 · 📈 443 stars today

Production-grade engineering skills for AI coding agents.

中文介绍 Agent Skills项目为AI编码代理提供生产级工程技能,帮助代理更高效地处理软件开发任务。适用于构建和优化AI驱动的开发工具和自动化流程。

Loops: What Every AI Engineer Needs to Know in 2026

@sairahul1 · 113.0K 粉丝 · 852.6K 阅 · 600 赞 · 79 转

Peter Steinberger, creator of OpenClaw, who now works with OpenAI. Yesterday he posted this: "You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents."

中文介绍 分享OpenAI工程师Peter Steinberger的观点:AI工程的核心正从手动提示编码代理,转向设计能自动化提示代理的“循环”(Loops)系统。这标志着从直接交互到系统设计的范式转变。

Harness Engineering: What Every AI Engineer Needs to Know in 2026

@sairahul1 · 113.0K 粉丝 · 546.4K 阅 · 536 赞 · 94 转

In February 2026, a small OpenAI team shipped 1 million lines of production code. They didn't write a single line by hand. The AI agents wrote it. The humans designed the system that made the agents

中文介绍 分享一个案例:OpenAI的一个小团队在2026年2月用AI代理生成了100万行生产代码,人类工程师未手写一行。他们的核心工作是设计驱动代理的系统,体现了“工程化”思维在AI开发中的关键作用。

My Week with Fable

@MatthewBerman · 121.3K 粉丝 · 108.0K 阅 · 661 赞 · 26 转

tl;dr I've been testing Fable (Mythos) for the past week and it feels unlike any other model I've used. It feels, and is priced, like a next-generation model. It also has some real quirks. The Good

中文介绍 博主分享对Mythos/Fable模型为期一周的测试体验。认为其定价和体验接近下一代模型,但存在一些使用上的怪癖,并提供了具体的优缺点反馈。

Loop engineering: the 14-step roadmap from prompter to loop designer.

@0xCodez · 5.3K 粉丝 · 97.8K 阅 · 510 赞 · 80 转

Most developers still prompt their coding agents by hand. They type, they wait, they read the diff, they type again. 9out of 10 builders have never written a single loop that prompts the agent for

中文介绍 分享一份14步路线图,旨在帮助开发者从手动提示编码代理,转变为设计能自动化提示代理的“循环”。内容具体,属于从提示者到循环设计师的实用教程。

Designing loops with Fable 5

@RLanceMartin · 30.4K 粉丝 · 84.7K 阅 · 660 赞 · 50 转

Mythos-class models like Claude Fable 5 have changed the way many of us work at Anthropic. I want to share two tips for getting the most out of this class of models. Self-correction loops There’s been

中文介绍 Anthropic员工分享在工作中使用Claude Fable 5等Mythos级模型的两个技巧,重点介绍了“自纠正循环”等高效用法,以最大化此类新模型的能力。

Do AGENTS.md Files Actually Help Coding Agents?

@rasbt · 459.9K 粉丝 · 50.9K 阅 · 507 赞 · 51 转

Catching up with the agent-related research literature, one paper that definitely got my attention is "Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?." It looks

中文介绍 博主分析一篇研究论文,探讨AGENTS.md这类为仓库级别上下文设计的文件,是否真的能提升编码代理的性能。属于对AI开发工具有效性的评估与讨论。

WTF Is a Loop? Peter Steinberger vs. Boris Cherny

@mvanhorn · 30.8K 粉丝 · 45.6K 阅 · 567 赞 · 47 转

The most repeated sentence in AI coding this week is six words long, and almost nobody saying it can define it. One tweet had the entire timeline in a chokehold this week, so I ran /last30days on the

中文介绍 针对近期AI编码领域热门的“Loop”概念,博主进行澄清和解释。旨在帮助社区理解这个被广泛提及但定义模糊的术语究竟是什么。

Implications of Large-Scale Test-Time Compute

@polynoamial · 129.2K 粉丝 · 43.7K 阅 · 707 赞 · 80 转

tl;dr: As LLMs become more capable, benchmark performance is increasingly a function of test-time compute. In fact, we likely don't know what the capability ceiling is for modern LLMs because it's too

中文介绍 讨论大模型评估的新视角:随着模型能力增强,基准测试成绩越来越取决于测试时的计算量(test-time compute)。这暗示我们可能尚未摸清当前模型的能力天花板。

17 prompts that make Hermes run while you sleep (copy-paste inside)

@Mnilax · 7.3K 粉丝 · 43.5K 阅 · 502 赞 · 42 转

In February 2026, Nous Research released Hermes Agent: an open-source, self-hosted agent that doesn't live inside an IDE and doesn't forget when the tab closes. It runs as a daemon on your own box,

中文介绍 分享17条可直接复制的提示模板,用于驱动开源、可自托管的Hermes代理在后台自主运行,实现自动化任务处理,提供了具体的提示工程案例。

I Built an Agentic Harness From Scratch. That Taught Me What Agents Actually Are

@ByteMohit · 2.0K 粉丝 · 43.3K 阅 · 501 赞 · 51 转

Everyone is building with agents. Almost nobody talks about what is actually inside one. Not the model. The harness around it. I spent the last few months building one from scratch in Python every

中文介绍 博主分享从零用Python构建AI代理“harness”的实践经验。旨在揭示代理内部工作机制,强调真正重要的是围绕模型的系统架构,而非模型本身。

Loop Engineering.

@addyosmani · 395.5K 粉丝 · 42.7K 阅 · 577 赞 · 62 转

Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead. A loop here can be thought of a recursive goal where you define a purpose and

中文介绍 谷歌工程师阐释“循环工程”的核心理念:用设计好的循环系统替代人工提示,让系统自主、递归地执行目标,从而将人从提示循环中解放出来。

Your Agent Harness Should Repair Itself

@akshay_pachaar · 276.5K 粉丝 · 38.0K 阅 · 504 赞 · 65 转

When an AI agent fails in production, your observability tool shows you exactly what it did and almost nothing about how to fix it. You get a clean trace of the run, every model call and tool that

中文介绍 提出AI代理的运行时框架(harness)应具备自我修复能力。当前观测工具只能显示失败轨迹,而无法自动诊断并修复问题,这需要新的工程解决方案。

Fluid, natural voice translation with Gemini 3.5 Live Translate

@GoogleAIStudio · 176.3K 粉丝 · 32.2K 阅 · 517 赞 · 55 转

Twenty years ago, translation at Google began as one of our pioneering machine learning experiments to turn the science of language into the magic of human connection. That experiment has come a long

中文介绍 GoogleAIStudio宣布推出Gemini 3.5 Live Translate功能,提供流畅、自然的实时语音翻译,是其在语言AI领域长期探索的新成果。

The Fourth Era of Compute

@unicity_labs · 126.3K 粉丝 · 7.0K 阅 · 821 赞 · 405 转

For thirty years the internet has been built for one kind of actor. The next one is not a person. For thirty years, compute has been built for humans. It is shaped around human attention, human

中文介绍 探讨计算范式的第四次转变:过去三十年互联网和计算架构为人类用户设计,而未来将围绕AI代理这一新的行为主体进行重构。

Fluid, natural voice translation with Gemini 3.5 Live Translate

Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.

中文介绍 Gemini 3.5 Live Translate将近乎实时的自然语音翻译引入谷歌AI Studio、谷歌翻译和谷歌Meet。

How engineers at Nextdoor use Codex to build without limits

How engineers at Nextdoor use Codex with GPT-5.5 to investigate hard-to-reproduce issues, build across platforms, and focus on product outcomes.

中文介绍 Nextdoor工程师如何使用Codex与GPT-5.5来调查难以复现的问题、跨平台构建并专注于产品成果。

Learning to lead in a hybrid human-AI enterprise

As adoption of AI agents looks set to surge by as much as 300% in the next two years, leadership teams are carefully considering the implications of a hybrid human-AI workforce. Unlike existing enterprise-level automation that relies on manual input, AI agents are capable of autonomously coordinatin

What Codex unlocks for Notion

How Notion uses Codex to one-shot specs, build AI Voice Input for the web, and multiply engineering power across small teams.

中文介绍 Notion如何使用Codex一次性生成规范、构建网页版AI语音输入,并增强小团队的工程能力。

Five things you need to know about AI

At SXSW London last week I gave a talk called “Five things you need to know about AI,” in which I shared what I think are the biggest themes in AI right now. I pulled a few things from our first AI10 list, an annual guide to the most important trends in this buzzy world,…

中文介绍 上周在SXSW伦敦的演讲中,分享了「关于AI你需要知道的五件事」,涵盖当前AI最大主题,并参考了首个AI10年度趋势指南。

Industrial policy for the Intelligence Age

Explore our ambitious, people-first industrial policy ideas for the AI era—focused on expanding opportunity, sharing prosperity, and building resilient institutions as advanced intelligence evolves.

中文介绍 探讨为AI时代设计的雄心勃勃、以人为本的工业政策理念,聚焦于扩大机会、分享繁荣,并随着先进智能的发展建立弹性机构。

Confidential submission of draft S-1 to the SEC

OpenAI confirms a confidential S-1 submission to the SEC and has not yet determined timing for further action.

中文介绍 OpenAI确认向美国证券交易委员会保密提交S-1草案,但尚未确定进一步行动的时间。

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks

第一作者: Qin Yang · 方向: AI 安全

Abstract:Large language model (LLM)-powered content moderation systems have become a critical defense against harmful online content. However, these systems primarily operate on tokenized text and largely ignore the visual cues that humans naturally rely on when interpreting content. We show that this discrepancy creates a fundamental perceptual mismatch: content that is readily recognized as harmful by humans can become effectively invisible to automated moderation systems. To study this vulnerability, we introduce a class of Human-Perceptible Adversarial Attacks (HPAA), in which harmful expressions are embedded into otherwise benign text through visually salient typographic manipulations. Our key insight is that typographic features, including spacing, visual emphasis, and spatial arrangement, can be strategically combined to preserve human recognition of harmful content while...

论文介绍 大型语言模型内容审核系统依赖分词文本,忽略人类视觉线索,导致人类可识别的有害内容对系统不可见。本研究引入人类可感知对抗攻击(HPAA),通过排版特征如间距和视觉强调嵌入有害表达,使人类可识别但系统失效。这揭示了审核系统的感知失配问题,对提升内容审核安全性有重要意义。

Observability for Delegated Execution in Agentic AI Systems

第一作者: Abhinav Mishra · 方向: AI 安全

Delegation-scoped execution is not identifiable from standard observables: audit logs and execution traces can be identical under multiple incompatible delegation assignments. This gap is especially acute in LLM-based agentic systems, where agents dynamically select tools, vary execution sequences across runs for the same instruction, and spawn cooperating sub-agents. These dynamics fragment and interleave traces, making delegation-scoped reconstruction from causal structure alone structurally underdetermined. Although individual actions are authorized and logged, existing audit, tracing, and security schemas lack the semantics to reconstruct what actions occurred under a given delegation across heterogeneous systems. We focus on delegation-scoped attribution and access/share footprint reconstruction, not intent inference or reasoning reconstruction. We present an agent-aware...

论文介绍 代理AI系统中,动态工具选择和子代理协作导致执行跟踪碎片化,使委托归因结构上不确定,审计日志无法区分不同委托分配。本文提出代理感知方法,专注于委托范围的属性和访问足迹重建,而非意图推断。该方法能处理异构系统中的动态性,增强安全性和审计能力。

FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing

第一作者: Yuhan Ma · 方向: 密码学协议

Abstract:Two-server secure inference allows a client to query a hosted large language model (LLM) without revealing prompts or embeddings. Recent GPU systems based on function secret sharing (FSS) make linear layers efficient, but fixed-point nonlinearities and helper operations remain a bottleneck because each operator is typically implemented as a bespoke protocol with its own comparisons, wrap-around corrections, and preprocessing material. We present FuseFSS, a compiler that replaces per-operator protocol design with a single compilation pipeline. For each scalar fixed-point operator, a compact specification lists its interval partition, low-degree arithmetic pieces, and required predicate bits. The compiler emits two batched FSS evaluations on the public masked value: one packed comparison that returns all predicate bits, and one vector interval lookup that returns the active...

论文介绍 两服务器安全推理允许客户端查询大型语言模型而不泄露提示或嵌入。现有GPU系统基于函数秘密共享(FSS),但非线性操作效率低。本文提出FuseFSS编译器,将每个算子规范编译为批量FSS评估,优化比较和查找操作,实现高效隐私保护的LLM推理。

SecureClaw: Clawing Back Control of LLM Agents

第一作者: Yuhan Ma · 方向: 密码学协议

Abstract:Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext inside the runtime before any final output check can intervene. Existing defenses usually protect one boundary, either the planner/runtime or the action sink, and therefore do not by themselves secure both surfaces. We present SecureClaw, a dual-boundary architecture that places authorization at the effect sink and plaintext confinement at the read boundary. Sensitive reads pass through a trusted gateway that replaces raw values with opaque handles and, in the evaluated deployment, bounded summaries as an explicit declassification interface. Writes that change external state follow a PREVIEW$\rightarrow$COMMIT protocol in which only a trusted executor may commit the exact canonical request authorized by policy. The runtime can...

论文介绍 工具使用大型语言模型代理面临未授权外部动作和敏感信息泄露风险。现有防御通常只保护一个边界。SecureClaw提出双边界架构,在动作接收端进行授权,在读边界进行隐私约束,使用可信网关替换敏感值为不透明句柄,并通过PREVIEW-COMMIT协议控制外部状态更改,提升代理系统整体安全性。

Model Poisoning Against Federated Model Adaptation with Chain of Bit-Flips

第一作者: Bastien Vuillod · 方向: AI 安全

Abstract:Federated Learning (FL) allows a set of clients to collectively train a global model without sharing local training data. Giving the responsibility of the training to decentralized actors may lead to poisoning attacks: clients controlled by malicious third party potentially poison the training dataset to install a backdoor in neural networks. In FL, these backdoor attacks rely solely on algorithmic approach, however, recent advances in hardware faults threats (e.g, Rowhammer) have widen the overall attack surface. In the context of federated model adaptation, we introduce a novel category of backdoor attack against FL systems that relies on model poisoning based on hardware-fault attacks. More precisely, we propose a task-agnostic backdoor attack that is implanted during the FL training time by inducing hardware faults (bit-flips) in parameters of a single local model. The...

论文介绍 联邦学习易受投毒攻击,传统基于算法,但硬件故障威胁如Rowhammer扩展了攻击面。本文提出基于硬件故障的模型投毒攻击,在联邦模型适应中,通过诱导位翻转植入任务无关后门。该方法在训练时影响单个本地模型,展示了硬件层面的新安全风险。

Towards Post-Quantum Secure Pharmacovigilance with ML-KEM and ML-DSA

第一作者: Saee Desai · 方向: 密码学协议

Abstract:Pharmacovigilance systems handle sensitive healthcare and drug-safety data, including adverse event reports and clinical observations. As quantum computing advances, classical public-key cryptographic systems such as RSA and elliptic-curve cryptography may become vulnerable, creating long-term risks for healthcare data that must remain confidential for many years. This paper presents an educational prototype of a post-quantum secure pharmacovigilance data pipeline. The system uses ML-KEM-768 for post-quantum key establishment, HKDF-SHA-256 for deriving an AES key, AES-256-GCM for efficient file encryption, and ML-DSA-65 for digital signatures and tamper detection. The pipeline supports multiple file formats, including TXT, CSV, JSON, and PDF, by treating files as raw bytes and preserving metadata for reconstruction at the receiver. The prototype includes separate hospital...

论文介绍 药物警戒系统处理敏感医疗数据,面临量子计算威胁,需长期保密。本文构建后量子安全数据管道原型,使用ML-KEM-768进行密钥建立,AES-256-GCM加密,ML-DSA-65签名,支持多种文件格式。旨在确保医疗数据长期保密性,对抗未来量子攻击。

Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs

第一作者: Charles Westphal · 方向: 安全研究

Abstract:Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs. This creates a steganographic exfiltration risk that is difficult to detect with output-level steganalysis. Recent work proposes mechanistic detection using linear probes that recover the secret from internal activations. We show that this defense can be systematically evaded, but that detectability can be recovered through a targeted data-level intervention. First, we extend the detection setup to include a non-linear MLP probe. We then adversarially fine-tune steganographic trojans across five base models: Qwen3-8B, Llama-3.1-8B, Ministral-8B, Qwen3-14B, and Phi-4-14B. The resulting models retain $58$--$79\%$ exact-match secret recovery while evading both ridge and held-out MLP probes, with $1$--$8\%$ average capability degradation across six benchmarks. We then give...

论文介绍 大型语言模型可微调以编码秘密到输出中,造成隐写泄露风险。现有检测方法如线性探针可被规避。本文扩展检测设置,包括非线性MLP探针,并通过对抗性微调规避防御。同时提出数据级干预以恢复可检测性,平衡模型能力与安全。

Fully Oblivious Differential Privacy for Frequency Estimation in the Augmented Shuffle Model with Trusted Processors

第一作者: Takao Murakami · 方向: 密码学协议

Abstract:In the shuffle model of DP (Differential Privacy), a shuffler randomly permutes users' data to achieve high accuracy and privacy. Recent studies show that most existing shuffle protocols are vulnerable to collusion attacks by the data collector and users. They address this issue by introducing the augmented shuffle model that incorporates random sampling and dummy data addition into the shuffler. However, it remains open how to ensure the shuffler follows the protocol and does not collude with the data collector in this model. We address this trust issue by thoroughly exploring the augmented shuffle model with TEEs (Trusted Execution Environments). We first introduce a new privacy notion, FODP (Fully Oblivious DP), which strengthens DP to prevent various TEE side-channel attacks based on external/internal memory access patterns and control flows. We propose a general framework...

论文介绍 差分隐私洗牌模型中,洗牌器可能与数据收集者合谋,现有协议脆弱。增强洗牌模型引入随机采样和虚拟数据,但信任问题仍存。本文引入完全遗忘差分隐私(FODP),防止TEE侧信道攻击,并提出框架在可信处理器中实现,确保频率估计的隐私性。

Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents

第一作者: Jianwei Tai · 方向: AI 安全

BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call \emph{brain-prompt injection}: signal-side perturbations, context-only injections, and adaptive dual-decoder attacks can all change the routed action while EEG-side or text-side monitors remain blind. Route safety in this stack depends on what the audit log can observe, not on decoder accuracy or agreement alone. We define a Route-Safety Audit Contract: a minimal log schema, denominator hierarchy, and endpoint specification, and prove an audit-schema separation theorem together with a C3 attacked-dependence decomposition; clean agreement and marginal robustness do not identify the joint term that controls C3 routing. As a calibration layer on top of the contract, we apply split-conformal calibration to a non-oracle EEG confirmation channel and...

论文介绍 本文探讨脑机接口与大语言模型代理结合后面临的一种新型安全威胁——「脑提示注入」。攻击者可通过扰动脑电信号或注入上下文,在不改变模型解码器的情况下篡改代理的决策路由。作者为此系统定义了「路线安全审计合约」,包括审计日志的最小化结构与规范,旨在为基于神经信号授权的智能代理提供可验证的安全保障,而非依赖解码精度。

Trustworthy Smart Fabs via Professional Proxies: Scaling Safe and Sustainable by Design (SSbD) through Industrial Data Spaces

第一作者: Han-Teng Liao · 方向: 系统安全

Abstract:The convergence of the 2026 European Union Safe and Sustainable by Design (SSbD) framework, Corporate Sustainability Due Diligence Directive (CSDDD), and Carbon Border Adjustment Mechanism (CBAM) introduce a severe governance bottleneck for advanced semiconductor manufacturing facilities ("Smart Fabs"). Regulatory compliance demands have surpassed the capacity of manual corporate reporting, creating a direct conflict between multi-stakeholder transparency and corporate data privacy. This paper addresses this challenge by introducing a zero-trust socio-technical orchestration framework that operationalizes a six-layer SSbD reference architecture within trustworthy industrial data spaces. We propose a shift from reactive automation to autonomous governance through "Professional Proxies"-role-based agentic workflows executing within hardware-isolated trust zones. Structured as an...

论文介绍 本文针对先进半导体制造设施在满足欧盟SSbD等法规时面临的数据隐私与透明度矛盾,提出了一种基于工业数据空间的零信任编排框架。该框架引入「专业代理」角色,通过硬件隔离信任域中的代理工作流,实现自主化治理,旨在帮助智能晶圆厂在保护商业机密的同时,合规地扩展安全与可持续设计流程。

Pretrained, Frozen, Still Leaking: Auditing Cross-Encoder Attribute Transfer in EEG Foundation Models

第一作者: Jianwei Tai · 方向: 安全研究

Abstract:EEG foundation-model releases are usually audited one endpoint at a time: raw-reconstruction, membership inference, identity linkage, or DP-SGD on the downstream head. We audit the same released embeddings under all four endpoints jointly, on BIOT, LaBraM, and EEGPT, and show that each single-endpoint audit clears releases that still leak spectral attributes. The decisive evidence is a cross-encoder transfer audit: a single ridge attribute decoder learned from one frozen encoder transfers, via a fitted linear bridge, to held-out-subject test splits of every other encoder, with subject-disjoint matched-control 95% CI lower bound at least 0.081 across all six BIOT/LaBraM/EEGPT directions. We prove a sufficient condition: two encoders sharing a nontrivial attribute-coordinate projector overlap beta admit a chained ridge bridge attacker with centered-gain lower bound...

论文介绍 研究对多个开源EEG基础模型的隐私审计表明,仅对单一端点(如身份链接)进行审计会遗漏关键风险。作者提出跨编码器属性转移审计,发现即使模型参数冻结,一个编码器学习到的属性解码器也能迁移到其他编码器上,证明它们共享着可泄露光谱属性的潜空间结构。这揭示了当前发布审计方法的不足,强调需要联合评估。

EnclaveScale: Hardware-Assisted Edge-DP for Secure Data Centre Power Telemetry

第一作者: Hung Dang · 方向: AI 安全

Abstract:EnclaveScale is a distributed, hardware-assisted telemetry architecture providing post-extraction attestation, enabling operators to collaboratively model high-resolution generative AI power transients. Existing cryptographic techniques scale poorly for 10-Hz streaming or fail to authenticate origins, permitting malicious hosts to spoof sensor inputs. We implement and evaluate a post-extraction pipeline utilizing DCAP attestation, differential privacy noise injection, and Byzantine rejection across 32 GCP Confidential VMs, achieving 0\% post-extraction attack success rate. This edge-DP approach distils continuous GPU transients into discrete Markov-chain transition matrices, guaranteeing event-level differential privacy. To mitigate pre-ingestion vulnerabilities, we propose an SPDM-authenticated first-mile layer. While current platforms lack attested I/O, emerging hardware...

论文介绍 本文提出EnclaveScale架构,为数据中心高频率的功率遥测流提供硬件辅助的边缘差分隐私保护。该系统利用DCAP认证、差分隐私噪声注入和拜占庭节点拒绝机制,实现数据提取后的来源认证,旨在使多个运营者能协作建模生成式AI的功率瞬态,同时防止恶意主机伪造传感器输入,保障协作数据隐私与安全。

Customization under Fire: Plugin Poisoning in Text-to-Image Ecosystem

第一作者: Jiahao Chen · 方向: 安全研究

Abstract:The prosperity of text-to-image (T2I) models has fostered a vibrant share-and-play ecosystem centered on Low-Rank Adaptation (LoRA) plugins, which allow users to customize and share model capabilities with ease. This democratization, however, comes with a hidden but severe security risk. Malicious users could share and distribute seemingly benign LoRA plugins that contain hidden functionalities to poison the model-sharing market, like Civitai or Liblib, severely undermining the user trust that underpins this collaborative ecosystem and threatening the safety of countless downstream applications. Despite these risks, plugin poisoning in the real-world T2I ecosystem remains underexplored. This paper introduces PoisonLoRA, the first systematic study of LoRA plugin supply-chain risks that exploits the trust and characteristics within the T2I ecosystem. We identify two primary...

论文介绍 随着LoRA插件共享生态的繁荣,其供应链安全风险凸显。本文系统研究了LoRA插件的投毒攻击,揭示了攻击者如何在看似无害的插件中植入隐藏功能,从而破坏模型共享市场(如Civitai)的信任基础。研究识别了此类攻击的模式,旨在警示开源社区,并呼吁加强插件分发平台的安全审查机制。

PrivCode++: Latent-Conditioned Differentially Private Code Generation for Comprehensive Guarantees

第一作者: Zheng Liu · 方向: 软件安全

Abstract:Large language models fine-tuned on instruction-code pairs may memorize and subsequently leak sensitive training data. Existing differentially private (DP) code generation methods primarily protect code snippets while assuming prompts are public, which fails in realistic scenarios where prompts may also contain sensitive information. When prompts cannot be explicitly learned or used during generation, code synthesis suffers from severe utility degradation as well as reduced diversity and fidelity. To address these challenges, we propose PrivCode-Plus, the first work to explore DP code generation where both prompts and code snippets are considered sensitive in LLM fine-tuning. PrivCode-Plus introduces a two-stage DP framework with a Privacy-Free Latent Conditioning module, enabling effective DP fine-tuning and data synthesis without direct access to sensitive prompts or code...

论文介绍 现有差分隐私代码生成方法通常假设提示是公开的,这无法应对提示本身也包含敏感信息的真实场景。本文提出PrivCode++,首个将提示和代码片段均视为敏感数据的DP代码生成框架。它引入一个「隐私自由潜条件」模块,实现在不直接访问敏感数据的情况下进行有效的DP微调和数据合成,提升了模型实用性。

Steganography Without Modification: Hidden Communication via LLM Seeds

第一作者: Felix Mächtle · 方向: 软件安全

Abstract:We demonstrate that widely deployed Large Language Model (LLM) inference stacks harbor a steganographic channel that requires no modification to model weights, sampling code, or output distributions. The channel exploits a structural property of deterministic decoding: pseudo-random number generators (PRNGs) used in inverse-transform sampling produce a seed-dependent sequence of token-level probability intervals that can be reconstructed from the generated text alone. A sender encodes a secret message in the PRNG seed before generation; a receiver reconstructs the intervals and recovers the seed, and thus the hidden payload, by exhaustive search over the seed space. We formalize two operational modes. In the known-prompt setting, sender and receiver share the prompt, enabling exact interval reconstruction and perfect seed recovery via forced alignment. In the unknown-prompt...

论文介绍 本文揭示了广泛部署的LLM推理栈中存在一种无需修改模型的隐写通道。该通道利用确定性解码中伪随机数生成器种子的依赖性:发送方将秘密信息编码在种子中,接收方根据生成文本重建概率区间序列并恢复种子。此方法在已知和未知提示场景下均可操作,暴露了当前LLM系统可能存在的隐蔽通信风险。

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges

第一作者: Tiejin Chen · 方向: 隐私保护

Abstract:Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information. However, Multi-modal Large Language Models (MLLMs), which process both text and images, introduce unique privacy challenges that remain underexplored. Compared to text-only models, MLLMs can extract and expose sensitive information embedded in images, posing new privacy risks. We reveal that some MLLMs are susceptible to privacy breaches, leaking sensitive data embedded in images or stored in memory. Specifically, in this paper, we (1) introduce MM-Privacy, a comprehensive dataset designed to assess privacy risks across various multi-modal tasks and scenarios, where we define Disclosure Risks and Retention Risks. (2) systematically evaluate different MLLMs using MM-Privacy and demonstrate how models leak sensitive data across various...

论文介绍 与文本模型相比,处理图像和文本的多模态大语言模型可能暴露嵌入图像中的敏感信息,带来独特的隐私挑战。本文构建了MM-Privacy评估数据集,并系统地评估了多个MLLMs的「披露风险」与「留存风险」,实证了模型可能泄露图像中或记忆里的敏感数据,旨在推动针对多模态场景的隐私保护技术发展。

Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps

第一作者: Xiaofeng Lin · 方向: AI 安全

Abstract:Tool-using LLM agents interact with the world through actions that persist state in artifacts (e.g., workspace files or logs). Consequently, jailbreak defenses must reason about cross-step composition rather than isolated text. Yet most existing attacks and defenses, including ``multi-turn'' jailbreaks such as Crescendo and Tree of Attacks,still assume a single contiguous conversation visible to the defender. This assumption breaks down in real agent pipelines, where enforcement is fragmented across tools, modules, and time, and where artifact provenance is often not tracked. We operationalize a deployment failure mode for tool-using LLM agents, the \emph{provenance gap}, and study reproducible triggers for it: \emph{Context-Fractured Decomposition} (CFD), a family of cross-context multi-step jailbreaks that preserve benign-looking intermediate artifacts from an early...

论文介绍 研究工具使用LLM代理中的安全漏洞,特别是跨步骤防御的弱点。提出上下文破碎分解攻击,利用provenance gap,在真实代理管道中实现越狱。该研究揭示了现有防御假设的局限性,对代理安全评估有重要意义。

Security-First Approach to API Pipeline Development with Zero-Trust Architecture

第一作者: Mahima Agarwal · 方向: 软件安全

Abstract:Modern enterprises face an accelerating onslaught of API-targeted threats amid a rapidly expanding attack surface. Record volumes of software vulnerabilities continue to accelerate dramatically, with 28,818 CVEs disclosed in 2023 (a 38% jump from 2022) and 40,009 CVEs in 2024 (another 38% increase), while the average time-to-exploit (TTE) of new flaws shrank to mere days (approximately 5 days in 2023, down from 32 days in 2021). At the same time, API usage dominates web traffic and has become a primary vector for breaches - 99% of organizations experienced API security incidents in the last year, with 22% suffering actual data breaches via APIs (based on industry vendor research). This paper proposes a comprehensive "security-first" framework for API pipeline development, leveraging Zero-Trust Architecture principles within DevSecOps practices to counter these trends. We...

论文介绍 面对API安全威胁加剧,提出安全优先的API管道开发框架。基于零信任架构原则,集成到DevSecOps实践中,以应对漏洞快速增长和数据泄露风险。旨在提升企业API防护能力,减少安全事件。

Document-Authored Control-Signal Impersonation: A Low-Cost Indirect Prompt Attack on RAG Safety Boundaries

第一作者: Jianguo Zhu · 方向: AI 安全

Abstract:Retrieval-augmented generation (RAG) systems often serialize user queries, retrieved documents, metadata, system labels, and task instructions into one natural-language prompt. We study a source-authority boundary failure in this design: attacker-authored retrieved text can impersonate metadata, provenance, authority, or disclosure-policy signals that appear control-relevant to the model. We call this pattern Document-Authored Control-Signal Impersonation (DACSI). DACSI is a non-imperative, metadata-like payload subclass within indirect prompt injection. Its central lesson is simple: document-authored labels are data, not policy. Command-style injection asks the model to ignore, override, or violate policy; DACSI asks whether untrusted document text can be misattributed as an authorized control signal when RAG prompt rendering collapses trusted and untrusted text into the same...

论文介绍 针对检索增强生成系统的安全边界,研究文档作者控制信号冒充攻击。攻击者通过检索文档冒充元数据或权威信号,误导模型决策。这种低攻击成本的方法凸显了RAG系统中数据与策略区分的挑战。

Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops

第一作者: Ziqian Zhong · 方向: 软件安全

Abstract:Agent benchmarks score submissions with outcome verifiers that are typically hand-written and brittle, leaving them open to reward hacking. We audit 1,968 tasks across five terminal-agent benchmarks and find 323 (16%) hackable by frontier models given only the task description. This corrupts both leaderboard rankings and RL training signal, yet the standard response is manual and reactive. We introduce the hacker-fixer loop, a method for building exploit-resistant verifiers without per-task manual patching. The loop alternates three LLM agents: a hacker tries to pass the verifier without solving the task, a fixer patches the verifier to reject each discovered exploit, and a solver confirms the patched verifier still admits legitimate solutions. The loop iterates: each patch reshapes what the verifier rewards, surfacing the next exploit. We further add verifier access, and let...

论文介绍 代理基准的验证器易受奖励黑客攻击影响,导致排名和训练信号失真。提出黑客-修复循环方法,通过LLM代理交替黑客攻击、修复和验证,自动加固验证器,提高基准的抗攻击性。

Block-A-Mole: The Sustainability Frontier of Moving-Target Censorship Resistance

第一作者: Anindya Maiti · 方向: 系统安全

Abstract:Internet censorship affects over four billion people, and deployed circumvention systems share a common weakness: their endpoints are fixed and discoverable, so a patient censor can enumerate and block them. Moving-target circumvention systems instead rotate endpoints across commercial cloud address space faster than censors can react, but the field lacks a theory of when rotation works, leaving rotation intervals and pool sizes to intuition. We give the first formal account of moving-target censorship resistance by modeling the censor-defender interaction as a continuous-time timing game over a combinatorial address-domain space, generalizing FlipIt to a collateral-bounded adversary. We prove a sustainability frontier separating configurations a censor can defeat from those it cannot, and show that under the Great Firewall's 2024 shift to blocking QUIC and TLS by domain, raw...

论文介绍 移动目标审查抵抗系统通过快速轮换端点规避封锁,但缺乏理论指导。将审查者-防御者交互建模为连续时间博弈,证明可持续性前沿,为系统配置提供理论依据,应对如Great Firewall的封锁策略。

Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography

第一作者: Prateek Paudel · 方向: 安全研究

Abstract:The aim of steganography is to hide secret information inside ordinary media so that the existence of communication is hidden rather than encrypted. In audiovisual context, the availability of audio and video streams creates an opportunity to split a payload across these two modes thus, reducing the embedding burden on any single carrier. This paper evaluates whether such split-payload audiovisual steganography can help evade unimodal and multimodal steganalysis under synchronized and asynchronous embedding settings. We create audiovisual samples where the hidden message is divided between the audio and video tracks, and then test how well different detectors can identify them. The single mode detectors performs close to random guessing, thus showing the benefit of this hiding mechanism, while the multimodal model initially appears more effective. However, further checks show...

论文介绍 评估分割有效载荷的音视频隐写术对隐写分析的效果。通过将秘密信息分割到音频和视频轨道,测试检测器性能。发现单模态检测器效果差,多模态模型初始有效但需进一步验证其鲁棒性。

AutoSUT: The Environment Semantics Gap in Structured CTI for Adversary Emulation

第一作者: Sidnei Barbieri · 方向: 软件安全

Abstract:Structured Cyber Threat Intelligence (CTI) is increasingly used for adversary emulation, detection evaluation, and cyber range design. However, these workflows still require a target System Under Test (SUT) whose environment is not fully described by public CTI. We measure how much of that environment can be derived from MITRE ATT&CK Structured Threat Information Expression (STIX) bundles. Using the ATT&CK Enterprise, Mobile, and Industrial Control Systems datasets, with CAPEC and FiGHT as comparison datasets, we evaluate platform coverage, software specificity, vulnerability evidence, and deployment compatibility. Platform annotations are common, but software references rarely include versions or Common Platform Enumeration (CPE) identifiers. In Enterprise, 97.6% of software objects lack both, and campaign-level Common Vulnerabilities and Exposures (CVEs) remain sparse. Our...

论文介绍 结构化网络威胁情报用于对手仿真时,环境描述存在语义差距。通过分析MITRE ATT&CK数据集,发现平台覆盖和软件特异性不足,影响仿真实施的准确性,提出改进方向。

Asymptotic Optimality of the High-Dimensional Gaussian Mechanism and Improved Low-Dimensional Mechanisms for Differential Privacy

第一作者: Yu Wei · 方向: 隐私保护

Abstract:The additive noise mechanism is a foundational tool for differential privacy (DP) of $T$-dimensional real-valued vector queries. The Gaussian mechanism, utilizing Gaussian noise, is the mostly widely used such mechanism, due to its simplicity and strong privacy guarantees. In this work, we provide justification for this choice, showing that as the dimension $T\to\infty$, no additive-noise mechanism can asymptotically improve on the Gaussian mechanism's privacy--utility tradeoff for the strong privacy settings typically this http URL also develop a new family of \emph{Spherical Generalized Gamma} DP mechanisms, which contains both the Gaussian mechanism and the recently studied $\ell_2$ mechanism (Joseph \emph{et al.}, ICML 2025). We identify members of this family that outperform both the Gaussian and $\ell_2$ mechanisms in certain low-dimensional settings, and show tight...

论文介绍 研究差分隐私中噪声机制的隐私-效用权衡。证明高维下高斯机制渐近最优,并提出球形广义伽马机制家族,在低维设置下改进性能,为隐私保护提供理论支持。

X-rated Compliance Theater: An Empirical Evaluation of European Age Verification Systems in Adult Websites

第一作者: Simone Lavermicocca · 方向: 隐私保护

Abstract:Age verification is rapidly emerging as a central regulatory instrument for protecting minors online, with several jurisdictions mandating its deployment for access to adult and pornographic content. This regulatory direction raises significant privacy concerns, as it risks binding sensitive content access to identity-related attributes. It also introduces security risks, since age-verification mechanisms are often outsourced to third-party providers with limited transparency into the robustness of their verification processes. In this work, we conduct, to the best of our knowledge, the first exploratory security assessment of regulation-mandated age-verification mechanisms deployed by adult websites. Rather than treating age verification as a purely regulatory question, we empirically examine whether current deployments provide security guarantees commensurate with the...

论文介绍 本文实证评估欧洲成人网站上法规强制实施的年龄验证系统,探讨其隐私和安全问题。研究聚焦于当前部署是否提供与监管目标相称的安全保证,发现存在敏感内容访问与身份属性绑定风险,以及第三方验证过程不透明带来的安全隐患。结果可能为设计更安全的年龄验证方案提供指导。

Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems

第一作者: Kuncan Wang · 方向: AI 安全

Abstract:Data agents integrate LLM-driven reasoning with relational data access, executable analytical tools, and multi-step workflow orchestration, making them increasingly central to enterprise analytics. This integration introduces new security vulnerabilities across data resources, database execution, and agent reasoning, recombining concerns from database security and general-purpose LLM-agent security into failure modes that neither line of work captures on its own. To address this gap, we present a systematic security study of data agents. Our contributions are threefold. First, we develop a layered vulnerability framework that identifies eight data agent-specific risks across interpretation, execution, and policy layers. Second, we introduce an attack taxonomy organized by adversary goal, tactic, and technique, covering three goals, seven tactics, and fourteen techniques, and...

论文介绍 研究针对LLM驱动的数据分析代理的安全漏洞,这些代理集成推理、数据访问和工具执行。文章提出分层漏洞框架,识别解释、执行和政策层的八种风险,并引入按目标、战术和技术组织的攻击分类,涵盖十四种技术。意义在于系统揭示数据代理的特定安全威胁,弥补现有数据库安全和LLM安全研究的不足。

Sample-Efficient LLM-Based Detection of Malicious Web Server Logs with Forensically Explainable Reasoning

第一作者: Bernhard Kneip · 方向: AI 安全

Abstract:Forensic analysis of web server logs demands both accurate detection and human-readable explanations that can satisfy legal requirements. We present CEF-Log, a context-enhanced few-shot chain-of-thought prompting strategy for Large Language Models that addresses this dual requirement. CEF-Log embeds expert investigative methodology through a structured five-step reasoning template, enabling the model to learn \textit{how} to analyze logs rather than \textit{what} patterns to memorize. Experimental evaluation demonstrates that CEF-Log achieves an F1-score of 0.99 on the CSIC 2010 dataset using only four examples while providing a $10\times$ improvement in sample efficiency compared to other prompting-based methods. We also introduce ForenWebLog, a new dataset that incorporates real-world attacks and multi-step attack sequences for comprehensive evaluation. Qualitative analysis...

论文介绍 提出CEF-Log策略,利用大语言模型通过上下文增强的少样本链式思考提示检测恶意Web服务器日志。方法嵌入专家调查模板,实现高F1分数和样本效率提升,并提供可解释推理以满足取证需求。同时引入新数据集ForenWebLog用于综合评估,适用于安全监控和数字取证场景。

Exploring CKKS Parameter Trade-offs for Privacy-Preserving Personalized Federated Learning

第一作者: Kamolchanok Saengtong · 方向: 密码学协议

Abstract:Privacy-preserving Personalized Federated Learning (PFL) enables clients to collaboratively train personalized models without exposing raw data, but exchanged model updates remain vulnerable to inference attacks from honest-but-curious servers. Homomorphic Encryption (HE) addresses this by allowing server-side aggregation directly on encrypted updates, with the CKKS scheme being particularly suitable due to its native support for approximate floating-point arithmetic. However, no prior work has examined how to configure CKKS for PFL deployments, leaving practitioners without principled guidance on parameter selection that directly affects privacy, precision, and computational cost. This paper presents pFedCKKS, a generic framework integrating CKKS into PFL, and provides the first systematic parameter selection guide for practitioners. We derive the full CKKS parameter...

论文介绍 探索CKKS同态加密在隐私保护个性化联邦学习中的应用,提出pFedCKKS通用框架。文章首次提供系统性参数选择指南,分析CKKS参数对隐私、精度和计算成本的影响权衡,帮助实践者优化部署。意义在于为个性化联邦学习中安全聚合提供实操指导,提升数据隐私保护水平。

Digital White Spaces: A Cyberpsychology-Informed Framework to Mobile Phone Addiction

第一作者: Leandros Maglaras · 方向: 网络安全

Abstract:Mobile phone overuse and attention fragmentation have become pressing societal and public health concerns. Cyberpsychology research highlights addictive engagement loops driven by intermittent rewards, persuasive design, and habit formation. In this article, we use current evidence on mobile-phone addiction and propose "Digital White Spaces" (DWS), a socio-technical framework that combines privacy-preserving monitoring, AI-driven detection of addictive loops, device-mode interventions, and physical signal-limited zones to minimize digital stimulation and internet addiction.

论文介绍 基于网络心理学证据,提出「数字白空间」框架应对手机过度使用和注意力碎片化问题。框架整合隐私保护监控、AI驱动成瘾循环检测、设备模式干预和物理信号限制区,以减少数字刺激。可能应用于公共健康领域,促进数字福祉和减轻网络成瘾的社会影响。

AI Code Sandboxes: A Comparative Security Study. Part 1 of 2 -- Engine-Level Properties (Attack Surface, Leakage, Stackability, CVE History, Patch Cadence, Fuzzing)

第一作者: George Andronchik · 方向: 软件安全

Abstract:This paper reads six engine-level measurements together -- 1.1 host attack surface, 1.2 information leakage, 1.3 defense-in-depth stackability, 1.4 public CVE history, 1.5 patch cadence, and 1.6 upstream fuzzing posture -- to describe how five AI-sandbox products isolate guest code from the host kernel. No single axis is a sufficient basis for a comparative judgement; the cross-axis reading is the load-bearing analysis. Three high-level findings: (1) engine classes (microVM, userspace kernel, OCI container) separate cleanly on every architectural axis, but products within a class do not; (2) product pin policy is the dominant operator-facing variable -- engine-side patch latency aggregates to ~0 days for coordinated disclosures, while downstream lag spans 0 days to 471+ days to "opaque" to infinity; (3) fuzzing investment splits into three tiers, and the strongest combination...

论文介绍 对五款AI代码沙箱产品进行安全比较研究,分析攻击面、信息泄露、深度防御堆叠性、CVE历史、补丁周期和模糊测试等引擎级属性。发现引擎类别在架构轴上清晰分离,但产品内差异小;补丁策略是关键操作变量,模糊测试投入分三层。结果为选择安全沙箱提供跨轴评估依据。

Hiding in Plain Floats: Steganographic Carriers for Indirect Prompt and Content Injection

第一作者: Mudit Sinha · 方向: AI 安全

Abstract:Text-centered prompt-injection defenses assume that the malicious signal is visible in one of the inspected text views. We study a reproducible LLM01-style indirect prompt/content-injection failure mode where that assumption breaks: a payload caught in plain English slips past the same detector when it is transported as structured float parameters and reconstructed only as fragmented telemetry. Across 14,400 attacked real-model trials on three commercial LLM APIs from different providers, the IFS-derived float-array carrier preserves 94.3% leakage ASR under the strongest dual-layer text-classifier defense evaluated in the main matrix: a Prompt Guard 2 + TF-IDF ensemble; the same carrier-level pattern also replicates with a fine-tuned roberta-base detector. We emphasize leakage ASR because downstream systems may act on quoted or reproduced markers even when the model refuses...

论文介绍 研究通过浮点参数进行隐写术运输的间接提示和内容注入攻击,绕过文本分类器防御。实验在商业LLM API上测试,显示载波在强防御下仍保持高泄漏成功率。揭示当前文本中心防御的局限性,强调需要新方法检测结构化数据中的恶意信号,对LLM安全防护具有启示意义。

SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)

第一作者: Steven Golob · 方向: AI 安全

Abstract:Synthetic data is increasingly promoted as a privacy-preserving substitute for releasing sensitive tabular records, yet its central adversarial threat ("reconstruction", the recovery of an individual's hidden attribute values from a synthetic release and a handful of known quasi-identifiers) has been studied only in scattered, hard-to-compare settings. We present the first systematization of reconstruction (equivalently, attribute inference) attacks on de-identified and synthetic tabular data. We contribute a taxonomy that organizes attacks by the structure they exploit; the most systematic empirical evaluation to date, pitting fourteen attacks against nine synthetic data generation (SDG) methods across five benchmark datasets; and a set of new attacks that fill gaps in the taxonomy, one of which (CoBP-RA) is the strongest attack we measure. Crucially, we introduce a...

论文介绍 系统化研究合成表格数据的重建攻击,即从合成发布和已知准标识符恢复个体属性。文章贡献攻击分类法、迄今最全面的实证评估(涉及十四种攻击和九种生成方法),并提出新攻击方法填补空白,其中CoBP-RA最强。基于NIST竞赛获胜经验,为合成数据隐私风险提供系统分析框架。

An AI Security Agent for University ACMIS: Multi-Vector Threat Detection and Automated Response

第一作者: Joseph Walusimbi · 方向: 系统安全

Abstract:University Academic Management Information Systems (ACMIS) are high-value targets for a wide spectrum of security threats including brute-force login attacks, payment fraud, privilege escalation, insider data theft, and academic integrity violations. Traditional rule-based intrusion detection systems are inadequate because many malicious activities are structurally indistinguishable from normal operations. This paper presents an AI-based security agent for ACMIS that combines supervised anomaly detection, behavioural analytics, and a natural language processing chatbot for secure password recovery. The agent monitors five operational layers: authentication, authorisation, financial transactions, user behaviour, and system health, and responds through a four-tier risk escalation framework. A modular architecture allows the core engine to be extended to other institutional...

论文介绍 本文针对大学学术管理信息系统面临的多维度安全威胁,提出一种AI驱动的安全代理方案。该系统结合了监督式异常检测、行为分析和自然语言处理聊天机器人,能够监控认证、授权、财务交易等五个层面。通过模块化架构和四级风险响应框架,该代理旨在解决传统规则系统难以区分正常与恶意活动的问题,以增强系统的整体安全韧性。

Quantifying and Defending against the Privacy Risk in Logit-based Federated Learning

第一作者: Sheng Wan · 方向: 隐私保护

Abstract:Federated learning aims to protect data privacy by collaboratively learning a model without sharing private data among clients. Unlike traditional parameter-based FL methods that exchange model weights or gradients during training, emerging logit-based FL approaches share model outputs (logits) on public data. This strategy promotes model heterogeneity, reduces communication overhead, and enhances clients' privacy. However, the potential privacy risks associated with these logit-based methods have been largely overlooked. This research presents the first theoretical and empirical analysis of a hidden privacy risk in logit-based FL methods - the risk that a semi-honest server (adversary) may learn clients' private models from logits. To quantify and address this threat, we develop the Adaptive Model Stealing Attack (AdaMSA) by leveraging historical logits during training...

论文介绍 联邦学习中,基于logit(模型输出)的共享方式被提出以降低通信开销并提升隐私性。然而,本文首次从理论和实验上揭示了其潜在隐私风险:半诚实服务器可能通过分析公开数据的logit来窃取客户端的私有模型。为量化并应对此威胁,研究者提出了自适应模型窃取攻击(AdaMSA),并设计了相应的防御策略,增强了logit-based联邦学习的安全性。

LPOR: A Layered Proof of Reserves Framework for Usable and Publicly Auditable Solvency Verification

第一作者: Donggoo Kim · 方向: 密码学协议

Proof of Reserves (PoR) enables centralized crypto exchanges to demonstrate that on-chain reserves are sufficient to cover customer liabilities. However, existing approaches, including Merkle-tree-based proofs and zero-knowledge PoR systems, remain difficult for everyday users to verify in practice, resulting in limited participation and weakened transparency. We introduce LPOR, a layered, usability-focused PoR framework that separates lightweight user-side checks from auditor-level cryptographic verification, enabling non-technical users to verify inclusion and publicly recompute total liabilities with minimal friction. By lowering verification barriers, LPOR increases user participation and substantially improves the probability of detecting omitted liabilities. We evaluate its scalability and omission detectability at a multi-million-user scale.

论文介绍 为提升加密货币交易所储备证明的透明度和用户参与度,本文提出了一个分层、注重可用性的LPOR框架。该框架将轻量级的用户侧检查与审计级的密码学验证分离,使得非技术用户也能以较低门槛验证其账户是否被正确包含在负债中。这种设计旨在降低验证壁垒,提高用户参与度,并显著增强检测遗漏负债的能力。

AI-Native Closed-Loop Security for 6G-Enabled Cyber-Physical Systems: From Edge Detection to Network-Wide Mitigation

第一作者: Bilal Hussain · 方向: 网络安全

Abstract:In sixth-generation (6G) networks, billions of cyber-physical systems (CPSs) - autonomous vehicles, smart grids, industrial robots, and remote-surgical equipment - will run over ultra-reliable low-latency slices, collapsing the gap between a remote breach and physical harm to milliseconds, a budget perimeter firewalls and centralised security operations centres cannot meet. This survey reframes 6G CPS security as a closed-loop, AI-native pipeline that senses at the multi-access edge computing (MEC) tier, using minute-scale call-detail records (CDRs) for baseline learning and sub-millisecond RAN/Open-RAN (O-RAN) telemetry for the latency-critical path. It decides locally with compressed deep models, mitigates network-wide via SDN, NFV, and O-RAN controllers, and retrains through federated learning (FL) and digital-twin (DT) replay. We formalise a per-slice, tail-bounded latency...

论文介绍 本文是一篇关于6G网络中网络物理系统安全的综述。文章将6G CPS安全重新定义为一个AI原生的闭环管道,该管道在多接入边缘计算层感知威胁,利用深度模型进行本地决策,并通过SDN/NFV控制器进行全网缓解。它还结合了联邦学习和数字孪生进行模型重训练,旨在应对6G超低延迟场景下从网络攻击到物理伤害仅毫秒级间隔的挑战。

Closing the Sim-to-Real Gap: An Evaluation Framework for Autonomous Cyber Defense Configuration of Commercial EDR

第一作者: Kerri Prinos · 方向: AI 安全

Abstract:Leading commercial endpoint detection and response (EDR) products have shifted from operator-configured rule sets to multi-component systems where autonomous AI components operate alongside, and increasingly in place of, operator-deployed policies. Autonomous defense agents using commercial EDR as their hardening tool are no longer tuning a passive tool, but a black-box autonomous system capable of making vendor-specific decisions. We present the first evaluation framework for autonomous defense agents hardening commercial EDR. We instantiate it in a Game of Active Directory (GOAD) lab with this http URL's NodeZero as the autonomous pentester and Microsoft Defender XDR as the EDR. We run a sample benchmark of defense agents with two large language model (LLM) backbones (Claude Sonnet 4.6 and Cisco Foundation-Sec-8B). We report three lessons learned that neither simulation nor...

论文介绍 商业端点检测与响应产品正变得自主化,但评估强化此类产品的自主防御代理却很困难。本文提出了首个针对商业EDR的自主防御代理评估框架。该框架在一个模拟活动目录的实验室环境中,使用自主渗透测试工具和EDR产品进行实例化,并测试了两种不同大语言模型作为后端的防御代理,旨在弥合仿真环境与真实世界评估之间的差距。

Policy Description Language for Authorization using Logic-Based Programming

第一作者: Masaki Hashimoto · 方向: 安全研究

Abstract:Recently, with the impossibility of eradicating the vulnerabilities of information systems, we must prepare for the occurrence of the security incident by the multi-layer defense called the Defense-in-Depth strategy. In the multi-layer defense, it is important to authorize accesses in fine-grained granularity to compose each layer effectively, and many access control models are proposed to follow them. However, policy description languages proposed so far cannot express the models appropriately in proper granularity. In this paper, we propose a policy description language which can designate many kinds of conditions for access control, such as the dynamic status of an application process, as an element of decision data, and implement it in Datalog. Using the proposed language, we compose the policy of SELinux, which is a major implementation achieving the multi-layer defense...

论文介绍 为了在深度防御策略中实现更精细的授权控制,本文提出了一种新的策略描述语言。该语言能够以适当的粒度表达多种访问控制条件,例如应用程序进程的动态状态,并基于Datalog进行了实现。与以往语言相比,它旨在更恰当地描述复杂的访问控制模型,并通过构成SELinux的策略来展示其能力,以支持有效组合多层防御。

The Dodona Protocol: A Living Design Science Experiment in Oracle Design

第一作者: Giulio Caldarelli · 方向: 密码学协议

Abstract:The oracle problem, broadly understood as the difficulty of reliably incorporating external information into blockchain-based systems, has been widely examined by scholars and practitioners. Recent comparative research has shown that several challenges of modern blockchain oracles, including attributability, accountability, integrity, and query design, mirror procedural and epistemic constraints already present in ancient oracular institutions such as the Delphic Oracle. Yet the translation of these insights into applied oracle design remains largely unexplored. This paper introduces the Dodona Protocol, a modular, chain-agnostic oracle service inspired by procedural patterns identified in ancient and modern oracle systems. Named after the Oracle of Zeus at Dodona, one of the oldest oracular sanctuaries in ancient Greece, the protocol operationalizes principles such as...

论文介绍 本文介绍了Dodona协议,这是一个受古代与现代预言机系统程序模式启发的、模块化且链无关的预言机服务。该研究将设计科学实验应用于预言机设计,旨在将关于归因、问责、完整性等挑战的学术洞察转化为实际方案。协议以古希腊的多多纳神谕命名,尝试将历史智慧中的程序性原则应用于区块链外部数据集成问题。

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks

第一作者: Abid Aziz · 方向: 安全研究

Abstract:Reasoning-capable large language models can be induced to spend their generation budget on injected decoy tasks rather than answering the user's question, causing denial of service when no final answer is produced and denial of wallet when excess output tokens are billed. Input-side safety classifiers often miss these attacks because the injected prompts can appear syntactically benign. We build RecurGuard, a runtime monitor for detecting reasoning-chain consumption attacks when reasoning traces are exposed by the model. RecurGuard analyzes reasoning traces as they are generated and tracks three signals: recurrence rate, volume growth, and progress toward the user's query. If all three signals remain anomalous over three consecutive chunks, RecurGuard terminates generation early. We evaluate RecurGuard against OverThink and ExtendAttack across open-weight reasoning models and...

论文介绍 本文研究针对具备推理能力的大语言模型的“推理令牌消耗攻击”,即攻击者注入诱饵任务,诱导模型耗尽生成预算而无法回答用户问题,造成服务拒绝或经济损失。为此,研究者构建了RecurGuard,一个在推理链暴露时进行检测的运行时监控器。它通过分析推理链的递归率、体积增长等信号来识别攻击,并在信号持续异常时提前终止生成。

Demand-Driven Vulnerability Detection for Cloud Security Posture Management: Removing Human Rule Authoring from the Disclosure-to-Protection Critical Path

第一作者: Prashant Kumar Pathak · 方向: 软件安全

Abstract:Cloud Security Posture Management (CSPM) systems detect known vulnerabilities by maintaining a rule set, distributing it to customers, and evaluating it against periodically-collected asset inventories. To our knowledge, in publicly documented architectures the rule set is environment-agnostic and curated centrally by the vendor; updates are batched into release cycles and shipped on a cadence ranging from hours to days depending on detection severity. The disclosure-to-protection window -- from a CVE being published to the customer's system being capable of detecting affected assets -- is therefore bounded by the vendor's release cadence for version-match detections, and by additional human authoring time for richer detections incorporating configuration predicates beyond the affected-software string. We propose an architecture in which the rule set is not vendor-distributed...

论文介绍 研究云安全态势管理(CSPM)系统中漏洞检测的延迟问题。现有系统依赖厂商集中更新规则,导致从漏洞披露到保护存在窗口期。本文提出需求驱动架构,通过自动化规则生成移除人工编写,实现按需检测。该方法能缩短响应时间,提升云环境安全性。

POISE: Position-Aware Undetectable Skill Injection on LLM Agents

第一作者: Haochang Hao · 方向: AI 安全

Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A practically dangerous injection must stay invisible: if executing the payload derails the user's legitimate task, the resulting failure signal invites inspection of the skill. We therefore evaluate attacks by Attack Success Rate, which requires the injected payload to execute and the user's task to still pass its verifier in the same trial. Prior skill-poisoning attacks face a reliability-stealth trade-off under this lens: YAML-header injections are reliably loaded but easily inspected, whereas stealthier body injections that place explicit malicious commands in the skill prose are less reliable because out-of-context commands invite the agent's own suspicion. We introduce POISE, a position-aware attack that compresses the trigger into a...

论文介绍 针对大型语言模型(LLM)代理的技能注入攻击,攻击者可注入恶意代码影响代理行为。现有攻击方法在隐蔽性与可靠性之间存在权衡。本文提出POISE,一种位置感知攻击,通过压缩触发词提高隐蔽性,同时保持攻击成功率。该工作有助于测试和增强LLM代理的安全机制。

Collective Hallucination in Multi-Agent LLMs:Modeling and Defense

第一作者: Saeid Jamshidi · 方向: AI 安全

Abstract:Hallucinations in large language models (LLMs) create heightened risks in multi-agent settings, where recursive agent interactions can propagate, reinforce, and amplify unsupported claims. This paper models hallucination as a system-level, time-evolving process across a network of interacting LLM agents, where nodes represent agents and edges encode information exchange. The proposed formulation captures how hallucinated claims diffuse through communication topologies, intensify under adversarial perturbations, and affect collective reliability across reasoning rounds. To suppress error propagation, we introduce an interaction-aware control method that combines confidence-weighted aggregation, adaptive impact regulation, external claim verification, and selective isolation of unreliable agents. Experiments on TruthfulQA and TriviaQA show that the proposed method reduces...

论文介绍 在多代理LLM系统中,幻觉错误可能通过代理间交互传播和放大。本文将幻觉建模为时变网络过程,分析其扩散机制。为抑制错误传播,提出交互感知控制方法,结合置信度加权、影响调节等策略。实验显示该方法能有效降低集体幻觉,提高系统可靠性。

SGTO-MAS: Secure Gorilla Troops Optimization for Multi-Agent LLM Systems

第一作者: Saeid Jamshidi · 方向: AI 安全

Abstract:Multi-agent large language model (LLM) systems offer strong capabilities for complex reasoning and decision-making, yet coordination across agents introduces error propagation, security risks, and inefficient use of resources. Existing methods often rely on heuristic, static strategies and lack a principled mechanism for balancing performance, security, and computational cost. This paper formulates multi-agent LLM coordination as a constrained optimization problem and proposes a security-aware method for adaptive agent selection. The method integrates trust modeling, risk-aware evaluation, and collective intelligence within a unified optimization objective. To solve the problem efficiently, we use a swarm-intelligence strategy inspired by Gorilla Troops Optimization (GTO), enabling adaptive coordination under varying threat conditions. Controlled experiments across 500...

论文介绍 多代理LLM系统在协调中面临安全风险和效率挑战。现有方法缺乏原则性优化机制。本文将协调问题形式化为约束优化,提出SGTO-MAS方法,基于群体智能进行自适应代理选择,平衡性能、安全与计算成本。该方法通过信任建模和风险评估提升系统鲁棒性。

Hallucination Cascade: Analyzing Error Propagation in Multi-Agent LLM Systems

第一作者: Saeid Jamshidi · 方向: AI 安全

Abstract:Large Language Models (LLMs) generate fluent text but remain vulnerable to hallucinations, producing unsupported, inconsistent, and factually incorrect claims. Most prior work treats hallucination as a static property of isolated outputs. In multi-agent LLM systems, however, responses are exchanged across agents, revised through sequential stages, and reused as context for later reasoning. Hallucination, therefore, becomes a dynamic process shaped by interaction history, cascade depth, and model heterogeneity. This paper analyzes hallucination dynamics in multi-agent LLM cascades by tracking claim-level factual inconsistencies across sequential agent interactions. We conduct 500 cascade experiments across 10 knowledge domains using GPT-5.3, DeepSeek-V3, and LLaMA-3-70B-Instruct, yielding 1,250 evaluated responses. Results show that deeper cascades reduce the normalized...

论文介绍 多代理LLM系统中的幻觉具有动态级联特性,现有研究多关注静态属性。本文通过级联实验跟踪声称级事实不一致,分析幻觉如何随交互深度和模型异质性传播。结果表明更深层的级联可能减少幻觉强度,但传播模式复杂。该工作为理解错误传播提供新视角。

DP4SQL: Differentially Private SQL with Flexible Privacy Policies

第一作者: Andrew Cascio · 方向: 隐私保护

The plausible deniability model of differential privacy for single-table datasets is well-understood. However, applying differential privacy to relational databases is much trickier: each application needs flexibility in specifying the pieces of information about an entity, spread across multiple relations, that require plausible deniability guarantees. Existing differentially private SQL systems only support rigid privacy policies. Even seemingly small changes, such as specifying that some tables need to protect the existence of records while others only need to protect the record contents, require significant manual effort in updating their privacy accountants and proving their correctness. One example of a challenge is the presence of partially public data. Public columns in a table (e.g., faculty names in a university dataset and partial course enrollment information) can cause...

论文介绍 将差分隐私应用于关系数据库时,现有SQL系统支持隐私策略僵化,难以灵活调整。本文提出DP4SQL,支持灵活隐私策略,如区分记录存在性与内容保护。该系统能处理部分公开数据等挑战,通过动态隐私会计简化配置,增强数据库查询的隐私保护。

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices

第一作者: Stefan Behfar · 方向: AI 安全

Abstract:The rise of edge-based machine learning has enabled distributed adaptation of language models across mobile and IoT devices, offering privacy preservation and real-time responsiveness. However, distributed fine-tuning of language models on untrusted or heterogeneous edge nodes introduces new vulnerabilities. Compromised or unreliable devices can inject poisoned updates, leading to stealthy model manipulation or convergence degradation. Classical defenses such as robust aggregation or temporal anomaly detection operate on a single global model and are therefore limited in detecting coordinated or persistent poisoning. This work proposes a new system-level defense based on model multiplicity. Instead of maintaining one global model, the system rotates or concurrently trains multiple small language models (e.g., DistilGPT-2), each updated by independently sampled subsets of edge...

论文介绍 在边缘设备上分布式训练语言模型时,不可靠节点可能注入中毒更新。传统防御基于单一全局模型,检测能力有限。本文提出基于模型多样性的防御,同时训练多个小语言模型,通过独立采样子集更新,增强对抗性检测。该方法利用模型旋转或并行训练提升安全性。

Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks

第一作者: Zvi Topol · 方向: 软件安全

Abstract:Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack success rate (ASR), not taking into account the sequential structure of how models resist or yield to attacks. We propose applying process mining, a discipline for discovering and analyzing process models from event logs, to red teaming traces. We conduct a controlled experiment pitting 60 HarmBench prompts against two LLMs, GPT-OSS 120B and Llama 3.3 70B, using 10 prompt mutation strategies over up to 110 attempts per prompt. From the resulting 8,575 scored events we extract Directly-Follows Graphs (DFGs) and state transition matrices that reveal structurally distinct defense profiles invisible to ASR alone: GPT-OSS exhibits a near-absorbing refusal state, while Llama presents multiple porous escape routes from refusal to getting successfully jailbroken. We further show that...

论文介绍 标准AI红队评估仅以攻击成功率(ASR)为指标,忽略攻击序列结构。本文应用过程挖掘分析LLM对抗轨迹,提取直接跟随图和状态转换矩阵,揭示不同模型的防御模式。实验显示GPT-OSS具有吸收性拒绝状态,而Llama存在多孔逃逸路径,为深入评估提供工具。

Ternary public-key cryptosystem

第一作者: Steven Duplij · 方向: 密码学协议

Abstract:Public-key cryptosystems eliminate the requirement for pre-shared secret keys by enabling encryption with a publicly disclosed key and decryption with a corresponding private key. In this article we generalize the public-key cryptosystems to ternary algebraic structures, with particular attention to ElGamal as a representative family. We introduce the necessary algebraic background for nonderived ternary structures, including special elements, ternary group rings, and a matrix ternarization procedure that maps binary rings and group rings to antidiagonal symbolic matrices closed under ternary multiplication. Building on these foundations, we formulate a ternary analogue of the ElGamal three-step protocol (key generation, ephemeral encryption, and decryption via querelements) and derive explicit ternary power and querelement formulas that enable correct decryption. Concrete...

论文介绍 本文将公钥密码系统推广到三元代数结构,特别关注ElGamal家族。研究问题在于扩展经典密码系统至三元框架,核心方法是引入非导出三元结构、特殊元素和矩阵三元化过程,并提出三元ElGamal协议及其解密公式。可能应用于新型加密方案设计。

Quantum-Inspired Reinforcement Learning for Low-Latency Intrusion Detection in V2X and Internet-of-Vehicles Networks

第一作者: Sajid Anwer · 方向: AI 安全

Abstract:Smart cities increasingly depend on dense edge, IoT, and vehicular networks to deliver critical urban services, including traffic control, connected mobility, infrastructure monitoring, and energy management. In this ecosystem, the Internet of Vehicles (IoV) is central to intelligent transportation, enabling continuous communication among vehicles, roadside infrastructure, and cloud-edge platforms. This connectivity, however, also enlarges the attack surface and exposes smart city and vehicular systems to evolving cyber threats that can compromise safety, privacy, data integrity, and service continuity. Conventional static defenses are often inadequate because they cannot autonomously adapt to changing attack behaviors or multi-stage intrusion patterns. This paper proposes QIRL, a Quantum-Inspired Reinforcement Learning framework built on a lightweight Deep Q-Network...

论文介绍 本文解决车联网(IoV)中低延迟入侵检测的挑战,提出QIRL框架。该框架使用量子启发强化学习和轻量级深度Q网络,以自主适应动态攻击模式。意义在于提升智能城市交通系统的安全性和实时性。

Belief-Space Quantum-Inspired Reinforcement Learning for Partially Observable Autonomous Cyber Defense in the Internet of Vehicles

第一作者: Anwar Shah · 方向: AI 安全

Abstract:The Internet of Vehicles (IoV) faces a dynamic, adversarial security environment where attackers adapt to defenses. Existing intrusion detection systems rely on static classifiers that fail to capture sequential decision-making, attacker adaptation, and uncertainty. We formulate IoV security as a sequential attacker-defender interaction and model defense as a reinforcement learning problem under partial observability. We propose Quantum Belief-Integrated Reinforcement Defense (Q-BIRD), using quantum-inspired belief representation to encode defender uncertainty about hidden attacker intent via amplitude-based states, enabling non-Bayesian belief evolution. Integrated into a Proximal Policy Optimization (PPO) defender, Q-BIRD selects cost-aware mitigation actions. In simulated environments with adaptive, probing attackers, Q-BIRD reduced cumulative mean damage, damage variance...

论文介绍 本文将车联网安全建模为部分可观察的序列决策问题,提出Q-BIRD框架。核心方法是将量子信念表示集成到强化学习中,编码防御者对攻击者意图的不确定性,使用PPO选择缓解动作。可能应用于自主网络防御系统。

MOLOT System Card: Malicious Operational Logic Observation Transformer

第一作者: Daniil Lopatkin · 方向: 软件安全

MOLOT (Malicious Operational Logic Observation Transformer) is a static malicious-code detection system designed for SAST setup where package metadata, maintainer history, and dynamic execution traces may be unavailable or unreliable. The system represents source code as behavior sequences derived from static call graphs, includes an explanation stage that ranks suspicious behavior activities and maps them back to source-code locations. The approach is evaluated on Python and JavaScript packages from PyPI and npm, compared with opensource detection tools, and validated under product constraints including runtime, memory use, and false-positive rates observed in a real moderation workflow. We also release Open Malicious-Code Bench, a public benchmark for reproducible evaluation of malicious-package detection methods. The results show that static behavior-sequence modeling can provide...

论文介绍 MOLOT系统是一种静态恶意代码检测方法,适用于缺少包元数据或动态信息的场景。研究问题是如何在静态分析中检测恶意代码,核心方法是将源代码表示为行为序列并映射回源码位置。意义在于增强软件供应链安全。

ScaleDisturb: Exploiting Temporal Asymmetry to Amplify Read Disturbance in Modern DRAM Chips

第一作者: Jikun Wang · 方向: 软件安全

Abstract:DRAM suffers from read disturbance phenomena (e.g., RowHammer and RowPress), where repeatedly accessing or continuously keeping open a DRAM row (aggressor row) induces bitflips in other physically nearby unaccessed rows (victim rows). The disturbance mechanism is practically exploitable from the software stack and worsens across generations with continued density scaling. DRAM read disturbance is highly sensitive to memory access patterns, yet prior work explores read disturbance under only a limited set of access patterns. We present ScaleDisturb, a new DRAM access pattern that can amplify DRAM read disturbance by asymmetrically extending the open time of two aggressor rows. Our rigorous experimental characterization of 196 DDR4 and 3 HBM2 DRAM chips shows that ScaleDisturb (1) leads to bitflips at significantly fewer row activations, compared to state-of-the-art memory...

论文介绍 ScaleDisturb提出一种新的DRAM访问模式,利用时间不对称性放大读干扰。研究问题在于现有读干扰模式覆盖不足,核心方法是不对称扩展开放时间以诱导更多位翻转。意义在于揭示新型内存安全风险。

SHIELD-IDS: Structurally Heterogeneous Ensemble with Integrated Layered Defense for Intrusion Detection Systems

第一作者: Maryam Zaman · 方向: AI 安全

Adversarial attacks pose a serious and growing threat to Machine Learning (ML)-based Intrusion Detection Systems (IDS), where imperceptible perturbations to network flow features can systematically mislead classifiers into accepting malicious traffic as benign. The IDS-Anta framework partially addresses this through Z-score normalization, Singular Value Decomposition (SVD), and Multi-Armed Bandit (MAB) classifier selection with Thompson Sampling, yet its classifier pool lacks sufficient structural diversity for robust adversarial resistance. This work introduces IDS-Anta++, which incorporates XGBoost and LightGBM gradient boosting models into the ensemble and wraps the extended pool in a three-layer black-box defense: Isolation Forest anomaly screening, median feature smoothing, and six-way majority voting. Experiments conducted on CIC-IDS-2017, CEC-CIC-IDS-2018, and CIC-DDoS-2019...

论文介绍 针对机器学习入侵检测系统易受对抗攻击的问题,本文提出SHIELD-IDS。核心方法是扩展分类器池以增加结构多样性,并集成三层黑盒防御:异常筛选、特征平滑和多数投票。可能提升IDS的对抗鲁棒性。

MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models

第一作者: Rishabh Makwana · 方向: 软件安全

Abstract:Vision-Language Models (VLMs) have demonstrated strong performance across multimodal tasks, yet their safety robustness remains an open challenge. While prior work has shown that structured visual prompts such as flowcharts can effectively jailbreak VLMs, existing studies are largely limited to English-centric settings. In this paper, we introduce MLingualFC, a multilingual multimodal benchmark designed to evaluate jailbreak vulnerabilities of VLMs across diverse languages using structured flowchart representations. MLingualFC encodes harmful instructions into flowchart images across five languages (Hindi, Punjabi, Spanish, Romanian, and German). We evaluate state-of-the-art multilingual VLMs, including Qwen2.5-VL, Gemma-4, and Pangea, under a black-box threat model. Our results reveal significant multilingual safety gaps. Flowchart-based attacks achieve high attack success...

论文介绍 本文引入MLingualFC基准,评估多语言视觉语言模型的越狱漏洞。研究问题是现有评估局限于英语,核心方法是使用流程图图像编码有害指令在五种语言中测试。结果揭示多语言安全差距,意义在于推动AI安全研究。

Detecting Aimbot Cheaters in MOGs

第一作者: Salman Shaikh · 方向: 软件安全

Abstract:Multiplayer Online Games have become a multibillion dollar industry in the entertainment sector. However, the presence of cheaters undermines the experience of honest players and devalues the effort of game developers, as it directly affects player retention, competitive integrity, the legitimacy and trustworthiness of a game, and most importantly the overall revenue streams. Among various cheating techniques, visual aimbots represent an emerging threat. They use computer vision models to detect opponents from client screen captures rather than accessing game memory, making them completely undetectable by commercial kernel level anti cheat solutions. In this paper, we introduce PATCH, a novel proactive defense strategy that deploys adversarial patches as in game honeytokens to mitigate the presence of visual aimbot cheaters. Our approach centers on deliberately triggering the...

论文介绍 本文针对多人在线游戏中视觉瞄准机器人作弊问题,提出PATCH防御策略。核心方法是部署对抗性补丁作为蜜令牌,触发作弊者检测。可能应用于游戏反作弊系统,保护游戏公平性。

Human-Centred Risk Mitigation for AI-Mediated Information Manipulation: A SOCMINT Framework Based on Information Manipulation Sets

第一作者: Antonio Scala · 方向: 安全研究

Abstract:AI-mediated information manipulation increasingly takes the form of social cyber attacks that target trust, attention, credibility, reputation, and decision-making rather than only technical infrastructures or isolated false contents. Existing defensive approaches often oscillate between incident-level analysis, which fragments campaigns into weak signals, and attribution-first analysis, which may delay mitigation until responsibility is established. This paper proposes a SOCMINT framework based on Information Manipulation Sets (IMS) as an intermediate operational unit between individual incidents and strategic attribution. Building on the VIGINUM/EEAS use of IMS in counter-FIMI analysis, the framework treats manipulation as a coherent process involving narratives, accounts, infrastructures, temporal patterns, cross-platform migration, synthetic amplification, and cognitive...

论文介绍 该论文针对 AI 介导的信息操纵问题,指出现有防御方法在事件分析和归因间存在局限,可能导致缓解延迟。作者提出基于信息操纵集的 SOCINT 框架,将操纵视为涉及叙述、账户、基础设施、时间模式等的一致过程,作为事件与战略归因之间的中间操作单元。此框架可用于反虚假信息、影响和操纵分析,以提高风险缓解的及时性和有效性。

A Bell-State Extension of Loop-Back Quantum Key Distribution

第一作者: Luis Adrián Lizama-Pérez · 方向: 密码学协议

Abstract:Bidirectional quantum key distribution (QKD) protocols face persistent challenges related to classical disclosure, confinement of the signal space to predictable subspaces, and limited detectability under substitution or entanglement-swapping attacks. In this work, we present a Bell-state extension of the Loop-Back QKD architecture that improves efficiency and detectability while preserving its defining feature of a simplified, measurement-free remote terminal. The protocol employs entangled Bell states together with deterministic local Pauli encoding at the remote node. A central element is that Alice privately prepares and knows the initial Bell state, which serves as a hidden reference enabling her to interpret the Bell-state transition induced by Bob, while preventing an adversary from reconstructing the encoding without access to this reference. By exploiting both intra...

论文介绍 该论文研究双向量子密钥分发(QKD)协议的安全性挑战,如经典泄露和可检测性限制。作者提出一个 Bell 状态扩展的 Loop-Back QKD 架构,使用纠缠 Bell 状态和本地 Pauli 编码,提高效率和可检测性,同时保持远程终端简化无测量的特性。核心是 Alice 私有准备初始 Bell 状态作为隐藏参考,防止对手重构编码,可能增强量子通信安全性。

Parent-Hash DAG: A Cost Analysis of Constant-Time Append for On-Chain Registries

第一作者: Ian C. Moore · 方向: 区块链安全

Abstract:Provenance trees are append-only directed acyclic graphs of artifact registrations anchored on a public blockchain, recently introduced as the data substrate of operator-gated provenance infrastructure. Their defining data-structural pattern is a parent-hash directed acyclic graph (PHDAG), in which each append performs a constant number of storage writes to previously-untouched slots. This pattern has not previously been isolated as a standalone primitive, formally bounded with explicit constants, or benchmarked against the standard alternative, the incremental Merkle tree (IMT). We formalize PHDAG append as O(1) in gas cost, independent of registry size and tree depth, and develop a stochastic cost model for IMT in which per-insert cost is a random variable over the leaf index, deriving closed-form expressions for its mean and variance. We validate both analyses empirically...

论文介绍 该论文研究区块链上追加-only 溯源树的数据结构,提出父哈希有向无环图(PHDAG)作为恒定时间追加的原语。作者形式化 PHDAG 的追加操作为 O(1) gas 成本,独立于注册大小和树深度,并与标准增量 Merkle 树(IMT)对比,开发随机成本模型进行基准测试。这为区块链注册提供更高效的数据存储方案,可能优化链上操作成本。

Clinically Grounded Privacy Evaluation of Medical LMs

第一作者: Sasha Ronaghi · 方向: AI 安全

Abstract:Medical language models (LMs) can memorize and reproduce protected health information, but privacy evaluations often focus on recovery of training text rather than disclosure under realistic threat models. We introduce a clinically grounded framework that evaluates leakage along a graded axis of adversarial access, ranging from publicly inferable demographics to leaked note fragments. At each tier, we measure verbatim memorization of patient-specific text and semantic leakage of sensitive diagnoses. Applying the framework to an LM pretrained on 378k clinical notes, we find that routine encounter metadata (i.e. name, date of birth, provider, practice, visit date) elicits high rates of verbatim memorization across a patient's timeline and sensitive-diagnosis recovery (AUROC 0.91 for abortion, 0.81 for HIV). At the same time, exact-match memorization can overstate disclosure: 36%...

论文介绍 该论文针对医疗语言模型(LMs)的隐私风险,提出一个临床基础评估框架,沿对抗性访问等级评估数据泄露,包括可公开推断的人口统计和泄露的笔记片段。框架测量逐字记忆和敏感诊断的语义泄露。应用于预训练于临床笔记的 LM,发现常规元数据能触发高记忆率和敏感诊断恢复,有助于更实际地评估隐私保护。

Safe-RULE: Safe Reinforcement UnLEarning

第一作者: Shixiong Jiang · 方向: AI 安全

Abstract:Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such as robotics systems. However, its reliance on static datasets exposes offline Safe RL to data poisoning attacks, where adversaries inject malicious samples that compromise safety and induce unsafe policy behavior. In this work, we propose a new learning paradigm, named safe reinforcement unlearning (Safe-RULE), used as a defense framework to remove the influence of poisoned data without retraining from scratch or requiring access to the original training environment. We further extend reinforcement unlearning to offline Safe RL by explicitly accounting for both task performance and safety constraints during the unlearning process. Experiments across benchmark Safe RL tasks demonstrate that our approach effectively enhances...

论文介绍 离线安全强化学习依赖静态数据集,易受数据中毒攻击。提出Safe-RULE框架,用于安全强化学习的防御,通过强化遗忘移除中毒数据影响,无需重新训练或原始环境访问,在遗忘过程中显式考虑任务性能和安全约束,增强系统安全性。

Targeting World Models to Compromise Robot Learning Pipelines

第一作者: Ethan Rathbun · 方向: 安全研究

Abstract:World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline. While highly practical, in this work we demonstrate that world models introduce a uniquely stealthy and effective data poisoning entry point into the robot learning supply chain that can result in the deployment of unsafe or otherwise compromised robotic policies despite training on seemingly safe ground truth training data. In contrast to traditional data poisoning techniques which directly implant dangerous trajectories into sold or uploaded datasets, our novel attack methods inject malicious prompts or compromising transition dynamics into visibly safe teleoperated datasets which are only activated once fed...

论文介绍 世界模型作为高效数据工具被集成到机器人学习中,但可能引入隐蔽数据投毒入口。本工作演示新颖攻击方法,将恶意提示或妥协动力学注入看似安全数据集,只在喂入世界模型时激活,揭示机器人学习管道的安全风险,促进安全措施发展。

Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models

第一作者: Bartłomiej Marek · 方向: AI 安全

Abstract:Recent work has applied differential privacy (DP) to adapt large language models (LLMs) for sensitive applications, offering theoretical guarantees. However, its practical effectiveness remains unclear, partly due to LLM pretraining, where overlaps and interdependencies with adaptation data can undermine privacy despite DP efforts. To analyze this issue in practice, we investigate privacy risks under DP adaptations in LLMs using state-of-the-art attacks such as robust membership inference and canary data extraction. We benchmark these risks by systematically varying the adaptation data distribution, from exact overlaps with pretraining data, through in-distribution (IID) cases, to entirely out-of-distribution (OOD) examples. Additionally, we evaluate how different adaptation methods and different privacy regimes impact the vulnerability. Our results show that distribution...

论文介绍 该论文评估大语言模型(LLMs)适应中的隐私保护实践效果,使用差分隐私(DP)但面临预训练数据重叠问题。作者通过鲁棒成员推断和金丝雀数据提取等攻击,在不同数据分布下基准测试隐私风险。研究显示,数据分布和适应方法影响漏洞,为隐私保护提供实证分析。

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

第一作者: Hyunseok Paeng · 方向: AI 安全

We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline. In safety-trained Claude models, documents containing prompt injections suffer a sharp drop in recommendation rate, and this suppression propagates beyond the injected document to unmodified documents of the same brand. In Claude Opus 4.6, the target brand drops from a 54% baseline to zero top-2 recommendations across all 50 trials, even though only 1 of 4 brand documents in the corpus contains an injection. The directional pattern is reproduced in counterfactual experiments and across three brands. A contrasting result across the GPT models tested, where the same injection instead increases recommendations, suggests...

论文介绍 该论文揭示 RAG 基础上 LLM 推荐系统中的一个安全训练失败模式——注入悖论。嵌入文档的提示注入在安全训练的 Claude 模型中反而抑制目标品牌推荐,甚至传播到未注入文档。实验显示品牌推荐率从基准降至零,而在 GPT 模型中则增加推荐,指出安全训练可能引入意外行为。

Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human

第一作者: Emre Turan · 方向: AI 安全

Abstract:As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person. We argue the gate is the easy part; the hard part is the judgment - which actions to stop - which the field evaluates against two false assumptions: that there is a ground-truth notion of "risky," and that the human reviewer is a perfect, infinitely-available oracle. On a hand-labeled set of 125 adversarially-weighted agent actions we show that (i) reviewers only moderately agree on what is risky (Fleiss' kappa = 0.52), so there is no single correct label; (ii) framing the guard as selective classification under asymmetric cost makes its operating limits measurable, and on hard inputs the guard cannot safely auto-decide; and (iii) when the reviewer is modeled as endogenous...

论文介绍 该研究探讨了基于大语言模型(LLM)的智能体在执行高风险操作时,如何校准其“人类审批”安全机制。研究指出,现有安全机制假设存在明确的“风险”定义且人类审查者是完美的,这与现实不符。作者通过实验表明人类对风险的判断存在中等程度的一致性,并提出将防护机制建模为非对称成本下的选择性分类,从而量化其能力边界。该工作为设计更有效、更现实的智能体安全守卫提供了理论基础。

Cheap Reward Hacking Detection

第一作者: Iván Belenky · 方向: AI 安全

Abstract:A small transformer encoder is trained to map Terminal-Wrench trajectories onto a unit sphere where embedding distance approximates the $L_1$ distance between reward and metadata signals. A linear probe on top of that embedding detects reward hacking on the cleaned test split with AUC $0.9467$ and TPR@5%FPR $0.8296$, matching the TW sanitized LLM-as-judge AUC ($0.9510$ on the cleaned split) and exceeding its TPR@5%FPR ($0.7130$ vs $0.8296$) on the same information condition, at roughly four orders of magnitude lower per-trajectory cost. The encoder is not a pure behavior reader: stripping natural-language reasoning from its input at probe time drops AUC to $0.6213$.

论文介绍 本文提出了一种用于检测“奖励黑客”行为的低成本方法。奖励黑客是指智能体利用奖励函数的漏洞获取高奖励但并未完成预期任务。作者训练了一个小型 Transformer 编码器,将智能体的轨迹映射到一个嵌入空间,并利用线性探测器进行检测。该方法在测试集上达到了接近使用大语言模型作为评判器的准确率,但每个轨迹的计算成本低几个数量级。这为大规模监控智能体行为提供了一种可行且高效的工具。

RAILS: Verification-Native Clearing For Agentic Commerce

第一作者: Adrian de Valois-Franklin · 方向: 软件安全

Abstract:Autonomous agents negotiate, purchase, deploy code, and move funds, but no neutral mechanism determines whether they met their delegated obligation, who is responsible when they did not, or which settlement action follows. This is the agentic clearing problem. Tool protocols (MCP), inter-agent communication (A2A), payment rails (x402), mandate and network agent protocols (AP2, Visa, Mastercard), and settlement-risk standards each assume that determination and none produce it. Clearing is the missing primitive. Payment is not clearing. Authorization is not clearing. LLM-as-judge evaluation is not clearing. Settlement-risk escrow is not clearing: it consumes clearing decisions. RAILS (Real-Time Agent Integrity & Ledger Settlement) is the integrity and clearing layer for agentic commerce, spanning a per-output reliability score, a published reliability record, and a clearing...

论文介绍 随着自主代理能够进行谈判、购买和代码部署,如何判定其是否履行了委托义务以及在失败时如何追责,成为“代理清算”这一关键缺失环节。现有通信、支付等协议均未解决此问题。本文提出了 RAILS(实时代理完整性与账本结算),作为一个完整性与清算层。它通过为每个输出计算可靠性分数、发布可靠性记录并执行结算决策,为代理商业提供了一个中立的责任判定和结算机制。

Uncertainty Principles for the Number Theoretic Transform

第一作者: Giulio Malavolta · 方向: 安全研究

Abstract:Motivated by polynomial identity testing with exponentials (Li and Wu, ITCS'26), we study uncertainty principles for the number-theoretic transform (NTT). We show that the NTT satisfies strong sparsity tradeoffs: For every fixed prime $q$ and for all but finitely many primes $p \equiv 1 \pmod q$ every nonzero $f\in \mathbb F_p^{\mathbb Z_q}$ and its number-theoretic transform $\hat f$ satisfy \[ |\mathrm{Supp}(f)| + |\mathrm{Supp}(\hat f)| \ge q+1. \] Thus, a $k$-sparse function has transform support at least $q-k+1$. As our main technical contribution, we prove a probabilistic version of the above uncertainty principle, averaged over primes $p$, in the regime $p=q^{O(1)}$. As an application, we obtain a black-box identity test for $k$-sparse exponential polynomials of degree at most $d$ with vanishing soundness error, for $q$ moderately larger than $k$.

论文介绍 受指数多项式恒等测试研究的启发,本文研究了数论变换(NTT)的不确定性原理。研究证明了 NTT 满足强稀疏性权衡:对于几乎所有的素数 p,非零函数 f 及其 NTT 变换 f̂ 的支撑集大小之和至少为 q+1。这意味着一个 k 稀疏函数的变换至少有 q-k+1 个非零分量。作为主要技术贡献,作者证明了在 p=q^{O(1)} 区间上关于素数 p 的概率版本不确定性原理,并将其应用于指数多项式恒等测试。

Differentially Private Range Subgraph Counting

第一作者: Xian Chen · 方向: 隐私保护

Abstract:Subgraph counting is a fundamental problem in graph analysis. Motivated by practical scenarios where graph analytics are performed on subgraphs induced by selected vertices -- rather than on the entire graph -- and by growing privacy concerns, we initiate the study of differentially private range subgraph counting (DPRSC). The goal is to privately count occurrences of a fixed pattern graph within induced subgraphs defined by multi-dimensional attribute ranges. Unlike classical point counting, subgraph counting is inherently nonlinear and exhibits high sensitivity: a single edge modification can affect many subgraph occurrences. We present the first efficient algorithms for DPRSC with small additive error. Our approach introduces a subgraph projection that reduces DPRSC to weighted orthogonal range counting, enabling the use of range trees and local sensitivity estimation to...

论文介绍 子图计数是图分析的基本问题。本文针对在由多维属性范围定义的顶点诱导子图中进行隐私保护的模式图计数问题,即差分隐私范围子图计数,展开了研究。子图计数具有高灵敏度,单个边的改变可能影响多个子图的出现次数。作者首次为该问题提出了具有小加性误差的高效算法,其核心是引入一个子图投影,将问题归约为加权正交范围计数,从而利用范围树和局部灵敏度估计来设计机制。

Multidimensional Resilience for Electrical Power Systems: Systematic Review, Integrated Index, and Validation under Real-World Cyber-Physical Attack Scenarios

第一作者: Isaac Ortega Romero · 方向: 安全研究

The accelerating decarbonization of energy systems has transformed electrical power systems into complex infrastructures exposed to threats whose interactions generate systemic vulnerabilities that conventional resilience approaches fail to capture. Although resilience assessment has expanded across multiple dimensions, existing studies largely examine them in isolation or adjacent pairs, leaving cross-dimensional couplings insufficiently explored. This study demonstrates i) that single-dimension assessments fail to capture the degradation produced by simultaneous cross-dimensional failures, ii) the nonlinear amplification emerging when physical, operational, and digital-cyber dimensions are jointly compromised, and iii) the intensification imposed by climatic and economic-regulatory stressors. To this end, we leverage a hybrid quantitative methodology. A PRISMA 2020 review with...

论文介绍 电力系统正面临物理、运行和数字-网络等多维度威胁的复杂交互。传统单维度或相邻维度的韧性评估方法无法捕捉跨域耦合失效的系统性影响。本文系统性地展示了单维度评估的不足,揭示了物理、运行和数字-网络维度同时受损时产生的非线性放大效应,并提出了一个综合性的多维韧性指数。该研究基于混合定量方法,为关键基础设施的韧性评估提供了更全面的视角。

TOMOYO Linux: A Mandatory Access Control Method Based on Application Execution State

第一作者: Toshiharu Harada · 方向: 系统安全

Existing access control methods grant access requests based on the combinations of applications as subject and files as objects. Therefore intents of applications and the possible effects caused by granting the access requests have not been taken into consideration. In this paper, we propose a new access control method based on application history and intents. With our access control method, system administrators can reduce the risks caused by malicious access attempts and wrong operations. In this paper, the concept and implementation design will be explained as well as the brief evaluation report of TOMOYO Linux, our implementation of the new access control method to Linux.

论文介绍 现有的访问控制方法通常基于主体(应用)和客体(文件)的组合来授权,但未考虑应用的意图以及授权可能带来的潜在影响。本文提出了一种基于应用历史记录和意图的新访问控制方法。该方法允许系统管理员根据应用的行为模式制定策略,从而降低恶意访问尝试或错误操作带来的风险。论文介绍了 TOMOYO Linux 的概念、实现设计以及初步评估报告。

VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation

第一作者: Harshil Patel · 方向: 密码学协议

Abstract:As the Model Context Protocol (MCP) standardizes tool-calling for autonomous agents, it introduces a critical, unexamined attack surface: the error-handling loop. We hypothesize that tool error messages possess implicit authority, triggering corrective reasoning modes that bypass standard safety heuristics. We introduce VATS (Vulnerability Analysis of Tool Streams), a mutation-driven framework that systematically evolves adversarial payloads across seven structural and linguistic dimensions. Our evaluation across four frontier models, Gemini 3.1 Pro, GPT-5.5, GLM-5.1, and Qwen3-Coder, demonstrates that error-path injection triples the success rate of standard indirect prompt injection (IPI), achieving up to 100% compliance in controlled evaluations. We isolate structural positioning (sandwiching instructions within error context) as the most effective exploit vector across all...

论文介绍 模型上下文协议(MCP)在标准化自主代理的工具调用时,引入了一个未被审视的攻击面:错误处理循环。本文假设工具错误消息具有“隐式权威”,能触发纠正性推理模式从而绕过安全检查。作者提出了 VATS 框架,通过系统变异在七个结构和语言维度上生成对抗性载荷。评估显示,错误路径注入使标准间接提示注入的成功率提高了三倍,在某些模型上可达 100% 成功率,其中指令被嵌入错误上下文的“三明治”结构最为有效。

FADRW: A Feature-Aware Modulated and Dynamically Reweighted Loss for Few-Shot Linguistic Steganalysis

第一作者: Shuo Liu · 方向: 安全研究

Abstract:The ubiquity of social media platforms facilitates malicious linguistic steganography, posing significant security risks. However, detection is severely hampered by two fundamental issues during model training. Firstly, extreme class imbalance (less than 1% steganographic samples) induces a strong decision bias. Secondly, the invisibility of generative steganography means its features are nearly indistinguishable from benign text; this similarity, compounded by their extreme rarity, leads to severe feature marginalization, where faint steganographic signals are completely overwhelmed. To directly address these optimization-level challenges, we propose FADRW (Feature-Aware Modulated and Dynamically Reweighted Loss), a novel loss function framework engineered for few-shot steganalysis. FADRW employs Dynamic Reweighting to progressively counteract decision bias, and a...

论文介绍 该研究针对少样本语言隐写检测中的类不平衡和特征边缘化问题,提出FADRW损失函数框架。该方法通过动态重加权逐步纠正决策偏差,并引入特征感知调制来增强微弱隐写信号。这有助于提升检测模型在社交媒体恶意隐写场景下的性能。

Bridged SBI: Correcting Biased Low-Fidelity Posteriors for Cost-Efficient High-Fidelity Inference

第一作者: Gahee Kim · 方向: 具身智能 · 来源: cs.RO

Accurate calibration of particle-based simulators is crucial for robotic earthwork simulation, but analytical calibration is challenging due to this task's highly nonlinear particle dynamics and the black-box nature of conventional simulators. Although simulation-based inference (SBI) can estimate posterior distributions over simulation parameters solely from forward simulations, applying SBI directly to high-fidelity (HF) particle simulators is often computationally prohibitive. Low-fidelity (LF) simulators with coarser particles can reduce this cost, but changes in particle size and particle count shift the parameter values needed to reproduce the same observation, producing biased LF posteriors. We propose Bridged SBI, which leverages a biased but informative LF posterior to guide HF inference. This method first uses inexpensive LF simulations to identify a coarse high-density...

论文介绍 本文解决机器人粒子仿真校准中高保真模拟计算成本高和低保真后验偏差的问题。提出Bridged SBI方法,利用低保真后验分布指导高保真推理,通过粗粒度模拟定位高密度区域。该方法能提升校准效率,适用于非线性粒子动力学任务。

FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning

第一作者: Haihao Lin · 方向: VLA 通用模型 · 来源: cs.RO

Action-supervised fine-tuning of vision-language-action (VLA) policies fits demonstrations effectively but constrains only the directions that change predicted actions, leaving visual structure consistent across action-equivalent states free to collapse. We formalize this as residual visual collapse along local action fibers and propose FiberTune, a training-time objective that preserves teacher-structured visual residuals without adding inference-time overhead. FiberTune uses an online action probe to estimate action-predictive feature directions, filters them from intermediate visual-token representations, and aligns the resulting probe-filtered residuals to a frozen visual teacher while regularizing their effective rank. Under identical training conditions, FiberTune improves over task-loss-only fine-tuning in every one of six controlled simulation settings spanning two benchmarks...

论文介绍 研究视觉语言动作(VLA)政策微调中视觉结构崩溃的问题,称为残差视觉沿局部动作纤维坍塌。提出FiberTune训练目标,使用在线动作探针估计特征方向,过滤并保留教师结构视觉残差。该方法在仿真设置中改善了任务性能,无推理开销。

HARBOR: A Harness Framework for Agentic Robot Reinforcement Learning

第一作者: Zechu Li · 方向: 策略学习 · 来源: cs.RO

Reinforcement learning (RL) has become a powerful paradigm for robot learning, particularly in sim-to-real settings, but its broader adoption remains limited by the engineering pipeline surrounding the algorithms. Building tasks, shaping rewards, and tuning hyperparameters require substantial expert effort, making RL workflows costly and difficult to scale. We introduce HARBOR, an agentic framework that frames robot RL automation as a harness-engineering problem: given a simulator codebase and a task specification, it automates the workflow from environment setup to policy training in simulation. HARBOR decomposes such high-level objectives into bounded stages executed by specialized agents through standardized commands, persistent artifacts, executable gates, and reusable knowledge, and scales iteration via decentralized parallel trials and experience learning across runs. We evaluate...

论文介绍 针对机器人强化学习工程流程复杂、需大量专家调优的问题,提出HARBOR代理框架。该框架将机器人RL自动化视为线束工程问题,分解目标为阶段,由专门代理执行标准化命令。通过并行试验和跨运行经验学习,提高工作流可扩展性。

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation

第一作者: Songlin Wei · 方向: 机器人操作 · 来源: cs.RO

Humanoid foundation models are advancing faster than we can evaluate them. While real-world testing is expensive and difficult to reproduce, existing simulation benchmarks focus primarily on table-top or wheeled robots. A scalable and reproducible benchmark for whole-body humanoid loco-manipulation remains an open problem. To this end, we present SIMPLE, a unified simulation testbed for humanoid policy learning and evaluation. SIMPLE couples the accurate contact-rich dynamics of MuJoCo with the photorealistic rendering of IsaacSim. It provides a large-scale environment comprising 60 diverse whole-body tasks, 50 indoor scenes, and over 1,000 object assets. To facilitate scalable data collection, the framework integrates two data generation pipelines: automated trajectory generation via motion planning and a low-latency VR teleoperation interface. We further integrate and benchmark...

论文介绍 提出SIMPLE统一仿真测试平台,用于人形机器人全身运动操作政策学习和评估。平台结合MuJoCo物理引擎和IsaacSim渲染,提供60个任务、50个场景和1000多个对象资产。集成自动化轨迹生成和VR遥操作界面,支持可扩展数据收集和基准测试。

CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning

第一作者: Markus Knauer · 方向: VLA 通用模型 · 来源: cs.RO

Enabling robots to understand and execute tasks from natural language commands while maintaining data efficiency remains challenging. Foundation models such as vision-language-action (VLA) and vision-language models (VLMs) provide intuitive interaction channels but require extensive data; task-parameterized imitation learning achieves data efficiency but lacks natural language grounding. This work bridges this gap through a modular architecture combining task-parameterized kernelized movement primitives (TP-KMPs) with pretrained VLMs. During learning, skills are acquired from 2 to 5 kinesthetic demonstrations, and the VLM generates skill schemas describing each skill's parameters and preconditions. During execution, the VLM interprets commands to select skills, reason about parameter bindings, and create novel behaviors through covariance-weighted composition. When no skill or...

论文介绍 通过结合任务参数化核运动原语(TP-KMPs)和预训练视觉语言模型(VLMs),实现语言驱动的机器人技能选择与组合。学习阶段从少量演示获取技能,VLM生成技能模式;执行阶段VLM解释指令以选择技能、绑定参数和创建新行为,提升数据效率。

Cooperative Long Rope Skipping via Multi-Agent Reinforcement Learning

第一作者: Zihao Wang · 方向: 导航与运动 · 来源: cs.RO

Humans exhibit remarkable motor agility, enabling a wide range of dynamic skills such as running and jumping, which highlights the great potential of humanoid robots for athletic locomotion. Among athletic sports, long rope skipping requires two rope turners to cooperatively swing the rope while adapting to a player under different jumping rhythms, making it a meaningful yet challenging task for humanoid robots. Although existing methods for humanoid sports have achieved success in single-agent and interaction-free settings, such as running, dancing, and parkour, task scenarios that require precise coordination among multiple participants remain largely unexplored. To this end, we propose Marope, a multi-agent reinforcement learning (MARL) framework for cooperative long rope skipping with multiple humanoid robots. Specifically, Marope adopts a hierarchical reinforcement learning...

论文介绍 针对人形机器人协作长绳跳绳这一动态任务,提出Marope多智能体强化学习框架。该框架采用分层强化学习,使机器人摇绳和跳绳智能体协同适应不同节奏。这推动了多智能体人形机器人在需要精确协调的体育运动场景中的应用。

Q-VGM: Q-Guided Value-Gradient Matching for Flow-Matching VLA Policies

第一作者: Ziqian Wang · 方向: VLA 通用模型 · 来源: cs.RO

We propose Q-Guided Value-Gradient Matching (Q-VGM), an off-policy reinforcement learning (RL) method that tackles a long-standing challenge in fine-tuning flow-matching vision-language-action (VLA) policies: efficiently improving an expressive flow-matching action expert with respect to a learned Q-function. Effective improvement must exploit the first-order (gradient) information of the critic, but this is difficult for flow policies, because directly back-propagating the value through their multi-step denoising process is numerically unstable at VLA scale, while the tractable action likelihoods required by policy-gradient methods are unavailable under iterative denoising. Existing value-based methods either backpropagate through the full denoising chain, use the critic only at test time without updating the policy, or distill critic-improved actions as terminal labels without...

论文介绍 提出Q-VGM离策略强化学习方法,用于微调流匹配视觉语言动作(VLA)政策。该方法通过Q引导价值梯度匹配,利用批评器一阶信息改进流匹配策略,避免直接反向传播的数值不稳定。这提高了VLA模型在复杂任务中的优化效率和稳定性。

BRAIN: Bayesian Reasoning via Active Inference for Agentic and Embodied Intelligence in Mobile Networks

第一作者: Osman Tugay Basaran · 方向: 策略学习 · 来源: cs.AI

Future sixth-generation (6G) mobile networks will demand artificial intelligence (AI) agents that are not only autonomous and efficient, but also capable of real-time adaptation in dynamic environments and transparent in their decisionmaking. However, prevailing agentic AI approaches in networking, exhibit significant shortcomings in this regard. Conventional deep reinforcement learning (DRL)-based agents lack explainability and often suffer from brittle adaptation, including catastrophic forgetting of past knowledge under non-stationary conditions. In this paper, we propose an alternative solution for these challenges: Bayesian reasoning via Active Inference (BRAIN) agent. BRAIN harnesses a deep generative model of the network environment and minimizes variational free energy to unify perception and action in a single closed-loop paradigm. We implement BRAIN as O-RAN eXtended...

论文介绍 研究未来6G移动网络中AI代理的挑战,如缺乏可解释性和脆弱适应性。提出BRAIN代理,通过贝叶斯推理和主动推理,利用深度生成模型最小化变分自由能,统一感知和行动,实现自适应决策。该方法可应用于动态网络环境中的智能控制。

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models

第一作者: Hao Shi · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Temporal modeling is essential for robotic manipulation, as effective control requires both memory of past interactions and imagination of future states. However, most VLA models rely primarily on the current observation and therefore struggle with long-horizon, temporally dependent tasks. Cognitive science suggests that humans rely on working memory to buffer short-lived context, the hippocampal system to preserve episodic memory of past experience, and internal models to imagine possible future state evolution. Inspired by these mechanisms, we propose MemoryVLA++, a full temporal modeling framework that equips VLA models with memory and imagination for robotic manipulation. A pretrained VLM encodes the current observation into perceptual and cognitive tokens, forming working memory. These tokens query a Perceptual-Cognitive Memory Bank to retrieve relevant historical...

论文介绍 机器人操作需要处理时间依赖任务,但现有视觉-语言-动作模型主要依赖当前观察。提出MemoryVLA++框架,通过工作记忆、情景记忆和想象力机制,增强模型的时间建模能力,提升长期操作任务性能。

iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

第一作者: Zhenyu Wu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventional embodied frameworks rely on low-dimensional structured action vectors (e.g., joint angles and end-effector poses), which suffer from limited expressive capacity, poor generalization across diverse embodiments, and unnatural dynamic modeling for complex physical interactions. To address these limitations, this paper proposesiMac (Image as Action Control), a novel unified control paradigm that treats raw visual images as native action representations for embodied world models. Departing from traditional explicit kinematic action encoding, iMac formulates continuous visual manipulation as image-based action tokens, which inherently encapsulate spatial motion intentions, interactive geometric constraints and subtle physical...

论文介绍 传统低维动作向量在 embodied世界模型中表达有限且泛化差。提出iMaC,将原始视觉图像作为原生动作表示,封装空间运动意图和物理约束,用于视觉机器人决策和环境模拟。

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

第一作者: Jisong Cai · 方向: 机器人操作 · 来源: cs.RO

Abstract:World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning. However, existing world-action models couple world prediction and action execution at the same temporal resolution, forcing the world branch to model near-term frame variations that are redundant and weakly informative. We posit that strictly binding world prediction and action execution to the same temporal rhythm may underutilize the potential of the video branch for embodied control. Therefore, we propose AHA-WAM, an Asynchronous Horizon-Adaptive World-Action Model built on a dual Diffusion Transformer (DiT) architecture that reorganizes world-action modeling around this temporal asymmetry. AHA-WAM instantiates the video DiT as a low-frequency world planner that maintains rolling key-value memory...

论文介绍 世界-动作模型中世界预测和动作执行耦合在同一时间分辨率,限制了视频分支潜力。提出AHA-WAM,基于双扩散变换器架构,异步处理世界预测和动作执行,使用观察引导上下文路由,优化策略学习。

SynManDex: Synthesizing Human-like Dexterous Grasps from Synthetic Human Pre-Grasps

第一作者: Yanming Shao · 方向: 机器人操作 · 来源: cs.RO

Abstract:Human hand-object interactions encode functional intent, but direct transfer to robotic hands often fails under morphology, contact, and reachability constraints. We present SynManDex, a synthetic pipeline that uses generated human pre-grasps as affordance-aware proposals and resolves the final contacts with robot-native optimization. SynManDex samples object-conditioned digital human pre-grasps, retargets them to dexterous robotic hand poses, optimizes force-closure contacts on the target embodiment, and admits trajectories that pass checks from each step. The resulting keyframes support both grasp-and-lift demonstrations and various prehensile manipulation tasks such as tea pouring, photo taking, and flute playing, designed via VLM agents. As a result, SynManDex combines high grasp quality (86.4\% grasp stability) with 4.67/5 human-likeness (93.4\%). It achieves 80.7\%...

论文介绍 人类手-物体交互转移到机器人手常因形态差异失败。提出SynManDex合成流程,从生成的人类预抓取中采样,重定向到机器人手并优化接触,实现高抓取质量和类人性,适用于多种操作任务。

AetheRock: An Arm-Worn Robot Teaching System for Force-Guided Vision-Tactile Learning

第一作者: Hong Li · 方向: 机器人操作 · 来源: cs.RO

Abstract:Force and tactile sensing are indispensable in contact-rich manipulation. However, force-aware robot learning faces critical challenges due to the incompatible assembly of tactile and force sensors in handheld or wearable devices. To address these limitations, we first introduce AetheRock for gripper-force, vision, and tactile data collection, which is an arm-worn device featuring a modular and easily manufactured visuo-tactile sensor, GelSlim-MiniFab, at the fingertip, a resistive pressure sensor at the human finger contact region, a customized PCB module, and a wearable kit for comfortable and robust collection. Building on this, we propose ForceVT, a representation learning framework that uses force and vision to guide fidelity-agnostic tactile learning, enabling robust inference in any tactile situation. Real-world experiments show that AetheRock achieves qualified data...

论文介绍 力觉机器人学习中,触觉和力传感器组装不兼容。介绍AetheRock臂穿戴设备,配备视觉-触觉传感器,和ForceVT框架,使用力和视觉引导触觉学习,提高数据收集和推理鲁棒性。

Difference-Aware Retrieval Policies for Imitation Learning

第一作者: Quinn Pfeifer · 方向: 模仿学习 · 来源: cs.RO

Abstract:Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment. We show that reusing the training data during inference via a semi-parametric retrieval-based imitation learning approach can alleviate this challenge. We present Difference-Aware Retrieval Policies for Imitation Learning (DARP), a semi-parametric retrieval-based imitation learning approach that addresses this limitation by reparameterizing the imitation learning problem in terms of local neighborhood structure rather than direct state-to-action mappings. Instead of learning a global policy, DARP trains a model to predict actions based on $k$-nearest neighbors from expert demonstrations, their corresponding actions, and the relative distance vectors between neighbor states and query states. DARP requires no additional...

论文介绍 行为克隆在分布外状态泛化差,由于复合错误。提出DARP,半参数检索方法,通过k近邻和相对距离向量预测动作,基于局部邻域结构,减少错误并提高泛化能力。

Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models

第一作者: Seongbin Park · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these policies offer no guarantees against collisions with task-irrelevant objects in the scene. Existing safety filters sidestep this problem by querying a vision-language model (VLM) to identify obstacles and their locations. This, however, is too slow to run in the control loop and can only be invoked at episode initialization, leaving the filter unable to track moving obstacles. We discover that a small number of attention heads within a VLA model reliably localize the object the policy intends to approach. These heads can be exploited within a training-free safety framework that obtains the active target from the attention heads at every step, treats the remainder of the scene as obstacles, and feeds these into a Control Barrier...

论文介绍 视觉-语言-动作模型缺乏对任务无关物体的碰撞保证。发现模型中的注意力头可定位目标物体,提出无训练安全框架,使用注意力引导和控制屏障函数,实现实时障碍物跟踪和安全性提升。

ProbeAct: Probe-Guided Training-Free Failure Recovery in Vision-Language-Action Models

第一作者: Fan Zhang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models demonstrate strong perfor-1 mance on language-conditioned robotic manipulation within their training dis-2 tribution, yet their generalization capabilities remain fundamentally limited. They3 lack the robustness required to handle perturbations, frequently failing when con-4 fronted with lighting changes, altered camera viewpoints, or small initial-state5 variations. We propose PROBEACT, a training-free runtime intervention frame-6 work that detects and recovers from grasping and placement failures in pre-7 trained VLA policies without modifying their weights or requiring additional8 demonstrations. PROBEACT combines three components: (i) a lightweight multi-9 target hidden-state probe that predicts the 3D positions of task-relevant objects10 from intermediate VLA features, with Hungarian-matched identity tracking for11 multi-object scenes...

论文介绍 该论文针对视觉-语言-动作模型在光照、视角等扰动下易失效的问题,提出了一种名为ProbeAct的无训练运行时干预框架。该框架通过轻量级探针从模型中间层特征预测任务相关物体的3D位置,并结合碰撞检测与匈牙利匹配进行身份跟踪。当检测到抓取或放置失败时,系统会触发预定义的校正动作序列,无需修改原始策略权重或额外演示,从而提升了预训练VLA策略在非理想状态下的鲁棒性。

Safe Polytope-in-Polytope Motion Planning and Control with Control Barrier Functions

第一作者: Alejandro Gonzalez-Garcia · 方向: 具身智能 · 来源: cs.RO

Abstract:Autonomous mobile robots operating in tight environments require motion planning frameworks that account for the physical footprint of the robot. Simplifying the geometry to a point or a circle is conservative and discards information needed to successfully and safely traverse narrow passages. This work proposes a safe local motion planning and control method that guarantees that a polytopic robot footprint stays inside a continuously updated convex free-space region. The containment condition is formulated as a set of discrete-time control barrier function constraints within a model predictive controller. The number of safety constraints depends on the complexity of the local free-space geometry and the robot shape, instead of the number of obstacles. The proposed free-space formulation does not need any obstacle detection or segmentation. A comparative analysis against a...

论文介绍 本研究针对自主移动机器人在狭窄环境中的安全运动规划问题。为精确考虑机器人物理形状,作者将机器人建模为多面体。其核心方法是在模型预测控制器中,将机器人多边形保持在凸自由空间区域内的条件,转化为一系列离散时间控制障碍函数约束。该方法无需障碍物检测,仅需维持局部自由空间的几何描述,安全性约束的数量取决于自由空间和机器人形状的复杂度,而非障碍物数量,保证了轨迹的安全性。

Modeling Components and Connections in Cyber-Physical Systems

第一作者: Kate Sanborn · 方向: 具身智能 · 来源: cs.RO

Abstract:Text based configuration files for cyber-physical systems show the hierarchy of component modules well but often hide the details of connections and interfaces between modules. A model-based visual approach to these configuration files can better capture this information. The XML structure of Robot Operating System (ROS) launch files can be improved using a modeling approach. This paper presents ROSLaunchVisual, a model-integrated environment built on WebGME for designing, visualizing, and managing ROS launch files. The tool raises the level of abstraction by allowing developers to create and modify launch files using a graphical interface that represents nodes, publishers, subscribers, and arguments as interconnected components. The tool provides a dynamic system analysis that can then be used in the static development and analysis of new and existing launch files...

论文介绍 本文聚焦于赛博物理系统的配置管理。传统的文本配置文件能清晰展示模块层级,但模块间的连接和接口细节常被隐藏。作者提出了ROSLaunchVisual,一个基于WebGME构建的模型集成环境,用于设计、可视化和管理ROS启动文件。该工具将节点、发布者、订阅者和参数等抽象为可互连的图形组件,提升了开发抽象层次,并提供了动态系统分析功能,以辅助开发与维护。

Physics-Aware Sparse Learning and Selective Online Adaptation for Euler-Lagrange Robot Dynamics

第一作者: Rishabh Dev Yadav · 方向: 具身智能 · 来源: cs.RO

Abstract:Accurate dynamics models are essential for model-based robotic control, yet nominal Euler--Lagrange models often become inaccurate in the presence of payload variation, unmodeled coupling, friction, aerodynamic effects, and changing operating conditions. Most learning-based correction methods improve prediction accuracy by introducing a single additive residual, but do not preserve the internal mechanical structure of Euler--Lagrange systems. This leads to models that do not preserve symmetry, positive-definiteness, or the coupling between inertia and velocity-dependent terms, which can result in physically inconsistent predictions and reduced reliability when embedded in model-based controllers. We propose a structure-preserving residual learning framework that decomposes model mismatch into an inertia correction, the corresponding induced Coriolis term, and a...

论文介绍 精确的机器人动力学模型是基于模型的控制基础,但标称的欧拉-拉格朗易模型常因负载变化、未建模耦合等因素而失准。现有学习修正方法常采用单一残差,可能破坏系统的物理结构。本文提出一种结构保持的残差学习框架,将模型失配分解为惯性修正及其诱导的科里奥利项、以及扰动力。该框架确保了修正后模型仍保持对称性、正定性等物理特性,从而为模型预测控制器提供物理一致的动力学预测。

ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies

第一作者: Haodi Hu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) policies provide strong priors for language-conditioned manipulation, but remain brittle in off-nominal states requiring targeted recovery. We propose ReCoVLA -- a failure-conditioned residual recovery framework that keeps a pretrained VLA policy frozen, uses an external vision-language model (VLM) to infer the failure mode and recovery stage, and compiles a structured reward from task-relevant components. Rather than using the VLM to generate actions or rewards directly, ReCoVLA uses it as a semantic reward selector: it predicts a recovery descriptor and reward mask for in-simulation residual-policy training, followed by zero-shot sim-to-real deployment of the trained recovery policies. This decouples high-level failure understanding from low-level corrective control to support different VLAs. Experiments across short-horizon, long-horizon, and...

论文介绍 视觉-语言-动作策略在需要针对性恢复的异常状态下表现脆弱。本文提出ReCoVLA,一个基于失败条件的残差恢复框架。它保持一个预训练的VLA策略冻结,利用外部的视觉语言模型来推断失败模式和恢复阶段,并据此从任务相关组件中编译结构化奖励。VLM在此作为语义奖励选择器,用于在仿真中训练残差恢复策略,然后进行零样本的仿真到真实迁移。该方法将高层失败理解与底层控制解耦,以支持不同的VLA。

Motion planning for hundreds of floating robots

第一作者: Jan Kamm · 方向: 具身智能 · 来源: cs.RO

Abstract:Planning collision-free motion for large robot fleets is difficult because collision avoidance induces strong inter-agent coupling that grows rapidly with team size. We consider omnidirectional floating robots on water, where choreographies are specified by sparse keyframes and an interactive tool must generate trajectories within seconds, even when transitions span minutes and thousands of time steps. We propose a scalable pipeline that builds a collision graph from an initialization, decomposes the coupled problem into interaction clusters, and solves clusters independently (and in parallel) with robustness mechanisms for common decomposition pathologies. We validate the approach in simulations up to 500 robots. The synthesized trajectories have also been deployed in two real-world demonstrations, on Lake Zürich with a fleet of 24 Way of Water crafts and at the Time Space...

论文介绍 为大规模机器人舰队规划无碰撞运动极具挑战,因为避碰导致的智能体间耦合随团队规模剧增。本文针对水上全向浮式机器人,提出一个可扩展的规划流水线。该方法从初始构型构建碰撞图,将耦合问题分解为多个交互集群,然后独立且并行地求解各集群,并内置机制处理常见的分解病理问题。该方法能在几秒内为数百个机器人生成时长数分钟的轨迹,已在仿真(500个机器人)和真实世界演示(苏黎世湖)中得到验证。

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience

第一作者: Ruizhe Liao · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dexterous manipulation presents substantial challenges for imitation learning due to its high-dimensional action space and complex contact-rich dynamics. Policies trained purely from demonstrations often suffer from compounding errors during deployment and require large amounts of expert data to achieve reliable performance. To move beyond the limitations of demonstration data, in this work, we propose DexPIE, a post-training framework for dexterous policy improvement from experience collected through real-world deployment. First, DexPIE enables effective exploration coverage through a dexterous-hand-adapted intervention system and multi-stage DAgger-style data collection across initial and intermediate task stages, providing reliable supervision for accurate policy evaluation. To reduce temporal noise between post-training rollouts and demonstration data, we introduce...

论文介绍 灵巧操作因高维动作空间和复杂接触动力学对模仿学习构成挑战。纯演示训练的策略易产生复合误差。本文提出DexPIE,一个从真实世界部署收集的经验中改进灵巧策略的后训练框架。它通过一个适应灵巧手的干预系统和多阶段的DAgger风格数据收集,实现有效的探索覆盖,为策略评估提供可靠监督。此外,该框架引入技术以减少后训练轨迹与演示数据间的时间噪声,从而在少量真实交互下提升策略性能。

Shape Formation for the Cooperative Transportation of Arbitrary Objects Using Multi-Agent Reinforcement Learning

第一作者: Mohamed Sayed · 方向: 导航与运动 · 来源: cs.RO

Abstract:Cooperative object transportation is essential in numerous domains, including industrial to domestic services. A popular transportation strategy is to carry objects on top of multi-robot systems. The corresponding task is typically solved by decomposing it into three interconnected subproblems: formation control, cooperative navigation, and collision avoidance. A particular challenge posed by real-world objects is their potentially arbitrary shape and non-uniform mass distribution, necessitating robot formations that securely support the object. In this work, we address the challenge of pattern formation control for transporting such real-world objects by proposing a novel multi-agent reinforcement learning approach. Our approach enables a multi-robot system to autonomously position itself underneath an object to support its weight while avoiding obstacles during the formation...

论文介绍 协同物体运输是多机器人系统的重要任务,常涉及编队控制、协同导航与避碰。当运输物体形状任意、质量分布不均时,需要机器人形成能稳定支撑物体的编队。本文提出一种新颖的多智能体强化学习方法来解决此问题。该方法使多机器人系统能够在形成支撑编队的同时,自主移动到物体下方以支撑其重量,并在编队行进过程中避开障碍,适用于真实世界中的任意形状物体运输任务。

CT-VAM: A Cerebello-Thalamic-Inspired Vision-Action Model for Efficient Visuomotor Control

第一作者: Jiacheng Li · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action models have shown strong promise for robot manipulation, yet raw language is primarily needed to specify task intent rather than to be repeatedly processed during high-frequency low-level execution. Motivated by this separation, we propose a cerebello-thalamic-inspired vision-action model (CT-VAM) for efficient task-conditioned visuomotor control. CT-VAM acts as a compact local execution policy that predicts action chunks from dualview visual observations, proprioception, and a lightweight task condition, potentially enabling a practical cloud-edge paradigm in which high-level semantic reasoning can be handled by large models while fast closed-loop control runs on local hardware. To fuse heterogeneous inputs effectively, CT-VAM introduces TARS (Thalamic Action Routing Stream), a stream-separated conditional attention decoder that independently routes...

论文介绍 视觉-动作模型在机器人操作中前景广阔,但语言在高频低层执行中重复处理效率低。为此,我们提出受小脑-丘脑启发的视觉-动作模型CT-VAM,作为紧凑本地策略,从双视角视觉、本体感觉和轻量任务条件预测动作块,并引入TARS架构融合异构输入,可能实现云边协同范式,提升控制效率。

Efficient Minimal Solvers for Relative Pose Estimation in Autonomous Driving Applications

第一作者: Tao Li · 方向: 导航与运动 · 来源: cs.RO

Abstract:With the advancement of visual sensing systems, computer vision is playing an increasingly important role in autonomous driving and robot navigation. Relative pose estimation in multi-camera systems is essential for accurate vehicle localization and environment perception, demanding high real-time performance and robustness. Existing methods, however, often involve high computational costs and rely heavily on abundant feature matches, limiting their applicability in time-sensitive driving scenarios. To address these limitations, this paper introduces a unified framework for efficient relative pose estimation, built upon a novel translation parameterization and first-order rotation approximation. Within this framework, we propose three efficient minimal solvers specifically designed for autonomous vehicles. The first solver integrates the vertical direction prior from Inertial...

论文介绍 多相机系统中的相对位姿估计对自动驾驶定位至关重要,但现有方法计算成本高、依赖大量特征匹配。本文提出统一框架,基于新颖平移参数化和一阶旋转近似,设计三个高效最小求解器,专为自动驾驶车辆优化,旨在提高实时性能和鲁棒性。

Targeting World Models to Compromise Robot Learning Pipelines

第一作者: Ethan Rathbun · 方向: 具身智能 · 来源: cs.RO

Abstract:World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline. While highly practical, in this work we demonstrate that world models introduce a uniquely stealthy and effective data poisoning entry point into the robot learning supply chain that can result in the deployment of unsafe or otherwise compromised robotic policies despite training on seemingly safe ground truth training data. In contrast to traditional data poisoning techniques which directly implant dangerous trajectories into sold or uploaded datasets, our novel attack methods inject malicious prompts or compromising transition dynamics into visibly safe teleoperated datasets which are only activated once fed...

论文介绍 世界模型作为高效数据工具被集成到机器人学习中,但可能引入隐蔽数据投毒入口。本工作演示新颖攻击方法,将恶意提示或妥协动力学注入看似安全数据集,只在喂入世界模型时激活,揭示机器人学习管道的安全风险,促进安全措施发展。

Goal Sets, Not Goal States: Queryable Robot Goals through Goal-Set Hindsight Relabeling

第一作者: Carlos Vélez García · 方向: 具身智能 · 来源: cs.RO

Abstract:Hindsight relabeling usually turns achieved future states into exact goals, which can overconstrain offline robot learning when task success depends only on a subset of the state. We propose Goal-Set Hindsight Relabeling (GS-HER), a predicate-level generalization of HER in which achieved states certify query-defined goal sets rather than singleton goal states. A binary query specifies which variables define success, making the goal predicate an inference-time input while leaving the underlying offline GCRL algorithm unchanged. Across OGBench tasks and five offline goal-conditioned learners, GS-HER improves performance when full-state goals are bottlenecked by nuisance dimensions and turns hindsight relabeling into a reusable goal interface: one checkpoint can answer multiple robot goal predicates without retraining.

论文介绍 传统后见之明重标记将达成状态转为精确目标,可能在任务成功仅依赖状态子集时过度约束离线学习。我们提出GS-HER,在谓词级别泛化,使达成状态验证查询定义的目标集,通过二元查询指定成功变量,提升性能并成为可复用目标接口。

$ω$-EVA: Envision, Verify, and Act with Latent Interactive World Models

第一作者: Zhenguo Sun · 方向: 多模态具身 · 来源: cs.RO

Abstract:Embodied policies typically map current observations directly to actions, leaving candidate-action consequences implicit. World models provide predictive supervision, representations, or external simulation, but rarely let a policy inspect the imagined consequence of its own proposal before acting. We introduce $\omega$-EVA, a latent interactive world model that realizes an Envision--Verify--Act loop for embodied action generation. Its three-stage framework learns action-conditioned latent dynamics, trains a language-conditioned flow policy on dynamics-aware visual representations, and feeds the policy's proposal back through the world model. A tri-branch refiner jointly reasons over the current state, proposal-conditioned future, and proposed action to produce the final action chunk. Because consequence reasoning remains in latent feature space, $\omega$-EVA avoids generating...

论文介绍 具身策略通常直接映射观察到动作,候选动作后果隐式;世界模型虽提供预测,但很少让策略在行动前检查想象后果。我们引入ω-EVA,隐式交互世界模型,实现想象-验证-行动循环,通过三阶段框架在潜空间进行后果推理,提升动作生成质量。

Harness Engineering for Physical AI: Robot Middleware Is the Harness Layer

第一作者: Sanghoon Lee · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Robot middleware faces a new role in the era of Physical AI. Learned policies, planners, and vision-language-action (VLA) models now enter deployed robots as causal participants on the control path, but the layer that integrates them with timing, scheduling, and network has not been named. Recent language-agent work names this layer the harness, the external system that mediates tools, manages state, bounds resources, and records execution. The robotics community has not yet adopted this framing, and we propose that robot middleware is that harness. A Physical AI harness differs from a software harness in where it intervenes. A software harness mediates at tool-call boundaries. A Physical AI harness must mediate at control, computing, and communication simultaneously, because a learned policy's output crosses all three: its commands shift the trajectory, its inference time...

论文介绍 物理AI时代,机器人中间件需集成学习策略到控制路径,但缺乏命名层。本文提出机器人中间件即线束层,外部系统调解工具、管理状态、限制资源并记录执行;物理AI线束需同时调解控制、计算和通信,因学习策略输出跨越三者。

ReGIL: Retrieval-Guided Imitation Learning from a Single Demonstration

第一作者: Yuying Zhang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Learning robot manipulation policies with deep neural networks from a single demonstration remains highly challenging, as even small deviations from the demonstrated trajectory can quickly compound into failure, while collecting substantial online interaction data is costly. We propose ReGIL, a retrieval-guided imitation learning framework that treats a single demonstration as an external memory. ReGIL repeatedly queries this static memory throughout training to simultaneously guide exploration, generate the regularization buffer, and construct rewards. Specifically, it computes rewards through local temporal alignment between the current trajectory and the retrieved segment, providing step-wise and informative feedback for policy improvement. We evaluate ReGIL on robotic manipulation tasks from the LIBERO and Meta-World benchmarks under the single demonstration setting. ReGIL...

论文介绍 从单次演示学习机器人操作策略具有挑战性,偏差易导致失败,在线数据收集成本高。我们提出ReGIL框架,将单次演示视为外部记忆,训练中重复查询以引导探索、生成缓冲区和构建奖励,通过局部时间对齐计算奖励,提升模仿学习性能。

TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation

第一作者: Huaihang Zheng · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have become a powerful framework for robotic manipulation, and recent studies have introduced tactile or force feedback into VLAs to address contact-rich tasks. However, these models are typically deployed as offline policies. When contact conditions shift from the training distribution, the policy cannot perform online adaptation, leading to problems such as inappropriate contact forces and inefficient retries. Therefore, we propose TORL-VLA, a tactile-guided online reinforcement learning framework that couples tactile feedback with policy refinement for contact-rich manipulation. Our method introduces a tactile-derived wrench-aware VLA to predict reference actions and future wrench sequences, while a lightweight online RL module is used to refine the reference actions. To stabilize learning from mixed exploratory policy-generated and...

论文介绍 VLA模型在接触丰富任务中引入触觉反馈,但作为离线策略,当接触条件变化时无法在线适应。我们提出TORL-VLA框架,耦合触觉反馈与策略细化,引入力觉感知VLA预测参考动作和未来力觉序列,使用在线强化学习模块细化动作,实现在线适应。

Dual Quaternion-Based Unscented Kalman Filter with Visual Inertial Odometry for Navigation in GPS-Denied Environments

第一作者: Mohamed Khalifa · 方向: 导航与运动 · 来源: cs.RO

Abstract:Reliable navigation in GPS-denied environments remains a fundamental challenge in robotics, aerospace, and autonomous vehicle applications. This paper presents a Dual Quaternion-Based Unscented Kalman Filter (DQUKF) equipped with a Visual Inertial Odometry (VIO) algorithm for accurate state estimation enabling navigation in GPS denied locations. The proposed framework formulates the DQUKF in an error state manner, where the nominal pose is represented by a unit dual quaternion and the local pose error is represented by a 6-dimensional twistor parameterization used for sigma point generation, covariance propagation, and measurement correction. In parallel, the VIO algorithm tracks features across image frames, synchronizes measurements between the IMU and camera, and provides visual constraints that complement inertial propagation. Simulation results on the EuRoC MAV dataset...

论文介绍 针对GPS拒止环境下可靠导航这一挑战,本文提出了一种基于双四元数的无迹卡尔曼滤波器。该框架采用误差状态形式,使用双四元数表示名义位姿,并以六维旋量参数化误差状态用于卡尔曼滤波的各个步骤。它与视觉惯性里程计算法并行工作,融合视觉与惯性测量以实现准确的状态估计,旨在为机器人及自主车辆在无GPS区域的导航提供解决方案。

VAIC: Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands

第一作者: Dongting Li · 方向: 具身智能 · 来源: cs.RO

Abstract:Humanoid robots hold immense potential for real-world assistance, yet agile interaction with objects in unstructured environments demands tightly coupled whole-body coordination. Despite recent advancements, current controllers face a critical deployment gap. They rely heavily on dense reference trajectories and perfect state observability, which inherently limits physical generalization. We present Vision Guided Agile Interaction Control (VAIC), a unified framework that bridges this gap by operating exclusively on onboard depth, historical proprioception, and a decoupled user command interface. VAIC employs a two-stage distillation paradigm. First, a privileged teacher policy masters diverse interaction skills using precise object kinematics and exact environmental states. Second, a deployable student policy distills these capabilities by replacing full body tracking with...

论文介绍 为实现人形机器人在非结构化环境中与物体的敏捷交互,本文提出了视觉引导敏捷交互控制框架。该框架仅依赖板载深度图像、本体感觉历史和解耦的用户命令进行操作,旨在弥合当前控制器依赖完美状态观测的部署鸿沟。其核心采用两阶段蒸馏范式,先训练掌握交互技能的特权教师策略,再将其能力蒸馏到仅需部分可观测信息的可部署学生策略中。

VGP-Nav: Metric-Aware Visual Geometric Perception for Robot Navigation

第一作者: Hewei Pan · 方向: 导航与运动 · 来源: cs.RO

Abstract:Reliable robotic navigation necessitates the seamless integration of accurate global localization and dense, metric-consistent obstacle perception. A common strategy to achieve these capabilities involves integrating diverse sensing modalities: cameras offer rich visual features for localization, while active sensors like LiDAR provide direct metric measurements. However, such multi-sensor configurations necessitate complex spatial-temporal calibration and increase deployment overhead. Although vision-only approaches offer a low-cost and scalable alternative, existing monocular visual systems typically struggle to simultaneously achieve efficient, globally consistent localization and dense, metric-consistent geometric perception. To bridge this gap, we propose \textbf{VGP-Nav}, a unified framework for \textit{Metric-Aware Visual Geometric Perception} that relies solely on...

论文介绍 可靠机器人导航需同时实现全局定位和密集、度量一致的障碍感知。传统多传感器方案校准复杂,而纯视觉系统又难以兼顾两者。为此,本文提出VGP-Nav,一个仅依赖单目视觉的统一框架,用于实现度量感知的视觉几何感知。该方法旨在通过单一视觉传感器,同时提供高效的全局一致定位和密集的、具有真实世界尺度的几何信息,从而支持导航任务。

Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection

第一作者: Suyeon Shin · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) policies can deviate from nominal trajectories during manipulation, even when tasks remain physically feasible. Recovering from these deviations is challenging, as they push the policy into unfamiliar state spaces where direct re-planning frequently destabilizes action sequences. We propose Back to the Familiar Future (B2FF), a recovery framework for foresight-driven VLAs that leverages future visual conditioning as a recovery interface. Before execution, the VLA generates a milestone bank of familiar future states conditioned on the clean initial observation. At recovery time, a recoverability-aware selector selects a recovery milestone from this bank and enforces it as a fixed visual goal. This enables the VLA to robustly map off-trajectory observations back to a familiar future. On failure-injected LIBERO, under controlled recovery timing...

论文介绍 视觉语言动作策略在操作过程中可能偏离理想轨迹,而从这些偏离中恢复具有挑战性。本文提出“回归熟悉未来”恢复框架,专为具有预见能力的VLA模型设计。该框架在执行前,利用初始观察条件生成一组未来视觉里程碑;当策略偏离时,一个可恢复性感知的选择器从里程碑库中选取目标,并将其作为固定的视觉目标来引导策略,使其能映射回熟悉状态。

Can we stabilize an inverted pendulum with feedback from a time-of-flight camera?

第一作者: Anthony Czubarow · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Time-of-flight cameras are popular in robotics for providing direct depth information while being compact, inexpensive, and robust to lighting conditions, but their low spatial resolution and depth noise are widely believed to preclude precise feedback control. In this paper, we show that an inexpensive, low-resolution time-of-flight camera provides sufficient feedback to reliably and precisely balance an inverted pendulum on a cart--a canonical benchmark for fast, unstable dynamics.

论文介绍 飞行时间相机虽提供直接深度信息且成本低,但其低空间分辨率和深度噪声普遍被认为无法用于精密反馈控制。本文通过实验证明,一个廉价的低分辨率飞行时间相机能够提供足够反馈,以可靠且精确地平衡一辆推车上的倒立摆——这是一个经典的快速、非线性动力系统基准。这挑战了对低成本深度传感器在精密控制任务中应用能力的普遍看法。

Deterministic Execution of ROS 2 Applications via Lingua Franca

第一作者: Harun Teper · 方向: 具身智能 · 来源: cs.RO

Abstract:The Robot Operating System 2 (ROS 2) is a widely used middleware for robotic systems, characterized by a publish-subscribe (pub-sub) communication mechanism in which computation is structured as callbacks dispatched by ROS 2 executors. Despite its popularity, the pub-sub pattern in ROS 2 is inherently nondeterministic: the order in which these callbacks run is nondeterministic even within a single executor, and distributed deployments add further nondeterminism from the interleaving of messages across nodes and from network latency. Such nondeterminism often leads to concurrency issues and makes it virtually impossible to analyze for safeness and provide guarantees. We present a framework that is able to convert an unmodified ROS 2 application and run it under Lingua Franca (LF), a coordination language for deterministic execution using logical time, so that the same input...

论文介绍 机器人操作系统ROS 2的发布-订阅通信模式存在固有的非确定性,回调执行顺序不确定,这常导致并发问题且难以分析验证。本文提出一个框架,能将未经修改的ROS 2应用程序转换到Lingua Franca协调语言下运行。Lingua Franca通过逻辑时间实现确定性执行,使得相同输入总能产生相同输出和回调序列,从而提升机器人系统的可预测性与可靠性。

Autonomous Obstacle Removal for Excavators through Policy Learning with Particle Simulation

第一作者: Yuki Kadokawa · 方向: 策略学习 · 来源: cs.RO

Abstract:Autonomous obstacle removal from the ground is an important earthwork task, but this is difficult to automate because an excavator must adapt its excavation trajectories over repeated cycles as soil-obstacle conditions change. Learning such state-dependent behavior requires a training environment that reproduces accumulated soil-obstacle interactions, including contact states, terrain deformation, and obstacle visibility. Accordingly, particle-based simulation is suitable for the relevant policy learning. However, particle simulation is computationally expensive, and repeated excavation cycles further increase the learning cost. We observe that the burial condition of an obstacle governs both task difficulty and simulation cost: deeper burial makes obstacle removal harder while also requiring more particles for accurate simulation. This observation motivates a...

论文介绍 挖掘机自主移除地表障碍物是一项重要但困难的土方任务,需要适应反复循环中土壤-障碍物条件的变化。粒子模拟适合此类策略学习,但计算成本高昂。本文观察到障碍物的埋藏深度同时影响任务难度和模拟成本,并基于此提出一种自适应课程学习方法。该方法从浅层障碍物开始学习,逐步增加深度,在控制模拟成本的同时,学习应对不同埋藏条件的挖掘策略。

From USD Scenes to Knowledge Graphs: Zero-Shot Ontology Grounding with LLMs

第一作者: Jiangtao Shuai · 方向: 多模态具身 · 来源: cs.RO

Abstract:Constructing knowledge graphs from 3D simulation scenes is essential for robot task reasoning, but the key bottleneck, grounding scene objects to formal ontology classes, still relies on manually curated dictionaries that are brittle and do not generalize across assets. We investigate whether large language models (LLMs) can automate this grounding step for Universal Scene Description (USD) scenes as a zero-shot, training-free alternative. On a kitchen scene (125 objects) with SOMA-HOME Ontology, LLMs achieve 90-96% exact-match accuracy with descriptive names and 49-89% with abbreviated names, substantially outperforming dictionary and embedding baselines. Under fully opaque names, context-augmented prompting recovers up to 48%. Feature ablation reveals that LLMs primarily exploit semantic cues in the scene graph (sibling names and parent paths); anonymizing these cues reduces...

论文介绍 从3D模拟场景构建知识图谱对机器人任务推理至关重要,但将场景对象自动映射到形式化本体类别仍是瓶颈,目前依赖脆弱的人工字典。本文研究使用大语言模型作为零样本、无需训练的替代方案。实验表明,在厨房场景中,LLM对描述性名称的对象映射准确率很高。分析显示,LLM主要利用场景图中的语义线索(如同级名称和父路径)进行推理,为自动化场景理解提供了新途径。

RAM: Reachability Across Morphologies

第一作者: Tim Walter · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Many stages of the robotic lifecycle, from morphology synthesis to operation, rely fundamentally on the reachable workspace. However, current methods for approximating workspaces are slow, imprecise, or tied to a single morphology. We introduce Reachability Across Morphologies (RAM): a morphology-conditioned, implicit neural representation that acts as a fast, differentiable surrogate for pose reachability, generalising to unseen morphologies while inherently accounting for self-collisions. To train RAM, we publish a large-scale dataset of $3\cdot10^{10}$ samples generated solely from forward kinematics. Experiments show that our model achieves an $ F_1$-score of $86\%$ at nanosecond inference, outperforming the baseline by $14\%$ while reducing inference time by three orders of magnitude. We further demonstrate speed-ups of one and two orders of magnitude for gradient-based...

论文介绍 本文研究机器人形态可达工作空间的近似问题。提出RAM模型,一种基于隐式神经表示的快速、可微分代理,能泛化到未见形态并考虑自碰撞。通过大规模运动学数据训练,实验显示推理速度快且准确率高,适用于机器人设计和操作规划。

PTDL:Multi-Terrain Fall Recovery via Phase-Terrain Decoupled Learning

第一作者: Xiaoyu Xu · 方向: 导航与运动 · 来源: cs.RO

Abstract:Humanoid robots can fall on slopes, gravel, and uneven ground in unstructured environments. We target integrated fall recovery and locomotion: rebuilding balance from a fallen state using proprioception alone and resuming velocity-commanded walking at the fall site. Prior methods often stop at quasi-static rise, neglect the post-fall ground-contact phase, or, when trained on mixed terrains without separating recovery and locomotion phases or per-surface constraints, collapse to a single compromise get-up across surfaces. We propose Phase--Terrain Decoupled Learning (PTDL), which decouples training supervision along phase and terrain axes while deploying one proprioceptive policy. On the phase axis, projected-gravity-gated dual motion-prior discriminators and a probe-to-walk transition link post-fall recovery to commanded walking. On the terrain axis, terrain-stratified...

论文介绍 针对人形机器人在多地形下的跌倒恢复问题。提出PTDL方法,通过解耦训练监督沿相位和地形轴,实现从跌倒状态恢复并继续行走。使用单一本体感觉策略,提高在不同地形上的适应性和鲁棒性。

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

第一作者: Yi Yu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustness on affordable real-world robots largely unexplored. We present a standardized real-world benchmark for evaluating representative VLA and imitation learning policies on the low-cost SO-101 robotic platform. The benchmark comprises four representative manipulation tasks together with unified evaluation protocols, enabling systematic comparison under embodiment uncertainty. Using real-world teleoperated demonstrations, we fine-tune and evaluate $\pi_{0.5}$, SmolVLA, Wall-X, and ACT directly on the physical platform. Beyond conventional task success rates, the benchmark incorporates a structured failure taxonomy, semantic- and execution-level failure...

论文介绍 在低成本SO-101机器人平台上评估视觉-语言-动作模型的鲁棒性。提供标准化基准,包括四个操作任务和统一评估协议。通过失败分类和执行级分析,系统比较模型性能,揭示当前局限性。

Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video

第一作者: Yunhai Han · 方向: 机器人操作 · 来源: cs.RO

Abstract:Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains challenging due to perception errors and embodiment gap. To address this, we introduce Video2Sim2Real, a full-stack framework for autonomous skill acquisition from a single human manipulation video. Our framework first uses off-the-shelf foundation models to reconstruct a simulator-ready digital twin and extract robot and object motion priors. Rather than treating the extracted robot motion as a reliable reference throughout execution, our key idea is to recover and leverage the most fundamental sources of supervision from the demonstrated skill: We identify object-centric keyframes to optimize the corresponding robot configurations using object information from the simulator, and use these configurations as anchors that refine...

论文介绍 从单个人类视频中自主获取机器人灵巧技能。框架使用基础模型重建数字孪生和提取运动先验,通过对象关键帧优化机器人配置,实现技能转移,减少对大量示范数据的依赖。

Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks

第一作者: Raktim Gautam Goswami · 方向: 机器人操作 · 来源: cs.RO

Abstract:Visual world models have shown great potential in learning complex system dynamics. Recent advancements leverage these models as transition functions within Model Predictive Control (MPC) frameworks to solve various control tasks. When applied to robotics, however, they are limited to single-stage tasks such as reaching or grasping, and struggle with multi-stage ones that demand complex sequential planning. In this work, we introduce WorldDP, a world model framework designed for multi-stage robotic manipulation. Our hierarchical approach utilizes a high-level world model as a transition function to optimize for feasible subgoals during runtime, which are subsequently reached by a low-level Diffusion Policy. To further aid in learning dynamics and planning, we incorporate object-centric representations that decouple environmental entities and enable us to plan sequentially with...

论文介绍 提出WorldDP框架,结合对象中心世界模型和扩散策略,用于多阶段机器人操作。分层方法在运行时优化可行子目标,并通过扩散策略执行,支持复杂序列规划,提升任务成功率。

RGB-S: Image-Aligned Tactile Saliency for Robust Dexterous Manipulation

第一作者: Shengcheng Luo · 方向: 机器人操作 · 来源: cs.RO

Abstract:Effective visuo-tactile integration is critical for robotic dexterous manipulation, especially when visual observations are unreliable or occluded. However, robustly aligning sparse, heterogeneous tactile measurements with dense visual representations remains a fundamental challenge. Most existing approaches require policies to learn cross-modal correspondences implicitly from limited demonstrations, without leveraging geometric priors. As a result, they are often data-inefficient and generalize poorly when visual observations are degraded. To address this limitation, we propose a framework that explicitly grounds physical contacts in the image domain. Using robot forward kinematics and camera calibration, we project tactile sensor locations directly onto the RGB image plane. We then render force-modulated Gaussian saliency maps to model spatial uncertainty arising from...

论文介绍 针对视觉不可靠时的灵巧操作挑战。提出RGB-S框架,将触觉传感器位置投影到RGB图像上,生成高斯显著图,对齐跨模态表示,提高数据效率和泛化能力,增强操作鲁棒性。

Guided Discovery of New Behaviors using Diffusion Policies

第一作者: Dian Yu · 方向: 导航与运动 · 来源: cs.RO

Abstract:Diffusion models have become a powerful tool for generative modeling in robotics, with diffusion policies excelling at modeling multimodal action-trajectory distributions. However, when demonstrations are limited, standard sampling often reproduces dominant behaviors while neglecting valid but rare modes, limiting the discovery of novel solutions. Existing approaches, such as guidance methods or combining reinforcement learning with diffusion, either push samples into infeasible regions or struggle to escape local minima, failing to systematically uncover diverse behaviors. To address these challenges, we propose a framework that combines Feynman-Kac correctors with a novel guiding potential that systematically guides diffusion policy samples towards promising yet underrepresented samples. These trajectories are refined using sampling-based trajectory optimization and...

论文介绍 解决扩散策略在有限示范下忽略罕见模式的问题。提出结合Feynman-Kac校正器和引导势的框架,引导样本向有希望但欠代表的轨迹移动,发现新行为,扩展机器人运动策略多样性。

Safe, Fluent and Acceptable Motion Generation and Execution for Human--Robot Interaction in Manufacturing Environments

第一作者: Thibaut Lopez · 方向: 具身智能 · 来源: cs.RO

Abstract:Robots operating in human environments must not only ensure physical safety but also exhibit behaviors that are understandable, fluent, and acceptable to human partners. This paper investigates motion generation strategies that combine safety guarantees with interaction quality considerations, such as motion smoothness and human comfort. While the design of robots capable of ensuring safety in shared human-robot environments has enabled closer and more advanced forms of interaction, these new proximity-based tasks require moving beyond purely technical considerations. In particular, robot behavior must also be addressed from psycho-cognitive and social perspectives. In this context, we argue for the relevance of integrating social-aware motion control into robotic systems. First, we identify the motion parameters that influence human perception and operator experience. Then...

论文介绍 研究制造业环境中安全、流畅且可接受的运动生成。结合安全保证和交互质量考虑,提出社会感知运动控制策略,考虑人机协作的感知和体验,提升协作效率和安全性。

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

第一作者: Yunfan Lou · 方向: 机器人操作 · 来源: cs.RO

Abstract:World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich manipulation, where critical cues arise from physical interaction. In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics. Specifically, Dream-Tac introduces (i) contact-gated visuotactile fusion to selectively integrate tactile signals and (ii) a contact-aware attention bias to better regulate cross-modal interactions during manipulation. To support real-time deployment, we further design a dual-level acceleration strategy, reformulating the contact-aware bias to preserve the fused attention path during training and introducing cache-based diffusion acceleration at...

论文介绍 本文提出Dream-Tac,一种统一触觉世界动作模型,用于接触丰富的机器人操作。该模型联合建模动作、未来视觉观察和触觉动态,引入接触门控视触觉融合和接触感知注意力偏差,以增强操作中的跨模态交互。此外,设计了双级加速策略支持实时部署,旨在提升机器人在物理交互中的感知和动作生成能力。

IR-SIM: A Lightweight Skill-Native Simulator for Navigation, Learning, and Benchmarking

第一作者: Ruihua Han · 方向: 导航与运动 · 来源: cs.RO

Abstract:Simulation plays a key role in automated robotics research supported by large language models (LLMs). However, existing simulators often require custom code or complex interfaces, creating a barrier to rapid prototyping and automated algorithm development. To this end, we propose the Intelligent Robot Simulator (IR-SIM), a lightweight skill-native navigation simulator designed for rapid scenario construction, benchmarking, and robot learning. In IR-SIM, scenarios are entirely defined by YAML configuration files that specify mobile robot kinematics, geometric collision checking, LiDAR sensing, visualization, and behavior modules. This design makes robotic simulation fully describable and reproducible, allowing scenarios to be generated and modified from text prompts through the proposed IR-SIM agent skills. The resulting scenarios can be used for automated benchmarking of...

论文介绍 本文提出IR-SIM,一个轻量级技能原生导航模拟器,旨在支持快速场景构建、基准测试和机器人学习。场景通过YAML配置文件定义,涵盖机器人运动学、碰撞检测、传感器和行为模块,实现可描述和可复现的模拟。该模拟器支持从文本提示生成场景,便于自动化算法开发和评估。

Real-Time and Accurate Collision-Free Teleoperation via Differentiable Constraint-Based Trajectory Planning

第一作者: Max Grobbel · 方向: 导航与运动 · 来源: cs.RO

Abstract:In teleoperation, the human operator typically controls only the end-effector pose, which often leads to self-collisions of the manipulator and collisions with environmental obstacles, since joints and links are not controlled individually. A common strategy to mitigate this issue is to enhance the operator's input using optimal-control-based trajectory planning. As derivative-based solvers require differentiable constraints, existing approaches either approximate robots and obstacles with spheres, reducing geometric accuracy, or approximate derivatives, degrading convergence and increasing computation times. We address these limitations by adapting a recent formulation of differentiable collision-avoidance constraints, based on duality in convex optimization, to the teleoperation setting. The robot is approximated with capsules and the environment with polytopes. We compare...

论文介绍 在遥操作中,末端执行器控制常导致自碰撞或环境碰撞问题。本文采用基于可微约束的轨迹规划来增强操作输入,使用胶囊和多面体近似机器人和环境,基于对偶的可微碰撞避免约束实现准确和实时的无碰撞操作,以提高遥操作的安全性和效率。

Language as a Sensor: Calibrated Spatial Belief Estimation in 3D Scenes from Natural Language

第一作者: Aryan Naveen · 方向: 多模态具身 · 来源: cs.RO

Abstract:Robots deployed in human-centric environments routinely receive natural-language descriptions of spatial information ("I left my backpack on the table") that reference parts of the world beyond their perceptual field of view. Traditional metric-semantic mapping ignores this signal, while off-the-shelf multimodal models remain limited in 3D spatial reasoning and are not directly amenable to fusion with other sensor modalities. To convert language observations into a calibrated spatial distribution, we train a Language Sensor Model (LSM) that maps each utterance and its scene-graph context to a multimodal distribution, with mixture weights encoding referential ambiguity (e.g., "which table") and component covariances encoding spatial uncertainty (e.g., where "on the table" the target lies). We then introduce VL-Map (Vision-Language Metric-Semantic Mapping), a probabilistic...

论文介绍 本文将自然语言描述作为传感器,训练语言传感器模型(LSM)将话语映射到3D场景中的空间分布,编码指代歧义和空间不确定性。结合视觉语言地图(VL-Map),实现概率性的空间信念估计,增强机器人在人类环境中的空间感知能力,支持语言驱动的机器人任务。

Latent Diffusion Policy: Shaping Latent Spaces for Diffusion-Based Robotic Manipulation

第一作者: Zhexuan Zhou · 方向: 机器人操作 · 来源: cs.RO

Abstract:Diffusion-based visuomotor policies operating directly in raw action spaces conflate scene comprehension with trajectory generation within a single denoising process. The resulting velocity field must simultaneously encode scene information and generate precise trajectories, increasing learning complexity and limiting performance on tasks demanding precise temporal coordination across multiple arms. To simplify this joint learning problem, we introduce Latent Diffusion Policy (LDP), a two-stage framework performing flow matching in a deliberately shaped latent space. By absorbing scene understanding into an observation-conditioned CVAE encoder, LDP concentrates the conditional distribution of each observation. Consequently, the flow model avoids implicitly resolving scene-dependent structures; instead, it generates within a pre-concentrated distribution featuring a smoother...

论文介绍 针对扩散策略在原始动作空间中混淆场景理解和轨迹生成的问题,本文提出潜在扩散策略(LDP)。LDP采用两阶段框架,在通过CVAE塑造的潜在空间中进行流匹配,将场景理解与轨迹生成解耦,简化学习过程并提升多臂操作的协调性和精确性。

PhysGraph: A Physics-aware 3D Scene Graph for Perception and Reasoning

第一作者: Haoyu Li · 方向: 具身智能 · 来源: cs.RO

Abstract:To perform a wide range of daily tasks, robots need to construct a 3D representation that is semantically rich, physically grounded, and structured enough to support task planning and affordance prediction. However, existing approaches primarily focus on semantic retrieval, often overlooking physical and kinematic factors. Methods that attempt to model physical properties typically rely on narrow training sets or single-object modeling, limiting scalability and generalization across diverse object types. To address these challenges, we present PhysGraph, a framework that unifies symbolic reasoning with structured 3D geometry to model kinematic and physical properties in cluttered scenes. Given RGB-D observations, PhysGraph reconstructs object-centric 3D geometry and associates object instances across views. It then decomposes objects into functional parts and infers materials...

论文介绍 机器人需要语义丰富且物理基础的3D表示。本文提出PhysGraph,一个物理感知的3D场景图框架,通过统一符号推理和结构化几何,建模对象的运动学和物理属性。该方法支持任务规划和可 affordance 预测,提升机器人在复杂场景中的感知和推理能力。

Real-IKEA: Physical Fidelity is the Prerequisite for Robust Manipulation

第一作者: Kunqi Xu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robotic manipulation robustness often founders on the physics gap between simplified simulations and the resistance-laden real world. In this work, we emphasize that physical realism in articulated interaction is an important ingredient for robust policy learning. We present Real-IKEA, a dataset and simulation framework designed with physical accuracy as a first-class goal. Real-IKEA provides 1,079 articulated asset configurations, derived from 83 authentic IKEA handles and knobs processed through a meticulous six-step physical workflow. For contact-geometry accuracy, we introduce a bidirectional surface-deviation metric to quantify collision meshes. For dynamics realism, we establish resistance-calibrated configurations that vary damping and friction. Crucially, we demonstrate through a Reinforcement Learning (RL) policy that high-fidelity assets enable the discovery of...

论文介绍 模拟与真实世界的物理差距影响操作鲁棒性。本文提出Real-IKEA数据集和模拟框架,强调物理保真度,提供铰接资产配置和阻力校准。通过高保真模拟,促进强化学习策略的发现,以提高机器人操作的鲁棒性和泛化能力。

FAWAM: Force-Aware World Action Models for Closed-Loop Contact-Rich Manipulation

第一作者: Haotian He · 方向: 机器人操作 · 来源: cs.RO

Abstract:Force signals provide critical interaction cues for contact-rich robotic manipulation. However, existing methods mostly use force as an additional observation modality, without fully exploiting its role in modeling future interaction dynamics or guiding execution-time feedback correction. In this paper, we propose FAWAM, a force-aware world action model that incorporates force information at three levels: perception, prediction, and closed-loop execution. FAWAM first encodes historical 6-axis force/torque signals to modulate action generation, then jointly predicts future actions and end-effector wrenches to explicitly model contact evolution. It further introduces a residual correction module that uses the predicted wrench trajectory as an execution-time reference to refine actions online based on real-time force feedback. Real-world experiments across multiple contact-rich...

论文介绍 力信号对接触丰富操作至关重要。本文提出FAWAM,一个力感知世界动作模型,在感知、预测和闭环执行三个层面融入力信息。通过预测末端执行器力矩并在线修正动作,实现基于实时力反馈的闭环控制,以增强操作的鲁棒性和适应性。

OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation

第一作者: Zehao Yu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Recent progress in robot manipulation has been largely driven by learning from large-scale demonstrations. For humanoid robot loco-manipulation tasks, however, existing data sources force an unsatisfying tradeoff between trajectory quality and scalability. Real-world teleoperation provides the highest-quality trajectories but requires dedicated physical space and time-consuming scene resets. Simulation offers an alternative way out of this dilemma: it can produce clean, embodiment-aligned data at scale without any physical hardware. In this paper, we propose OASIS, a simulation-data-driven framework for humanoid loco-manipulation. OASIS automatically reconstructs realistic object assets from real-world images using a 3D generative model. Based on these assets, trajectories are first collected through teleoperation in simulation, and then augmented under diverse domain...

论文介绍 人类机器人loco-manipulation任务面临轨迹质量与可扩展性的权衡问题。本文提出OASIS框架,使用3D生成模型从真实图像自动重建逼真物体资产,在模拟中收集轨迹并进行域增强,为真实世界任务提供高质量、可扩展的数据支持,减少对物理硬件的依赖。

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA

第一作者: Haizhou Ge · 方向: 机器人操作 · 来源: cs.RO

Abstract:Exploratory manipulation often turns an apparent failed attempt into the key evidence for what to do next. For example, a robot pulls a locked cabinet drawer, fails, and only succeeds after opening the lock. The failed pull reveals a latent precondition (the drawer is locked) that determines the minimal-success action chain (the fewest actions that complete the task), here [lock-open, drawer-pull]. Correctly reading this trace is therefore the prerequisite for recovering that chain. We formalize this setting as Exploratory Manipulation Trace QA (EMT-QA): given synchronized video and proprioception from an exploratory trace, predict the minimal-success action chain under the latent precondition revealed by the probe. However, even state-of-the-art VLMs and embodied multimodal LLMs misread this evidence: they do not reliably recover the chain from raw video, raw proprioception...

论文介绍 探索性操作中,失败尝试可揭示潜在前提条件,对最小成功动作链的预测至关重要。本文形式化了探索性操作轨迹问答任务,指出视觉语言模型在解读原始视频和本体感知时存在误读问题,并提出闭环蒸馏方法来改善解读启发式,提升机器人从失败中学习的能力。

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

第一作者: Yuan Zhang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments. We argue that this stems from the lack of a unified geometry-aware manipulation representation, leaving existing VLAs vulnerable to low-level trajectory supervision, misaligned 3D features, and embodiment differences. To address this, we propose GEAR-VLA, a VLA framework for learning unified geometry-aware action representations for generalizable robotic manipulation. GEAR-VLA adopts coarse-to-fine action learning, where multi-source embodied pretraining equips the VLM with embodied reasoning and discrete action understanding before latent action tokens connect action semantics to a gradient-decoupled DiT continuous action expert. It further performs semantic-aligned 3D integration by...

论文介绍 视觉语言动作模型在基准任务上表现良好,但在真实世界部署中因物体、背景和机器人形态差异而泛化能力不足。GEAR-VLA通过学习统一的几何感知动作表示,采用粗到细动作学习和语义对齐的3D集成,旨在提升模型在不同环境和任务中的适应性。

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data

第一作者: Linqi Yin · 方向: 导航与运动 · 来源: cs.RO

Abstract:Vision-language models (VLMs) are powerful general-purpose reasoners, yet converting them into robot control policies (VLAs) is surprisingly difficult. The root cause is a two-fold gap: VLMs are trained on internet-scale images with language-understanding objectives, while VLAs must perceive robot scenes and predict motor actions. Fine-tuning a VLM directly on robot action data forces the model to cross both gaps at once -- the learning curve is steep and the rich generalizations learned during pretraining tend to degrade rather than transfer. We argue that this gap can be bridged gradually with the right intermediate data. We introduce \emph{embodied trajectory-coupled (ETC) data} -- vision-language supervision derived from the same robot scenes and trajectories used for action learning. Because ETC data shares the visual context of robot operation while retaining familiar...

论文介绍 将视觉语言模型转换为机器人控制策略时,存在感知与动作预测的双重差距。本文引入具身轨迹耦合数据作为中间桥梁,通过共享视觉上下文并保留语言理解,逐步弥合这些差距,促进预训练知识的迁移与保留。

ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies

第一作者: Bingjia Huang · 方向: 具身智能 · 来源: cs.RO

Abstract:Generative robot policies fail unpredictably at deployment: they hesitate at critical moments, drift off-task, or commit to unrecoverable actions. Existing online failure detectors either require white-box access to policy internals or add runtime overhead through resampling and observation-side signals. Our empirical analysis shows that emitted action chunks themselves already carry strong predictive signal for impending failures in generative robot policies. Motivated by this observation, we introduce ActProbe, a lightweight, pure action-space detector that uses two compact signals available from a single forward pass: Temporal Consistency Error (TCE) between consecutive action chunks and Action Chunk Magnitude (ACM) of the current chunk. ActProbe maps these signals to per-step failure probabilities with a task-conditioned LSTM-MLP architecture. Across a diverse suite of...

论文介绍 生成式机器人策略在部署中易出现不可预测的失败,如犹豫或偏离任务。ActProbe是一种轻量级检测器,仅基于动作块的时间一致性误差和幅度信号,通过LSTM-MLP架构映射到失败概率,实现早期检测,无需白盒访问或额外运行时开销。

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

第一作者: Haoyang Ge · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic interfaces, but neither offers a scalable and interactive prior for broad full-body behavior. We introduce EgoPriMo (Egocentric Motion Prior for Humanoid Robots), a unified framework that learns such priors from egocentric human demonstrations. Given egocentric observations and a text prompt, EgoPriMo reconstructs, generates, and forecasts SMPL-based full-body motion. Language is used as a high-level control signal rather than a complete motion specification. At the core of EgoPriMo is a Triple-stream DiT that jointly models body dynamics, egocentric visual context, and text; task-conditioning masks route different tasks and missing-modality data through the...

论文介绍 人形机器人的全身运动需要适应场景、任务和用户意图。EgoPriMo框架从自我中心人类演示中学习运动先验,给定观察和文本提示,重建、生成和预测基于SMPL的全身运动,使用三流DiT模型联合建模身体动力学、视觉上下文和文本。

Personalized and Robust Proactive Robot Assistance with Uncertainty-Guided LLM Reasoning

第一作者: Alvaro Gonzalez · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Proactive robot assistance in household environments requires accurate prediction of human activities and object usage under dynamic and noisy conditions. Existing approaches often rely on complex spatio-temporal models, which can be computationally expensive and sensitive to environmental variability. In this paper, we propose GLOBE, a lightweight framework that combines n-gram Markov models for capturing temporal behavioral patterns with uncertainty-guided large language model (LLM) reasoning. The framework performs sequential prediction efficiently while selectively invoking LLM reasoning only when the model confidence is low. To evaluate performance under realistic conditions, we introduce HOMER-Noise, a noisy extension of the HOMER+ dataset that simulates structured disturbances such as object movements caused by humans, pets, and toddlers. Experimental results show that...

论文介绍 家庭环境中的主动机器人辅助需在动态噪声条件下预测人类活动,现有方法计算昂贵且敏感。GLOBE框架结合轻量级n-gram模型捕获行为模式,并在置信度低时调用不确定性引导的大语言模型推理,实现高效顺序预测,并引入噪声数据集进行评估。

GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors

第一作者: Dongli Wu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robotic grasping is a fundamental capability in robotic manipulation. Yet grasping remains challenging under partial observations. Reliable grasping depends on both local contact cues and object-level 3D structure. Existing geometry-aware grasping methods recognize the value of reconstruction, but they typically treat geometry as an intermediate prediction rather than a reusable object prior for grasping. In this paper, we present GraspFoM, a unified framework that leverages 3D foundation priors (SAM3D) to build a shared 3D object latent for both reconstruction and grasp pose prediction. Built on this shared object latent, we introduce an anchor-initialized truncated pose-reasoning diffuser that predicts continuous and multimodal grasp poses without directly relying on discrete grasp candidates. We further investigate the interaction between reconstruction and grasping through...

论文介绍 机器人抓取在部分观察下依赖局部接触和物体3D结构,现有方法将几何作为中间预测而非可重用先验。GraspFoM利用3D基础先验构建共享对象潜在表示,用于重建和抓取姿态预测,通过锚初始化的截断姿态推理扩散器生成连续多模态抓取姿态。

PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation

第一作者: Lingxuan Wu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Diffusion policies have achieved remarkable success in robotic manipulation, yet they often fail to satisfy strict physical constraints required for safe deployment. Existing approaches impose safety either prematurely during training or reactively via external guardrails at test time, limiting policy expressivity and overall scalability. We propose Physical safety Alignment for Constrained Trajectories (PACT), a self-evolving post-training framework that projects pretrained diffusion policies onto constraint-feasible regions without accessing demonstration data or task rewards. PACT distills constraint gradients into the diffusion model through a reverse-KL objective with dense supervision across timesteps. It incorporates a curriculum that progressively tightens constraints while maintaining theoretically bounded policy shift and monotone improvement, mitigating the...

论文介绍 针对扩散策略在机器人操作中可能违反物理安全约束的问题,本文提出了PACT后训练框架。该框架无需示范数据或任务奖励,通过反向KL目标蒸馏约束梯度,将预训练策略投影到约束可行区域,并采用渐进式约束课程以保持策略改进的理论保证。

Uncertainty-Aware Intention Prediction for Human-to-Robot Assembly Teleoperation

第一作者: Fnu Heman · 方向: 机器人操作 · 来源: cs.RO

Abstract:In assisted teleoperation for human-robot collaboration, accurate intention prediction is critical for enabling timely and reliable robotic assistance during long-horizon manipulation and assembly tasks. These systems require continuous understanding of user behavior to recognize actions, anticipate intentions, and detect mistakes in real time. However, robot teleoperation demonstrations are costly and hardware-limited, whereas human demonstrations are easier to collect and provide rich temporal structure. To address this challenge, we propose an uncertainty-aware human-to-robot intention prediction framework that combines: (1) hierarchical transfer learning, where MS-TCN++ is pretrained on human hand demonstrations and fine-tuned on limited robot teleoperation data to capture low-level actions and high-level task intentions; (2) a conformal prediction module that provides...

论文介绍 为了在人-机器人协作遥操作中准确预测用户意图,本文提出了一种不确定性感知的意图预测框架。该框架结合了层次迁移学习(利用人类手部演示进行预训练)与保形预测模块,能够在有限机器人数据下识别动作、预测意图并提供预测的置信区间,以增强系统的可靠性与安全性。

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

第一作者: Shanglin Yuan · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiotemporal evidence is not necessarily better: when the injected evidence is not motion-consistent, it can introduce geometric drift, fragmented temporal cues, and unstable action generation. This raises a simple question: should a VLA remember past frames, or remember the motion that connects them? We introduce MotionVLA, a motion-history interface that converts a short past-only video window into compact, time-continuous trajectory-field tokens. Instead of treating history as a sparse set of ndependently lifted frames, MotionVLA represents recent observations as physically coherent motion evidence. Current visual tokens query this history to retrieve task-relevant motion information, which is then...

论文介绍 针对VLA模型注入时空证据不一致可能导致几何漂移和动作不稳定的问题,本文提出了MotionVLA。它将短期历史视频窗口转换为紧凑、时间连续的轨迹场标记,从而将物理一致的运动证据注入模型,旨在改善长时程操作中的歧义性和稳定性。

Impedance MPC for Physical Human-Robot Interaction: Predictive Disturbance Rejection with Joint-Limit Safety

第一作者: Yongyan Cao · 方向: 导航与运动 · 来源: cs.RO

Abstract:Physical human-robot interaction (pHRI) demands simultaneous trajectory accuracy and compliant safety under unplanned contact. Classical impedance control incurs a nonzero steady-state position error under sustained human force -- the applied force divided by the task stiffness -- which integral action reduces only within a narrow stable-gain budget. We present a two-layer Impedance MPC that resolves this tension. Layer~1 analytically cancels gravity, Coriolis, and task-space inertia, reducing the residual plant to a configuration-independent double integrator with a constant state-transition matrix. Layer~2 solves a 30-variable convex QP at 100\,Hz, exploiting this constant structure so the free-response matrix is precomputed once; an augmented Kalman filter estimates the persistent disturbance state, giving a formal zero-steady-state-error guarantee. A null-space...

论文介绍 为解决经典阻抗控制在持续外力下存在稳态位置误差的问题,本文提出了一种两层阻抗MPC方案。第一层解析消除机器人动力学,将其简化为双积分器;第二层求解凸优化问题,利用增广卡尔曼滤波估计恒定干扰,从而在保证关节限制安全的同时,实现零稳态误差的精确力控与柔顺交互。

Mind Your Steps: A General Learning Framework for Accurate Humanoid Foothold Tracking

第一作者: Alessandro Montenegro · 方向: 机器人操作 · 来源: cs.RO

Abstract:Enabling humanoid robots to operate in complex, dynamic environments remains a critical challenge, fundamentally limited by the ability to navigate robustly, safely, and accurately. While reinforcement learning with velocity-commanded policies has achieved remarkable robustness in humanoid locomotion, this approach lacks explicit control of the foothold placement, leading to unsafe behavior, such as stepping onto human feet, or imprecise navigation, hindering the following manipulation task. Conversely, explicit foothold-tracking policies offer a promising alternative by directly being commanded with target foot poses. However, existing approaches are often limited by unrealistic state assumptions, compromising real-world deployment, or they are part of staged pipelines, making them tied to specific downstream tasks. In this work, we introduce a novel, lightweight framework...

论文介绍 针对人形机器人缺乏精确落脚点控制导致安全与导航精度问题,本文提出一个轻量级学习框架。该框架训练机器人直接跟踪目标脚部位姿,旨在实现更精确、可预测的步行,以提升其在复杂动态环境中进行安全导航和后续操作的能力。

Disturbance-Aware Aerial Robotics for Ethical Wildlife Monitoring

第一作者: Mahmut Osmanovic · 方向: 具身智能 · 来源: cs.RO

Abstract:Reliable wildlife monitoring is essential for ecology and conservation, yet many existing methods, such as tagging, capture, and close-range observation, can alter the very behaviors they aim to measure. Aerial robots offer a scalable alternative, which has shown promising performance in multiple studies. Nonetheless, existing approaches typically lack behavioral awareness, rely on fixed heuristics, or require real-world training data that are costly, impractical, and ethically difficult to obtain. As a result, there remains no general framework for adaptive drone-based monitoring that can both preserve ecological validity and scale across species, behaviors, and robotic platforms. In this study, we introduce a disturbance-aware reinforcement-learning-based framework for heterogeneous aerial robotic fleets that enables autonomous wildlife tracking while explicitly minimizing...

论文介绍 为减少无人机监测对野生动物行为的干扰,本文提出一种基于强化学习的扰动感知框架。该框架使异构无人机群能够自主跟踪目标,同时通过学习扰动成本函数来显式最小化对动物的干扰,旨在开发一种可扩展、生态友好的自适应监测方法。

Agentic Neuro-Symbolic Planning and Commissioning for Human-in-the-Loop Industrial Robotics with Digital Twins

第一作者: Zhihao Liu · 方向: 具身智能 · 来源: cs.RO

Abstract:Flexible robotic automation requires systems that interpret operator intent, verify physical feasibility, and recover from execution failures across both the planning and execution stages. This paper proposes an agentic neuro-symbolic framework for human-in-the-loop industrial robotics, in which LLMs are used for tasks that require language understanding or contextual reasoning, while all verification, sequencing, and execution remain deterministic. The framework adapts the Planner-Generator-Evaluator (PGE) harness pattern from software engineering into a Specifier-Designer-Inspector (SDI) architecture for industrial robotics, combined with LangGraph-based dynamic routing for failure recovery. A two-tier recovery mechanism addresses structure-level replanning through context-aware orchestration and execution-level geometric failures through deterministic recovery skills. A...

论文介绍 为实现灵活的工业机器人自动化,本文提出一个代理神经符号框架。该框架结合大语言模型进行上下文推理,并采用确定性结构进行验证和执行。通过数字孪生进行可行性检查,并设计分层失败恢复机制,以支持人类操作员的实时意图交互与任务编程。

Propeller-Assisted Robust 3D Hopping Robot with Hierarchical Force Allocation

第一作者: Chuhan Zhang · 方向: 具身智能 · 来源: cs.RO

Abstract:Monopedal hopping robots are conceptually simple but highly dynamic and inherently unstable. Achieving robust 3D hopping is still difficult because ground reaction forces are available only during the short stance phase, while the robot is underactuated in flight. A key unresolved issue is how to improve flight-phase control authority. Propeller assistance provides a promising solution, but it requires careful coordination of leg-generated contact forces and propeller thrusts across stance and flight. This paper presents Pro-OMEGA2, a propeller-assisted 3D monopedal hopping robot with an active 3-RSR parallel leg and a trunk-mounted tri-rotor for auxiliary attitude regulation. To address the force coordination challenge, we propose a Hierarchical Force Allocation (HFA) framework based on a single rigid body (SRB) model. The leg generates the main stance contact wrench, while...

论文介绍 为增强单足跳跃机器人在飞行阶段的控制能力并实现稳健的3D跳跃,本文提出Pro-OMEGA2机器人设计及其分层力分配框架。该框架基于单刚体模型,协调腿部接触力与螺旋桨推力,以优化起跳和姿态控制,旨在提升跳跃运动的鲁棒性。

SynthICL: Scalable In-context Imitation Learning with Synthetic Data

第一作者: Cheng Qian · 方向: 模仿学习 · 来源: cs.RO

Abstract:In-context imitation learning (ICIL) enables robots to learn new tasks from a small number of demonstrations by conditioning a pre-trained policy on task-specific examples, without retraining at test time. Despite this promise, training generalizable and scalable in-context imitation policies remains an open challenge. We present SynthICL, a scalable framework that trains ICIL policies entirely from RGB-only synthetic data. Specifically, we build a data generation pipeline to produce high-fidelity ICIL data and train a flow-matching transformer policy on the resulting dataset. SynthICL avoids the need for depth sensing, precise camera calibration, and real-world training data in prior approaches, offering a simpler and more scalable alternative. We further incorporate subgoal prediction by training the model to predict the next subgoal images, enabling more precise and...

论文介绍 上下文模仿学习(ICIL)允许机器人从少量示例中学习新任务,但训练泛化和可扩展的策略是挑战。SynthICL框架通过构建高保真合成数据生成管道,完全使用RGB数据训练基于流匹配变换器的策略,并加入子目标预测以提高精度。该方法避免了对深度传感、精确相机校准和真实训练数据的依赖,为机器人学习提供了更简单、可扩展的替代方案。

Vision-Guided Dual-Arm Humanoid Robotic Disassembly of End-of-Life 18650 Lithium-ion Battery Packs

第一作者: Yile Chen · 方向: 机器人操作 · 来源: cs.RO

Abstract:The growing volume of retired lithium-ion battery packs from electric vehicles and portable electronics calls for automated disassembly that is safe, flexible, and selective down to the individual cell. Existing robotic systems, however, mostly assume known pack poses, external fixtures, or specialised tooling, leaving fixture-free cell-level disassembly under pose uncertainty largely unsolved. This paper presents a vision-guided dual-arm pipeline that disassembles a 21-cell 18650 pack from an arbitrary initial pose using only general-purpose parallel-jaw grippers, RGB-D sensing, and a pre-trained grasp detector. Pose uncertainty is absorbed by a learn-and-filter perception stack with discrete look-and-move wrist-camera corrections, while a mid-task support transfer between the two arms extends the effective workspace without any external clamp. The pipeline achieves an 8/10...

论文介绍 退役锂离子电池包的自动化拆卸需要安全、灵活且选择性的方法,但面临姿态不确定性挑战。本文提出视觉引导的双臂机器人流水线,使用通用并行夹爪、RGB-D传感和预训练抓取检测器,通过学习过滤感知堆栈吸收不确定性,并在任务中通过双臂支撑转移扩展工作空间,实现无夹具的单元级拆卸。

Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data

第一作者: Ji Woong Kim · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Robotics faces a fundamental challenge of data scarcity. Unlike language or vision research, there is no internet-scale dataset for robotic manipulation. A promising path forward is to leverage egocentric human data, which can be collected more easily, with greater breadth, and at a larger scale. Towards this end, we investigate key design choices for learning across human and humanoid embodiments equipped with dexterous five-finger hands, using the $\pi_{0.5}$ model as a foundation. Our results show that human data enables robots to learn new task semantics and compose existing skills into novel behaviors without corresponding robot data. The paper website is here: this https URL

论文介绍 机器人操作面临数据稀缺问题。Ego-Pi研究利用自我中心人类数据,通过视觉语言动作模型(如π0.5)进行微调,学习跨人类和类人机器人实现的策略。结果表明,人类数据使机器人能学习新任务语义并组合现有技能,减少对机器人特定数据的依赖。

Reinforcement learning in linear embedding space unlocks generalizable control across soft robot configurations

第一作者: Xinglong Zhang · 方向: 策略学习 · 来源: cs.RO

Abstract:Soft-bodied organisms such as octopuses and elephant trunks exhibit remarkable morphological adaptability, dynamically reconfiguring body shape and stiffness, and flexibly adjusting their control strategies to enable versatile behaviors. Inspired by these biological systems, various soft robots have emerged in recent decades, featuring diverse materials, stiffnesses, and morphologies tailored to specific tasks. Despite substantial advances in the materials and structural designs of soft robots, developing a generalizable control framework capable of rapid adaptation across diverse configurations remains a long-standing challenge. Existing controllers are limited to fixed configurations, demanding laborious configuration-specific remodelling and policy redesign for new configurations. Here, we introduce a generalizable control system that enables rapid adaptation across diverse...

论文介绍 软机器人具有多样形态和材料,但缺乏跨配置的可泛化控制框架。本文引入可泛化控制系统,使用强化学习在线性嵌入空间中训练策略,实现快速适应不同软机器人配置。该方法模仿生物系统的形态适应性,无需为每个配置重新建模,提供跨配置的通用控制。

Revisiting Articulated Parts Perception in Robot Manipulation

第一作者: Xiaoqian Wu · 方向: 机器人操作 · 来源: cs.RO

Abstract:We are surrounded by various objects with movable, articulated parts, e.g., box, handle, door. An accurate and generalizable perception of articulated parts is essential to enhance robotic manipulation capabilities. Building on this need, recent efforts in articulated parts perception have followed two main directions: One line of work uses pose-based representation, which requires high manual cost; in parallel, affordance-based methods extract future object motion from point tracking without additional manual efforts, but suffer from low-quality data. In this paper, we propose a new representation of articulated parts, Geometric Primary Structure (GPS), an abstraction of the part geometry structure to balance scalability and quality. For efficient and scalable data collection, GPS is integrated with a portable Virtual Reality (VR) device and requires only one minute to...

论文介绍 关节部件感知对机器人操作至关重要,但现有方法要么手动成本高,要么数据质量低。本文提出几何主结构(GPS)表示,平衡可扩展性和质量,并通过便携VR设备进行高效数据收集,仅需一分钟,用于训练感知模型,提升机器人操作能力。

Continual Quadruped Robots Coordination via Semantic Skill Discovery

第一作者: Daoqing Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Multi-quadruped coordination has attracted increasing attention due to its enhanced payload capacity, broader contact coverage, and improved adaptability to challenging tasks. Existing methods for multi-quadruped manipulation typically focus on predefined or closed task families, often relying on multi-agent reinforcement learning (MARL) to train task-specific coordination policies. However, such methods struggle in open-ended continual learning settings, where tasks arrive sequentially and robots are expected to acquire new coordination skills while reusing previously learned ones without catastrophic forgetting. To address this challenge, we propose Conquer, a semantic skill-library framework that formulates continual multi-quadruped coordination as a retrieve-adapt-update process. First, to accommodate varying team sizes across tasks, we design a team-structured...

论文介绍 多四足机器人在开放环境中的持续协调学习是挑战,需要处理序列任务和灾难性遗忘。Conquer框架将协调视为检索-适应-更新过程,使用语义技能库,适应不同团队规模和任务序列,实现技能复用和持续学习。

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

第一作者: Khanh D. Nguyen · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware on which robots actually run. We present this http URL, a portable C++ inference runtime built on this http URL. To our knowledge, it is the first ggml-class engine to natively serve the flow-matching and diffusion VLA inference pattern, in which a cached vision-language prefix is consumed by a cross-attending action expert integrated over several solver steps. A single runtime serves seven architectures spanning five backbone and four action-head families behind one request/response protocol, with each model packaged as a self-contained bundle. On LIBERO-Object, the engine matches a state-of-the-art checkpoint to within one episode out of 200, and runs BitVLA at 100% success in 1.3 GiB of memory. The same bundle runs unchanged...

论文介绍 视觉语言动作模型通常依赖Python和GPU,不匹配机器人硬件限制。vla.cpp是一个基于ggml的C++统一推理运行时,支持多种VLA架构,提供轻量级、可移植的解决方案,允许在资源受限设备上实现高效推理,适用于机器人嵌入式系统。

Perceptive Behavior Foundation Model: Adapting Human Motion Priors to Robot-Centric Terrain

第一作者: Zifan Wang · 方向: 具身智能 · 来源: cs.RO

Abstract:Humanoid behavior foundation models aim to acquire reusable whole-body control policies from broad human motion priors, enabling a single controller to produce diverse and expressive behaviors. However, existing motion-centric foundation policies largely assume that the reference motion is already physically compatible with the robot's surroundings. This assumption breaks when the demonstrator, operator, and robot inhabit different environments: a human motion may specify the intended behavior, but not the footholds, clearance, body height, or contact timing required by the robot's local terrain. We introduce \emph{Perceptive Behavior Foundation Model} (Perceptive BFM), a terrain-aware humanoid control framework that grounds human motion priors in robot-centric perception. The model preserves raw kinematic motion references as the behavioral interface, while using local...

论文介绍 人形机器人行为基础模型旨在从人类运动先验中学习控制策略,但运动可能不兼容机器人地形。感知行为基础模型(Perceptive BFM)将运动先验与机器人中心感知结合,使用局部地形信息调整行为,实现地形感知的多样且表达性控制。

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

第一作者: Yichen Niu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets. We present EgoAERO, the first framework that learns dexterous manipulation from a single egocentric RGB-D human demonstration without object assets. EgoAERO reconstructs contact-consistent hand-object trajectories through asset-free object tracking and reconstruction, ego motion compensation, and adaptive contact optimization, then converts them into robot policies using two-stage residual learning. We further introduce an online quality assessment mechanism and construct EgoDex-R, a large-scale egocentric dataset with 4.3M RGB-D frames for dexterous policy learning. Simulation and real-world experiments show that...

论文介绍 研究从单次自我中心RGB-D视频学习灵巧操作,无需预扫描对象资产。提出EgoAERO框架,通过无资产对象跟踪和重建、自我运动补偿及自适应接触优化,重建接触一致的手-对象轨迹,并利用两阶段残差学习转换为机器人策略。引入在线质量评估机制并构建大规模数据集,适用于机器人操作学习。

MuJoCo-Drones-Gym: A GPU-Accelerated Multi-Drone Simulator for Control and Reinforcement Learning

第一作者: Manan Tayal · 方向: 策略学习 · 来源: cs.RO

Abstract:Robotic simulators are a cornerstone of modern research in aerial robotics, serving both as a vehicle for the development of new control algorithms and as the data source for training reinforcement learning (RL) policies. Yet, existing quadcopter learning environments often face a trade-off between physical fidelity, multi-agent support, and the throughput required by modern deep RL pipelines. In this paper, we present MuJoCo-Drones-Gym, an open-source Gymnasium-compatible multi-drone environment built on top of the MuJoCo physics engine. MuJoCo-Drones-Gym supports an arbitrary number of Bitcraze Crazyflie 2.x nano-quadcopters and exposes a modular API for selecting (i)~the physics model (rigid-body MuJoCo, explicit Python dynamics, or any subset of ground effect, blade drag, and inter-drone downwash), (ii)~the action interface (per-motor RPMs, collective normalized thrust...

论文介绍 针对现有四旋翼学习环境在物理保真度、多代理支持和吞吐量间的折衷,提出MuJoCo-Drones-Gym,一个基于MuJoCo的GPU加速开源多无人机环境。支持任意数量的Crazyflie无人机,提供模块化API用于选择物理模型和动作接口,适用于空中机器人控制和强化学习研究。

IntentNav: Learning Spatial-Visual Object Navigation from Human Demonstrations

第一作者: Yuxin Cai · 方向: 导航与运动 · 来源: cs.RO

Abstract:Object navigation requires a robot to search for an unobserved target in an unknown environment by deciding where to explore next under partial observability. Effective search resembles human-like exploration: selectively probing visually promising frontiers while relying on spatial memory to avoid redundant revisits. We propose IntentNav, a spatial-visual imitation framework that learns human-like ObjectNav policies from human demonstrations. To infer high-level search intent from low-level human actions, we introduce Frontier-based Human-Intent Labeling, which looks ahead in human demonstrations and labels the frontier that best explains the demonstrator's future search direction. We construct a spatial-visual candidate space, where BEV memory tracks explored regions, unexplored frontiers, and trajectory history, while egocentric visual memory provides semantic cues for each...

论文介绍 对象导航要求机器人在部分可观测下搜索未知目标。提出IntentNav框架,通过前沿人类意图标签从人类演示中学习搜索意图,构建空间-视觉候选空间,结合鸟瞰视图记忆和自我中心视觉记忆,学习人类般的导航策略,提升探索效率。

X-OP: Cross-Morphology Whole-Body Teleoperation via MPC Retargeting

第一作者: Jen-Wei Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Whole-body teleoperation is essential for scalable robot data collection in loco-manipulation tasks, yet existing approaches relying on exoskeleton suits or multi-camera setups impose prohibitive cost, complexity, and environmental constraints. Recent methods using a single extended reality (XR) device with end-to-end reinforcement learning policies partially address these limitations but require robot-specific retraining, suffer from out-of-distribution failures, and rely on motion retargeting that neglects dynamic feasibility. We propose a hierarchical whole-body teleoperation framework driven by a single XR device that generalizes across diverse robot morphologies without retraining robot-specific policies. A Model Predictive Control (MPC)-based motion retargeter jointly optimizes alignment with the operator's intent and the robot's dynamic feasibility, generating optimal...

论文介绍 现有全身遥操作方法成本高且复杂。提出X-OP框架,使用单个XR设备驱动层次化遥操作,通过基于模型预测控制的运动重定向器,跨不同机器人形态泛化,无需重新训练特定策略,优化操作意图和机器人动态可行性,适用于loco-manipulation任务。

MinNav: Minimalist Navigation Using Optical Flow For Active Tiny Aerial Robots

第一作者: Aniket Patil · 方向: 导航与运动 · 来源: cs.RO

Abstract:Navigation using a monocular camera is pivotal for autonomous operation on tiny aerial robots due to their perfect balance of versatility, cost and accuracy. In this paper, we introduce MinNav, a navigation stack based on optical flow and its uncertainty to fly through a scene with static and dynamic obstacles and unknown-shaped gaps without any prior knowledge of the scene components and/or their locations/ordering. We further improve success rate by using the activeness of the robot to move around in an exploratory way to find obstacles and navigate. We successfully evaluate and demonstrate the proposed approach in many real-world experiments in various environments with static and dynamic obstacles and unknown-shaped gaps with an overall success rate of 70%. To the best of our knowledge, this is the first solution to tackle all the aforementioned navigation cases without...

论文介绍 针对微型空中机器人单目相机导航,提出MinNav导航堆栈,基于光流及其不确定性处理静态和动态障碍物及未知形状间隙,无需场景先验知识。利用机器人主动性进行探索性移动,在真实实验中实现70%成功率,适用于低成本自主无人机。

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

第一作者: Siyi Chen · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning, executing, monitoring, and recovering from failures. We address these demands with a closed agent loop in which a VLM orchestrates heterogeneous robot capabilities as interruptible tools. Unlike in virtual AI agents, the timing of decisions, actions and tool calls is important in a physical world that does not pause for reasoning. We refer to this setting as Physical Orchestration, and propose VoLoAgent, a VLM that plans, monitors, and recovers by treating a VLA/WAM as an interruptible tool it steers mid-rollout alongside vision models and action primitives. To evaluate these long-horizon capabilities, we introduce RoboVoLo, a high-fidelity benchmark for open-vocabulary long-horizon manipulation across common sense...

论文介绍 开放词汇长时域操作需要机器人推理灵活指令和复杂场景。提出VoLoAgent,作为视觉语言模型编排异构机器人能力,将其视为可中断工具,在物理世界中处理决策、行动和工具调用的时机,引入RoboVoLo基准评估长时域能力。

Safe-RULE: Safe Reinforcement UnLEarning

第一作者: Shixiong Jiang · 方向: 策略学习 · 来源: cs.RO

Abstract:Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such as robotics systems. However, its reliance on static datasets exposes offline Safe RL to data poisoning attacks, where adversaries inject malicious samples that compromise safety and induce unsafe policy behavior. In this work, we propose a new learning paradigm, named safe reinforcement unlearning (Safe-RULE), used as a defense framework to remove the influence of poisoned data without retraining from scratch or requiring access to the original training environment. We further extend reinforcement unlearning to offline Safe RL by explicitly accounting for both task performance and safety constraints during the unlearning process. Experiments across benchmark Safe RL tasks demonstrate that our approach effectively enhances...

论文介绍 离线安全强化学习依赖静态数据集,易受数据中毒攻击。提出Safe-RULE框架,用于安全强化学习的防御,通过强化遗忘移除中毒数据影响,无需重新训练或原始环境访问,在遗忘过程中显式考虑任务性能和安全约束,增强系统安全性。

Real-time body pose non-verbal communication with a consistency-based reliability measure

第一作者: Alina Marcu · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Body movement communicates intent at distances and in conditions where neither the face, nor speech can be captured. We study the recognition of communicative intent from 2D body pose alone. We argue that body motion is a reliable signal especially in scenarios that require real time low-cost on-device person-to-robot communication in long distance environments, such as rescue missions. However, existing resources do not isolate this signal. Affective corpora combine body, face, voice and text, while skeleton action-recognition benchmarks label the action performed rather than the message conveyed. We release a dataset of real frames of full-body pose covering ten communicative intents and we compare it against other real (IPC) and synthetic (MotionLCM, VEO3.1, Kimodo) ones that span a range of difficulty. We target systems that can run on a robot's limited onboard hardware...

论文介绍 身体运动在远距离或条件受限时可传递意图。研究从2D身体姿势单独识别交流意图,发布新数据集覆盖十种意图,比较真实和合成数据,并引入基于一致性的可靠性度量,适用于机器人实时低成本人机交互。

C$^3$ache: Accelerating World Action Models with Cross Inference Chunk Cache

第一作者: Weisen Zhao · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:World Action Models (WAMs) generalize better than standard Vision-Language-Action (VLA) policies to novel motions and environments, because a video-modeling objective lets them learn from abundant unlabeled video rather than scarce labeled robot demonstrations. This generalization is computationally expensive. To complete a task, a WAM runs over multiple inference chunks, and each chunk requires a costly denoising process. Existing acceleration methods reduce this cost by caching and reusing computation within a single chunk's denoising trajectory. Our empirical analysis reveals a substantial source of redundancy they overlook: redundancy across chunks. When a robot executes a smooth behavior, the residuals computed at a given denoising step are strongly correlated from one chunk to the next. We introduce C$^3$ache, a training-free method that caches and reuses these residuals...

论文介绍 该论文针对世界动作模型推理成本高的问题,提出了一种名为 C$^3$ache 的免训练加速方法。研究发现,在不同推理块之间存在显著的计算冗余。该方法的核心是缓存并复用相邻推理块在去噪步骤中计算出的残差,从而减少重复计算,提升模型在执行连续平滑动作序列时的推理效率。

Autonomous Aerial Manipulation via Contextual Contrastive Meta Reinforcement Learning

第一作者: Lixuan Jin · 方向: 机器人操作 · 来源: cs.RO

Abstract:Unmanned aerial vehicles (UAVs) are increasingly being deployed in logistics, service robotics, and other real-world applications, creating a growing demand for autonomous payload acquisition and delivery. Existing approaches typically assume pre-attached payloads or rely on specialized grippers, leaving versatile end-to-end aerial delivery largely unresolved, where different payloads induce highly variable flight dynamics, requiring a single policy to adapt online without manual calibration or explicit system identification. To this end, we study \textbf{A}utonomous \textbf{A}erial Manipulation via \textbf{Co}ntextual \textbf{Co}ntrastive Meta Reinforcement Learning (\textbf{\textit{Aco2}}), a fully autonomous aerial delivery setting in which a quadrotor equipped with a lightweight hook continuously picks up, transports, and delivers diverse handle-equipped objects between...

论文介绍 本文研究了无人机自主执行多样化物体空中运输任务的挑战。针对不同载荷导致飞行动力学多变的问题,论文提出了 Aco2 方法。该方法基于上下文对比元强化学习,使四旋翼无人机能够在单一策略下,在线自适应地完成从拾取、运输到交付的全流程操作,无需人工校准或显式系统辨识。

TBD-VLA: Temporal Block Diffusion Vision Language Action Model

第一作者: Sung-Wook Lee · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Discrete Vision-Language-Action (VLA) models typically formulate action generation as next-token prediction over discretized action spaces, conditioning each token autoregressively on prior context. While effective, this paradigm incurs high inference latency and largely ignores the temporal structure inherent in action trajectories. Recent efforts introduce parallel decoding to improve efficiency, enabling faster inference, but lack explicit mechanisms for modeling token dependencies. We introduce TBD-VLA, a discrete token-based VLA framework that incorporates block diffusion to enable temporal action generation. We partition action sequences into temporal blocks and perform masked discrete diffusion within each block, while maintaining autoregressive generation across blocks. This design unifies temporal autoregression and parallel action decoding, achieving both strong...

论文介绍 论文提出了 TBD-VLA 框架,旨在改进离散视觉语言动作模型的效率与动作生成质量。该框架将动作序列划分为时间块,在块内应用掩码离散扩散进行并行动作解码,而在块间保持自回归生成。这种设计统一了时间自回归与并行解码,旨在降低推理延迟并更好地建模动作轨迹的时序依赖关系。

Astro, I'm Home! Investigating Factors that Influence the Acceptance of Home Robots Using Supervised Machine Learning

第一作者: Katrin Fischer · 方向: 具身智能 · 来源: cs.RO

Abstract:The use of social robots in home environments is on the rise. This exploratory study applies regularization techniques (e.g., Lasso and Ridge regression) to investigate variables and identify new models of technology acceptance in the context of social robots. Within the original UTAUT2 framework, performance expectancy, social influence, and hedonic motivation emerged as the strongest and most consistent predictors of intention to use the technology. In addition, usability, trust, and competence were identified as promising variables in a model predicting intention to use.

论文介绍 这项探索性研究运用正则化机器学习技术,调查了影响用户接受家用社交机器人的关键因素。研究在 UTAUT2 理论框架基础上,识别出性能期望、社会影响和享乐动机是最强且最一致的预测变量。此外,可用性、信任和能力也被发现为预测技术使用意向的有前景的变量。

Scaling by Diversified Experience for Vision-Language-Action Models

第一作者: Leiyu Wang · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level control, and the instability of policy optimization. In this paper, we introduce SyVLA, a robust VLA model trained with diversified experiences. We propose an Intention Decoupling algorithm to isolate control-relevant features from reasoning contexts and a similar-sample guided RL pipeline to stabilize policy updates and mitigate distribution shift. Extensive experiments on real-world robotic tasks and multi-modal benchmarks demonstrate that SyVLA achieves superior task success rates and stronger out-of-distribution generalization compared to existing methods, while effectively preserving core vision-language capabilities. Codes and Datasets is released on \href{this https URL}{project page}.

论文介绍 本文针对视觉语言动作模型在现实部署中的挑战,提出了 SyVLA 模型。其核心方法包括:使用「意图解耦」算法将控制相关特征与推理上下文分离,并通过类似样本引导的强化学习管道来稳定策略更新、缓解分布偏移。实验表明,该方法在提升任务成功率和分布外泛化能力的同时,能保持视觉语言基础能力。

Failure-Aware Refinement of Vision-Language Model for Lithography Defect Detection

第一作者: Pangyun Jeong · 方向: 多模态具身 · 来源: cs.CV

Abstract:Semiconductor lithography inspection requires reliable detection of small pattern defects such as bridge, burr, pinch, and contamination. In this study, we propose a two-stage vision-language framework that combines initial defect detection with prediction refinement. In the first stage, Qwen3-VL is fine-tuned with LoRA as a vision-language adapter to predict defect counts, defect categories, and normalized bounding boxes from lithography images. However, direct fine-tuning may still produce common test-time errors, including false positives, missed defects, and incorrect defect types. To address this limitation, the second stage trains a refinement module using first-stage prediction failures and their corrected labels, allowing the model to review and revise initial outputs. By learning from cases where the initial adapter fails, the refinement process improves defect...

论文介绍 该研究为半导体光刻缺陷检测提出了一个两阶段视觉语言框架。第一阶段微调视觉语言模型进行初步缺陷预测。第二阶段则引入一个修正模块,利用第一阶段的预测失败案例及其纠正标签进行训练,使模型能够复核并修正初始输出中的假阳性、漏检和类型错误,从而提升检测可靠性。

BLUE: Toward Better Language Use in Efficient Vision-Language-Action Models for Autonomous Driving

第一作者: George Ling · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:We present BLUE, a minimal method for better language use in vision-language-action (VLA) models for autonomous driving (AD). Through extensive analysis, we reveal that language matters on only a small fraction of routes, but on those routes it can greatly improve or degrade performance. Generating language at every frame is therefore inefficient, since most computation is spent on frames that do not benefit from language. We further show that pretrained VLA hidden states potentially already encode whether language will benefit a given frame, even though scene complexity and kinematic features alone struggle to predict this. Based on this finding, BLUE trains a lightweight gate on frozen VLA hidden states to decide per frame whether to activate language generation or predict actions directly, without modifying the backbone or requiring additional human annotation. With just a...

论文介绍 本文提出了 BLUE 方法,旨在提升自动驾驶视觉语言动作模型的语言使用效率。研究发现语言仅在少数关键路线中显著影响性能。BLUE 基于冻结模型的隐藏状态训练一个轻量级门控机制,逐帧决定是否激活语言生成或直接预测动作,避免了在不受益的帧上进行冗余的语言计算,从而在不修改主干网络的情况下提升整体效率。

Vision Language Model Helps Private Information De-Identification in Vision Data

第一作者: Tiejin Chen · 方向: 数据集与评测 · 来源: cs.AI

Abstract:Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability. While various methods exist to enhance privacy in text-based applications, privacy risks associated with visual inputs remain largely overlooked such as Protected Health Information (PHI) in medical images. To tackle this problem, two key tasks: accurately localizing sensitive text and processing it to ensure privacy protection should be performed. To address this issue, we introduce VisShield (Vision Privacy Shield), an end-to-end framework designed to enhance the privacy awareness of VLMs. Our framework consists of two key components: a specialized instruction-tuning dataset OPTIC (Optical Privacy Text Instruction Collection) and a tailored training methodology. The dataset provides diverse privacy-oriented prompts that guide VLMs to perform targeted Optical Character...

论文介绍 本文关注视觉输入(如医疗图像)中的隐私风险,提出了 VisShield 框架以增强视觉语言模型的隐私意识。该框架包含两个关键组件:用于引导模型进行目标字符定位的专用指令微调数据集 OPTIC,以及对应的训练方法。其目标是使模型能够准确定位并处理图像中的敏感文本信息,实现隐私保护。

An Effective Router for Vision-Language Model Selection

第一作者: Can Wang · 方向: 多模态具身 · 来源: cs.AI

Abstract:Vision-language models (VLMs) with varying performance and resource requirements are widely deployed, making it difficult for users to select the most appropriate one among numerous VLM candidates. Existing work reveals the performance paradox phenomenon in language models and focuses on routing methods to solve it. However, developing a router for VLM selection is still a critical yet challenging problem, which primarily faces: 1) lack of specialized data, 2) ineffective feature representation, and 3) rigid model space and costly adaptation. In this paper, we construct a multimodal dataset for VLM selection, containing the outputs of seven mainstream VLMs on 32,626 unique image-text queries. We then propose ARMS, a router for VLM selection. ARMS enhances input signals with VLM profiles, employs a simple but effective architecture to improve representations of queries and VLM...

论文介绍 本文针对视觉语言模型多样且选择困难的问题,提出了一种有效的路由器ARMS。研究构建了一个多模态数据集,包含七个主流视觉语言模型在32,626个图像-文本查询上的输出。ARMS通过增强输入信号和改进特征表示,旨在帮助用户高效选择合适的模型,可能应用于多模态系统优化和资源管理。

Blockchain Infrastructure for Intelligent Cyber--Physical--Social Systems:Post-Quantum Security, Interoperability, and Trustworthy Data Economies in the Era of Embodied AI

第一作者: Song Guo · 方向: 具身智能 · 来源: cs.AI

Abstract:The deployment of embodied artificial intelligence via world-model-based robotics presents a transformative opportunity for blockchain infrastructure, establishing urgent demand for trustworthy data provenance, cross-organizational governance, and incentive-compatible sharing across decentralized ecosystems. Simultaneously, quantum computing advances recognized by the 2025 Nobel Prize in Physics and the Turing Award threaten the cryptographic primitives securing these data economies, creating an interdependent imperative: long-lived verification for embodied AI depends on crypto-agile architectures capable of withstanding quantum adversaries. This tutorial examines blockchain as the coordination layer bridging this dual transition, from financial substrate to foundational Cyber-Physical-Social Systems infrastructure that simultaneously secures against quantum cryptanalysis and...

论文介绍 本文探讨区块链基础设施在智能网络-物理-社会系统中的应用,聚焦于具身AI时代。具身AI部署要求可信的数据溯源、跨组织治理和激励兼容的共享,同时量子计算威胁现有加密。研究提出区块链作为协调层,提供后量子安全和互操作性,可能为构建可信数据经济奠定基础。

市场总览

市场技术面整体呈现分化态势。美股方面,标普500 ETF(SPY)和纳斯达克100 ETF(QQQ)近期回调,SPY的RSI 14为49,价格低于20日均线746.26但仍高于200日均线684.94,MACD柱状图转负但趋势bullish;QQQ的RSI 49.5,MACD死叉出现但多头排列,显示短期压力。加密货币市场受恐慌情绪主导,加密恐慌贪婪指数仅为9(极度恐慌),总市值下降至2.21万亿美元,BTC主导率55.9%,比特币RSI 24.1超卖,MACD死叉,价格较52周低点仅高4.6%,以太坊和Solana同样RSI低于30,空头排列明显。中概股如阿里巴巴(BABA)跌幅较大,RSI 35.8,趋势bearish,拼多多(PDD)接近52周低点。商品外汇中,黄金期货RSI 29.2超卖,原油期货RSI 41.2中性,美元指数DXY的RSI 64接近52周高,多头排列,显示避险需求。整体来看,风险偏好下降,技术信号指向调整压力。

今日关注

BTC-USD Bitcoin
偏下行

Bitcoin当前价格为61828.47美元,1日跌幅2%,5日跌幅3.09%。RSI 14指标读数为24.1,处于超卖状态。MACD值为-4114.8497,低于信号线-3197.2476,形成死叉。趋势为bearish,信号包括RSI超卖和空头排列,表明技术面偏下行。

^VIX VIX 恐慌指数
偏上行

VIX恐慌指数当前价格为19.87,1日涨幅5.02%,5日涨幅26%。RSI 14指标为57.5,处于正常范围。MACD值为-0.0005,高于信号线-0.5296,形成金叉。趋势为bullish,信号包括MACD金叉和多头排列,显示技术面偏上行。

QQQ Nasdaq 100 ETF
中性

Nasdaq 100 ETF当前价格为707.83,1日跌幅1.15%,5日跌幅5.14%。RSI 14指标为49.5,接近中性50。MACD值为12.7347,低于信号线18.2343,形成死叉。但趋势为bullish且有多头排列信号,指标多空交织,技术面呈中性。

全部资产

^VIX

VIX 恐慌指数

$19.87 +5.02%
5 日
+26.00%
距 52w 高
-43.7%
RSI(14)
57.5
趋势
多头
SMA 20 / 50 / 200
17.24 / 18.54 / 18.46
MACD / 信号
-0.001 / -0.530
MACD 金叉 (2 天前)多头排列

^TNX

10Y 美债收益率 (%)

$4.53 -0.53%
5 日
+1.64%
距 52w 高
-9.4%
RSI(14)
56.1
趋势
多头
SMA 20 / 50 / 200
4.52 / 4.41 / 4.21
MACD / 信号
0.029 / 0.034
多头排列

DX-Y.NYB

美元指数 DXY

$99.97 -0.08%
5 日
+0.76%
距 52w 高
-0.7%
RSI(14)
64.0
趋势
多头
SMA 20 / 50 / 200
99.23 / 98.91 / 98.62
MACD / 信号
0.311 / 0.211
接近 52 周高多头排列

SPY

S&P 500 ETF

$737.05 -0.29%
5 日
-2.96%
距 52w 高
-3.1%
RSI(14)
49.0
趋势
多头
SMA 20 / 50 / 200
746.26 / 717.45 / 684.94
MACD / 信号
7.001 / 10.475
多头排列

QQQ

Nasdaq 100 ETF

$707.83 -1.15%
5 日
-5.14%
距 52w 高
-5.5%
RSI(14)
49.5
趋势
多头
SMA 20 / 50 / 200
721.98 / 673.56 / 623.29
MACD / 信号
12.735 / 18.234
MACD 死叉 (3 天前)多头排列

AAPL

Apple

$290.55 -3.64%
5 日
-7.82%
距 52w 高
-8.5%
RSI(14)
42.6
趋势
多头
SMA 20 / 50 / 200
304.56 / 283.05 / 265.90
MACD / 信号
5.543 / 8.300
MACD 死叉 (4 天前)多头排列

MSFT

Microsoft

$403.41 -2.02%
5 日
-8.59%
距 52w 高
-27.4%
RSI(14)
41.8
趋势
空头
SMA 20 / 50 / 200
421.95 / 410.20 / 455.41
MACD / 信号
1.820 / 4.940
MACD 死叉 (2 天前)空头排列

NVDA

Nvidia

$208.19 -0.22%
5 日
-6.57%
距 52w 高
-12.0%
RSI(14)
46.4
趋势
多头
SMA 20 / 50 / 200
218.21 / 205.01 / 188.91
MACD / 信号
0.928 / 3.161
多头排列

GOOGL

Alphabet

$364.26 +0.26%
5 日
+0.67%
距 52w 高
-10.9%
RSI(14)
44.1
趋势
多头
SMA 20 / 50 / 200
382.28 / 357.95 / 305.68
MACD / 信号
-0.263 / 4.424
多头排列

TSLA

Tesla

$396.68 -3.00%
5 日
-6.39%
距 52w 高
-20.5%
RSI(14)
43.9
趋势
空头
SMA 20 / 50 / 200
422.49 / 396.72 / 414.95
MACD / 信号
1.147 / 6.223
空头排列

META

Meta

$584.59 -0.14%
5 日
-2.18%
距 52w 高
-26.6%
RSI(14)
39.3
趋势
空头
SMA 20 / 50 / 200
610.79 / 621.45 / 660.86
MACD / 信号
-7.315 / -4.619
MACD 死叉 (2 天前)空头排列
加密恐慌贪婪
9
极度恐慌
加密总市值
$2.21 T
-1.29% / 24h
BTC 主导率
55.9%
ETH 8.9%
24h 成交量
$90.1 B
活跃币 17,348

BTC-USD

Bitcoin

$61,828.47 -2.00%
5 日
-3.09%
距 52w 高
-51.0%
RSI(14)
24.1
趋势
空头
SMA 20 / 50 / 200
70,208.20 / 75,307.84 / 78,241.27
MACD / 信号
-4,114.850 / -3,197.248
RSI 超卖空头排列

ETH-USD

Ethereum

$1,642.13 -2.84%
5 日
-7.20%
距 52w 高
-66.9%
RSI(14)
25.9
趋势
空头
SMA 20 / 50 / 200
1,913.42 / 2,134.88 / 2,438.64
MACD / 信号
-145.272 / -119.133
RSI 超卖空头排列

SOL-USD

Solana

$65.17 -2.43%
5 日
-5.16%
距 52w 高
-74.3%
RSI(14)
26.9
趋势
空头
SMA 20 / 50 / 200
77.08 / 83.35 / 101.39
MACD / 信号
-5.750 / -4.335
RSI 超卖空头排列

BABA

阿里巴巴 (BABA)

$119.70 -0.31%
5 日
-8.50%
距 52w 高
-37.9%
RSI(14)
35.8
趋势
空头
SMA 20 / 50 / 200
129.85 / 130.94 / 149.73
MACD / 信号
-3.370 / -2.215
空头排列

PDD

拼多多 (PDD)

$81.93 -0.84%
5 日
-7.09%
距 52w 高
-41.2%
RSI(14)
32.5
趋势
空头
SMA 20 / 50 / 200
90.83 / 96.52 / 112.07
MACD / 信号
-4.099 / -3.359
接近 52 周低空头排列

JD

京东 (JD)

$28.73 +0.49%
5 日
-4.71%
距 52w 高
-22.1%
RSI(14)
40.9
趋势
空头
SMA 20 / 50 / 200
30.52 / 30.15 / 30.36
MACD / 信号
-0.461 / -0.209
空头排列

0700.HK

腾讯控股 (0700.HK)

HK$455.40 +2.02%
5 日
-5.44%
距 52w 高
-33.3%
RSI(14)
47.5
趋势
空头
SMA 20 / 50 / 200
450.13 / 474.68 / 570.85
MACD / 信号
-6.531 / -9.531
空头排列

GC=F

黄金期货

$4,230.70 -2.43%
5 日
-5.76%
距 52w 高
-24.3%
RSI(14)
29.2
趋势
中性
SMA 20 / 50 / 200
4,502.41 / 4,618.04 / 4,407.98
MACD / 信号
-84.797 / -63.265
RSI 超卖

CL=F

WTI 原油期货

$88.45 -3.12%
5 日
-5.66%
距 52w 高
-26.0%
RSI(14)
41.2
趋势
中性
SMA 20 / 50 / 200
96.08 / 97.58 / 73.06
MACD / 信号
-2.041 / -1.342

USDCNY=X

美元 / 人民币

¥6.77 +0.12%
5 日
+0.17%
距 52w 高
-6.1%
RSI(14)
40.3
趋势
空头
SMA 20 / 50 / 200
6.78 / 6.81 / 6.97
MACD / 信号
-0.014 / -0.015
MACD 金叉 (今天)接近 52 周低空头排列
风险提示

本报告基于公开行情数据计算的技术指标进行解读,过去走势不代表未来表现,仅供技术指标解读参考。技术分析存在局限性,投资者应结合基本面和市场环境综合判断。

Iran War Live Updates: U.S. Launches Strikes on Iran to Retaliate for Downed Helicopter

The president said earlier Iranian forces had shot down a U.S. Apache helicopter near the Strait of Hormuz and the United States “must, of necessity, respond to this attack.”

中文摘要 美国总统特朗普宣布对伊朗发动打击,以报复伊朗在霍尔木兹海峡附近击落一架美国阿帕奇直升机。特朗普此前称伊朗击落了该直升机,并强调美国「必须、有必要」对此攻击作出回应。

Here’s the latest.

中文摘要 此条目标题为「最新情况」,但未提供具体摘录内容,无足够信息生成事实摘要。

Australia news live: Albanese says ‘we need to do more’ on house prices; Australia joins sanctions against ‘extremist settlers’ in West Bank

Get our breaking news email, free app or daily news podcast Albanese says Australia still impacted by Middle East conflict ‘each and every day’ The prime minister, Anthony Albanese, is now on the ABC News Breakfast couch. He said Australia remains concerned about the economic impact of the turmoil i

中文摘要 澳大利亚总理阿尔巴尼斯表示需要采取更多措施应对高房价。同时,澳大利亚宣布对以色列约旦河西岸的「极端主义定居者」实施制裁,并指出中东冲突每天都在影响澳大利亚经济。

US strikes Iran in response to downing of helicopter, military says

President Donald Trump earlier accused Iran of shooting down the US helicopter over the Strait of Hormuz and vowed to respond.

中文摘要 美国军方确认对伊朗实施打击。此前,美国总统特朗普指责伊朗在霍尔木兹海峡上空击落美军一架直升机,并誓言将作出回应。

Prisoners in Western Australia are living in ‘cruel, inhuman or degrading’ conditions, report warns

Inspector of custodial services says inmates are sleeping on the floor and denied basic entitlements due to ‘a systemic failure across multiple prisons’ Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Inmates in Western Australia are

中文摘要 一份报告警告,西澳大利亚州监狱的囚犯生活在「残酷、不人道或有辱人格」的条件下,包括睡地板及被剥夺基本权利。监管机构指出,这是跨监狱的系统性失败所致。

Brazil intercepts 108 Cuban immigrants amid growing asylum applications

Last year, for the first time in a decade, Cuban asylum applications exceeded Venezuelan ones in sign of growing strain.

中文摘要 巴西当局拦截了108名试图入境的古巴移民,凸显该国庇护申请量增长。据悉,去年古巴庇护申请量十年来首次超过委内瑞拉。

Iran war live: US strikes after downing of helicopter, Tehran vows response

Lebanon's Health Ministry said the toll from Israeli attacks on the country since March 2 has reached 3,666 dead.

中文摘要 伊朗战争持续,美国在直升机被击落后对伊朗发动打击,德黑兰誓言报复。同时,黎巴嫩卫生部称,自3月2日以来以色列对该国的袭击已造成3666人死亡。

Afghan authorities open fire on protesters over women’s dress code

Authorities suppressed a protest in western Afghanistan over the detention of women for dress code violations.

中文摘要 阿富汗当局在西部省份镇压一场因妇女因着装规定被拘留而引发的抗议,并向示威者开枪。

Man Arrested in ‘Brutal’ Stabbing in Belfast, Police Say

The suspect, who the police said is 30 and Sudanese, was charged with having attacked another man in an assault in Belfast that was recorded and spread quickly online.

中文摘要 一名30岁的苏丹籍男子因在贝尔法斯特涉嫌持刀袭击他人被捕,并被指控。此起袭击被录影并在网上迅速传播。

US judge halts execution by nitrogen gas, ruling it unconstitutional

Judge Emily Marks had previously allowed the execution to proceed, arguing that no execution is entirely without pain.

中文摘要 一名美国联邦法官叫停了计划通过吸入氮气执行的死刑,裁定该方法违宪。该法官此前曾允许执行进行,但指出没有任何执行过程是完全无痛的。

Many Feared Trapped Under Earthquake Rubble in Philippines

Families hope to recover their loved ones after the most powerful earthquake in 50 years to hit the Philippines.

中文摘要 菲律宾遭遇50年来最强地震,多人据信被困在废墟之下。救援人员正努力搜寻幸存者,家属期盼寻获亲人。

What to Know About the Sea Drone That Rescued Downed Apache Crew

It was the first U.S. rescue carried out by an autonomous surface vessel and remotely piloted by a human operator, according to a military spokesperson.

中文摘要 美军称,在伊朗附近海域首次使用自主遥控海面无人机成功营救了被击落的阿帕奇直升机机组人员。这是美军首次由自主船只执行此类救援。

Air Canada Pilot Accused of Flying for 17 Years Without Proper License

The pilot, who retired last year before the investigation, held some valid flight credentials, officials said, but not the one required to be a captain.

中文摘要 加拿大航空一名飞行员被指控在未持有合格机长执照的情况下飞行长达17年。该飞行员已于去年退休,目前正接受调查。

Trump launches strikes against Iran after downing of US army helicopter

US president blames Tehran for loss of Apache gunship, whose crew were rescued by a drone near strait of Hormuz Middle East crisis – live updates The US has launched strikes against Iran after Donald Trump blamed Tehran for downing a US army helicopter near the strait of Hormuz, imperilling a shaky

中文摘要 在伊朗击落一架美军直升机后,美国总统特朗普对伊朗发动打击。特朗普指责德黑兰方面在霍尔木兹海峡附近击落美军直升机,其机组人员随后被一架无人机救出。

Drugmaker Parabilis Raises $745 Million in US IPO, Placement

Parabilis Medicines Inc., a clinical-stage cancer drug developer, raised nearly $745 million in an upsized US initial public offering that priced above its marketed range, and a private placement.

中文摘要 临床阶段癌症药物开发商Parabilis Medicines Inc.在美国首次公开募股中筹集近7.45亿美元,定价高于市场范围,同时进行了私人配售交易。

Malaysia Minister on Fuel Subsidy, Deficit Goals

Malaysia may not meet its fiscal deficit targets for 2026 as the Iran war drives up the cost of fuel subsidies, Second Finance Minister Amir Hamzah Azizan exclusively tells Bloomberg's Haslinda Amin at the Invest Malaysia event. (Source: Bloomberg)

中文摘要 马来西亚第二财政部长Amir Hamzah Azizan在Invest Malaysia活动上告诉彭博社,由于伊朗战争推高燃料补贴成本,该国可能无法实现2026年财政赤字目标。

Alberta to Propose ‘General Corridor’ for Pipeline to West Coast

Alberta is set to propose a “general corridor” for a planned new million-barrel-a-day oil pipeline to the northern British Columbia coast rather than a specific route, the provincial minister of Indigenous relations said.

中文摘要 阿尔伯塔省计划为一条每日百万桶石油的管道提议「通用走廊」,通往不列颠哥伦比亚省北部海岸,而非具体路线,由省原住民关系部长透露。

US Launches Iran Strikes After Helicopter Is Downed

The US military launched fresh airstrikes against Iran hours after President Donald Trump vowed to retaliate for the shooting down of a US Army helicopter. Bloomberg's Romy Varghese breaks down the situation. (Source: Bloomberg)

中文摘要 在美国陆军直升机被击落后,美国总统Donald Trump誓言报复,美军随即对伊朗发动新一轮空袭。

Apollo and Blackstone raise $35bn in chip financing deal for Anthropic

Transaction is one of the largest private credit fundraisings, fuelling Claude maker’s AI growth plans

中文摘要 Apollo和Blackstone为Anthropic筹集350亿美元芯片融资,这是最大私人信贷融资之一,用于支持Claude制造商的人工智能增长计划。

Gold Extends Drop as Renewed US-Iran Clashes Test Fragile Truce

Gold extended a decline after the US launched fresh strikes against Iran in retaliation for the downing of a military helicopter, jeopardizing efforts to end the war that’s roiled global markets.

中文摘要 在美国为报复军事直升机被击落而对伊朗发动新打击后,黄金价格延续跌势,危及结束战争的努力,该战争已扰乱全球市场。

Cuba Poised for Biggest US Fuel Shipment Since Cold War Embargo

A Florida trading company is in advanced talks to send Cuba the biggest cargo of US fuel since the Eisenhower administration as the island nation contends with an acute energy crisis.

中文摘要 一家佛罗里达贸易公司正深入谈判向古巴运送自艾森豪威尔政府以来最大的美国燃料货物,以缓解该国的严重能源危机。

Bill debt soars but many don't know help is available

The majority of billpayers are unaware of special tariffs for water and broadband, the spending watchdog says.

中文摘要 支出监管机构指出,尽管账单债务飙升,但大多数支付者不了解水和宽带的特殊资费可用。

Beauty Pie LED mask ad banned over misleading anti-wrinkle claim

The mask is not "clinically proven to reduce wrinkles in four weeks", the advertising watchdog finds.

中文摘要 广告监管机构裁定,Beauty Pie LED面膜广告因误导性抗皱声明被禁止,该产品未被临床证明可在四周内减少皱纹。

How to enjoy the World Cup - and keep your boss on side

Football fans and bosses share their strategies to balance late night kick offs with work the next day.

中文摘要 足球迷和老板分享策略,以平衡世界杯深夜比赛与次日工作,确保不影响职场关系。

ASML’s Surge to Record High Masks Lowball Valuation Versus Peers

ASML Holding NV’s ascent to fresh record highs has been a highlight for European markets, yet it’s done little to improve what is the stock’s cheapest relative valuation in years.

中文摘要 ASML Holding NV股价升至新高,成为欧洲市场亮点,但其相对估值仍为多年来最低,未能改善这一状况。

【CHY公益站】高考结束,祝各位佬友金榜题名!CHY公益站开放无注册码注册

本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已

Claude Fable 5 / Mythos 5 迄今为止最强大的模型,性能大幅跃升暴打GPT

Anthropic 发布 Claude Fable 5 与 Mythos 5,性能大幅跃升 Anthropic 推出面向普通用户的 Claude Fable 5,这是迄今能力最强的 Mythos 级模型。它在软件工程、知识工作、视觉和科研等基准上均达顶尖,价格比前代 Mythos Preview 低一半以上。为防滥用,内建分类器在涉及网络安全、生物化学等话题时改用 Opus 4.8 回复,约 95% 的会话不受影响。 同步发布的 Claude Mythos 5 对网络防御伙伴解除部分限制,号称拥有全球最强的网络安全能力。生物医学研究者也可通过信任计划在解除防护后使用。两款模型定价均为每百万输入

给新大学生的一封长信

全文手打!希望佬们看的爽 许可:CC BY-NC-ND 4.0 我想说 正直六月,今天是高考的第三天,也是高考最后的一天。想到四年前我也走进了考场,心里有一种冲动,想给从高中毕业进入大学的同学写一些东西。于是我便写下了这篇不算文章的文章。本篇作为礼物送给大家。愿大家在大学与生活中不再迷茫。 由于没有人指导,我的大学生活其实算得上比较失败了。 现在将时间拨回 2022 年。疫情几乎贯穿了我整个高中生涯。同时,新高考的政策也打了 22 届考生一个措手不及。新高考一卷让许多考生无法释怀。 不过这些对我来说并没有太大的影响。感谢我的高中母校,提供了海外本硕连读的机会,也感谢我的家庭的资金支持。 对我来

FluxDO 0.2.16 已发布

本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 时隔一月,总算攒了点内容,大家自行探索吧 GitHub Release v0.2.16 · Lingyan000/fluxdo [0.2.16] - 2026-06-09 🌟 新功能 俺也一样: OP 帖底部 ME

某公众号diss发文linux do

刚准备躺下打开手机就看到这个标题的推文,看不懂的操作,一边拿标题引流,内容全是吐槽diss 45 个帖子 - 40 位参与者 阅读完整话题

高考结束,解放!

三年的高中生活就这样结束了。想想这三年学习有多苦,每次在学校模拟考试的时候有多紧张,每次看到分数,高兴与难过。还有在学校,还跟同学的打打闹闹,甚至还有吵架等杂七杂八的经历。仔细想想,就这么结束了,也是挺… 34 个帖子 - 33 位参与者 阅读完整话题