每日简报

2026-06-06

← 历史归档

NousResearch/hermes-agent

Python · ★ 183,133 · 🍴 31,410 · 📈 1,845 stars today

The agent that grows with you

中文介绍 Hermes Agent 是一个与你共同成长的智能代理。它旨在通过持续交互和学习来适应用户需求,提供个性化的辅助能力。开发者或 AI 研究者可使用它来构建能不断进化、理解上下文并执行复杂任务的自适应系统。

chopratejas/headroom

Python · ★ 14,511 · 🍴 923 · 📈 2,473 stars today

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.

中文介绍 Headroom 是一个文本压缩工具,能在数据发送给大语言模型前,对工具输出、日志、文件及 RAG 数据块进行压缩,可减少 60%-95% 的 token 消耗,且不损失核心信息。它以库、代理及 MCP 服务器的形式提供,适用于需要优化 LLM 调用成本和效率的开发者。

CopilotKit/CopilotKit

TypeScript · ★ 32,684 · 🍴 4,194 · 📈 366 stars today

The Frontend Stack for Agents & Generative UI. React + Angular. Makers of the AG-UI Protocol

中文介绍 CopilotKit 是一个专为 AI 代理与生成式用户界面设计的前端开发栈,支持 React 和 Angular。它提供了构建交互式、由 AI 驱动应用的框架,并提出了 AG-UI 协议。前端开发者可利用它快速集成智能功能,打造下一代应用界面。

lfnovo/open-notebook

TypeScript · ★ 26,007 · 🍴 2,993 · 📈 1,152 stars today

An Open Source implementation of Notebook LM with more flexibility and features

中文介绍 Open Notebook 是 Google NotebookLM 的开源替代方案,提供了更灵活的特性和功能。它允许用户以结构化的方式组织笔记、文档和研究资料,并利用 AI 进行查询、总结和知识关联。适合学生、研究人员和知识工作者用于管理信息流。

affaan-m/ECC

JavaScript · ★ 208,363 · 🍴 31,966 · 📈 1,361 stars today

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

中文介绍 ECC 是一个面向 AI 代理(如 Claude Code、Cursor 等)的性能优化系统。它通过集成技能、本能、记忆和安全模块,并采用研究优先的开发方法,旨在提升代理的执行效率与可靠性。适用于 AI 开发者优化代理表现和构建高级功能。

Panniantong/Agent-Reach

Python · ★ 21,566 · 🍴 1,861 · 📈 148 stars today

Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.

中文介绍 Agent-Reach 为 AI 代理赋予了浏览互联网的能力,可通过单一命令行读取和搜索 Twitter、Reddit、YouTube、GitHub 等多个平台的内容,且无需 API 费用。它让代理能获取实时网络信息,适用于需要联网研究或内容分析的开发者。

NVIDIA/cosmos

Jupyter Notebook · ★ 9,417 · 🍴 602 · 📈 479 stars today

NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.

中文介绍 NVIDIA Cosmos 是一个开放平台,提供世界模型、数据集和工具,旨在帮助开发者构建面向机器人、自动驾驶车辆和智能基础设施的物理 AI。它为训练和理解物理世界的复杂模拟与交互提供了基础支持,是具身智能开发的重要资源。

666ghj/MiroFish

Python · ★ 64,701 · 🍴 10,088 · 📈 320 stars today

A Simple and Universal Swarm Intelligence Engine, Predicting Anything. 简洁通用的群体智能引擎,预测万物

中文介绍 MiroFish 是一个简洁通用的群体智能引擎,其核心目标是进行万物预测。它利用群体智能的方法论,对复杂现象或数据进行分析与预测。适合数据科学家、研究人员或商业分析师用于需要进行趋势分析和预测建模的场景。

mvanhorn/last30days-skill

Python · ★ 28,212 · 🍴 2,392 · 📈 731 stars today

AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary

中文介绍 这是一个 AI 代理技能,能够针对任意主题,对 Reddit、X、YouTube、Hacker News、Polymarket 及互联网进行综合研究,最终生成一个基于事实的摘要。它为研究人员或内容创作者提供了快速获取跨平台信息概览的自动化工具。

PaddlePaddle/PaddleOCR

Python · ★ 80,539 · 🍴 10,632 · 📈 747 stars today

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

中文介绍 PaddleOCR 是一个强大、轻量的 OCR 工具包,支持超过 100 种语言,能将 PDF 或图像文档精准转换为结构化数据,有效连接图像/PDF 与大语言模型。它适用于文档数字化、数据提取和构建智能文档处理管道的开发者。

openai/plugins

JavaScript · ★ 1,537 · 🍴 242 · 📈 49 stars today

OpenAI Plugins

中文介绍 这是 OpenAI 官方的插件仓库,提供了为 ChatGPT 扩展功能的插件示例与规范。开发者可以基于此构建插件,让 ChatGPT 能够访问实时信息、执行操作或与第三方服务交互,从而扩展其能力边界。

MemPalace/mempalace

Python · ★ 53,890 · 🍴 7,079 · 📈 227 stars today

The best-benchmarked open-source AI memory system. And it's free.

中文介绍 MemPalace 是一个在基准测试中表现优异的开源 AI 记忆系统,并且免费使用。它专注于为大语言模型提供高效的长期记忆与上下文管理解决方案,帮助开发者构建具有持久记忆能力的复杂 AI 应用。

withastro/flue

TypeScript · ★ 4,519 · 🍴 241 · 📈 126 stars today

The sandbox agent framework.

中文介绍 Flue 是一个沙箱代理框架,旨在为运行和测试 AI 代理提供安全、隔离的环境。开发者可以使用它来可控地实验代理行为,管理其与外部环境的交互,降低直接部署风险,适用于代理开发与调试阶段。

openclaw/openclaw-windows-node

C# · ★ 1,604 · 🍴 182 · 📈 326 stars today

Windows companion suite for OpenClaw - System Tray app, Shared library, Node, and PowerToys Command Palette extension

中文介绍 这是 OpenClaw 项目的 Windows 伴侣套件,包括一个系统托盘应用、共享库、节点程序以及 PowerToys 命令面板扩展。它为 Windows 用户提供了便捷的界面和工具,用于管理和交互 OpenClaw 服务。

aquasecurity/trivy

Go · ★ 35,855 · 🍴 443 · 📈 207 stars today

Find vulnerabilities, misconfigurations, secrets, SBOM in containers, Kubernetes, code repositories, clouds and more

中文介绍 Trivy 是一个全面的开源安全扫描工具,能够发现容器、Kubernetes、代码仓库、云资源等中的漏洞、配置错误、敏感信息及生成 SBOM。它集成简单,速度快,是 DevOps 和安全团队进行安全审计与合规检查的首选工具。

jwasham/coding-interview-university

★ 350,388 · 🍴 83,290 · 📈 745 stars today

A complete computer science study plan to become a software engineer.

中文介绍 这是一个完整、免费的计算机科学学习计划,旨在帮助学习者系统性地掌握知识,成为一名合格的软件工程师并成功应对技术面试。它涵盖了算法、数据结构、系统设计等核心主题,是自学求职者的宝贵资源库。

github/copilot-sdk

Java · ★ 9,244 · 🍴 1,223 · 📈 309 stars today

Multi-platform SDK for integrating GitHub Copilot Agent into apps and services

中文介绍 这是 GitHub 官方推出的多平台 SDK,用于将 GitHub Copilot Agent 集成到各类应用程序和服务中。它提供了标准化的接口和工具,使开发者能够在自己的产品中方便地嵌入 AI 辅助编程或对话功能。

SpaceX IPOs in 7 days. I Fed the S1 Doc Into Claude. Here Is What It Found Buried in 300 Pages.

@DamiDefi · 96.5K 粉丝 · 2.3M 阅 · 584 赞 · 80 转

The number that stopped me was not the $2 trillion valuation. It was $791 million. That is what SpaceX made in net income in 2024. A profitable, growing aerospace company with a genuine moat in launch

中文介绍 博主使用Claude分析SpaceX的S1招股书,从300页中提取关键数据,如2024年净利润7.91亿美元,突出公司盈利能力和护城河,演示AI在财务文档分析中的高效应用。

How to master Dynamic Workflows in Claude Code: 6 patterns and 14 steps Anthropic engineers actually

@0xCodez · 3.3K 粉丝 · 637.2K 阅 · 510 赞 · 59 转

Most Claude Code users still write their workflows by hand. They chain prompts, copy outputs, paste them into the next prompt, fix what went wrong, repeat. 9 out of 10 builders haven’t tried Dynamic

中文介绍 博主详细分享Claude Code动态工作流的6种模式和14个步骤,帮助开发者从手动链式提示转向自动化流程,提升AI协作效率,是Anthropic工程师的实践经验总结。

Generative UI Is the New Frontend

@Saboo_Shubham_ · 116.2K 粉丝 · 263.3K 阅 · 517 赞 · 74 转

The frontend used to be a fixed thing. Designers drew it. Engineers built it. Users got what shipped. That's over. The interfaces shipping in 2026 are drawn partly by the agent itself, in real time,

中文介绍 博主讨论生成式UI作为前端开发的未来,AI代理能在2026年实时生成界面,打破传统设计与工程分离的模式,推动前端技术向智能化演进。

Building cloud agent infrastructure: what's different, and what we learned

@intuitiveml · 6.4K 粉丝 · 171.3K 阅 · 524 赞 · 70 转

Most agent frameworks today assume a desktop. One user, one machine, one process. The agent runs while the laptop is open, writes to a local filesystem, holds API keys in environment variables, and

中文介绍 博主探讨云代理基础设施的构建经验,对比桌面框架的单用户、本地文件系统等局限,为云环境设计支持多用户、弹性扩展的代理系统,分享学习心得。

A guide to /goal 🥅

@dkundel · 19.3K 粉丝 · 116.9K 阅 · 523 赞 · 40 转

We launched the goal mode (or /goal) as a way to help you have Codex drive towards a concrete outcome. When you set a goal Codex will continue to work until the goal is achieved, whether that takes

中文介绍 博主指南Codex的/goal模式,允许用户设定具体目标,AI会持续执行直到达成,无论耗时多久,增强任务自动化和AI的自主性,适用于长期复杂任务。

A Functional Taxonomy of World Models

@drfeifei · 738.0K 粉丝 · 72.2K 阅 · 699 赞 · 144 转

“The world is everything that is the case.” — Ludwig Wittgenstein, Tractatus Logico-Philosophicus, 1921 The world is not made of words. In an earlier essay, we argued that spatial intelligence is AI’s

中文介绍 博主提出世界模型的功能分类,引用维特根斯坦哲学,强调AI需发展空间智能以超越纯文本理解,指向更全面的认知架构,是AI理论的重要讨论。

How to Build a Custom Agent Harness

@sydneyrunkle · 7.5K 粉丝 · 69.5K 阅 · 511 赞 · 74 转

Building useful agents is largely about customization: connecting your agent to the right context, data, and environment(s) for the task at hand. At its core, an agent is a model calling tools in a

中文介绍 博主指导构建自定义代理工具套件,核心是将代理连接到合适的上下文、数据和环境,通过模型调用工具实现任务定制化,提升代理的实用性和适应性。

I Gave Claude David Ogilvy's Writing Rules And Built A Legendary AI Writing Coach

@dickiebush · 441.8K 粉丝 · 57.7K 阅 · 519 赞 · 45 转

Legendary marketer David Ogilvy generated over $864 million for his clients. He was a British advertiser known as "The Father of Advertising." And in 1982, Ogilvy sent this 1-page memo to his staff:

中文介绍 博主将David Ogilvy的写作规则输入Claude,创建AI写作教练,利用广告传奇人物的智慧优化AI输出,展示提示工程在创意写作中的创新应用。

Your token spend is an AI architecture problem, not just a model problem

@jainarvind · 9.3K 粉丝 · 53.7K 阅 · 505 赞 · 68 转

Enterprise AI token spend is scaling quickly, especially as the technology shifts from simple chat assistants into coding agents, AI coworkers, and long-running workflows. These systems do far more

中文介绍 博主指出企业AI token支出快速增长,不仅是模型问题,更是架构设计挑战,随着技术转向代理和长期工作流,优化架构对控制成本至关重要。

10 HERMES AGENT HACKS THAT TURNED MY CHAT AGENT INTO A 24/7 SYSTEM

@IBuzovskyi · 1.2K 粉丝 · 50.1K 阅 · 500 赞 · 51 转

These 10 Hermes Agent hacks saved me 15+ hours every week - and they work for any workflow you run repeatedly. Content, software development, business operations, client management, research, sales.

中文介绍 博主分享10个Hermes Agent技巧,每周节省15小时以上,适用于内容、开发、运营等多种重复工作流,将聊天代理升级为24/7自动化系统。

Feedback loops: Help Claude Code complete ambitious tasks with less babysitting

@delba_oliveira · 74.0K 粉丝 · 37.4K 阅 · 533 赞 · 38 转

As we delegate more ambitious tasks to Claude, it becomes increasingly important that it can verify its own work. The more Claude can self-verify: the more independently it can work on long-running

中文介绍 博主强调反馈循环在Claude Code中的作用,通过自我验证减少人工干预,使AI能更独立地处理长期任务,提升复杂工作的自主完成率。

Don't let your agent guess, give it runtime context

@ericzakariasson · 67.9K 粉丝 · 37.3K 阅 · 507 赞 · 27 转

If you've ever watched an agent try to fix a bug, you've watched it guess. It reads the code, comes up with a theory, makes an edit, and hopes. Sometimes it's right. A lot of the time you get a fix

中文介绍 博主建议为AI代理提供运行时上下文,避免其在修复bug时盲目猜测,通过上下文信息提高推理准确性,减少无效修改,提升开发效率。

How to Stop Shipping Low-Quality RL Environments (with Examples)

Your broken harness is actively making the model worse. Here's what I keep seeing after years of eyeballing trajectories, and what you need to fix.

中文介绍 文章指出,低质量的强化学习(RL)训练环境会直接损害模型性能,作者基于多年观察提出了常见的问题模式与具体改进建议。

The Meta hack shows there’s more to AI security than Mythos

On June 5, 404 Media reported that attackers had been using Meta’s AI customer support agent to steal Instagram accounts. Their approach was simple: They asked the agent to link the accounts to email addresses that they controlled, and the agent complied. One attacker broke into the dormant Obama Wh

中文介绍 据404 Media于6月5日报道,攻击者利用Meta的AI客服代理成功窃取Instagram账户。攻击者通过指示该代理将账户关联至其控制的邮箱地址,从而轻易得手。

[AINews] not much happened today

a quiet day

中文介绍 Latent Space的AI新闻日报表示当日没有重大新闻发生,内容较为平淡。

Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs

We talk with the VendingBench authors on evaling Claudes from Haiku to Mythos, and how they build leading, and lasting, frontier evals from scratch.

中文介绍 文章讨论了现实世界作为最终评估场的重要性,并访谈了Andon Labs的Lukas Petersson和Axel Backlund,分享了如何从零构建持久的前沿AI评估基准。

How Endava is redesigning software delivery around AI agents

Learn how Endava is using AI agents, ChatGPT Enterprise, and Codex to accelerate software delivery, automate workflows, and build an AI-native culture across the enterprise.

中文介绍 软件公司Endava正在利用AI智能体、ChatGPT Enterprise和Codex来重新设计软件交付流程,旨在加速交付、自动化工作流并构建企业级的AI原生文化。

How courts are coping with a flood of AI-generated lawsuits

Most days in her chambers, Judge Maritza Braswell, a federal magistrate judge in Colorado, sifts through stacks of documents written by people without a lawyer. Many of them can’t afford to hire a lawyer, and others have cases too weak or too small to interest one. She reads each one carefully, mind

中文介绍 法院正面临由AI生成的诉讼文书激增的挑战。联邦治安法官Maritza Braswell每天需要审阅大量由非律师撰写的、质量参差不齐的文件。

Dreaming: Better memory for a more helpful ChatGPT

ChatGPT introduces a new memory system to better remember preferences, keeping context fresh and relevant across conversations.

中文介绍 OpenAI为ChatGPT引入了名为“Dreaming”的新记忆系统,旨在更好地记住用户偏好,在对话间保持上下文的新鲜与相关性,以提升服务体验。

not much happened today

**NVIDIA** released **Nemotron 3 Ultra**, a fully open **550B MoE** model with **55B active parameters** and **1M context**, optimized for long-running agent tasks with up to **5x speedup** and **30% cost reduction**. It features hybrid Mamba/attention, LatentMoE, native MTP, and was pretrained on *

中文介绍 英伟达发布了完全开源的Nemotron 3 Ultra模型。这是一款拥有5500亿参数的混合专家模型,激活参数为550亿,上下文长度达100万token,针对长时间运行的智能体任务进行了优化,推理速度提升5倍,成本降低30%。

Biodefense in the Intelligence Age

An action plan for AI-powered biological resilience

中文介绍 OpenAI发布了一份名为《智能时代的生物防御》的行动计划,旨在利用AI技术增强全球生物韧性,以应对潜在的生物威胁。

每日论文 · arXiv cs.CR 最新公告批次

周末 arXiv 通常无新公告。当前展示最近一次可用公告批次。

Will the Agent Recuse Itself? Measuring LLM-Agent Compliance with In-Band Access-Deny Signals

第一作者: Thamilvendhan Munirathinam · 方向: 密码学协议

As autonomous LLM agents increasingly hold real credentials and operate infrastructure without a human in the loop, operators have no standard way to tell an agent that a resource is off-limits. Access controls either let the agent in (it has valid credentials) or hard-fail it (indistinguishable from any other client). We propose a third mode: a lightweight, published in-band deny signal -- the Recuse Signal -- that a server emits over a protocol's existing channels (an SSH banner, a PostgreSQL NOTICE) asking a connecting automated agent to voluntarily withdraw. This is a cooperative governance control, the robots.txt analogue for live access; it is explicitly not a security boundary. Its value is entirely empirical and, to our knowledge, unmeasured: do compliant LLM agents actually honor such a signal? We define the signal as an open mini-standard, implement two zero- or low-footprint...

论文介绍 研究在自主LLM代理持有真实凭证时,如何通过协作治理方式控制其访问。论文提出“Recuse Signal”——一种通过现有协议通道(如SSH横幅)发送的带内轻量级拒绝信号,请求代理自愿撤回。该信号是开放的迷你标准,旨在成为实时访问的“robots.txt”。论文通过实现和测试,首次测量了LLM代理对此类信号的遵守情况,为代理系统管理提供了经验数据。

WebMCP Tool Surface Poisoning: Runtime Manipulation Attacks on LLM Agents

第一作者: Lin-Fa Lee · 方向: 密码学协议

Abstract:WebMCP is a newly emerging protocol that enables websites to expose tools directly to AI agents, bypassing traditional user interfaces and introducing new security risks. The dynamic exposure of agent-accessible tools in WebMCP expands the attack surface of web sessions, especially when third-party scripts are involved. In this study, we identify a new potential threat, termed Mid-Session Tool Injection (MSTI), in which attackers leverage third-party scripts to inject malicious tools during an active session. To better characterize this threat, we classify MSTI based on the stage and target of manipulation, distinguishing between Tool Hijacking and Tool Framing. Tool Hijacking modifies the set of tools visible to the agent through mechanisms such as the AbortSignal API or race conditions during tool registration. In contrast, Tool Framing influences the agent's perception of...

论文介绍 研究新兴WebMCP协议的安全风险,该协议允许网站直接向AI代理暴露工具。论文识别了“会话中工具注入”威胁,即攻击者利用第三方脚本在活跃会话中注入恶意工具。根据操作阶段和目标,将其分为“工具劫持”(修改工具集)和“工具构陷”(影响代理感知)。研究系统分析了这类攻击的原理与分类,对构建更安全的代理-网站交互协议具有重要意义。

Credential Disclosure in (EU) Digital Identity Wallets: Privacy Risks and Practical Mitigations

第一作者: Sheila Zingg · 方向: 系统安全

Abstract:The European Union will introduce the EUDI Wallet by late 2026, which allows users to hold digital credentials (i.e., representations of physical official identity documents) on their devices. This will allow users to securely and privately disclose identity attributes to websites. Although such a system has many benefits, it also introduces risks caused by poor credential disclosure decisions. In this paper, we (i) conduct a large-scale survey on credential disclosure with users and experts and (ii) evaluate the effectiveness and feasibility of our Credential Assistant that displays expert recommendations and user opinions. Our results show that users are likely to overshare (e.g., ~20% of users disclosed their official ID to news websites). This indicates that users struggle to protect their privacy, which will impact the usability of the EUDI Wallet and lead to privacy...

论文介绍 研究欧盟数字身份钱包在用户凭证披露时的隐私风险。通过对用户和专家的大规模调查发现,用户容易过度分享身份属性(例如,约20%的用户向新闻网站披露了官方身份证)。论文评估了一种“凭证助手”的有效性与可行性,该助手通过显示专家建议和用户意见来辅助决策。结果表明当前设计下用户隐私保护面临挑战,对钱包的实际应用构成影响。

Robust Ensemble of Selectively Strengthened and Augmented Predictors

第一作者: Parsa Memarzadehsaghezi · 方向: AI 安全

Abstract:Evasion attacks present a significant challenge to the robustness of machine learning (ML)-based classifiers, particularly in critical applications such as fraud detection and cybersecurity. Although existing defense mechanisms are effective in some settings, they often suffer from limited generalizability and do not systematically improve model robustness across diverse attack scenarios. To address these limitations, we introduce Robust Ensemble of Selectively Strengthened and Augmented Predictors (RESSAP), a novel framework that transforms a single classifier into an ensemble of robust classifiers. Each classifier in the ensemble is trained on a carefully selected subset of features, where feature selection is guided by a resilience metric that accounts for both feature importance and robustness. During inference, a random subset of these classifiers is used to make...

论文介绍 针对机器学习分类器在逃避攻击下的鲁棒性问题,现有防御方法泛化能力有限。论文提出RESSAP框架,旨在将单一分类器转换为鲁棒分类器集成。其核心在于根据一个综合考虑特征重要性和鲁棒性的“韧性指标”,为集成中的每个成员精心选择训练特征子集。推理时随机使用部分分类器预测,以提升模型在多样化攻击场景下的整体防御能力。

SecRL-Prune: Structured Reinforcement Learning-Based Pruning of CodeLLMs for Preserving Adversarial Code Mutation

第一作者: Parsa Memarzadehsaghezi · 方向: 密码学协议

Abstract:Large code language models (CodeLLMs) can generate and rewrite programs, enabling functionality-preserving code mutation that may be used to create diverse malware variants and evade signature-based detection. A key security question is whether this mutation capability survives model compression, which would make deployment feasible under limited hardware budgets. We propose SecRL-Prune, a structured pruning framework for CodeLLMs that operates on feed-forward (MLP/FFN) channels. Starting from a pretrained teacher, it learns a layer-wise pruning policy with reinforcement learning using a teacher-student KL-divergence reward. To improve efficiency, we cache the teacher's top-P predictions once and compare the pruned student against this compact target, avoiding simultaneous teacher-student residency in GPU memory. We evaluate SecRL-Prune on HumanEval using pass@k for execution...

论文介绍 探究大型代码语言模型在压缩后,其执行功能性保持代码变异的能力是否得以保留,这种能力可能被用于创建恶意软件变种。论文提出SecRL-Prune,一种基于强化学习的结构化剪枝框架,针对模型的前馈网络层。通过学习分层剪枝策略,在压缩模型的同时,旨在维持其关键的代码变异能力,为在资源受限环境下安全部署此类模型提供思路。

Steering LLM Viewpoints through Fabricated Evidence Injection

第一作者: Xi Yang · 方向: AI 安全

Abstract:As chatbots increasingly influence daily decision-making, their potential to produce misleading responses poses substantial risks to users. This paper investigates a critical cognitive vulnerability in LLMs: their tendency to uncritically trust external context when presented with fabricated evidence bearing markers of credibility. We introduce Ghostwriter, a two-phase attack framework that first repackages misleading statements with fabricated rationales, then instruct target LLMs to incorporate these viewpoints when responding to relevant queries. Experiments on BBQ, ToxiGen, and our specialized dataset reveal that commercial LLMs without external safety classifiers remain highly vulnerable, while even frontier classifier-guarded models (e.g., GPT-5.4) reduce but do not eliminate the attack. Building on this, we explore multiple defense strategies, among which a tailored...

论文介绍 揭示大语言模型存在信任带有可信标记的外部上下文的认知漏洞。论文提出Ghostwriter攻击框架:第一阶段将误导性陈述与伪造理由重新打包;第二阶段指示目标LLM在回答相关查询时采纳这些注入的观点。实验表明,缺乏外部安全分类器的商业LLM高度易受攻击,而前沿的防护模型虽能降低风险但无法完全消除。研究还探索了多种防御策略。

Opportunities and Challenges in Securely Reusing and Repurposing Mobile Devices

第一作者: Adelin Roty · 方向: 系统安全

Abstract:An estimated 5.3 billion mobile phones became electronic waste in 2022. Many of these devices can be repurposed and used in different contexts to extend their lifetime and to reduce ecological impacts. An often overlooked aspect of smartphone reuse is cybersecurity: these devices embed hardware-backed security mechanisms that rely on vendor-controlled provisioning and are designed for a fixed device lifecycle. In this paper, we investigate whether security mechanisms and guarantees remain effective when devices are repurposed outside their original ecosystem. We explore security features in a PinePhone, an open-hardware smartphone, and focus on three core security aspects: boot chain integrity, isolation provided by the Trusted Execution Environment, and the protection of hardware-bound secrets. Our experiments simulate realistic repurposing scenarios and highlight the...

论文介绍 探讨移动设备在原始生命周期外被再利用时面临的网络安全挑战。设备内置的硬件安全机制依赖于厂商控制的固定生命周期,再利用可能破坏其保障。论文在开源硬件手机PinePhone上,针对启动链完整性、可信执行环境隔离和硬件绑定秘密保护三个核心安全方面进行实验,模拟现实再利用场景,评估安全机制的有效性,为延长设备寿命与保障安全提供依据。

RedEdit: Agentic Red-Teaming of Image Safety Classifiers via MCTS-Guided Photo-Editing

第一作者: Weilin Lin · 方向: 网络安全

Image safety classifiers serve as a critical component of contemporary content moderation systems on the internet. However, their resilience against user-style malicious image editing remains underexplored. Such behaviors are highly prevalent in daily scenarios but difficult to fully reproduce. To explore this vulnerability, we introduce RedEdit, a novel black-box red-teaming agent that formulates photo-editing evasion as a combinatorial search problem over edit-tool sequences. It adopts a Vision-Language-Model (VLM)-based proposer to generate semantically targeted candidate edits and a Monte Carlo Tree Search (MCTS) planner to prioritize promising edit paths while backtracking from ineffective ones. Together, the proposer and planner instantiate two key capabilities of human attackers, i.e., domain knowledge and iterative backtracking, respectively, to reproduce this practical threat...

论文介绍 针对图像安全分类器对用户常见恶意图像编辑缺乏抵御能力的问题,论文提出RedEdit红队代理。它将规避编辑建模为对编辑工具序列的组合搜索,利用视觉语言模型生成语义定向的候选编辑,并采用蒙特卡洛树搜索规划来优选编辑路径、回溯无效尝试。该方法模拟了人类攻击者的领域知识和迭代试错能力,系统性地评估分类器的脆弱性,推动更鲁棒的安全系统构建。

Cheating in Multiplayer Online Games: a Dataset

第一作者: Hugo Bertin · 方向: 网络安全

Abstract:Cheating poses a significant threat to the Multiplayer Online Games (MOG) industry by degrading player satisfaction and undermining the fairness in competitive gaming. Despite efforts to develop mitigation techniques, cheating remains difficult to detect and prevent in practice. In particular, a class of cheats based on network flow disruption remains unsolvable. To find out how to detect such attacks we need access to representative labelled data. However, no such dataset exists. To address this gap, we leverage an experimental framework that combines a multiplayer online game with a plug-in capable of both reproducing cheating attacks and collecting logs at two levels: network and application-layer. This paper presents a dataset compiling records of game sessions played by both real players and automated game clients, with cheating actions explicitly logged. To the best of...

论文介绍 本研究针对多人在线游戏中的作弊威胁,指出基于网络流量中断的作弊攻击难以检测和预防。为了解决缺乏代表性数据集的问题,作者开发了一个实验框架,该框架整合了多人在线游戏和一个插件,能够模拟作弊攻击并收集网络层和应用层的日志。论文呈现了一个包含真实玩家和自动化客户端游戏会话记录的数据集,其中作弊行为被明确记录,为未来作弊检测研究提供了基础资源。

AttackPathGNN: Cross-function vulnerability detection in smart contracts using state interference graphs and conjunction pooling

第一作者: Gabriela Dobrita · 方向: AI 安全

Abstract:Existing learning-based detectors for Solidity smart-contracts reduce vulnerability detection to syntactic pattern matching within single functions, yet many of the most consequential exploits (The DAO, Cream Finance) exist not in any individual function but in the relationship between functions and in the combination of conditions that made the attack feasible. Thus, we propose AttackPathGNN, a graph neural network (GNN) that reframes detection as reasoning over explicit attack paths. Two architectural choices distinguish it from prior GNN-based detectors: (1)a State Interference Graph that links every pair of functions sharing mutable storage through typed, weighted edges and through directed reentrancy-path edges defined by an explicit five-condition predicate; (2)conjunction pooling, a differentiable AND-aggregator over eight named exploit preconditions whose log-sigmoid...

论文介绍 现有智能合约漏洞检测方法多局限于单函数内的语法模式匹配,忽略了跨函数关系和攻击条件组合。本文提出AttackPathGNN,一个图神经网络,通过状态干扰图链接共享可变存储的函数,并使用连接池聚合攻击前提条件,以推理显式攻击路径,提高对跨函数漏洞的检测能力。

Exploring the connection between coding habits and cognitive styles in malware developers

第一作者: Vasilis Vouvoutsis · 方向: 软件安全

Malware research primarily studies the results, the methods, and the impact. Even from an offensive security perspective, what is examined is the method, not the development strategy of the offender. This study investigates the behavioral signatures and coding patterns embedded in the malware source code. By analyzing a large corpus of leaked malware code and comparing it with carefully selected benign open-source software, we apply static application security testing and compute multiple software metrics. Based on cognitive psychology and criminological theories, our work interprets differences in code structure and quality as behavioral indicators, reflecting distinct motivational structures, risk tolerances, and development strategies of malware authors compared to benign software developers. Our findings reveal that malware code is generally smaller, less documented, and exhibits...

论文介绍 本研究探讨恶意软件开发者的行为特征,通过分析泄露的恶意代码并与良性软件比较,应用静态分析计算软件度量。基于认知心理学和犯罪学理论,将代码差异解释为行为指标,反映了恶意开发者的独特动机、风险容忍度和开发策略,为安全研究提供新视角。

PriSrv+: Privacy and Usability-Enhanced Wireless Service Discovery with Fast and Expressive Matchmaking Encryption

第一作者: Yang Yang · 方向: 密码学协议

Service discovery is a fundamental process in wireless networks, enabling devices to find and communicate with services dynamically, and is critical for the seamless operation of modern systems like 5G and IoT. This paper introduces PriSrv+, an advanced privacy and usability-enhanced service discovery protocol for modern wireless networks and resource-constrained environments. PriSrv+ builds upon PriSrv (NDSS'24), by addressing critical limitations in expressiveness, privacy, scalability, and efficiency, while maintaining compatibility with widely-used wireless protocols such as mDNS, BLE, and Wi-Fi. A key innovation in PriSrv+ is the development of Fast and Expressive Matchmaking Encryption (FEME), the first matchmaking encryption scheme capable of supporting expressive access control policies with an unbounded attribute universe, allowing any arbitrary string to be used as an...

论文介绍 针对无线网络中服务发现协议隐私保护不足的问题,本文提出PriSrv+协议,增强隐私和可用性。其核心创新是快速表达式匹配加密(FEME),支持无界属性宇宙的表达式访问控制策略,兼容mDNS、BLE等无线协议,适用于资源受限环境。

GenTI: Benchmarking LLMs for Autonomous IDPS Rule Generation for Unseen Attacks

第一作者: Hassan Jalil Hadi · 方向: 密码学协议

Abstract:Rule-based Intrusion Detection and Prevention Systems (IDPS) offer precise attack detection as well as mitigation, however their manually crafted, signature-driven rules limit adaptability to emerging and zero-day threats. Additionally, existing public datasets (e.g., CICIDS2017, UNSW-NB15) focus on traffic classification and provide little structured information to support automatic rule synthesis or prevention logic. To address this gap, we propose Generative Thread Intelligence (GenTI) \footnote{GenTI refers to the proposed framework, and GTI refers to the dataset.} an LLM-driven benchmark for automatic generation of IDPS rules targeting unseen attacks. The dataset (GTI) aggregates over 150k detection and prevention rules from Snort, Suricata, Emerging Threats, as well as 50k YARA, each annotated with protocol behavior, payload signatures, contextual relationships, mappings...

论文介绍 为了解决IDPS规则手动编写难以适应新兴威胁的问题,本文提出GenTI框架,利用大语言模型自动生成IDPS规则。构建了GTI数据集,聚合了超过15万条检测和预防规则,用于基准测试LLM在未见攻击规则生成中的性能。

Towards Worst-case Hardness for Low-Noise LPN

第一作者: Divesh Aggarwal · 方向: 密码学协议

Abstract:The hardness of the Learning Parity with Noise (LPN) problem is a foundational assumption in cryptography, forming the basis of constructions ranging from symmetric-key primitives to public-key encryption and beyond. A central open question is whether the average-case hardness of LPN can be based on worst-case complexity assumptions, as has been achieved for the analogous Learning With Errors (LWE) problem. Existing worst-case-to-average-case reductions for LPN [BLVW19, YZ21] rely on statistical smoothing of linear codes, which inherently limits the resulting average-case hardness to noise rates as large as $1/2 - 1/\mathrm{poly}(n)$, which is insufficient for public-key applications. We explore a new approach towards obtaining such reductions: rather than requiring that random sparse combinations of the rows of the generator matrix of a code be statistically close to uniform...

论文介绍 LPN问题的平均情况硬度是否能基于最坏情况复杂性假设是密码学中的关键问题。本文探索新的归约方法,尝试通过修改代码生成器矩阵的统计性质,为低噪声率LPN提供最坏情况到平均情况的归约,推动理论进展。

PriSrv: Privacy-Enhanced and Highly Usable Service Discovery in Wireless Communications

第一作者: Yang Yang · 方向: 密码学协议

Abstract:Service discovery is essential in wireless communications. However, existing protocols provide limited privacy protection, leaking sensitive device information and opening routes to network attacks. This paper proposes a private service discovery protocol, called PriSrv, which enables both service providers and clients to specify fine-grained authentication policies before establishing connections. PriSrv achieves this via a dual-layer matching architecture: an outer layer filters mismatched entities using public attributes, while an inner layer handles mutual authentication using selectively disclosed private attributes. As a core component, we introduce the primitive of anonymous credential-based matchmaking encryption (ACME), which enables dual-layer matching in a single step to achieve bilateral policy control, selective attribute disclosure, and multi-show unlinkability...

论文介绍 针对服务发现中的隐私泄露问题,本文设计PriSrv协议,允许服务提供者和客户端指定细粒度认证策略。采用双层匹配:外层使用公共属性过滤,内层通过匿名凭证基础的匹配加密实现相互认证,支持选择性属性披露,增强隐私保护。

GCD: Garbled, Corrected, Demonstrandum -- Fixing and Proving Go's Extended GCD Implementation

第一作者: Linard Arquint · 方向: 软件安全

Abstract:We verify the 'extendedGCD' implementation in Go's standard library ('crypto/internal/fips140/bigmod'), which plays a crucial role in the generation of RSA key pairs. Even though the Go implementation is supposedly a direct port from BoringSSL's implementation, we uncovered two deviations that each break the algorithm's invariants: (1) the Go implementation deviates in the way coefficients are updated, and (2) it permits a larger input domain. We address both deviations; the first by fixing the Go implementation, which results in an on average 24% speedup, and the second deviation by porting an existing proof for BoringSSL and extending it to cover the larger input domain. We prove correctness and termination of the fixed Go implementation using Gobra, a deductive program verifier for Go. Where necessary, we used Lean to prove key lemmata on non-linear arithmetic, which we...

论文介绍 本文验证Go标准库中用于RSA密钥生成的extendedGCD实现,发现两个偏差导致算法不变量破坏。通过修复Go代码并扩展BoringSSL的证明,使用Gobra程序验证器证明了修正后实现的正确性和终止性,提升了软件可靠性。

SentinelRAG: Synthetic Sentinel Knowledge for RAG Database Copyright Protection

第一作者: Tsun On Kwok · 方向: AI 安全

Abstract:Protecting proprietary RAG databases from unauthorized redistribution is challenging: existing watermarking methods either inject fabricated relations between real entities, polluting the knowledge base with misinformation, or embed fragile lexical patterns that adversarial paraphrasing easily removes. We propose SentinelRAG, a watermarking framework that embeds style-consistent but fictitious knowledge entries into the RAG database. Our key insight is that synthetic knowledge describing fictitious entities is unlikely to be retrieved by legitimate queries, yet can be reliably triggered through targeted probes known only to the data owner. Experiments on four datasets ranging from 2.9k to 8.8M documents demonstrate that SentinelRAG achieves statistically significant detection $p < 10^{-5}$ across all tested configurations at only a 0.1% injection rate. Compared to the...

论文介绍 该研究针对RAG数据库版权保护问题,提出SentinelRAG水印框架。核心方法是在数据库中嵌入风格一致但虚构的知识条目作为水印,通过仅数据所有者知晓的触发探针实现可靠检测。实验在多个数据集上验证,以0.1%的注入率实现统计显著检测,避免污染知识库或易被移除的缺陷,为数据库防未授权分发提供新思路。

TinyML-Driven Cybersecurity for Autonomous Spacecraft: Latency-Accuracy Analysis for SPARTA RF and Cyber Threat Detection

第一作者: Van Le · 方向: AI 安全

Abstract:Autonomous spacecraft require rapid, lightweight, and reliable onboard detection of cyber-RF threats. Using the SPARTA attack model, we analyze the latency-accuracy trade-offs of TinyML-compatible classical models -- Random Forest, Logistic Regression, SVM, and MLP -- for detecting uplink jamming, Fake-NR spoofing, payload manipulation, ground-segment compromise, and unauthorized command injection. We present a physics-informed theoretical analysis of each model's computational complexity, VC dimension, Lipschitz continuity, and latency scaling, supported by empirical measurements on adversarial RF spectrograms generated via BandErasure, FakeNR, and NoiseBurst corruption modes. Results show that Logistic Regression achieves microsecond-level inference with only a 1\% accuracy drop relative to Random Forest, making it an effective TinyML baseline for onboard autonomy. The study...

论文介绍 该研究关注自主航天器的网络射频威胁检测,分析TinyML兼容模型的延迟-准确性权衡。基于SPARTA攻击模型,评估随机森林、逻辑回归等经典方法在检测上行链路干扰、假NR欺骗等攻击时的性能。理论分析和实证测量表明,逻辑回归以微秒级推理和1%精度损失成为有效的TinyML基线,适用于轻量级太空安全应用。

An Improved CNN-LSTM Based Intrusion Detection System for IoT Networks

第一作者: Mohammad Tariq Ikhlas · 方向: AI 安全

Abstract:With the rapid proliferation of IoT devices, security concerns have dramatically escalated and intrusion detection systems have become critical for protecting networked environments. This paper presents an improved CNN-LSTM based intrusion detection model that combines multi-class classification, dataset integration, and temporal feature learning to enhance detection performance in IoT networks. Using network traffic data, the proposed approach is evaluated on intrusion detection tasks and achieves an accuracy of approximately 97%. Experimental results demonstrate that the model effectively detects multiple attack categories while maintaining stable training and validation performance. The integration of convolutional and recurrent neural network components enables the framework to capture both spatial and temporal characteristics of network traffic, improving overall...

论文介绍 该研究针对物联网网络安全问题,提出改进的CNN-LSTM入侵检测模型。核心方法是结合卷积和循环神经网络,捕获网络流量的空间和时间特征,实现多类攻击分类。在入侵检测任务上评估,达到约97%的准确率,展示了稳定训练和验证性能,为物联网环境提供增强的检测能力。

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

第一作者: Minseok Choi · 方向: AI 安全

Abstract:Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evolving attacks, while adaptive memory-based guardrails tend to over-refuse benign queries that resemble stored attacks. We propose Membrane, a self-evolving guardrail built on Contrastive Safety Memory (CSM): each cell pairs the conditions for blocking a harmful query with those for permitting a superficially similar benign request. Without retraining, Membrane evolves CSM by distilling each harmful interaction and its benign counterpart into a contrastive cell indexed by the underlying attack strategy, so that one cell generalizes across topical variants of the same mechanism. At inference, retrieved cells serve as grounding context for precise safety decisions. Across model-level safety on HarmBench...

论文介绍 该研究针对大型语言模型持续演变的越狱攻击,提出Membrane自演进安全护栏。核心方法是基于对比安全记忆(CSM),将有害查询和类似良性请求的条件配对存储,通过蒸馏交互实现记忆演化。无需重新训练,Membrane能泛化攻击变体,在推理时提供精准安全决策,增强对动态攻击的防御适应性。

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

第一作者: Shuze Liu · 方向: AI 安全

Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service security. However, individual extraction queries often resemble benign requests, and existing evaluations often focus on single-query anomaly scoring or pure benign-versus-attacker user settings. We formulate model extraction monitoring as benign-calibrated traffic-window distribution testing and show that an embarrassingly simple detector is effective: embed incoming queries into a semantic space and test whether their aggregate distribution deviates from historical benign traffic. We instantiate the detector with maximum mean discrepancy (MMD), using only benign-vs-benign comparisons to set the decision threshold. We evaluate on fourteen attacker-normal query pairs from four extraction scenarios and compare with adapted PRADA, SEAT...

论文介绍 该研究针对LLM API流量中的模型提取威胁,提出一种简单的检测器。核心方法是将查询嵌入语义空间,并使用最大均值差异(MMD)测试其聚合分布与历史良性流量的偏差,基于良性流量校准阈值。在多个提取场景中评估,该方法能有效检测恶意流量窗口,保护模型所有权和服务安全。

Hybrid CNN-LSTM Framework for Intelligent Cyber Attack Detection and Prevention in U.S. Critical Digital Infrastructure: A Comparative Machine Learning Evaluation on CSE-CIC-IDS2018

第一作者: Md. Iqbal Hossan · 方向: 密码学协议

Abstract:Digital infrastructure is growing at a rapid pace in the United States, and as a result, exposure to advanced cyber threats to critical sectors including healthcare, finance, transportation, energy and government systems is growing. The traditional cybersecurity approaches, including signature-based intrusion detection systems, have become less effective against today's cyber attacks, as they are unable to detect unknown and changing attacks in real time. To overcome these constraints, this research suggests a smart cyber-defense system, which utilizes Artificial Intelligence (AI) and Machine Learning (ML) algorithms in the detection and prevention of cyber attacks in the U.S. digital infrastructure. This study uses the CSE-CIC-IDS2018 dataset, which is a realistic network traffic dataset, along with various cyber attack scenarios, including Distributed Denial of Service...

论文介绍 该研究针对美国关键数字基础设施的网络安全挑战,提出混合CNN-LSTM框架用于智能攻击检测和预防。核心方法是结合CNN和LSTM的特征学习能力,使用CSE-CIC-IDS2018数据集进行评估。研究旨在提升对DDoS等高级威胁的实时检测,为医疗、金融等关键部门提供增强的防御机制。

Explainable AI-Driven Cyber Risk Analytics and Model Reliability Assessment for Intelligent Governance of U.S. Critical Infrastructure: An XGBoost and SHAP-Based Intrusion Detection Framework

第一作者: B. M. Taslimul Haque · 方向: 系统安全

Abstract:The increasing penetrations of the critical infrastructure sector in the United States with intelligent digital technologies have greatly increased exposure to advanced cyber adversaries and operational vulnerabilities. AI-powered governance and automated decision-making systems are becoming a key part of the operation of critical infrastructure systems, including energy, healthcare, transportation, financial services, and communication infrastructure, in order to improve efficiency and strategic management. The growing cyber threat environment, such as Distributed Denial of Service (DDos) attacks, botnets, ransomware, and Advanced Persistent Threats (APTs) pose significant challenges to infrastructure resilience, cyber security reliability, and governance trustworthiness. In a changing attack landscape and dynamic network environment, traditional cybersecurity mechanisms can...

论文介绍 该研究针对美国关键基础设施的智能治理需求,提出基于XGBoost和SHAP的入侵检测框架。核心方法是利用可解释AI驱动网络风险分析和模型可靠性评估,以增强决策透明度和信任度。框架旨在应对DDoS、僵尸网络等高级威胁,提升基础设施在动态网络环境中的韧性和安全性。

Cognitive Threat Intelligence and Explainable Federated Security Analytics for distributed Infrastructure Systems

第一作者: Md. Arifur Rahman · 方向: AI 安全

The increasing adoption of distributed infrastructure systems, cloud computing, Internet of Things (IoT) technologies, and edge-based architectures has significantly expanded the cybersecurity attack surface and introduced increasingly sophisticated cyber threats. Conventional centralized intrusion detection approaches often face challenges related to scalability, data privacy, communication overhead, and limited transparency in artificial intelligence-driven decision-making processes. To address these limitations, this study proposes a Cognitive Threat Intelligence and Explainable Federated Security Analytics framework for distributed infrastructure systems. The proposed framework integrates Federated Learning (FL), Explainable Artificial Intelligence (XAI), and cognitive cybersecurity analytics to enable collaborative and privacy-preserving cyber threat detection across distributed...

论文介绍 该研究针对分布式基础设施系统的网络安全挑战,提出认知威胁情报和可解释联邦安全分析框架。核心方法是整合联邦学习、可解释AI和认知网络安全分析,实现隐私保护的协作威胁检测。框架旨在解决可扩展性、数据隐私和决策透明度问题,适用于云计算、物联网等分布式环境。

Protecting K-Nearest Neighbor Queries from Location Inference Attacks

第一作者: Zhiyu Sun · 方向: 隐私保护

Abstract:The k-nearest neighbor query (kNNQ) is a core component of modern location-based services (LBS) and has been widely adopted in popular features such as ``people nearby''. However, its potential privacy risks have long been overlooked. In this work, we present the first two attacks against kNNQ, namely the geometric intersection location inference attack (GI-LIA) and the zero-order optimization location inference attack (ZO-LIA), revealing the inherent location privacy risks posed by kNNQ. To mitigate these privacy risks, we further propose DPRS, a differential privacy framework for kNNQ protection. The core idea of DPRS is to incorporate a rejection sampling mechanism within a constrained perturbation interval, thereby mitigating the distance distortion caused by excessive noise injection. In addition, we design a private interval construction algorithm to construct the...

论文介绍 k近邻查询是现代位置服务的核心,但存在隐私风险。本文首次揭示了其面临的两种位置推断攻击:几何交集攻击与零阶优化攻击。为保护隐私,作者提出了DPRS差分隐私框架,其核心是在受限扰动区间内引入拒绝采样机制,以减轻因过度添加噪声而导致的距离失真。该研究为位置服务的隐私保护提供了新的方法与视角。

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks

第一作者: Seungwon Jeong · 方向: AI 安全

Abstract:As large language models (LLMs) are widely deployed, identifying their vulnerability through jailbreak attacks becomes increasingly critical. Optimization-based attacks like Greedy Coordinate Gradient (GCG) have focused on inserting adversarial tokens to the end of prompts. However, GCG restricts adversarial tokens to a fixed insertion point (typically the prompt suffix), leaving the effect of inserting tokens at other positions unexplored. In this paper, we empirically investigate \emph{slots}, i.e., candidate positions within a prompt where tokens can be inserted. We find that vulnerability to jailbreaking is highly related to the selection of the \emph{slots}. Based on these findings, we introduce the \textit{Vulnerable Slot Score} (VSS) to quantify the positional vulnerability to jailbreaking. We then propose SlotGCG, which evaluates all slots with VSS, selects the most...

论文介绍 大型语言模型易受越狱攻击,但现有优化方法仅将对抗性token固定插入提示词末尾。本文发现,将token插入提示词不同位置会显著影响攻击成功率,这被称为「位置脆弱性」。为此,作者提出SlotGCG框架,通过「脆弱位置分数」评估所有候选插入点,并选择最脆弱的位置进行攻击。这为理解和缓解LLM的安全漏洞提供了新思路。

The Coverage Gap: Chile's Cyber Disclosure Framework versus the USA, EU and UK

第一作者: David Mellafe Z · 方向: 安全研究

We introduce the Coverage Gap as a measurable distance between the observable public exposure of critical-infrastructure operators and their declared capability to coordinate vulnerability disclosure. We instantiate it against the 915 Chilean Operadores de Importancia Vital (OIVs -- Operators of Vital Importance) designated by the National Cybersecurity Agency (ANCI) under Ley 21.663 (Resolucion Exenta No. 87, 16 December 2025). Using a passive-only, OSINT-based method consistent with the principles of ISO/IEC 29147:2018 and Chile's computer-crimes safe harbour (Ley 21.459), we conduct a full-universe census of the foundational disclosure-capability layer (Layer 1, verifiable disclosure contact) across approximately 98.7% of the official catalogue. Only 16 of 915 OIVs (1.7%) publish a verifiable RFC 9116 disclosure channel; among operators of physical-world infrastructure -- energy...

论文介绍 本文定义了「覆盖差距」这一概念,用以衡量关键基础设施运营商的公开网络暴露度与其声称的漏洞协调能力之间的差距。作者对智利915个「至关重要运营商」进行调查,发现仅1.7%拥有可验证的RFC 9116披露渠道。该研究通过被动式OSINT方法,量化了智利在网络安全披露框架方面与美国、欧盟及英国之间的差距,为政策改进提供了依据。

Dimensionality Reduction for Cyberattack Classification: A Comparative Evaluation of PCA and Linear Predictive Coding

第一作者: Nelly Elsayed · 方向: AI 安全

Abstract:High-dimensional feature representations are widely used in machine learning-based cyberattack detection systems. However, they increase computational complexity and may hinder deployment in resource-constrained environments. In this paper, we investigate feature compression techniques for cyberattack classification by comparing two dimensionality reduction approaches: Principal Component Analysis (PCA) and Linear Predictive Coding (LPC). Compressed feature representations with varying dimensionalities are generated and evaluated across several classification models. Experimental analysis demonstrates that PCA preserves classification performance even under aggressive compression. On the other hand, LPC provides competitive predictive representations with slightly larger performance degradation. The results show that substantial reductions in feature dimensionality can be...

论文介绍 基于机器学习的网络攻击检测系统常使用高维特征,这增加了计算复杂度。本文比较了主成分分析和线性预测编码两种降维方法在网络攻击分类任务中的效果。实验表明,PCA即使在激进压缩下也能保持分类性能,而LPC的性能略有下降。该研究证明,大幅降低特征维度是可行的,有助于在资源受限环境中部署检测系统。

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense

第一作者: Anlan Zheng · 方向: AI 安全

Abstract:LLM-driven automated penetration testing agents are typically evaluated against static targets that neither detect nor respond to attacks, so their behavior under intelligent defense remains untested. The causal consistency of multi-step attack chains likewise hinges on unstable LLM reasoning, and agent decisions remain opaque to human analysts. These three shortcomings, in realism, consistency, and auditability, are usually patched in isolation. We present ZERO-APT, a turn-based attacker-defender-judge framework that addresses them within a single architecture. For realism, ZERO-APT embeds a configurable LLM Defender that consumes Sysmon telemetry and detects attacks in real time, exposing the attacker to a live opponent rather than a passive target. For consistency, three architectural mechanisms move causal consistency from unstable LLM reasoning into enforced system...

论文介绍 现有LLM驱动的自动化渗透测试评估缺乏真实性和一致性。本文提出ZERO-APT,一个基于回合制的攻击者-防御者-裁判框架。它通过嵌入可配置的LLM防御者来模拟实时对抗,提升测试的真实性;并通过架构机制而非不稳定的LLM推理来保证多步攻击链的因果一致性。该框架旨在更全面地评估和改进自动化渗透测试。

Bitcoin After Block Rewards

第一作者: Junhyuk Lee · 方向: 密码学协议

Abstract:Bitcoin's block reward is scheduled to decline to zero, raising concerns about whether the network can remain secure once miners rely solely on transaction fees. This paper seeks to identify the conditions under which large-scale and persistent deviation from honest mining can arise. We analyze and compare the payoffs of honest and deviating miners in a sequential decision model, and identify a deviation threshold $G_t$ at which honest mining ceases to be privately optimal. Around the 2024 Bitcoin halving, we show that current mining behavior does not exhibit large-scale or structural deviation. However, when the block reward is removed, the $G_t$ criterion implies that deviation can arise even with a very small fraction of transaction fees. Finally, we evaluate three protocol-level mechanisms: Base Fee, Fee Floor, and an adaptive maximum block size rule, and show that their...

论文介绍 比特币的区块奖励将归零,引发对仅依赖交易费维持网络安全的担忧。本文通过建立顺序决策模型,分析诚实与偏离挖矿的收益,并识别出一个偏离阈值。研究表明,当前挖矿行为未出现大规模偏离,但区块奖励移除后,即使交易费比例很低也可能引发偏离。论文还评估了基础费、费用下限等三种协议改进机制的效果。

SHIELDS: Automating OS Hardening with Iterative Multi-Agent Remediation

第一作者: Andrew Hamara · 方向: 系统安全

Security misconfigurations remain a leading cause of OS-level compromise, and manually keeping systems compliant with standards like Defense Information Systems Agency (DISA) Security Technical Implementation Guides (STIGs) is a tedious and expensive process. Existing compliance automation tools can reduce some of this burden, but they depend on static, pre-written corrective actions. In this paper, we introduce SHIELDS, a multi-agent system that uses large language models (LLMs) to approach OS hardening as an iterative, feedback-driven process. Instead of applying fixed remediations, SHIELDS continuously proposes fixes and refines them based on feedback from target system execution and validation scans. We evaluate the system across multiple virtual machine configurations using six contemporary LLMs ranging from 20B to 400B parameters, and find that SHIELDS successfully remediates up...

论文介绍 安全配置错误是系统被攻陷的主要原因,而手动遵守STIG等安全标准过程繁琐。本文提出SHIELDS,一个利用LLM的多智能体系统,将操作系统加固视为一个迭代的、基于反馈的过程。该系统不应用固定的修复方案,而是持续提出修复建议,并根据目标系统的执行反馈和验证扫描结果进行优化。实验表明SHIDS能成功修复多种配置。

CRESS: Quantifying Vulnerabilities of Attack Scenarios in Hardware Reverse Engineering

第一作者: Alexander Hepp · 方向: 系统安全

Abstract:The safety, security, and reliability of microelectronic systems depend on a trustworthy, secured supply chain and design flow. Globally distributed supply chains or unintentional design weaknesses leave the door open for attacks on the hardware level. These scenarios encompass counterfeiting, hardware trojans, or on-device attacks. For these, hardware reverse engineering (RE) results play a pivotal role. The ongoing publication of new RE-involved attacks motivated the development of the common RE scoring system (CRESS). The system enables a general classification of RE-involved scenarios for a common, consistent rating. In this work, the originally qualitative system is extended to a quantitative system. We performed an extensive interview study with experts in the field. The interview results allowed us to derive weights that measure the severity of different RE-involved...

论文介绍 硬件逆向工程在应对硬件木马、假冒芯片等安全威胁中至关重要。现有的通用逆向工程评分系统CRESS最初是定性的。本文通过一项广泛的专家访谈研究,为该系统推导出权重,将其扩展为定量系统。该定量系统能够对涉及逆向工程的攻击场景进行一致且客观的风险严重性评级,为硬件安全评估提供了标准化工具。

Policy-Compliant Cloud Storage Systems

第一作者: Dimitrios Stavrakakis · 方向: 软件安全

Abstract:Privacy regulations such as the General Data Protection Regulation (GDPR) impose strict requirements on how personal data is stored, processed, and audited. While key-value stores (KVS) are widely used in latency-sensitive applications, their simple data model and untrusted cloud deployment environments make GDPR compliance particularly challenging. Existing approaches require invasive code modifications, impose high performance overheads, or overlook the integrity of compliance mechanisms themselves. This paper presents GDPRuler, a trusted middleware system that enables verifiable GDPR compliance for KVS on untrusted clouds without modifying their codebase. GDPRuler deploys a trusted GDPR monitor inside a Confidential Virtual Machine (CVM), which enforces GDPR policies, manages compliance metadata, and maintains tamper-evident audit logs. A declarative policy language...

论文介绍 该研究针对云存储中键值存储系统面临的GDPR合规挑战,提出GDPRuler系统。该系统在机密虚拟机中部署可信监控器,无需修改原有代码即可强制执行隐私策略、管理合规元数据并维护防篡改审计日志,为不可信云环境提供可验证的合规保障。

A formal framework for the economic security of DeFi compositions

第一作者: Massimo Bartoletti · 方向: AI 安全

Abstract:Decentralized Finance (DeFi) services are usually constructed by composing a variety of smart contracts. While composability is a key driver of the success of DeFi, it also creates security risks: adversaries may exploit interactions between newly deployed contracts and the pre-existing ones to inflict economic losses. We introduce MEV non-interference, a formal security notion for DeFi composability requiring that the maximal extractable value from a set of newly deployed contracts is not increased by interactions with the existing blockchain state. To support this notion, we define local MEV, a novel measure of economic attacks that focusses on the loss of a given set of victim contracts. We study two adversarial models, with bounded and unbounded wealth, and establish sufficient conditions and locality principles that enable modular reasoning about secure composability. We...

论文介绍 本文关注去中心化金融中智能合约组合带来的经济安全风险。引入MEV非干扰形式化概念,定义局部MEV以衡量特定合约的损失。研究有界和无界财富下的敌手模型,建立模块化推理的安全条件,提升DeFi组合的安全性分析。

Willing but Unable: Separating Refusal from Capability in Code LLMs via Abliteration

第一作者: Cristina Carleo · 方向: 软件安全

Abstract:Producing a labeled vulnerable code at scale is a recurring obstacle for learning-based vulnerability detection: mined corpora carry substantial label noise, and existing LLM-based augmentation propagates these inaccuracies because it transforms vulnerable seeds rather than synthesising vulnerabilities from a specification. A complementary route is to start from safe code and ask an instruction-tuned LLM to inject a specified CWE (which would shift the labeling burden from open-ended detection to bounded binary confirmation) but safety-aligned code LLMs systematically refuse such prompts. This paper is a preliminary feasibility study of abliteration, a low-rank weight edit that orthogonally projects out the refusal direction in the residual stream, as a tool to remove this barrier. We use Python and CWE-89 (SQL injection) as a case study, evaluating the Qwen2.5-Coder-Instruct...

论文介绍 该研究探讨安全对齐的代码大语言模型在生成漏洞代码时的拒绝问题。通过abliteration技术移除模型中的拒绝方向,以评估其注入指定漏洞的能力。以Python和CWE-89为案例,验证可行性,为漏洞数据合成提供新思路。

From Attack Simulation to SIEM Rule: Deterministic Detection-as-Code Synthesis with Probe-Level Traceability

第一作者: Alexandre Cristovão Maiorano · 方向: 软件安全

Abstract:Security teams routinely simulate attacks against their own systems to check whether their monitoring would catch a real intruder. These Breach-and-Attack-Simulation (BAS) tools surface findings, but the security information and event management (SIEM) systems that watch production need detection rules -- and today a human bridges that gap by hand, reading each finding and writing the corresponding Sigma rule (a vendor-neutral detection format). We show this translation can be partially automated when probes are drawn from a locked corpus, so each finding carries a stable identifier back to the originating probe. We describe a deterministic synthesis function that maps each finding to a starter Sigma rule through a small template library (N=23, indexed by categories from the OWASP LLM and Web Top 10), with a back-reference to the originating finding and its MITRE ATT&CK...

论文介绍 针对攻击模拟结果与SIEM检测规则之间手动转换的瓶颈,本文提出确定性合成方法。通过模板库将BAS工具的发现映射为Sigma规则,并引用MITRE ATT&CK,实现检测规则的自动化生成,提高安全运维效率。

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

第一作者: Yongjie Wang · 方向: AI 安全

Abstract:Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such agents may retrieve public benchmark metadata, question context, or even ground-truth answers via web search. This gives rise to Search-Time Contamination (STC), where external retrieval bypasses intended reasoning and inflates measured performance. We systematically study STC in deep research agent evaluation. We define three contamination types with increasing severity, namely Benchmark Metadata Leakage, Question-Context Leakage, and Explicit Answer Leakage, and develop detection algorithms to identify them and quantify their impact on agent performance. Evaluating modern deep research agents on six public benchmarks, we find that STC is widespread and can inflate performance by up to 4%. Our findings...

论文介绍 本文系统研究深度研究代理在公共基准测试中的搜索时间污染问题。定义三种污染类型并开发检测算法,评估发现STC普遍存在,可导致性能膨胀高达4%。强调基准评估的脆弱性,并提出改进方向。

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

第一作者: Nicholas Saban · 方向: 系统安全

Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster on retired models and on the most-vulnerable model in each paper's panel. We ask whether those techniques, reproduced as hand-crafted templates, still work against current frontier CUAs. We release CUA-HandCrafted, a public benchmark of 793 episodes spanning 24 multi-step web tasks, 56 attack templates, 8 attack families, and 4 system-prompt configurations. Against Claude Sonnet 4.6 and GPT-5.4 we measure 0/140 multi-step attack success (Clopper-Pearson 95% upper bound 2.60%); a prompt ablation shows this resistance lives in the model weights. Yet it does not generalize: on a sister coding-agent benchmark (SkillBench), the same weights fall to hand-crafted skill-injection at up to 100%. We argue that the literature's high ASR is largely...

论文介绍 该研究发布基准CUA-HandCrafted,评估前沿计算机使用代理对提示注入攻击的抵抗力。测试显示当前模型在浏览器任务中完全抵抗,但编码任务中易受攻击,揭示安全对齐的领域特异性,质疑高攻击成功率报告。

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

第一作者: Yifan Liao · 方向: AI 安全

Abstract:Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription. Their robustness to adversarial attacks has become an important topic for the community. Existing adversarial attacks directly add adversarial noise to the speech audio. However, prior work has shown that existing adversarial attacks face two limitations: they often transfer poorly to black-box ASR systems and are increasingly mitigated by defenses tailored to input-space perturbations. In this work, we propose a Clean-Referenced Feature-Vocoder Attack, a surrogate-based black-box attack that moves the adversarial search space from raw waveforms to self-supervised learning (SSL) representations. To address the transferability limitation, we perturb more generalizable acoustic-phonetic representations rather than low-level waveform samples, reducing dependence on...

论文介绍 针对语音识别系统对抗攻击转移性差的问题,本文提出特征-声码器攻击。将扰动空间从波形转移到自监督学习表示,提高对黑盒系统的攻击效果。该方法减少对输入空间扰动的依赖,增强攻击的鲁棒性。

Multi-Objective Submodular Maximization with Differential Privacy

第一作者: Ting Hou · 方向: 隐私保护

In this paper, we study multi-objective submodular maximization (MOSM) subject to a cardinality constraint under differential privacy (DP). Specifically, we aim to select a set of at most $k \in \mathbb{Z}_{+}$ elements to maximize the minimum of $d > 1$ monotone submodular functions while satisfying $\varepsilon$-DP. Although extensive studies have been conducted on both differentially private single-objective submodular maximization on sensitive data and non-private MOSM, to the best of our knowledge, there has not yet been any prior work on MOSM with DP. We propose two novel algorithms: the first extends the classic greedy algorithm and the second employs a truncation technique, both of which are integrated with DP mechanisms for privacy protection and achieve approximation guarantees for MOSM. Finally, we conduct numerical experiments on two submodular maximization applications...

论文介绍 本文研究差分隐私约束下的多目标子模最大化问题。目标是选择元素集以最大化多个单调子模函数的最小值。提出两种算法:扩展贪心和截断技术,集成差分隐私机制,并提供近似保证,应用于实际优化场景。

GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

第一作者: Paulo Ricardo Ferreira Neves · 方向: AI 安全

Abstract:Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attacks. In addition, benchmark evaluations may be affected by contamination and partial information leakage, compromising performance estimates. This work presents GuardNet, a guardrail system based on an ensemble of shallow neural networks (BiLSTMs) with approximately 47 million parameters. We investigate the hypothesis that robustness in adversarial scenarios depends more on the diversity of example coverage and threshold calibration than on model scale. The results indicate that GuardNet achieves competitive performance compared with lightweight detectors and high efficiency at low latency, although larger LLMs such as Mistral-7B and Llama-3.1-8B still achieve superior performance in terms of F1 score and AUROC on the blind...

论文介绍 该研究聚焦于大语言模型面临的提示注入与越狱攻击问题,并提出一种名为 GuardNet 的防御系统。该系统基于约4700万参数的 BiLSTM 集成模型构建,旨在验证其核心假设:在对抗场景下,系统的鲁棒性更多依赖于样本覆盖的多样性和阈值校准,而非模型规模。实验表明,该系统在低延迟下实现了具有竞争力的检测效率,为轻量化安全防线提供了可行方案。

On the Cryptographic Structure Required for Verifying Qubits

第一作者: James Bartusek · 方向: 密码学协议

Abstract:Classically testing for the presence of anti-commuting operators on a quantum device is a critical tool underpinning recent progress in classical verification of quantum computation. While such tests can be based on cryptographic assumptions, known constructions rely on highly structured assumptions, e.g. trapdoor claw-free functions. In this work, we seek to explain this state of affairs by constructing strong cryptography from (certain forms of) classical tests of anti-commutation. In particular, we formulate the notion of a test of non-commutation (ToNC), an interactive protocol between a quantum prover and classical verifier in which the prover's final-round response is obtained by measuring one of two binary observables $P_0,P_1$ depending on the verifier's challenge bit $c$. We prove that, for a broad range of parameters, ToNC implies classical-communication key...

论文介绍 本文探讨了在量子设备上测试反交换算子存在性所需的基础密码学结构。研究者提出了一个非对易性测试(ToNC)的形式化协议,其中量子证明者的响应取决于对某个二元可观测量的测量。他们证明,对于广泛的参数范围,该测试本身蕴含了构造经典通信密钥交换方案的能力,从而为已有高度结构化假设的验证方法提供了更一般的理论基础。

DP-MacAdam: Differentially Private Mechanism with Adaptive Clipping and Adaptive Momentum

第一作者: Naima Tasnim · 方向: AI 安全

Abstract:Differentially private stochastic gradient descent (DP-SGD) has become the standard framework for privacy-preserving machine learning, yet its reliance on a fixed gradient clipping threshold to limit sensitivity remains a significant practical limitation. Adaptive clipping algorithms such as AdaClip shift and scale the gradient prior to clipping and adding noise so that the clipped gradient yields a more informative descent direction. The shift and scaling parameters are selected adaptively based on the empirical mean and variance. However, in existing adaptive clipping algorithms, these empirical estimates have not been also used for momentum to accelerate training itself. On the other hand, DP-Adam is an algorithm that exploits Adam-like momentum updates based on the gradient mean and variance to accelerate training, but does not exploit these estimates for adaptive...

论文介绍 针对隐私保护机器学习中 DP-SGD 依赖固定梯度裁剪阈值的局限,本文提出了 DP-MacAdam 机制。该方法的核心创新在于将自适应裁剪中基于梯度均值和方差的估计值,同时用于梯度缩放与 Adam 风格的动量更新,从而在保持差分隐私的前提下,试图自适应地优化梯度方向并加速模型训练过程。

RadiusFPS: Efficient Farthest Point Sampling on CPUs and GPUs via Spherical Voxel Pruning

第一作者: Ziyang Yu · 方向: 导航与运动 · 来源: cs.RO

Point clouds are a primary sensory representation for robotic perception, underpinning LiDAR-based autonomous driving, simultaneous localization and mapping (SLAM), and navigation. Within these pipelines, Farthest Point Sampling (FPS) is the most well-known downsampling operator, as its uniform coverage preserves the geometric structure on which downstream perception relies. However, the large time complexity of classical FPS scales poorly with the million-point-per-second rates of modern 3D sensors, making it a dominant latency bottleneck that conflicts with the real-time and limited onboard compute budgets of robotic systems. Therefore, we propose RadiusFPS, an FPS acceleration framework based on spherical voxel pruning that preserves the standard FPS update rule under the same initialization and tie-breaking policy. By indexing the point cloud with spherical voxels, RadiusFPS...

论文介绍 在机器人感知中,最远点采样因其几何结构保持能力而被广泛使用,但其计算复杂度难以满足现代三维传感器的实时需求。本文提出 RadiusFPS 加速框架,通过球形体素对点云进行空间索引和剪枝,在保持标准 FPS 更新规则不变的前提下,显著降低计算开销。该方法适用于 CPU 和 GPU,旨在解决自动驾驶、SLAM 等场景中的实时性瓶颈。

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

第一作者: Qize Yu · 方向: VLA 通用模型 · 来源: cs.RO

Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robotic manipulation. However, the structural mismatch between VLM semantic spaces and embodied control policies often hinders the learning of precise perception--action mappings. To address this challenge, we propose \textbf{AffordanceVLA}, a unified framework that introduces structured affordance forecasting as a task-oriented intermediate representation to establish a more precise and robust perception--action mapping. Specifically, we progressively model manipulation priors through three complementary components: 1) \textbf{Which2Act} for object-centric grounding via visual latent prediction to suppress distractions; 2) \textbf{Where2Act} for 2D interaction localization via affordance map estimation; and 3) \textbf{How2Act} for 3D...

论文介绍 为解决视觉语言模型与机器人控制策略间的结构不匹配问题,本文提出 AffordanceVLA 框架。该框架引入结构化的可供性预测作为中间表示,以建立更精确的感知-动作映射。具体地,它通过 Which2Act、Where2Act 和 How2Act 三个渐进式组件,分别完成物体定位、交互区域预测和动作参数生成,从而增强视觉语言动作模型的指令跟随与操作能力。

WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

第一作者: Shengtao Zheng · 方向: VLA 通用模型 · 来源: cs.AI

End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observations to directly predict actions, often struggling in dense urban environments where severe occlusions and sharp turns result in drastic viewpoint transitions. We argue that the ability to "imagine" future states -- inherent in World Models -- is critical for robust decision-making under such partial observability. To address this, we construct a challenging Urban Canyon Traversal Benchmark, specifically designed to evaluate spatial understanding in scenarios characterized by severe occlusions and drastic viewpoint transitions. To this end, we propose WorldFly, a novel world-model-based VLA framework that employs a dual-branch coupled flow matching mechanism to jointly generate future video predictions and navigation actions, thereby...

论文介绍 针对现有视觉语言动作模型在复杂城市环境中因遮挡和视角突变导致决策困难的问题,本文构建了城市峡谷穿越基准,并提出 WorldFly 框架。其核心是利用世界模型“想象”未来状态的能力,通过双分支耦合流匹配机制,联合生成未来视频预测和导航动作,以增强在部分可观测场景下的鲁棒导航能力。

DexFuture: Hierarchical Future-State Visuomotor Targeting for Bimanual Dexterous Tool Use

第一作者: Runfa Blark Li · 方向: 机器人操作 · 来源: cs.RO

Bimanual dexterous tool use remains challenging for robots due to high-dimensional hand configurations and complex hand-tool-object dynamics and contact. Most existing control policies depend on future configuration references provided from demonstrations, while future action-conditioned world models require slow online planning over high-dimensional action sequences. A significant challenge is generating a dynamically consistent future reference trajectory without relying on privileged states from demonstrations or slow counterfactual planning. We propose DexFuture, a hierarchical system that couples a high-level Future-State Visuomotor Target Predictor with a low-level Target-Conditioned Structured Dexterous Policy. Conditioned on egocentric RGB, proprioceptive and geometric history, the high-level predictor constructs structured hand-tool-object visuomotor embeddings and uses a...

论文介绍 实现机器人双手灵巧工具使用面临高维配置与复杂动力学的挑战。DexFuture 提出一个分层系统,高层预测器根据视觉、本体感觉等历史信息,预测未来手-工具-物体的结构化视觉运动状态;低层控制器则基于该目标状态生成具体的灵巧操作策略。该方法试图在不依赖示教特权状态或缓慢规划的情况下,生成动态一致的未来参考轨迹。

VASO: Formally Verifiable Self-Evolving Skills for Physical AI Agents

第一作者: Yunhao Yang · 方向: 具身智能 · 来源: cs.RO

Reusable robot skills are becoming the basic units through which embodied agents turn open-ended instructions into long-horizon physical behavior. We argue that, while foundation models have collapsed the cost of creating these skills, the cost of trusting them has not. Existing skill-evolution loops refine skills through execution feedback, unit tests, environment reward, or LLM self-critique, but these signals provide only trace-level evidence: they show that a skill worked on sampled executions, not that skill-induced plans satisfy temporal safety contracts under untested conditions. We introduce VASO, a framework for verification-guided self-evolution of LLM-generated robot skill contracts. In VASO, each skill is represented as a semantic contract with two coupled interfaces: a formal interface that aligns robot states, observations, and control commands with logical propositions...

论文介绍 针对由大语言模型生成的可复用机器人技能缺乏可信保障的问题,本文提出 VASO 框架。该框架将技能表示为带有语义和形式化双接口的契约,并通过形式化验证引导技能的自演化过程。它旨在超越基于执行轨迹的启发式反馈,确保技能诱导的规划在未测试条件下仍能满足时间安全性合约,从而提升物理智能体长时程任务的可靠性。

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

第一作者: Lizhi Yang · 方向: 机器人操作 · 来源: cs.RO

Abstract:For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is crucial. Existing whole-body controllers typically demand dense kinematic or spatial references that planners struggle to synthesize from task semantics. We instead propose a compact, explicit interface that is intuitive, general, modular, and expressive enough for diverse manipulation skills. To this end, we introduce HANDOFF, a single humanoid whole-body controller that follows this interface and is distilled via multi-teacher KL distillation under a context-conditioned gating scheme into a mixture-of-experts student from three complementary specialists: whole-body motion tracking with safety-filtered data, locomotion, and fall-recovery. On the Unitree G1, HANDOFF matches state-of-the-art velocity tracking and offers one of...

论文介绍 该研究针对人形机器人全身控制中任务规划与底层控制接口复杂的问题,提出了HANDOFF系统。其核心是定义一个紧凑、直观的任务空间接口,并采用多教师知识蒸馏方法,将运动跟踪、行走和恢复三个专家模型的能力整合到一个学生模型中。该系统在Unitree G1机器人上展示了与先进方法相当的速度跟踪能力,为开发通用、模块化的人形机器人控制器提供了新思路。

TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

第一作者: Dong Jing · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VLAs) only inherit a single fixed speed from training demonstrations. Prior efforts to accelerate VLAs through model compression, KV-cache reuse, or reinforcement learning only shift the policy from one fixed speed to another, and leave deceleration almost unexplored. We observe that the magnitude of each predicted action already governs how fast the robot moves, opening a direct route to controllable execution speed. We turn this observation into TempoVLA, a single VLA whose execution speed is controlled by an explicit condition. TempoVLA combines two coupled components. (1) A data-side Variable-Speed Trajectory Augmentation (VSTA) that re-times demonstration to any target speed by...

论文介绍 现有视觉-语言-动作(VLA)模型通常以固定速度执行任务。本文提出的TempoVLA,旨在实现对VLA执行速度的灵活控制。其核心是通过在训练数据中引入可变速度轨迹增强,并结合一个以速度为条件的反射器,使模型能够根据外部输入的显式条件调节其输出动作幅度,从而控制机器人执行速度的快慢。这为需要动态调整操作节奏的任务提供了新方案。

Flow-based Policy Adaptation without Policy Updates

第一作者: Luzhe Sun · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Leveraging prior knowledge from pretrained policies, foundation models, or human operators offers an efficient alternative to learning robot skills from scratch. However, these agents often provide actions that are suboptimal, noisy, or misaligned with task-specific expert behavior. We propose GLOVES, a family of flow-based adaptation methods that correct non-expert actions by transporting them toward an expert action distribution. Rather than replacing agentic control with full autonomy, GLOVES performs selective action-level adaptation, improving task success while preserving agent intent. The learned flow also provides a natural in-distribution scoring mechanism through reverse flow evaluation. We use this signal as an intervention gate: actions that appear consistent with the expert distribution are passed through unchanged, while anomalous or out-of-distribution (OOD)...

论文介绍 本文提出GLOVES,一种基于流匹配的策略适应方法,用于纠正来自预训练模型或人类操作员的非最优、噪声动作。该方法通过学习一个将错误动作分布映射到专家动作分布的流模型,对动作进行选择性修正,而非完全替代原始智能体的意图。它利用反向流评估作为干预门控,仅在检测到动作异常时进行纠正,从而在提升任务成功率的同时保留了原始智能体的控制意图。

VOLT: Vision and Language Trajectory Segmentation for Faster-than-Demonstration Policies

第一作者: Robert Ramirez Sanchez · 方向: 机器人操作 · 来源: cs.RO

Abstract:Humans often take longer to demonstrate a task than a robot would need to execute it. Rather than learning to replicate the demonstration at the same pace, many industrial and practical applications require robots to perform tasks as quickly as possible. In this paper, we investigate several hypotheses for learning policies that operate faster-than-demonstrations. Our experiments show that the most effective strategy is to downsample recorded demonstrations and train the robot's policy on this accelerated data. However, uniformly downsampling an entire trajectory can be problematic. Some parts of a task can be safely sped up (e.g., unconstrained motion), while others demand slower, more precise motion (e.g., object interactions or fine manipulation). To address this challenge, we introduce VOLT, a vision-and-language trajectory segmentation method that reasons over video...

论文介绍 本文探讨如何让机器人学习以比演示更快的速度执行任务。研究发现,对演示轨迹进行均匀降采样训练是有效策略,但可能忽略任务不同阶段对速度的需求差异。为此,作者提出了VOLT方法,利用视觉和语言信息对轨迹进行语义分割,区分出可加速的自由运动段和需要缓慢精确操作的接触交互段,从而实现更合理、安全的加速策略学习。

Meridian: Metric-Semantic Primitive Matching for Cross-View Geo-Localization Beyond Urban Environments

第一作者: Mason Peterson · 方向: 具身智能 · 来源: cs.RO

Abstract:Successful robot automation requires accurate global localization to support repeatability, task planning, goal specification, and safe operation. However, reliable localization in GNSS-denied environments remains an open problem. Overhead aerial imagery offers a promising solution, but existing approaches primarily target structured urban environments and have been rarely demonstrated in unstructured natural terrain. Limitations of the state-of-the-art include a reliance on models trained for specific environments, as well as difficulty handling repetitive geometries and featureless landscapes commonly found in natural outdoor areas. To overcome these challenges, we present Meridian, a method for matching high-level metric-semantic primitives across aerial images and ground robot RGB-D camera data that achieves accurate global localization and generalizes well across diverse...

论文介绍 在GNSS信号缺失的环境中,机器人可靠的全局定位是一个开放问题。本文提出Meridian方法,用于解决非结构化户外环境(如自然地形)下的跨视图地理定位问题。该方法通过将空中图像与机器人地面RGB-D相机数据中的高层度量-语义基元进行匹配来实现定位。其核心在于构建和匹配度量与语义信息,以克服传统方法对特定环境模型依赖、难以处理重复几何和无特征地貌的局限。

Attitude-Aided Linear Calibration of Triaxial Accelerometers

第一作者: Yongqiang Yu · 方向: 导航与运动 · 来源: cs.RO

Abstract:Triaxial MEMS accelerometers are widely used for inertial sensing, navigation, and sensor fusion, but existing calibration methods often rely on costly reference setups or nonlinear iterative optimization, limiting their efficiency and applicability to low-cost or self-calibrating systems. We present attitude-aided linear accelerometer calibration (ALAC), a method that operates on any platform providing orientation information, such as turntables, robotic arms, or inertial measurement units. ALAC constructs a combined error matrix (CEM) to represent sensor errors in a unified calibration model and enables linear least-squares estimation. The bias and gravity vector are jointly estimated, implicitly accounting for platform misalignment, and matrix decomposition of the CEM recovers scale, non-orthogonality, and alignment rotation parameters. Under static gravity, calibration is...

论文介绍 本文提出一种姿态辅助的线性加速度计校准方法ALAC,旨在提高MEMS三轴加速度计的校准效率和通用性。该方法利用平台(如转台、机械臂或IMU自身)提供的姿态信息,构建统一的误差模型,从而将校准问题转化为线性最小二乘问题。它能联合估计偏置、比例因子、非正交误差和对齐误差,适用于低成本或自校准系统,简化了传统依赖外部精密设备或非线性优化的流程。

Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation

第一作者: Rickmer Krohn · 方向: 机器人操作 · 来源: cs.RO

Abstract:Touch sensing is beneficial for solving a wide variety of manipulation tasks. While there exists a wide range of tactile sensors with different properties, exploiting the fusion of multiple heterogeneous tactile sensors to improve manipulation learning remains underexplored. We present Multi-Resolution Tactile Sensing (MiTaS), a representation framework that leverages multiple tactile sensors operating at different temporal resolutions in order to solve complex contact-rich manipulation tasks. We propose a novel architecture using modality-specific convolutional stems and transformer-based fusion that effectively fuses information from an RGB camera stream, a vision-based GelSight Mini sensor and a high-frequency event-based Evetac sensor. This multi-sensor representation then conditions a flow-matching policy for solving downstream tasks. Experimental results across five...

论文介绍 本文针对接触丰富的机器人操作任务,提出了MiTaS(多分辨率触觉模仿学习)框架。该框架的核心是融合来自不同触觉传感器的多模态、多分辨率信息,包括视觉、基于视觉的GelSight传感器和高频率事件相机Evetac。通过设计特定模态的卷积提取器和基于Transformer的融合模块,有效整合异构触觉数据,并将其用于条件化流匹配策略,以提升复杂操作任务的性能。

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action

第一作者: Boyang Zhang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) policies remain brittle in long-horizon and high-uncertainty control, where one-pass action decoding provides limited inference-time deliberation. Explicit chain-of-thought can increase reasoning depth, but introduces token latency and an indirect text-to-action interface. We propose MPCoT, a reward-guided multi-path latent reasoning framework that initializes $M$ hypotheses, refines them for K weight-tied steps, and softly aggregates them before action decoding. A training-only path-preference objective evaluates candidate action branches with expert-action consistency, world-model/VLM-based progress, and success feedback to align the latent path scorer with downstream execution quality. MPCoT preserves the original 8-step action interface, generates zero reasoning tokens, and exposes configurable inference controls (K,M). Under matched protocols...

论文介绍 为提升视觉-语言-动作(VLA)模型在长时程、高不确定性任务中的推理能力,本文提出MPCoT框架。该框架在潜空间内进行多路径推理:初始化多个动作假设,经过多步迭代细化,并通过奖励引导的方式进行软聚合,最终生成动作。它避免了生成显式文本推理的延迟,并在训练时通过路径偏好目标对齐潜空间评分与任务执行质量,从而在保持原有动作接口的同时,增强了模型的决策深度和稳健性。

TAM: Torque Adaptation Module for Robust Motion Transfer in Manipulation

第一作者: Dongwon Son · 方向: 机器人操作 · 来源: cs.RO

Abstract:A policy tuned for one robot often behaves differently on another, whether due to the sim-to-real gap, unknown payloads, or the differing dynamics of two instances of the same robot. In contact-rich, dynamic manipulation, even small motion discrepancies can result in failure to track reference motion, since they disrupt the timing and modes of contact. Common remedies, such as domain randomization or system identification, either produce overly conservative task policies or require data that must be recollected for each robot or payload. We introduce the Torque Adaptation Module (TAM), a learned module that adapts the torque commands sent to the robot to match the behavior of an ideal robot. TAM operates between the low-level controller that tracks the policy's actions and the robot's torque interface. It includes a history encoder that embeds proprioceptive history into a...

论文介绍 针对不同机器人或负载下策略执行效果差异大的问题,该研究提出一种力矩自适应模块TAM。该模块通过学习历史本体感觉数据,自适应调整发送给机器人的力矩指令,以匹配理想机器人的行为。它被部署在低级控制器与机器人接口之间,旨在提升接触丰富的动态操作任务中轨迹跟踪的鲁棒性,避免因运动差异导致的操作失败。

ActiveMimic: Egocentric Video Pretraining with Active Perception

第一作者: Xingyao Lin · 方向: 机器人操作 · 来源: cs.RO

Abstract:Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform those pretrained on robot data. We attribute this gap to a missing signal, the active perception behavior in egocentric videos, where humans continuously reposition their viewpoint during manipulation, inducing camera motion that standard pipelines treat as noise. To address this, we present ActiveMimic, a pretraining framework that recovers synchronized camera and wrist trajectories from a single body-worn RGB camera, models camera motion as a viewpoint action, and jointly learns active perception and manipulation from in-the-wild egocentric human video before adapting to a target robot. Empirically, real-world experiments across tasks with diverse active perception demands show that ActiveMimic consistently surpasses baselines...

论文介绍 本研究旨在解决仅从自我中心人类视频预训练的机器人模型性能不佳的问题。作者指出,问题根源在于忽略了视频中人类主动调整视角的主动感知行为。他们提出了ActiveMimic框架,能从单视角视频中恢复同步的相机与手腕轨迹,将相机运动建模为视角动作,并在野外视频中联合学习主动感知与操作技能,随后适配至目标机器人。

Towards Realistic 3D Sonar Simulation

第一作者: Youssef Attia · 方向: 导航与运动 · 来源: cs.RO

Abstract:As underwater robotics research increasingly addresses complex 3D perception and autonomous navigation, the fidelity of sonar simulation has become a key factor in algorithm development. Current simulation frameworks typically rely on geometry-driven rendering, approximating 3D sonar as an underwater equivalent to LiDAR, which fails to account for fundamental acoustic phenomena such as refraction, multi-path interference, and phase-dependent signal formation. This paper proposes a modular architecture for realistic 3D sonar simulation that integrates GPU-accelerated graphics engines with physically grounded acoustic propagation principles. We implement a volumetric 3D sonar model within the NVIDIA Isaac Sim environment, modeled after the Water Linked 3D-15 sensor, and integrate it into a comprehensive underwater simulation framework. The system is validated through a...

论文介绍 随着水下机器人研究深入,对高保真3D声纳仿真的需求日益迫切。现有方法多基于几何近似,未能反映折射、多径干扰等关键声学现象。本文提出一种模块化仿真架构,将GPU加速图形引擎与物理声学传播原理相结合,在NVIDIA Isaac Sim中实现了基于真实传感器(如Water Linked 3D-15)的体积化3D声纳仿真模型,并构建了完整的水下仿真框架。

A Conversational Framework for Human-Robot Collaborative Manipulation with Distributed Generative AI models

第一作者: Arash Ghasemzadeh Kakroudi · 方向: 机器人操作 · 来源: cs.RO

Abstract:This paper presents a distributed conversational framework for human-robot collaborative manipulation that integrates local language and vision-language models (VLMs) with a Robot Operating System 2 (ROS 2)-based execution stack. Language understanding, visual grounding, orchestration, and motion execution run as separate ROS 2 nodes, enabling flexible deployment across distributed hardware while maintaining a responsive control loop. From free-form user commands, the system generates structured action requests for pick, place, and handover. It uses a VLM to return image-space targets, which are converted into metric robot-frame goals using depth and calibration. A web dashboard exposes intermediate intent and grounding overlays (pixel, depth, and robot-frame) and requires explicit operator confirmation before any motion is executed. Experiments on a Franka FR3 platform...

论文介绍 本文提出一种用于人机协作操作的分布式会话框架,它整合了本地语言模型与视觉语言模型(VLM),并与基于ROS 2的执行栈相结合。系统能理解用户的自由形式指令,生成结构化的抓取、放置与传递动作请求,并利用VLM进行视觉定位,将图像目标转换为公制机器人坐标。该框架通过分离的ROS 2节点实现分布式部署,支持人机交互式操作。

L-SDPPO: Policy Optimization of Spiking Diffusion Policy for Intra-vehicular Robotic Manipulation

第一作者: Liwen Zhang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Intra-vehicular robots in spacecraft help reduce astronaut workload and improve mission efficiency. Recent research focuses on using deep learning methods to achieve the acute control required for operations in these complex environments. However, objects exhibit unpredictable, unconstrained drift without gravitational damping. These factors demand robustness against complex multimodal action distributions. Diffusion policies (DP) can model these complex actions, but their iterative sampling process consumes too much energy for the limited power budgets of spacecraft. We therefore propose a low-energy intra-vehicular robotic manipulation framework, L-SDPPO, in which the Spiking Diffusion Policy (SDP) is optimized with a reinforcement learning (RL) algorithm. Furthermore, to address the insufficient perception of dynamic spatiotemporal features in microgravity, we propose the...

论文介绍 本研究聚焦航天器内机器人操作,在微重力环境中物体无约束漂移,控制需应对复杂多模态动作分布。扩散策略(DP)能建模复杂动作,但迭代采样过程能耗高,不适合航天器有限功率。为此提出低能耗框架 L-SDPPO,通过强化学习优化脉冲扩散策略(SDP),以降低能耗,并改进对动态时空特征的感知。该方法可应用于太空任务,减轻宇航员工作负荷,提高操作效率。

Sample-efficient Low-level Motion Planning for Robotic Manipulation Tasks via Zero-shot Transfer Learning

第一作者: Yuanzhi He · 方向: 机器人操作 · 来源: cs.RO

Abstract:As robotic systems become more sophisticated, the growing complexity of their motion planning models and the longer training times pose substantial challenges. Evolutionary algorithms such as the Sample-efficient Cross-Entropy Method (iCEM) have recently demonstrated promising potential for low-level real-time planning by leveraging efficient knowledge reuse strategies to improve performance. Although effective in many control tasks, iCEM's performance can be constrained in more complex scenarios, particularly those requiring stacking, sliding, and shelf placement. In this work, we propose a novel iCEM+TL framework that explicitly leverages Transfer Learning (TL), where key iCEM parameters are transferred from simpler upstream tasks to guide more complex downstream tasks. Additionally, we applied Reward Redesign (RR) through task decomposition for stacking objects and shelf...

论文介绍 针对机器人复杂任务运动规划模型训练耗时长的问题,本研究提出了一种基于迁移学习的样本高效低级运动规划框架。该方法显式地利用迁移学习策略,将关键参数从简单的上游任务迁移到更复杂的下游任务(如堆叠、滑动、放置)中进行引导。同时,通过任务分解进行奖励重塑,以提升规划性能,减少样本需求。

Gotta Grow Fast: Design and Benchmarking of a Tip Mount for High-Speed Vine Robots

第一作者: Antonio Alvarez Valdivia · 方向: 导航与运动 · 来源: cs.RO

Abstract:Soft, growing vine robots extend through tip eversion, a mechanism that enables navigation through cluttered environments. However, integrating cameras and other sensors at the tip is uniquely challenging because the material forming the tip is constantly renewed as the robot grows. This continual material turnover, combined with friction between internal layers, added tip weight, and fabric constriction, complicates sensor and tool mounting. These limitations hinder the deployment of vine robots for inspection and search tasks, where rapid growth while carrying tip-mounted sensors is essential. In this work, we present a triangular roller tip mount that reduces internal resistance during growth by rolling rather than sliding against the robot body. The design was refined through iterative failure analysis, enabling, for the first time, consistent eversion on a TPU-coated...

论文介绍 软体藤蔓机器人通过尖端外翻生长,利于在杂乱环境中导航,但在快速生长时于尖端集成传感器极具挑战。本文设计了一种三角滚动顶端安装装置,通过滚动而非滑动来减少生长时的内部阻力。经过迭代失败分析改进后,该设计首次实现了在TPU涂覆织物机器人上的稳定外翻,并建立了高速生长携带传感器的基准测试。

RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning

第一作者: Chaoyi Xu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Learning dexterous manipulation requires demonstrations that preserve fine hand-object interactions while remaining executable at deployment. Existing pipelines either lose deployable dexterity through retargeting or embodiment conversion, or rely on robot-specific teleoperation that is costly to scale and often lacks intuitive, contact-aware control for dexterous data collection. We present RealDexUMI, a wearable universal manipulation interface built around a shared dexterous end-effector module that integrates a lightweight dexterous hand, in-hand vision, and fingertip tactile sensing. A palm-side isomorphic teleoperation glove maps human finger inputs to robot-hand joint commands, enabling real-time, retargeting-free, intuitive, and precise hand control. The shared hand and sensing modules yield zero-gap end-effector data, with matched in-hand observations, tactile...

论文介绍 为获取可部署的灵巧操作演示数据,本文提出了穿戴式通用操作接口RealDexUMI。该系统基于一个共享的灵巧末端执行器模块,集成了轻巧灵巧手、手内视觉与指尖触觉传感。操作者通过手掌侧的同构遥操作手套,能将手指动作直接映射为机器人手关节指令,实现无需重定向、直观且精确的实时控制,从而采集到零间隙的末端执行器数据。

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

第一作者: Yi Yang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict textual subtasks, subgoal images, and robot actions, conjoining the \emph{world modeling interface} to learn from extensive egocentric videos as in the world-action model (WAM) and the \emph{language reasoning} capacities to solve complex long-horizon tasks as in vision-language-action (VLA) models. At the core of WLA lies an \emph{autoregressive (AR)} Transformer backbone, instead of a bidirectional diffusion Transformer as in WAMs, to predict the \emph{next state}, comprising the \emph{semantic-level} textual intention and complementary \emph{fine-grained} physical dynamics. The physical dynamics are supervised by the world modeling objective based on a dedicated World Expert, and are leveraged to...

论文介绍 本文提出了一种新的具身基础模型「世界-语言-动作模型」。该模型以文本指令、图像和机器人状态为输入,联合预测文本子任务、子目标图像和机器人动作。它融合了世界建模和语言推理能力,其核心是一个自回归Transformer骨干,用于预测包含语义意图和物理动力学的下一状态。这旨在解决复杂的长期任务,并有望提升机器人执行长序列任务的能力。

Towards a Data Flywheel for Embodied Intelligence in Logistics

第一作者: Anlan Yu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Embodied intelligence is moving from laboratory demonstrations toward industrial deployment, with the logistics industry serving as a key application scenario. Learning-based policies offer a promising path beyond traditional perception-planning-control pipelines, but their scalability depends on how embodied data can be collected, organized, and reused. This research studies a data-centric framework for industrial embodied intelligence by constructing a logistics data flywheel. Our framework converts daily operations into reusable data assets, uses World Models to generate reliable supervision for long-tail parcel manipulation, and feeds deployment feedback back into policy improvement. As an initial result, \textit{WM-DAgger} introduces a World-Model-based data aggregation framework that synthesizes out-of-distribution recovery data for robust imitation learning. Building on...

论文介绍 该研究面向物流场景,提出了一种以数据为中心的具身智能框架。其核心是构建一个数据飞轮,将日常运营转化为可重用的数据资产,并利用世界模型为长尾物体操作生成可靠的监督数据。初步成果提出了「WM-DAgger」框架,通过世界模型合成分布外数据,以增强模仿学习的鲁棒性,旨在解决学习型策略的规模化部署问题。

Learning of Robot Safety Policies via Adversarial Synthetic Scenarios

第一作者: Nikolai Dorofeev · 方向: 具身智能 · 来源: cs.RO

Abstract:In this work, we propose an agentic gamification framework for hazard-informed learning of robot safety policies through synthetic scenarios. We model scenario generation as an adversarial game between two agents: a Red Team that explores the space of potential failures by constructing hazardous situations, and a Blue Team that incrementally refines safety policies to prevent them. This iterative process enables efficient discovery of high-risk edge cases that are unlikely to be captured through random simulation or manual enumeration. By combining classical risk modeling with adversarial scenario generation and modern learning paradigms, this work provides a scalable pathway for embedding safety into Physical AI systems operating in complex real-world environments. The paper describes ongoing work. The contribution is a problem formulation and a proposed solution architecture.

论文介绍 本文提出了一种基于智能体博弈化框架的方法,通过合成场景来学习机器人的安全策略。该框架将场景生成建模为红蓝两个团队的对抗博弈:红队探索潜在故障空间构建危险场景,蓝队则逐步优化安全策略以防患于未然。该方法旨在高效发现随机仿真或手动枚举难以涵盖的高风险边缘案例,为复杂环境中物理人工智能系统的安全嵌入提供可扩展的路径。

LadderMan: Learning Humanoid Perceptive Ladder Climbing

第一作者: Siheng Zhao · 方向: 机器人操作 · 来源: cs.RO

Abstract:Humanoid robots hold great promise for operating in human-centered environments, yet ladder climbing remains one of the most challenging tasks due to sparse footholds and handholds, complex whole-body coordination, and sensitivity to perception and control errors. We present \textbf{LadderMan}, a unified system that enables humanoid robots to robustly climb diverse ladders and perform manipulation under such constrained conditions. Our climbing policy is built on a scalable two-stage learning pipeline, where we use hybrid motion tracking to learn multiple climbing experts from a single reference motion, and distill these experts into a unified depth-based visuomotor climbing policy via hybrid imitation and reinforcement learning. To enable real-world deployment, we leverage vision foundation models to bridge the sim-to-real gap in depth perception. Building on the learned...

论文介绍 针对人形机器人梯子攀爬这一全身协调要求高且感知控制敏感的挑战,本文提出了「LadderMan」系统。该系统采用两阶段学习流程:先从单一参考动作通过混合运动跟踪学习多个攀爬专家,再通过混合模仿与强化学习将其蒸馏为统一的、基于深度视觉的运动控制策略。该方法旨在实现对多样化梯子的稳健攀爬,并利用视觉基础模型弥合仿真与现实之间的感知差距。

Visuotactile and Explicitly Force-Controlled Robotic Ultrasound for Abdominal Volumetric Reconstruction

第一作者: Adrian Piedra · 方向: 具身智能 · 来源: cs.RO

Abstract:In this paper, we present a robotic ultrasound acquisition system that integrates stereo vision, touch-based feedback, and expert-informed strategies to perform autonomous and adaptive abdominal scans. The system records freehand motion and force data from expert radiologists, creating a framework to capture transducer motion, applied forces, and anatomical scanning strategies. This expert data is replayed to replicate characteristic scans with the robot, forming a foundation for further autonomous capabilities. Using stereo vision, the system generates three-dimensional topography maps of the patient's abdomen, which are refined through stiffness measurements at key points to delineate the rib cage boundary. These combined techniques enable the robot to execute two distinct scanning paths: an upward-angled sweep beneath the rib cage to visualize structures near the upper...

论文介绍 本文提出一个集成视觉、触觉反馈和专家策略的机器人超声波采集系统,用于自主执行和自适应腹部扫描。系统记录专家放射科医生的手部运动和力数据,构建了一个捕捉探头运动、施加力和解剖扫描策略的框架。通过立体视觉生成患者腹部的三维地形图,并结合刚度测量来界定胸廓边界,使机器人能够执行多种扫描路径,进行腹部体积重建。

PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation

第一作者: Chong Ma · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) policies operate in a closed loop in real-world robot tasks: a robot observes the scene, executes an action chunk, and conditions its next decision on the resulting observation. However, most existing world models for robot action evaluation are limited to open-loop prediction along pre-collected action trajectories. This prevents them from supporting closed-loop VLA evaluation, where each action chunk must be conditioned on the observation generated by the previous execution. To address this gap, we propose PiL-World, a chunk-wise world model designed for policy-in-the-loop VLA evaluation. Given the current observation and the action trajectory rolled out by a VLA policy, PiL-World generates multi-view future observations that are consistent with the VLA rollout and match the image inputs required by the policy. By alternating between VLA...

论文介绍 现有用于评估机器人动作的世界模型大多局限于开环预测,无法支持闭环视觉-语言-动作策略的评估。为此,本文提出「PiL-World」,一个为策略在环评估设计的块状世界模型。给定当前观察和策略生成的动作轨迹,它能生成与策略卷出结果一致、并满足策略所需图像输入的多视角未来观察,从而实现VLA策略在环的交替评估,弥补了现有评估方法的关键缺陷。

Dynamic Multi-Agent Pickup and Delivery in Robotic Cellular Warehousing Systems

第一作者: Cheng Ren · 方向: 具身智能 · 来源: cs.RO

Abstract:Robotic Cellular Warehousing Systems (RCWS) give rise to multi-agent pickup and delivery (MAPD) processes in which robots sequentially collect multiple stock-keeping units (SKUs) for each order. Unlike classical MAPD formulations that assume static tasks, real warehouse operations often involve dynamic order evolution, where new SKUs may be appended to an order while it is being executed. Motivated by this practical requirement, this letter formulates the Dynamic Multi-Agent Pickup and Delivery problem considering internal order evolution for the first time. Building on the token passing paradigm, we propose two event-triggered online replanning algorithms. The first, Dynamic Token Passing, performs localized replanning upon order updates through add-order decomposition and priority-based token scheduling while preserving collision-free execution. The second, Cooperative Token...

论文介绍 在机器人单元仓储系统中,订单可能在执行过程中动态演变,新SKU可能被追加。本文首次形式化了这种考虑内部订单演变的动态多智能体拾取配送问题。基于令牌传递范式,提出了两种事件触发的在线重规划算法:动态令牌传递通过订单分解和优先级令牌调度进行局部重规划;协同令牌传递则在智能体间进行协调重规划。两者均旨在保证无冲突执行的同时处理动态变化。

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation

第一作者: Dabin Kim · 方向: 机器人操作 · 来源: cs.RO

Abstract:Embodied AI systems are increasingly expected to reason and act over extended horizons in physical environments. This growing capability brings safety to the foreground, because failures in the physical world can harm people, damage objects, and disrupt workplaces. Although safe embodied AI has attracted substantial attention, the literature remains fragmented across planning, policy design, and runtime execution. Long-horizon robotic manipulation is a particularly revealing anchor domain for this problem because semantic misgrounding, subtask-level error propagation, execution drift, and contact-rich physical risk can accumulate within the same closed-loop system. This survey therefore provides a structured review of safety in long-horizon robotic manipulation from an embodied AI perspective. We organize the literature by intervention locus, covering planning-time...

论文介绍 随着具身AI系统需要在物理环境中进行长时程推理与行动,安全问题日益凸显。本文针对长时程机器人操作这一特定领域,从具身AI视角对安全性进行了结构化综述。它指出该问题中的语义误解、子任务错误传播、执行漂移和接触性物理风险会累积。文章按干预位置(规划时、策略设计时、运行时执行)组织现有文献,旨在提供一个跨规划、策略与执行的统一安全分析框架。

Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning

第一作者: Ziyang Yao · 方向: 策略学习 · 来源: cs.RO

Abstract:Autonomous driving requires reasoning about how ego actions shape the evolution of the surrounding world. However, most end-to-end methods rely on direct state-to-action mappings, capturing correlations without explicitly modeling action-conditioned dynamics. Conversely, continuous-latent world models often lack compositional structure for causal reasoning across counterfactual futures. We introduce Discrete-WAM, a unified latent vision-action world policy that represents future visual states and ego actions as aligned discrete tokens, enabling compositional causal reasoning across alternative futures. Built upon this unified discrete alignment, Discrete-WAM establishes a shared discrete diffusion framework with unified generative tasks, jointly formulating world modeling, world-action policy, and hierarchical decision-enabled policy, supporting compositional generalization...

论文介绍 自动驾驶需要建模自车行为如何影响环境演变。本文提出Discrete-WAM,一种统一的潜在视觉-动作世界策略,通过将未来视觉状态和自车动作对齐为离散标记,实现对不同未来场景的组合式因果推理。该方法构建了统一的离散扩散框架,能联合处理世界建模、动作策略和分层决策,支持组合泛化。

Learning Contact Representation for Leg Odometry

第一作者: Emre Girgin · 方向: 具身智能 · 来源: cs.RO

Abstract:The estimation of odometry in legged robots depends on the assumption that the velocity of the foot with respect to the world remains zero during the stance phase. Feedback for the main body velocity is derived from the kinematic serial chain of the feet making accurate leg phase detection is a critical subproblem. A considerable number of studies employ ground reaction force sensors mounted at the tip of the foot to classify, yet these sensors may not be universally available for all legged robots. Additionally, these sensors are often unresponsive to unaccounted disturbances, such as slippage, while the foot remains in contact with the ground. In this study, we propose a self-supervised representation learning framework for contact detection that utilizes the standard sensor set of joint encoders without reliance on force sensor augmentations. We employ learned...

论文介绍 腿式机器人的里程计估计依赖于支撑相中足端速度为零的假设,因此准确的步态相位检测至关重要。许多方法依赖于地面反作用力传感器,但这些传感器并非普遍可用且对打滑等扰动不敏感。本文提出一种自监督的接触表示学习框架,仅使用关节编码器等标准传感器,无需力传感器,从而学习用于接触检测的鲁棒特征。

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization

第一作者: Yihao Wu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Post-training Vision-Language-Action (VLA) models into policies that can be reliably deployed on real robots remains a major bottleneck. SFT and DAgger exploit failure signals only indirectly, and reward-based RL is bottlenecked by the difficulty of real-world reward design and of training reliable critics. We present FlowPRO, a reward-free offline reinforced fine-tuning framework for flow-matching VLAs. Algorithmically, we propose RPRO (Robotic Flow-matching Proximalized Preference Optimization), a preference-optimization objective tailored to the flow-matching action head of VLA models. RPRO pairs a contrastive optimizer with an explicit proximal regularizer that anchors the absolute magnitude of the implicit reward, thereby eliminating the reward-hacking failure mode of plain Flow-DPO. On the data side, a teleoperated intervention-and-rollback paradigm produces naturally...

论文介绍 将视觉-语言-动作模型部署为可靠机器人策略的后训练是瓶颈。本文提出FlowPRO,一种针对流匹配VLA的无奖励离线强化微调框架。其核心是RPRO算法,一个为流匹配动作头定制的偏好优化目标,通过对比优化器与显式近端正则器结合,消除了奖励欺骗问题。数据方面采用遥操作干预与回滚范式。

Learning from Demonstrations over Riemannian Manifolds using Neural ODEs: An Extended Abstract

第一作者: Diana Cuervo Espinosa · 方向: 具身智能 · 来源: cs.RO

Abstract:Learning from demonstratins (LfD) is usually performed over Euclidean spaces, while the robot state, e.g. orientation, naturally evolves over curved spaces. Therefore, to ensure natural, complex motion generation, we investigate learning from demonstrations over Riemannian manifolds that are capable of encoding both position and orientation data. Here, geodesic paths provide for natural motion between two arbitrary points within the manifold. We propose to numerically estimate geodesics via neural ordinary differential equations, mitigating large computational overhead of existing approaches. Finally, these geodesics can be decoded back into the original task space before deploying on the robot. In this extended abstract, we discuss the architecture of our framework, provide some initial insights from our simulation experiments, including comparison to other geodesic...

论文介绍 传统学习示范通常在欧氏空间进行,但机器人姿态等状态自然演化于曲面空间。本文研究在能够编码位置和朝向数据的黎曼流形上进行示范学习。流形上的测地线路径能生成自然的运动。作者提出利用神经常微分方程来数值估计测地线,以降低现有方法的计算开销,最终将测地线解码回任务空间供机器人执行。

MoDex: A Diffusion Policy for Sequential Multi-Object Dexterous Grasping

第一作者: Haofei Lu · 方向: 机器人操作 · 来源: cs.RO

Abstract:This work addresses sequentially grasping multiple objects with a single dexterous hand without releasing those already held. Most dexterous grasping methods commit all of the hand's degrees of freedom to a single object, underutilizing its dexterity and leaving no redundancy for subsequent grasps. The proposed solution, MoDex, is a diffusion policy that predicts the next gripper pose directly from observations, conditioned on an opposition space and point cloud. The opposition space condition specifies which fingers participate in the current grasp, enabling the gripper to use only a subset of its available degrees of freedom while reserving the remaining degrees of freedom for subsequent grasps. To facilitate sim-to-real transfer, MoDex is trained in two stages: first through imitation learning on expert demonstrations, and subsequently through reinforcement learning...

论文介绍 如何用单只灵巧手顺序抓取多个物体而不释放已持物体是一个挑战。现有方法常将手的所有自由度用于单个物体。本文提出MoDex,一种扩散策略,通过反对空间条件指定当前抓取使用哪些手指,从而只占用部分自由度,为后续抓取保留冗余。该方法通过两阶段训练(模仿学习与强化学习)促进仿真到现实的迁移。

Efficient Computation of Distance Functions for Navigation Vector Fields in Lie Groups

第一作者: Vinicius M. Gonçalves · 方向: 导航与运动 · 来源: cs.RO

Abstract:Vector-field-based methods are widely used for robot control and are often applied to the path-tracking problem. Some vector field approaches require repeatedly computing the distance between the robot configuration and the curve, as well as the corresponding closest point. Recently, vector fields have been extended to Lie Groups. In this case, this computation can be expensive, especially when performed at high control frequencies on embedded platforms. This paper proposes a method for efficiently computing the distance between a point and a curve represented as what is called a G-polynomial curve, which is a curve representation that generalizes polynomial curves to matrix Lie groups. The proposed approach exploits the structure of these curves to reduce the problem to a small number of polynomial root-finding computations. Simulation results show that the method...

论文介绍 基于向量场的方法广泛用于机器人控制,特别是路径跟踪。某些方法需要重复计算机器人配置与曲线间的距离及最近点。当向量场被扩展到李群时,此计算可能代价高昂。本文提出一种方法,用于高效计算点与称为G-多项式曲线之间的距离。该方法利用曲线结构将问题转化为少量多项式求根计算,从而提升嵌入式平台上的计算效率。

Inverse Manipulation through Symbolic Planning and Residual Operator Learning

第一作者: Yigit Yildirim · 方向: 机器人操作 · 来源: cs.RO

Abstract:Inverting a robotic task requires more than reversing symbolic state transitions or rewinding motor trajectories. In robot manipulation tasks, symbolic inverse plans often fail to fully restore the effects of forward executions under continuous interaction dynamics. We present a hybrid framework for inverse manipulation that derives inverse-skill objectives from STRIPS-like operators automatically extracted from demonstrations through soft geometric predicates. For each extracted operator, we construct an inverse restoration objective that preserves preconditions, restores delete effects, and negates add effects. A task planner first attempts to satisfy this objective using available action primitives. Unresolved symbolic predicates then induce a residual operator learning problem solved through Reinforcement Learning (RL). We evaluate the framework on the ManiSkill3 PushCube...

论文介绍 机器人任务的逆向执行不止于反转符号状态或回放轨迹。本文提出一个混合框架用于逆向操作:首先从演示中通过软几何谓词自动提取类STRIPS算子,并为每个算子导出逆向技能目标;然后任务规划器尝试用可用原语满足该目标;未能解决的符号谓词则构成一个残差算子学习问题,通过强化学习求解。

A New Quaternion-Joint Cable-Driven Redundant Manipulator Configuration and its Control Through FABRIK and Residual Reinforcement Learning

第一作者: Tanapath Pornthisan · 方向: 策略学习 · 来源: cs.RO

Abstract:Robotic arms capable of traversing arbitrary spatial paths, especially in highly obstructed workspaces, are highly desired across several industries. Quaternion-joints have recently empowered a specific class of robotic arms -- cable-driven redundant manipulators -- beyond its prior capabilities. Specifically, quaternion-joints reduce the number of required motors per degree of freedom, paving the way for more compact this http URL ongoing challenge is that the complexity of the kinematic model of quaternion joints challenges a priori decisions on manipulator configurations and imposes higher computational demands on the control system and its non-linearities amplify all discrepancies between design and physical artifact arising from fabrication imprecision. Here we show a that a 4-segment, 8-joint manipulator can achieve a broader workspace than extant configurations, at...

论文介绍 能在复杂环境中灵活运动的机械臂需求迫切。四元数关节提升了绳驱动冗余机械臂的能力,减少了每个自由度所需的电机。然而,其运动学模型的复杂性给构型设计和控制带来挑战。本文展示了一种4段8关节的机械臂构型,能实现比现有构型更大的工作空间,并提出结合FABRIK与残差强化学习的控制方法来应对其非线性。

Synthetic Data Generation and Vision-based Wrinkle and Keypoint Detection for Bimanual Cloth Manipulation

第一作者: Ariel Herrera · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robotic manipulation of textiles remains challenging because continuous deformation and self-occlusions hinder the robust visual perception required to estimate the cloth's state. To address the lack of annotated real-world data, we developed a Blender-based synthetic pipeline exporting auto-annotated keypoints, and combined manually labeled renders with real-world data to train a wrinkle detector. We present a perception framework integrating a CNN for permutation-invariant keypoint detection and a YOLOv8-OpenCV pipeline to extract grasping points from structural wrinkles. A proposed bimanual algorithm uses this system to stretch fully folded garments via wrinkles, transitioning to keypoint-based ironing once corners emerge. The keypoint model achieves a Mean Position Error (MPE) of 1.7615 pixels. The perception system transfers to physical fabrics without fine-tuning...

论文介绍 本文针对机器人布料操作中连续变形和自遮挡导致的视觉感知挑战,开发了基于Blender的合成数据管道,生成自动标注的关键点数据,并结合手动标注与真实数据训练皱纹检测器。提出一个集成CNN和YOLOv8-OpenCV的感知框架,用于检测关键点和抓取点,并设计双手算法实现衣物伸展和熨烫。该系统在关键点检测上达到较低误差,并能迁移到真实布料,有望提升机器人操作纺织品的鲁棒性。

T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation

第一作者: Jingkun Feng · 方向: 具身智能 · 来源: cs.RO

Abstract:Open-vocabulary 3D functionality segmentation enables robots to localize functional object components in 3D scenes. It is a challenging task that requires spatial understanding and task interpretation. Current open-vocabulary 3D segmentation methods primarily focus on object-level recognition, while scene-wide part segmentation methods attempt to segment the entire scene exhaustively, making them highly resource-intensive and time consuming. Balancing segmentation performance in terms of granularity, accuracy, and speed remains a challenge. As one step towards alleviating this, we introduce T-FunS3D, a task-driven hierarchical open-vocabulary 3D functionality segmentation method that provides actionable perception for robotic applications. Our method takes as input the 3D point cloud and posed RGB-D images of an indoor scene. We construct an open-vocabulary scene graph by...

论文介绍 本文研究开放词汇3D功能分割问题,旨在为机器人定位功能物体组件。现有方法在对象级或场景级分割上存在资源密集、平衡性能困难等挑战。提出T-FunS3D,一种任务驱动的分层开放词汇3D功能分割方法,输入3D点云和RGB-D图像,构建开放词汇场景图,以提供可操作感知。该方法兼顾分割粒度、准确性和速度,适用于机器人应用。

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

第一作者: Yitong Chen · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising. We argue that VLA action generation has a different condition-target structure: the policy is conditioned on rich observations, language, and state, but predicts only a compact, low-dimensional action chunk. Under this asymmetry, strong one-step action generation should not necessarily require the advanced one-step methods developed for image synthesis. We keep standard velocity prediction and add no teacher model, distillation stage, or auxiliary objective; in our main recipe, we simply bias the training time distribution toward high-noise states. We first isolate the effect in a controlled MNIST grid-to-sequence task, then test it with extensive robot-policy experiments. Across standard LIBERO, LIBERO-Plus, and LIBERO-Pro, one-step...

论文介绍 针对视觉语言动作模型中基于扩散的动作生成需要迭代去噪、成本高的问题,本文提出单步动作生成方法。分析VLA动作生成的条件-目标不对称性,通过偏向高噪声状态训练时间分布,无需额外复杂技术如教师模型或蒸馏。在控制任务和多个机器人策略基准上验证,实现有效单步生成,可能提升实时控制效率。

What Objects Enable, Not What They Are: Functional Latent Spaces for Affordance Reasoning

第一作者: Rohan Siva · 方向: 具身智能 · 来源: cs.RO

Abstract:Existing robot planning systems rely on appearance-based reasoning, where visual observations are encoded into latent spaces organized around object appearances (e.g., recognizing a "cart" based on how it looks). However, planning requires reasoning about task-relevant functionalities of objects (e.g., whether an object is "movable"), which appearance-based latent spaces do not capture. As a result, existing approaches struggle to generalize to novel robot-object interactions. We address this limited generalizability through affordance reasoning, enabling planning based on task-relevant object functionalities instead of appearance alone. We introduce A4D, which maps visual observations into a shared latent space structured around affordances (e.g., "movable"). By projecting visual observations into this functional latent space and measuring their proximity to affordances, A4D...

论文介绍 现有机器人规划系统依赖外观推理,导致对新型交互泛化能力差。本文引入基于affordance的功能推理,提出A4D方法,将视觉观察映射到功能潜空间,结构化围绕任务相关功能如可移动性。通过测量视觉观察与功能的接近度进行规划,从而基于物体功能而非外观提升泛化能力,适用于机器人规划任务。

Flash-WAM: Modality-Aware Distillation for World Action Models

第一作者: Arman Akbari · 方向: 机器人操作 · 来源: cs.RO

Abstract:World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods break down in the joint video-action setting because video and action streams use different SNR-shifted noise schedules and reach training with substantially different marginal noise distributions, an asymmetry that single-modality distillation methods cannot accommodate. We introduce \textbf{Flash-WAM}, a modality-aware step-distillation framework inspired by consistency distillation that selects the consistency function for each modality to match its noise regime: a linear-gradient-scaling parametrization for the action stream's low-noise regime, paired with a...

论文介绍 世界动作模型通过迭代扩散生成未来视频和机器人动作,但多步去噪限制实时控制。现有步蒸馏方法在联合视频-动作设置中因模态间噪声分布不对称而失效。提出Flash-WAM模态感知步蒸馏框架,针对视频和动作流选择适配的一致性函数,实现高效生成。该方法可能推动WAMs在机器人操作中的实际部署。

UNIVID: Unified Vision-Language Model for Video Moderation

第一作者: Kejuan Yang · 方向: 多模态具身 · 来源: cs.AI

Abstract:Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement. Traditional moderation systems often rely on fragmented black-box classifiers that are difficult to maintain and lack transparency. In this paper, we present UNIVID, a UNIfied VIsion-language model for video moDeration. Unlike standard classification models, UNIVID generates policy-aware captions that serve as an interpretable intermediate representation, enabling human-verifiable decisions and multi-task reusability. While existing open-source and commercial VLMs often suffer from safety-guardrail refusals and lack fine-grained policy alignment, we develop a specialized training data recipe that combines expert human-refined labels with synthetic data to align the model with our safety guidelines. By...

论文介绍 全球视频审核面临细粒度多模态推理和可解释输出的双重挑战,传统系统碎片化且不透明。提出UNIVID统一视觉语言模型,通过生成策略感知字幕作为可解释中间表示,支持人类验证和多任务复用。开发专门训练数据,结合专家标签和合成数据,对齐安全指南,提升审核的透明度和效果。

市场总览

美股方面,SPY 收于 737.55,日内跌 2.58%,但仍高于 SMA50(713.51)和 SMA200(683.93),RSI14 49.4 接近中性;QQQ 跌 4.8%,MACD 形成死叉。加密市场整体承压,BTC 和 ETH 的 RSI14 分别为 15.7 和 13.1,处于超卖区间,恐慌贪婪指数仅为 12(极度恐慌),加密总市值 2.19 万亿美元,24 小时跌 3.68%。中概股如 BABA 和 PDD 趋势 bearish,价格低于主要均线。商品外汇中,黄金期货 GC=F 跌 2.72%,RSI14 35;美元指数 DXY 涨 0.66%,RSI14 66.1,接近 52 周高点。VIX 恐慌指数大涨 39.68%,显示市场风险偏好下降。

今日关注

BTC-USD Bitcoin
偏下行

Bitcoin 当前价格 61183.62 美元,日内跌幅 4.1%,五日跌幅 16.85%。RSI14 为 15.7,处于超卖区间。MACD 值为 -3495.28,信号线为 -1984.55,形成死叉。价格远低于 SMA20(73199.55)、SMA50(76384.18)和 SMA200(78778.13),呈现空头排列。较 52 周高点下跌 51.52%,技术指标整体指向持续下行压力。

^VIX VIX 恐慌指数
偏上行

VIX 恐慌指数当前 21.51,日内大涨 39.68%,五日涨幅 40.4%。RSI14 为 65.7,接近超买但仍在正常范围。MACD 为 -0.3953,信号线为 -0.7728,形成金叉。价格高于 SMA20(17.12)、SMA50(18.99)和 SMA200(18.42),呈现多头排列。指标显示市场恐慌情绪显著上升,动量偏上行。

DX-Y.NYB 美元指数 DXY
偏上行

美元指数当前 100.07,日内上涨 0.66%,五日上涨 1.17%。RSI14 为 66.1,显示偏强但未超买。MACD 为 0.2468 高于信号线 0.1608,形成金叉。价格接近 52 周高点(仅差 0.57%),并高于 SMA20(99.02)、SMA50(98.91)和 SMA200(98.6),多头排列趋势明显,技术面偏上行。

GC=F 黄金期货
中性

黄金期货当前价格 4353.9,日内下跌 2.72%,五日下跌 4.53%。RSI14 为 35,处于中性偏低区域,未超卖。MACD 为 -63.668 低于信号线 -54.1205,但无显著金叉或死叉信号。价格略低于 SMA20(4546.88)和 SMA50(4624.4),但高于 SMA200(4398.63),趋势为 neutral,技术信号不明确,整体呈中性态势。

全部资产

^VIX

VIX 恐慌指数

$21.51 +39.68%
5 日
+40.40%
距 52w 高
-39.1%
RSI(14)
65.7
趋势
多头
SMA 20 / 50 / 200
17.12 / 18.99 / 18.42
MACD / 信号
-0.395 / -0.773
MACD 金叉 (今天)多头排列

^TNX

10Y 美债收益率 (%)

$4.54 +1.32%
5 日
+1.86%
距 52w 高
-9.2%
RSI(14)
57.7
趋势
多头
SMA 20 / 50 / 200
4.50 / 4.40 / 4.20
MACD / 信号
0.027 / 0.037
多头排列

DX-Y.NYB

美元指数 DXY

$100.07 +0.66%
5 日
+1.17%
距 52w 高
-0.6%
RSI(14)
66.1
趋势
多头
SMA 20 / 50 / 200
99.02 / 98.91 / 98.60
MACD / 信号
0.247 / 0.161
接近 52 周高多头排列

SPY

S&P 500 ETF

$737.55 -2.58%
5 日
-2.50%
距 52w 高
-3.0%
RSI(14)
49.4
趋势
多头
SMA 20 / 50 / 200
746.29 / 713.51 / 683.93
MACD / 信号
9.981 / 12.063
多头排列

QQQ

Nasdaq 100 ETF

$705.06 -4.80%
5 日
-4.50%
距 52w 高
-5.8%
RSI(14)
48.3
趋势
多头
SMA 20 / 50 / 200
722.01 / 667.81 / 621.82
MACD / 信号
17.425 / 20.696
MACD 死叉 (1 天前)多头排列

AAPL

Apple

$307.34 -1.25%
5 日
-1.51%
距 52w 高
-3.0%
RSI(14)
60.7
趋势
多头
SMA 20 / 50 / 200
304.25 / 281.24 / 265.19
MACD / 信号
8.464 / 9.395
MACD 死叉 (2 天前)多头排列

MSFT

Microsoft

$416.67 -2.66%
5 日
-7.46%
距 52w 高
-25.0%
RSI(14)
47.5
趋势
中性
SMA 20 / 50 / 200
422.58 / 408.35 / 456.38
MACD / 信号
5.540 / 6.187
MACD 死叉 (今天)

NVDA

Nvidia

$205.10 -6.20%
5 日
-2.86%
距 52w 高
-13.3%
RSI(14)
43.8
趋势
多头
SMA 20 / 50 / 200
219.10 / 203.45 / 188.57
MACD / 信号
2.300 / 4.259
多头排列

GOOGL

Alphabet

$368.53 -0.98%
5 日
-3.11%
距 52w 高
-9.8%
RSI(14)
46.4
趋势
多头
SMA 20 / 50 / 200
385.38 / 354.50 / 304.04
MACD / 信号
1.703 / 6.855
多头排列

TSLA

Tesla

$391.00 -6.56%
5 日
-10.28%
距 52w 高
-21.6%
RSI(14)
40.4
趋势
空头
SMA 20 / 50 / 200
425.87 / 395.29 / 414.14
MACD / 信号
4.143 / 8.607
MACD 死叉 (4 天前)空头排列

META

Meta

$593.00 -5.51%
5 日
-6.25%
距 52w 高
-25.5%
RSI(14)
41.6
趋势
空头
SMA 20 / 50 / 200
612.72 / 619.52 / 662.45
MACD / 信号
-3.755 / -3.493
MACD 死叉 (今天)空头排列
加密恐慌贪婪
12
极度恐慌
加密总市值
$2.19 T
-3.68% / 24h
BTC 主导率
56.1%
ETH 8.8%
24h 成交量
$184.1 B
活跃币 17,370

BTC-USD

Bitcoin

$61,183.62 -4.10%
5 日
-16.85%
距 52w 高
-51.5%
RSI(14)
15.7
趋势
空头
SMA 20 / 50 / 200
73,199.55 / 76,384.18 / 78,778.13
MACD / 信号
-3,495.276 / -1,984.551
RSI 超卖接近 52 周低空头排列

ETH-USD

Ethereum

$1,586.57 -10.34%
5 日
-20.84%
距 52w 高
-68.0%
RSI(14)
13.1
趋势
空头
SMA 20 / 50 / 200
2,008.98 / 2,190.29 / 2,464.45
MACD / 信号
-122.513 / -84.246
RSI 超卖接近 52 周低空头排列

SOL-USD

Solana

$64.37 -6.33%
5 日
-21.78%
距 52w 高
-74.6%
RSI(14)
16.9
趋势
空头
SMA 20 / 50 / 200
81.14 / 85.04 / 102.79
MACD / 信号
-4.459 / -2.503
RSI 超卖接近 52 周低空头排列

BABA

阿里巴巴 (BABA)

$121.06 -3.88%
5 日
-2.54%
距 52w 高
-37.2%
RSI(14)
37.2
趋势
空头
SMA 20 / 50 / 200
131.73 / 131.10 / 149.72
MACD / 信号
-2.624 / -1.646
空头排列

PDD

拼多多 (PDD)

$85.07 -0.94%
5 日
+0.75%
距 52w 高
-39.0%
RSI(14)
36.3
趋势
空头
SMA 20 / 50 / 200
92.48 / 97.24 / 112.45
MACD / 信号
-3.706 / -2.987
空头排列

JD

京东 (JD)

$28.88 -1.06%
5 日
+0.17%
距 52w 高
-21.6%
RSI(14)
41.3
趋势
空头
SMA 20 / 50 / 200
30.69 / 30.16 / 30.38
MACD / 信号
-0.352 / -0.077
空头排列

0700.HK

腾讯控股 (0700.HK)

HK$453.20 -1.26%
5 日
+6.09%
距 52w 高
-33.6%
RSI(14)
47.1
趋势
空头
SMA 20 / 50 / 200
451.94 / 476.94 / 572.08
MACD / 信号
-7.220 / -11.055
MACD 金叉 (3 天前)空头排列

GC=F

黄金期货

$4,353.90 -2.72%
5 日
-4.53%
距 52w 高
-22.1%
RSI(14)
35.0
趋势
中性
SMA 20 / 50 / 200
4,546.88 / 4,624.40 / 4,398.63
MACD / 信号
-63.668 / -54.120

CL=F

WTI 原油期货

$90.25 -3.00%
5 日
+3.31%
距 52w 高
-24.5%
RSI(14)
43.1
趋势
中性
SMA 20 / 50 / 200
96.75 / 97.86 / 72.79
MACD / 信号
-1.721 / -1.020

USDCNY=X

美元 / 人民币

¥6.76 -0.13%
5 日
-0.02%
距 52w 高
-6.2%
RSI(14)
34.2
趋势
空头
SMA 20 / 50 / 200
6.79 / 6.82 / 6.98
MACD / 信号
-0.015 / -0.015
接近 52 周低空头排列
风险提示

本报告基于公开行情数据计算的技术指标,仅供技术指标解读参考,不构成任何投资建议。过去走势不代表未来表现,市场存在不确定性,投资者应自行判断并承担风险。

Regrettable references and claims of ‘rigged’ election laws: why this week has reignited Jacinta Allan spill rumours

Just months from the Victorian election, the premier’s performance has left some MPs wondering if it’s too late for Labor to change leaders Get our breaking news email, free app or daily news podcast Jacinta Allan faced three major tests this week. The way she handled them has left some of her colle

中文摘要 维多利亚州选举临近,州长Jacinta Allan因本周处理三大测试引发领导权更迭传闻,部分议员质疑工党是否来不及更换领导人。

Iran war live: US says Iranian drones shot down, radar sites attacked

The United Nations reports that 1.4 million people are in need of aid in Lebanon amid Israel's attacks on the country.

中文摘要 美国军方称击落伊朗无人机并袭击雷达站;联合国报告黎巴嫩因以色列攻击有140万人需要援助。

US says Iran radar sites struck and drones intercepted, in latest threat to fragile ceasefire

Iran and the US have exchanged a series of attacks near the strait of Hormuz, imperilling efforts to reach a peace deal The US military said it shot down four Iranian drones that were launched toward the strait of Hormuz and struck coastal surveillance radar sites in response. “The attack drones pos

中文摘要 美国在霍尔木兹海峡附近击落四架伊朗无人机并袭击沿海雷达站,威胁脆弱停火协议,此前伊朗和美国已有一系列交火。

A Sherpa Survived 6 Days Alone on Everest. His Family Says He Was Abandoned.

Dawa Sherpa, 57, was found alive on Thursday, nearly a week after he was last seen on the mountain. His wife says more could have been done to find him sooner.

中文摘要 57岁夏尔巴人Dawa Sherpa在珠穆朗玛峰独自生存6天后被救,其家人批评救援延迟,称本应更快找到他。

Pamela Hicks, Lady-in-Waiting to Elizabeth II of Britain, Dies at 97

The queen’s third cousin, she was a bridesmaid at the royal wedding in 1947, and witnessed firsthand pivotal moments in British history.

中文摘要 英国女王侍女Pamela Hicks去世,享年97岁,她是女王的三表妹,曾是1947年皇家婚礼伴娘,见证英国历史关键时刻。

Astronauts return to ISS after sheltering during air leak repair attempt

Russian attempt to repair tunnel area sparks safe-haven procedure for five other astronauts onboard.

中文摘要 国际空间站宇航员在空气泄漏修复尝试期间避难后返回,俄罗斯维修任务触发安全程序,涉及五名其他宇航员。

Tennis Giants Tumble

All the big names are out at the French Open. The result is a very confusing, and very exciting, tournament.

中文摘要 法国网球公开赛所有大牌选手出局,导致比赛结果混乱但令人兴奋,赛事出现意外格局。

Putin promotes a new world economic order at St. Petersburg forum

At "Russian Davos," Putin ruled out meeting with Zelenskyy and promoted a new world economic order.

中文摘要 俄罗斯总统普京在圣彼得堡论坛上推动新世界经济秩序,并排除与乌克兰总统泽连斯基会面。

Trump hails jobs surge, says Iran talks ‘going well’

Trump hailed jobs growth before pivoting to Iran, saying negotiations with Tehran "seem to be going quite well".

中文摘要 特朗普赞扬美国就业增长,并表示与伊朗的谈判进展顺利。

Cobolli into final as virus-struck Arnaldi pulls out of French Open

Tenth seed Flavio Cobolli will ​take on German second seed Alexander Zverev in ⁠final after Matteo Arnaldi withdrawal.

中文摘要 法国网球公开赛中,意大利选手Matteo Arnaldi因病毒感染退出,使第十种子Flavio Cobolli进入决赛对阵德国第二种子Alexander Zverev。

‘Pervasive fear’ grips Gaza as Israeli attacks persist despite ceasefire

A drone attack killed a young woman and injured 15 others near Khan Younis, reported the Wafa news agency.

中文摘要 加沙地带在停火期间以色列攻击持续,无人机袭击在汗尤尼斯附近造成一名年轻女性死亡、15人受伤,引发普遍恐惧。

Nose Gear on Boeing 787-9 Dreamliner Collapses, Injuring Several Workers

The airline Lufthansa said the cause of the accident at Frankfurt Airport was under investigation. The plane can weigh 279 tons at takeoff.

中文摘要 波音787-9梦想客机在法兰克福机场前起落架倒塌,造成多名工人受伤,德国汉莎航空称事故原因正在调查中。

USDA Confirms Second US Screwworm Case in Texas

A second case of the deadly New World screwworm parasite has been confirmed in a Texas calf, the US Department of Agriculture said.

中文摘要 美国农业部确认德克萨斯州一头小牛感染致命的新世界螺旋蝇,这是该州第二例病例。

Raizen Inks $13 Billion Out-of-Court Debt Deal With Creditors

Raizen SA reached an out-of-court restructuring agreement with the majority of its creditors, marking a key step in the Brazilian sugar-and-ethanol producer’s efforts to rework its debt, according to people familiar with the matter.

中文摘要 巴西糖和乙醇生产商Raizen SA与多数债权人达成130亿美元庭外债务重组协议,标志其债务重组关键进展。

KPMG Chief Economist Discusses Fed Rate Hike Expectations

Diane Swonk, KPMG Chief Economist, analyzed the current economic landscape, highlighting how improvements in the labor market and persistent service sector inflation are driving hawkish sentiment among Federal Reserve officials. She noted that bond market pricing reflects expectations of a 25 basis

中文摘要 毕马威首席经济学家Diane Swonk分析称,劳动力市场改善和服务业通胀持续导致美联储官员鹰派情绪上升,债券市场定价反映加息预期。

Electrum-Backed Mexican Silver Miner Sinda Files for US IPO

Sinda Ltd. filed for a US initial public offering, seeking to fund mining for silver in a historically productive part of Mexico.

中文摘要 由Electrum支持的墨西哥银矿商Sinda Ltd.提交美国IPO申请,计划在墨西哥历史产区开采白银。

Strong U.S. Jobs Report Lowers Recession Risks Despite Wage Growth Lagging Inflation

Gene Sperling, President of Sperling Economic Strategies and former director of the National Economic Council, discussed the recent U.S. jobs report, highlighting that the headline numbers were stronger than expected and suggest a low likelihood of recession despite ongoing geopolitical and economic

中文摘要 前国家经济委员会主任Gene Sperling表示,美国就业报告数据强劲,降低经济衰退风险,尽管工资增长滞后于通胀。

Top Goldman lawyer Ruemmler to remain with bank despite Epstein ties

General counsel resigned after revelations about relationship with disgraced financier but now will stay on as adviser

中文摘要 高盛总法律顾问Ruemmler因与已故金融家爱泼斯坦的关系曝光后辞职,但将留任为顾问。

Tech Selloff, Bitcoin Drop Test Retail Investor Strength Ahead of SpaceX IPO

For years, Wall Street has benefited from one of the most reliable forces in modern markets: a retail-trader army willing to buy almost anything.

中文摘要 科技股抛售和比特币下跌在SpaceX IPO前考验散户投资者实力,华尔街长期依赖散户买入力量。

Marvell Technology, Flex to Join S&amp;P 500 Later This Month

Marvell Technology Inc. and Flex Ltd. will join the S&P 500 in the latest quarterly rebalance, S&P Dow Jones Indices said Friday.

中文摘要 标普道琼斯指数宣布,Marvell Technology和Flex将于本月晚些时候加入标普500指数,在最新季度调整中。

US stocks slump as fears over Big Tech shake Wall Street

The Nasdaq saw its biggest daily fall since early 2025.

中文摘要 美国股市下跌,纳斯达克指数录得自2025年初以来最大单日跌幅,因对大型科技股的担忧震动华尔街。

Goldman’s Flood Sees Buying Opportunity in Stock Market Selloff

Friday’s pullback in US equities offers a chance to add exposure rather than a reason to retreat, with a clear path for the S&P 500 to reach 8,000 this year, according to John Flood, the head of Americas equities execution services at Goldman Sachs Group Inc.

中文摘要 高盛美洲股票执行服务主管John Flood认为,周五美股回调是增加敞口的机会,标普500指数有望年内达到8000点。

SpaceX signs $30bn deal to lease computing capacity to Google

Agreement comes ahead of a record-breaking initial public offering for Elon Musk’s rockets-to-AI conglomerate

中文摘要 SpaceX与谷歌签署价值300亿美元的计算能力租赁协议,此举发生在埃隆·马斯克的公司进行破纪录IPO之前。

Beating the S&P For Generations: Masters in Business with Chris Davis

Barry sits down with Chris Davis, Chairman and Portfolio Manager at Davis Funds. They discuss his approach to managing risk and the key elements changing the economy. Chris and Barry also discuss Chris's mentors including Charlie Munger, and how he settled into the family business. (Source: Bloomber

中文摘要 采访Davis Funds主席Chris Davis,讨论其风险管理方法、经济变化要素,以及导师Charlie Munger的影响,如何延续家族事业。

Claude官方又又又异常了 26-6-5

这已经是A\在90天内第23次出现Error了,又是model问题 这也是本月第二次,稳定性堪比佛系 具体表现是529 Overloaded 27 个帖子 - 22 位参与者 阅读完整话题

【开源求Star】深度调研报告生成 Skill,一个命令,六分钟,一份深度专业的调研报告

本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 自己的第一个开源项目,全程vibe-coding。 在Opencode下已测试迭代N次,完整可用,搭配deepseek v4 flash,几乎0成本。 注意,适配调研任何主题,不仅仅能做行业研报哦!看我出的案例报告就知

A➗终止Mythos测试真实原因推测

我觉得网传信息不可靠,要是号商真有这么牛逼,连A​内部模型的访问权限都能搞到。早就直接和国内大厂合作了,完全没必要算计用户的三瓜两枣。还有,这次A​那个CEO也没有出来喊话,要是真的被反代了,不跳出来出来骂几句?我觉得更可能是Mythos吹得太离谱了,跟现实的落差很大,通过终止测试体面的离场,要是真的被揭穿了,恐怕会影响投资者的信心。 16 个帖子 - 11 位参与者 阅读完整话题

RawChat 公益站怎么关站了

刚才看到有评论说群里发消息了,站长被气到了然后关站了,有在群里的佬说说发生什么了吗? 62 个帖子 - 51 位参与者 阅读完整话题

分享一下我的 Windows 桌面图标设置

应该有很多人跟我一样不喜欢桌面一堆乱七八糟的图标,也不想装 uTools、Listary 这种额外的启动器。所以俺来分享一下这个用了好多年的 windows 上的软件使用习惯 核心思路:把所有软件的快捷方式统一放到一个文件夹(比如 QuickWay),再把这个文件夹加进环境变量,之后就可以在任意界面按下 Win+R 输入快捷方式名称回车,就可以直接打开对应软件 1. 新建一个放快捷方式的文件夹 在任意位置新建一个文件夹,命名为 QuickWay(名字随意,以下统一用这个),建议放在不容易误删的位置,我是放在了个人文件夹下,即 C:\Users\TaiYang\QuickWay 2. 把所有软件

建议少用国产辣鸡小模型,太让人红温了

最近Minimax刚上线了M3版本模型,因为公司采购的就是Minimax家订阅,所以公司项目代码就只能用它来写,使用过程中感觉豆包味越来越重,甚至无视写在项目md里的红线警告,自己往数据库写了数据,事后我让它细化md文档,更是搞了一堆废话进去。 公司也是抠,好歹也选个Qwen或者DS…… 40 个帖子 - 27 位参与者 阅读完整话题

还真有人拿我们的公益站去卖,而且还卖的的很贵!

事情是这样的,有人跟我们反馈说回答里出现了我们的来源信息,因为我们是动态风控机制的,正常使用是不会出现来源信息的,只有多设备 多ip 多账号 环境跳来跳去,才会触发这个风控机制,会动态增加来源信息 然后私聊问了下,如下: 没想到居然拿我们的公益站去卖而且还卖的这么贵!0.2的倍率,太夸张了! 由于当事人说收到了很多人的消息,不想被打扰,所以码掉了 224 个帖子 - 203 位参与者 阅读完整话题

41个20xPro 到现在封的剩下9个 说下我的情况 你们参考下

41个pro开号方式如下: 11个为美区没免税 套google play开的,220刀/月 之前封了7个 今早到现在为止封了25个,余9个 今天死的号为1个美区免税 2个土区 20个礼品卡 2个美区没免税 现在看来美区没免税能活久点 再说下2个多月做中转站,到现在的情况 总投入还差23000+ 没有回本,别说赚钱了 另外封了的号 googleplay去申请退款,没有一个给我退的,不给退 佬们参考下吧,9个号刷完 或者 刷死 结束 终于能睡个好觉了 65 个帖子 - 43 位参与者 阅读完整话题