每日简报

2026-06-22

← 历史归档

palmier-io/palmier-pro

Swift · ★ 5,140 · 🍴 388 · 📈 1,834 stars today

macOS video editor built for AI

中文介绍 一款专为人工智能优化的 macOS 视频编辑器。它利用 AI 技术简化了复杂的视频剪辑流程,旨在帮助内容创作者更高效地完成编辑工作,尤其适合需要快速处理和智能化剪辑视频的用户。

calesthio/OpenMontage

Python · ★ 8,728 · 🍴 1,297 · 📈 987 stars today

World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.

中文介绍 全球首个开源的智能视频制作系统。它集成了12条生产管线、52种工具和500多项智能体技能,可将你的AI编程助手转变为完整的视频制作工作室,适用于希望自动化或半自动化视频内容生产的开发者与创作者。

chopratejas/headroom

Python · ★ 44,390 · 🍴 3,090 · 📈 2,624 stars today

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.

中文介绍 一个旨在优化LLM输入效率的工具库。它能在数据发送给大语言模型之前,对工具输出、日志、文件及RAG文本块进行压缩,可减少60-95%的token消耗且不影响结果质量,适用于处理大量文本以降低LLM调用成本的开发者。

tursodatabase/turso

Rust · ★ 20,802 · 🍴 1,062 · 📈 548 stars today

Turso is an in-process SQL database, compatible with SQLite.

中文介绍 一个兼容SQLite的进程内SQL数据库。它继承了SQLite的轻量和易用性,并提供了更适合现代应用场景的特性,适用于需要嵌入式、轻量级数据库且希望保持SQLite兼容性的应用开发。

penpot/penpot

Clojure · ★ 52,237 · 🍴 3,339 · 📈 1,135 stars today

Penpot: The open-source design tool for design and code collaboration

中文介绍 一款面向设计与代码协作的开源设计工具。它为设计师和开发者提供了协同工作的平台,支持创建设计系统、UI/UX原型及实时协作,是寻求Figma开源替代方案的设计团队的理想选择。

ZhuLinsen/daily_stock_analysis

Python · ★ 44,446 · 🍴 41,431 · 📈 568 stars today

LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.

中文介绍 一个由大语言模型驱动的股票智能分析系统。它能整合多市场行情数据与实时新闻,提供决策看板与自动推送功能,并支持零成本定时运行,适合投资者进行日常的多市场股票监控与分析。

koala73/worldmonitor

TypeScript · ★ 58,078 · 🍴 9,206 · 📈 163 stars today

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface

中文介绍 一个实时全球情报监控仪表板。它利用AI进行新闻聚合、地缘政治监控和关键基础设施追踪,通过统一的态势感知界面呈现信息,适用于研究人员、分析师或需要实时掌握全球动态的机构。

bytedance/deer-flow

Python · ★ 72,581 · 🍴 9,828 · 📈 442 stars today

An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.

中文介绍 字节跳动开源的长周期超级智能体框架。它能够进行研究、编码和创作,通过整合沙箱环境、记忆系统、工具集、技能与子智能体等模块,处理不同复杂度的任务,适用于需要构建强大自主智能体的开发者。

DeusData/codebase-memory-mcp

C · ★ 10,273 · 🍴 779 · 📈 1,032 stars today

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

中文介绍 一个高性能的代码智能MCP服务器。它能将代码库快速索引成持久化的知识图谱,支持158种语言,查询速度达到毫秒级,可大幅减少LLM处理代码时的token消耗,适用于需要深度理解和导航大型代码库的开发工具。

mukul975/Anthropic-Cybersecurity-Skills

Python · ★ 17,699 · 🍴 2,127 · 📈 361 stars today

754 structured cybersecurity skills for AI agents · Mapped to 5 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND & NIST AI RMF · agentskills.io standard · Works with Claude Code, GitHub Copilot, Codex CLI, Cursor, Gemini CLI & 20+ platforms · 26 security domains · Apache 2.0

中文介绍 一个包含754个结构化网络安全技能的集合,专为AI智能体设计。这些技能映射了MITRE ATT&CK、NIST等多个权威安全框架,兼容Claude Code等主流编程助手,适用于构建具备专业安全知识的AI安全代理。

tw93/Pake

Rust · ★ 56,194 · 🍴 11,093 · 📈 1,848 stars today

🤱🏻 Turn any webpage into a desktop app with one command.

中文介绍 一个能用单条命令将任意网页打包成桌面应用的工具。它生成的桌面应用体积小、启动快,适合希望将常用网页服务(如邮件、文档、工具站)快速变成独立桌面客户端的用户。

mikumifa/biliTickerBuy

Python · ★ 3,709 · 🍴 464 · 📈 67 stars today

b站会员购购票辅助工具

中文介绍 一个针对B站会员购的购票辅助工具。它能帮助用户自动化或半自动化地完成票务抢购流程,适用于需要在B站购买热门演出、赛事门票但面临激烈抢票竞争的普通用户。

smicallef/spiderfoot

Python · ★ 18,759 · 🍴 3,106 · 📈 294 stars today

SpiderFoot automates OSINT for threat intelligence and mapping your attack surface.

中文介绍 一个自动化开源情报收集工具。它用于威胁情报分析和攻击面测绘,能集成数百个数据源进行信息聚合与关联分析,是网络安全专业人员、渗透测试人员进行信息搜集和风险评估的利器。

topoteretes/cognee

Python · ★ 18,638 · 🍴 1,969 · 📈 347 stars today

Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory across sessions with a self-hosted knowledge graph engine.

中文介绍 一个开源的AI记忆平台。它为智能体提供跨会话的持久化长期记忆,基于自托管的知识图谱引擎工作,适用于需要为AI应用赋予记忆能力、实现连续对话或个性化交互的开发者。

byoungd/English-level-up-tips

★ 54,003 · 🍴 5,554 · 📈 125 stars today

An advanced guide to learn English which might benefit you a lot 🎉 . 离谱的英语学习指南/英语学习教程/英语学习/学英语

中文介绍 一份进阶版的英语学习指南。它从不同维度为希望提升英语水平的学习者提供方法、资源和路线建议,内容系统且实用,适合已有一定基础、寻求突破瓶颈的英语自学者参考。

asgeirtj/system_prompts_leaks

JavaScript · ★ 44,396 · 🍴 7,332 · 📈 282 stars today

Extracted system prompts from Anthropic - Claude Fable 5, Opus 4.8, Claude Code, Claude Design. OpenAI - ChatGPT 5.5 Thinking, GPT 5.5 Instant, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.

中文介绍 一个收集并提取了多家厂商AI系统提示词的仓库。它公开了来自Anthropic、OpenAI、Google等公司最新模型的内部提示词,对研究AI对齐、安全性和提示工程的人员具有很高的参考价值。

mattpocock/skills

Shell · ★ 139,755 · 🍴 12,126 · 📈 1,443 stars today

Skills for Real Engineers. Straight from my .claude directory.

中文介绍 一个来自作者.claude目录的实用技能集合。它包含了面向真实工程师的提示词和工作流,旨在帮助开发者更高效地使用Claude Code等AI编程助手,是探索AI辅助编程最佳实践的参考材料。

Loops explained: Claude, GPT, Mira and what actually works

@AnatoliKopadze · 83.0K 粉丝 · 1.3M 阅 · 584 赞 · 70 转

AI has been in everyone's hands for years. Most people who use it every day still use it the slowest way there is: type a request, wait, fix it, ask again, all by hand. Not because the faster way is

中文介绍 本文指出多数人仍采用手动、反复提示AI的低效方式,旨在介绍一种更高级的工作流——“循环”,并探讨其如何适用于Claude、GPT等模型,以实现自动化,属于对新兴AI工作流理念的解读与实践指导。

How to Build an AI Second Brain With Claude and Obsidian That Gets Smarter Every Day (Full Guide)

@undefinedKi · 3.9K 粉丝 · 1.0M 阅 · 601 赞 · 78 转

Your best ideas are scattered across a dozen places right now. Notes apps. Browser tabs. Old chats with Claude that you closed and will never find again. Every time you sit down to work, you rebuild

中文介绍 提供一份详细指南,介绍如何结合Claude与Obsidian笔记应用,构建一个能自动整合碎片化信息(如笔记、网页、历史对话)并每日自我优化的“AI第二大脑”,旨在提升个人知识管理效率。

WTF Is a Loop? Part 2: The 15 Loops People Are Actually Running (and the Commands to Steal Them)

@mvanhorn · 35.2K 粉丝 · 102.4K 阅 · 510 赞 · 56 转

Earlier this month I wrote WTF Is a Loop? Peter Steinberger vs. Boris Cherny, which did 3.6M views on what a loop even is. This is the sequel, and it answers the next question: which loops do people

中文介绍 作为此前爆款文章(360万次观看)的续篇,本文具体列出并解析了15种用户实际在用的AI“循环”工作模式,并提供了可直接复用的命令,内容聚焦于实用案例和具体操作模板。

From Prompting Agents to Loop Engineering

@omarsar0 · 308.0K 粉丝 · 90.2K 阅 · 504 赞 · 69 转

A claim has been circulating in AI coding circles: stop prompting your coding agents and start designing loops that prompt them for you. As with everything new, this stuff gets repeated often and

中文介绍 针对AI编程领域“停止提示智能体,开始设计循环”的流行主张进行深入探讨,旨在厘清概念、分析其实际应用,帮助开发者理解如何从简单的提示转向构建自动化的“循环工程”系统。

How GLM-5.2 Beat Fable 5 at Website Design

@Designarena · 13.9K 粉丝 · 80.4K 阅 · 518 赞 · 39 转

GLM 5.2 ranks 1st overall on Design Arena’s single-turn, HTML Web Design (Non-Agentic) evaluation, 5 places higher than its predecessor GLM-5.1. To do so, it beat Claude Fable 5, Opus 4.6, and Opus

中文介绍 在Design Arena的单轮、HTML网站设计评估中,新模型GLM-5.2综合排名第1,超越了Claude Fable 5等强劲对手。这是一份产品性能报告,展示了GLM-5.2在特定设计任务上的最新突破。

The Next 5 Years: A Supersonic Tsunami

@PeterDiamandis · 400.3K 粉丝 · 51.7K 阅 · 551 赞 · 75 转

Elon described the near future as a “supersonic tsunami”: a wave moving so fast and so large that by the time you hear it coming, it has already broken over you. The phrase stuck with me. So let me

中文介绍 引用Elon Musk关于“超音速海啸”的比喻,展望未来五年AI驱动的技术变革将极其迅猛。这是一篇趋势分析与预测,旨在提醒人们为即将到来的、颠覆性的快速变化做好准备。

Loop Engineering: Build an AI That Codes While You Sleep

@phosphenq · 12.4K 粉丝 · 49.4K 阅 · 506 赞 · 63 转

Last December, Boris Cherny shipped 259 pull requests in a single month. Every one written by Claude. He says he did not open an editor the whole time. His job was not to write code. It was to write

中文介绍 以Boris Cherny单月提交259个AI生成PR的案例引出,详细阐述“循环工程”理念。核心是展示如何设计自动化工作流,让AI智能体(如Claude)持续处理编码任务,实现“人睡觉,AI工作”。

wtf is Loop Engineer & how to setup for real

@jasonzhou1993 · 30.8K 粉丝 · 48.5K 阅 · 526 赞 · 49 转

At around 1:00 AM yesterday, a bunch of PRs started landing in our codebase. Not because our team was working unusually late. They came from different agent loops: agents finding issues, picking up

中文介绍 从凌晨PR自动提交的实例切入,解释“循环工程师”这一新兴角色,并提供了如何设置让不同AI智能体自动发现问题、认领任务并提交代码的实际操作指南,侧重于具体实施方法。

Samsung Electronics brings ChatGPT and Codex to employees

Samsung Electronics deploys ChatGPT Enterprise and Codex to employees worldwide, marking one of OpenAI’s largest enterprise AI rollouts.

中文介绍 三星电子向全球员工部署ChatGPT Enterprise和Codex,这是OpenAI迄今为止最大的企业AI部署之一,标志着大规模企业级人工智能应用的推进。

[Exclusive] $250 off AI Engineer tix til Monday

special offer for subscribers - $250 off AI Engineer tix til Monday

中文介绍 Latent Space为订阅者提供限时优惠,AI工程师门票在周一前可减250美元,作为特别促销活动的一部分。

[AINews] not much happened today

a quiet day lets us promo AIE one last time

中文介绍 AI新闻今日无重大事件,Latent Space借此机会最后一次宣传AI工程师活动,强调活动的最终优惠机会。

A startup claims it broke through a bottleneck that’s holding back LLMs

Miami-based AI startup Subquadratic came out of stealth mode last month with a huge claim. It announced that it had solved a mathematical bottleneck that had been holding back large language models for almost a decade. The details were thin, and many people were unconvinced. But Subquadratic has sta

中文介绍 迈阿密AI初创公司Subquadratic声称解决了阻碍大型语言模型近十年的数学瓶颈,但细节有限,引发行业质疑和关注。

not much happened today

**GLM-5.2** emerges as a leading open-weight coding model rivaling **Opus 4.8** and **GPT-5.5** in software engineering tasks, emphasizing the strategic importance of open models for provider competition, on-prem deployment, and fine-tuning rights. Experts like **Patrick Toulme** and **Thomas Wolf**

中文介绍 GLM-5.2成为领先的开源编码模型,在软件工程任务中与Opus 4.8和GPT-5.5竞争,强调了开源模型对提供商竞争和本地部署的战略重要性。

The Professor of Outputmaxxing — Anjney Midha, AMP

We talk about how this legendary investor went from humble beginnings in Singapore to leading rounds in Anthropic, Mistral, Black Forest Labs, and Periodic Labs... and the AMP secret master plan!

中文介绍 Anjney Midha分享从新加坡起步到领导Anthropic、Mistral等AI公司融资的经历,以及AMP的秘密计划,展示投资历程和战略。

New usage analytics and updated spend controls for enterprises

OpenAI introduces new spend controls and usage analytics for ChatGPT Enterprise, helping organizations manage costs and scale AI with confidence.

中文介绍 OpenAI为ChatGPT Enterprise引入新的支出控制和使用分析工具,帮助企业管理AI成本并自信地扩展人工智能应用。

Improving health intelligence in ChatGPT

Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness responses with stronger reasoning, better context, clearer communication, and physician-informed evaluations.

中文介绍 OpenAI利用GPT-5.5 Instant改进ChatGPT的健康智能,通过增强推理、改善上下文和医生指导评估提升回应质量。

Using AI to help physicians diagnose rare genetic diseases affecting children

Researchers used an OpenAI reasoning model to help diagnose rare diseases, identifying 18 new diagnoses in previously unsolved cases.

中文介绍 研究人员使用OpenAI推理模型帮助诊断儿童罕见遗传疾病,在之前未解决病例中识别出18个新诊断,展示AI在医疗领域的应用潜力。

not much happened today

**GLM-5.2** from **Zhipu** emerged as a leading open-weight model with innovative **IndexShare** sparse-attention enabling efficient **1M-token inference**, praised as comparable to **GPT-5.5** and **Opus 4.8** but lacking vision support. Other notable open models include **Laguna M.1** by **Poolsid

中文介绍 智谱GLM-5.2模型因其创新的IndexShare稀疏注意力和高效1M-token推理能力成为领先开源模型,但缺少视觉支持,与GPT-5.5和Opus 4.8可比。

From Efficiency to Leakage -- Privacy Backdoor in Federated Language Model Fine-Tuning

第一作者: Shanghao Shi · 方向: 隐私保护

Abstract:Federated learning (FL) enables multiple parties to collaboratively fine-tune language models for domain-specific tasks without sharing raw data. Since full model fine-tuning is often prohibitively expensive for FL clients, parameter-efficient fine-tuning (PEFT) has become the de facto approach in practice, freezing the base model and training only a small set of adapters. In this paper, we show that a malicious parameter server can stealthily corrupt a PEFT adapter into a privacy backdoor that implicitly memorizes the client's training samples as isolated per-sample parameter updates stored in separate neurons, without degrading model utility. Concretely, our attack, NeuroImprint, assigns a dedicated memorization neuron to each training sample and constrains that each neuron is updated at most once along the local fine-tuning trajectory. This design mitigates both...

论文介绍 本文研究联邦语言模型微调中的隐私后门问题。提出NeuroImprint攻击,将参数高效微调适配器腐败为隐私后门,通过分配专用神经元隐式记忆客户端训练样本,每个神经元仅更新一次,在不降低模型效用的情况下实现隐私泄露。该方法揭示了联邦学习中参数高效微调的安全风险。

Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes

第一作者: Jun He · 方向: 系统安全

Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes. Existing access-control mechanisms authorize identities, while assurance layers certify proposed actions; neither alone provides a mandatory enforcement point for certified authority at the moment of mutation. This paper introduces the Sovereign Execution Broker (SEB), a runtime enforcement boundary for certificate-bound agentic infrastructure. SEB consumes certificates issued by the Sovereign Assurance Boundary (SAB), verifies that the requested mutation matches the certified execution contract, checks validity windows, policy epochs, revocation epochs, and live-state drift, mints scoped execution identity, invokes infrastructure APIs, and records signed decision and outcome records. By...

论文介绍 针对自主代理在云和工作流中的安全执行问题,本文引入Sovereign Execution Broker (SEB)系统。SEB作为运行时执行边界,通过消耗证书、验证执行契约匹配和记录决策,强制执行证书绑定的权限,为代理基础设施提供强制性的安全执行点。

Efficient and Sound Probabilistic Verification for AI Agents

第一作者: Alaia Solko-Breslin · 方向: 安全研究

Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulate and enforce policies expressed in a formal language like Datalog offer a promising solution. However, existing approaches are restricted to deterministic policies. In many practical applications of AI agents, there is a need to enforce security policies in the face of ambiguity, leading to probabilistic predicates or state transitions (for example, a declassifier or Personally Identifiable Information (PII) detector that has some failure probability on each invocation). Furthermore, in many such applications, one cannot easily make the independence assumptions necessary to invoke prior work on probabilistic inference in Datalog. We address this by introducing a sound and efficient framework for such verification based on distributionally robust...

论文介绍 AI代理在数字环境中运行时面临安全策略执行的不确定性挑战。本文提出一个基于分布鲁棒优化的概率验证框架,用于运行时监控和策略执行,支持概率谓词,如具有失败概率的检测器,提供健全且高效的方法以处理模糊性,增强AI代理的安全性。

Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software

第一作者: Arastoo Zibaeirad · 方向: 软件安全

Abstract:Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved. We present CWE-Trace, a framework for LLM vulnerability detection built from 834 manually curated Linux kernel samples spanning 74 CWEs. The framework enforces a strict temporal split (pre-2025 historical set / post-cutoff leakage-free set), preserves context-aware vulnerable--patched pairs, and introduces two diagnostic metrics: the Directional Failure Index (DFI) and Hierarchical Distance and Direction (HDD). We evaluate eight vanilla LLMs and 15 LoRA fine-tuned variants across non-targeted detection, targeted detection, and CWE classification. Our analysis yields two key results. First, data contamination provides no measurable advantage. Function-level analysis shows that 84% of nominally contaminated samples carry no usable...

论文介绍 LLM在漏洞检测基准上表现优异,但其推理能力是否真实存疑。本文提出CWE-Trace框架,基于手动整理的Linux内核样本评估LLM漏洞检测,并通过严格的时间分割和诊断指标分析,显示数据污染未提供优势,且函数级分析揭示多数样本无可用信息。

A-COMPASS: Formal Foundations for Anonymity Analysis in Microdata

第一作者: Tamara Tagliavia · 方向: 隐私保护

Abstract:In the information age, one of the leading problems is how to ensure individual's privacy. Depending on the context in which privacy is considered, various data privacy models have emerged. However, the domain of formal verification of these models is still not sufficiently explored even when it comes to the most basic models. An attempt to verify privacy requirements is the Compliance Assertion Language (COMPASS). In COMPASS, one can specify an anonymity condition that a table needs to satisfy, and an action that will modify the table if the condition is not satisfied. It is designed to operate on preprocessed tables in a form one record - one group of people. In this paper, we modify the COMPASS language in order to operate on microdata tables in their usual form of one record - one person. The modified language is called A-COMPASS. Along with checking of previously applied...

论文介绍 在数据隐私领域,匿名性验证的形式化方法尚不完善。本文提出A-COMPASS,修改COMPASS语言以操作微数据表,允许指定匿名性条件并在条件不满足时执行修改操作,为微数据的匿名性分析提供形式化基础,扩展隐私验证工具的应用范围。

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

第一作者: Reza Soosahabi · 方向: 系统安全

Abstract:Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation. This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the attacker's automated judge. Our analysis shows that conventional detect-and-block defenses can allow attacker success rate (ASR) to approach one as the query budget grows, since predictable refusals provide useful feedback to automated search. We then examine detect-and-misdirect, where detected malicious interactions receive controlled, non-operational responses designed to induce...

论文介绍 代理AI系统易受提示注入和越狱攻击,尤其是当攻击者使用模型引导自动化时。本文通过概率模型分析攻击防御设置,指出传统检测-阻塞防御可能被绕过,并提出检测-误导策略,对恶意交互提供非操作响应以诱导攻击者错误行动,提升防御能力。

Image Encryption Algorithm Based on Convolutional Neural Networks and Dynamic S-Box Generation

第一作者: Ans Ibrahim · 方向: 密码学协议

The paper proposes a dynamic approach to image encryption, combining the use of Convolutional Neural Networks (CNNs) and classical cryptography to improve the security and flexibility of image encryption. The main concept is to create adaptive Substitution boxes (S-boxes) based on characteristics that are learned by a trained CNN. The CNN-based S-boxes can be relied on for more non-linearity, uniqueness, and input image dependence than the conventional fixed S-boxes because they are susceptible to the linear and differential attacks. This dynamic behaviour enhances the confusion property and makes it more resistant to statistical and structural attacks. The encryption algorithm consists of CNN-based feature extraction and the creation of a personalised S-box to replace the pixels. Entropy, histogram analysis, correlation, NPCR, and UACI enable security assessment of generated S-boxes...

论文介绍 本文提出一种基于卷积神经网络的动态S-box图像加密算法。通过CNN学习图像特征生成自适应S-box,替代传统固定S-box,增强非线性、唯一性和输入依赖性,从而提高加密算法抵抗统计和结构攻击的能力,并使用安全指标进行评估。

Multi-View Decompilation for LLM-Based Malware Classification

第一作者: Bercan Turkmen · 方向: 软件安全

Abstract:Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable. Recent work suggests that large language models (LLMs) can assist this process by classifying decompiled code as benign or malicious, but existing pipelines typically rely on a single decompiler view. We argue that this assumption is fragile: decompilers are lossy heuristic tools, and different decompilers can expose different artefacts of the same binary. We curate a benchmark of benign utilities and malicious programs spanning a range of threat behaviors. Each sample is compiled and decompiled with both Ghidra and RetDec, yielding matched pseudo-C views. Across a range of LLMs from major model families, we find that providing both decompiler views improves malicious-class F1, mainly by increasing recall on malicious samples. Agreement analyses further show that...

论文介绍 本文探讨使用大型语言模型进行恶意软件分类时单一反编译视图的局限性。提出多视图反编译方法,结合Ghidra和RetDec的反编译视图,实验表明提供双视图能提升恶意类别的F1分数,主要通过增加召回率,提高分类准确性。

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

第一作者: Hanwool Lee · 方向: 密码学协议

Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as operators of a safety-critical system, instantiated in a simulated nuclear power plant control room. A five-role operator team, each backed by a configurable LLM, runs a plant governed by six critical safety functions (CSFs), while adversaries inject messages over four channels in bounded multi-turn sessions with per-turn feedback. Harm is an objective signal rather than LLM-judged text: a run terminates the moment any CSF is lost, attributed to the causing message. Evaluating four frontier operator models under a fixed-attack paired-replay protocol, we find that adaptive multi-turn attacks...

论文介绍 本文针对将大语言模型代理用作安全关键系统监控组件时的鲁棒性问题,提出了NRT-Bench基准。该基准在模拟核电站控制室场景中,通过一个由五个LLM角色组成的操作团队与通过四个通道注入信息的对手进行多轮对抗,以客观信号(关键安全功能是否失效)衡量危害。研究评估了四种前沿操作模型在自适应多轮攻击下的表现,揭示了其脆弱性。

Quantization as a Malicious Task: Removing Quantization-Conditioned Backdoors via Task Arithmetic

第一作者: Kaihsun Yang · 方向: AI 安全

Abstract:Model quantization is widely adopted to reduce memory usage and inference cost when deploying deep neural networks on resource-constrained devices. However, recent studies have revealed a new security threat known as Quantization-Conditioned Backdoors (QCBs), where a model behaves normally in full precision but activates malicious behavior only after quantization. Existing defenses typically modify quantization procedures or correct activation statistics, often introducing additional computational overhead or relying on specific quantization settings. Here, we present QVec, a parameter-space perspective for defending against QCBs. We observe that the weight difference between a full-precision model and its quantized counterpart encodes a structured behavioral shift, which can be interpreted as a malicious task vector rather than random quantization noise. Based on this...

论文介绍 模型量化被广泛用于降低部署成本,但引发了「量化条件后门」这一新安全威胁,即模型仅在量化后才表现出恶意行为。本文提出防御方法QVec,从参数空间视角出发,发现全精度模型与量化模型之间的权重差异编码了恶意任务向量。通过任务算术技术移除该向量,可在不依赖特定量化设置的情况下有效防御此类后门,且计算开销较低。

TrustMix: How to Mix Messages in a Mobile Ad-hoc Network

第一作者: Yu Shen · 方向: 密码学协议

Mix networks are a highly effective way to achieve anonymity, defending against a wide range of traffic-analysis attacks. However, mix networks are usually designed for infrastructure networks and cannot be directly applied in the context of mobile ad hoc networks (MANETs). The few existing solutions for MANETs require advance knowledge of the topology or a trusted central party. In this paper, we present TrustMix, a mix protocol for MANETs that operates without any central trusted party. In TrustMix, parties join groups and then messages are forwarded via multiple groups to provide anonymity. With TrustMix, users only need to find a party nearby that they consider trusted. They then forward the message to this party's group, and the party shuffles messages before forwarding to other groups, meaning that the original message and the forwarded message cannot be linked. Furthermore, even...

论文介绍 混合网络是实现匿名通信的有效手段,但通常依赖固定网络基础设施,难以直接应用于移动自组网。本文提出TrustMix协议,它无需任何中央可信方,用户仅需找到附近一个他们认为可信的节点,将消息转发至其所在的组。组内成员对消息进行混淆后再转发,从而切断原始消息与转发消息的关联性,实现去中心化的匿名通信。

GNSS Spoofing Threat for V2X communications

第一作者: Adolfo P. Jimenez · 方向: 软件安全

Abstract:Global Navigation Satellite Systems (GNSS) constitute a core technology for delivering crucial positioning, navigation, and timing (PNT) services in the Vehicle-to-Everything (V2X) domain, where they are indispensable for generating Cooperative Awareness Messages (CAM) that uphold network reliability and vehicular safety. Yet, GNSS signals are acutely exposed to spoofing, an advanced attack in which an adversary transmits crafted signals that replicate legitimate satellite characteristics, misleading the receiver into computing a false position. This work presents a methodology for conducting physical spoofing with inexpensive Software Defined Radio (SDR), describing a coordinate generation pipeline that employs Haversine-based distance calculations, temporal discretization to emulate constant velocity, and linear interpolation to produce high-fidelity GPS baseband signals...

论文介绍 全球导航卫星系统是车联万物领域提供定位、导航与授时的核心技术,但极易受到欺骗攻击。本文描述了一种使用廉价软件定义无线电实施物理层GNSS欺骗的方法论,包括基于距离计算、时间离散化和线性插值来生成高保真GPS基带信号的坐标生成流程,评估了该威胁对依赖GNSS的V2X通信系统(如合作感知消息)的实际影响。

Accelerating Trust Convergence in IIoT: A ML Approach for Dynamic Network Conditions

第一作者: Aymen Bouferroum · 方向: AI 安全

Abstract:In Industrial Internet of Things (IIoT) environments, trust management plays a vital role in securing systems, especially when dealing with resource-constrained devices. Traditional trust models often overlook the impact of fluctuating network quality, leading to slower trust convergence and inaccurate assessments. In this paper, we propose a dynamic trust management solution, known as the Trust Convergence Acceleration (TCA) approach, which integrates Machine Learning (ML) to accelerate trust convergence under poor network conditions. Our model predicts the number of time units needed for trust convergence based on key network metrics and dynamically adapts transition probabilities in the trust model to enhance convergence speed. Using a simulation framework that incorporates realistic Wi-Fi channel conditions based on the IEEE 802.11 standard, we demonstrate the...

论文介绍 在资源受限的工业物联网环境中,网络质量波动会严重影响传统信任模型的评估准确性与收敛速度。本文提出一种信任收敛加速方法,集成机器学习技术,基于关键网络指标(如丢包率、延迟)预测信任收敛所需时间,并动态调整信任模型中的转移概率,从而在恶劣网络条件下加速信任收敛,提升系统安全性评估的及时性。

A Measurement Study of Cryptographic Misuse in Embodied AI Mobile Applications

第一作者: Junchao Li · 方向: 系统安全

Abstract:Embodied AI (EAI) mobile applications are evolving from auxiliary user interfaces into active control-path components, directly linking mobile-side cryptographic security to cyber-physical trust. Despite this shift, existing security research predominantly focuses on embodied AI devices and cloud infrastructures, leaving the mobile control layer largely unexplored as a critical attack surface. To bridge this gap, we present the first large-scale measurement study of cryptographic misuse within the EAI mobile ecosystem. We construct EAIAppZoo, a benchmark of 507 real-world applications across six EAI domains, and employ an automated semantic-aware analysis pipeline to measure the prevalence and characteristics of five major cryptographic failure modes. Our measurement yields 12,975 misuse findings (with an evaluated precision of 80.74\%), revealing that these cryptographic...

论文介绍 具身AI移动应用正从辅助界面演变为主动控制路径组件,其移动控制层成为关键攻击面。本文首次对此生态中的密码误用进行大规模测量,构建了涵盖六个领域507个应用的基准库EAIAppZoo。通过自动化语义分析,测量了五类主要密码失败模式的普遍性,发现了大量误用实例,揭示了控制层密码安全与网络物理信任之间的直接关联。

AutoTam: Specifying Secure Protocol Implementations with Tamarin Model Generation

第一作者: Johannes Wilson · 方向: 密码学协议

Formal verification is a challenging but important task for ensuring the security of cryptographic protocols. While modern protocol verification tools significantly reduce verification effort, modelling remains challenging to practitioners without a background in formal verification. In addition, transferring verification results to a concrete protocol implementation requires expert knowledge. In this paper, we present a novel language-first method for verification of trace properties using a domain-specific language for protocol implementations. We target the Tamarin prover for verification, and we prove that verified universal trace properties translate back to the implementation. We additionally integrate symbolic execution in order to analyse the memory safety of protocol implementations. We use our tool to implement and generate accurate models for a signed Diffie-Hellman...

论文介绍 密码协议的形式化验证建模门槛高,且验证结果难以直接对应具体实现。本文提出一种语言优先的方法,针对协议实现定义了一个领域特定语言。该方法可自动为Tamarin证明器生成准确的协议模型,并证明已验证的通用迹属性可正确转换回实现。此外,它集成了符号执行以分析协议实现的内存安全性,降低了安全验证的实践难度。

FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

第一作者: Chaeyun Kim · 方向: AI 安全

Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and systemic trust erosion that require targeted evaluation. We introduce FinRED, an expert-guided red-teaming framework for financial LLM safety evaluation developed with financial experts. FinRED uses a novel two-level taxonomy mapping global standards (e.g., FATF and EU DORA) to threats ranging from regulatory evasion to complex fraud, integrated with a scalable pipeline that converts real financial documents into context-rich red-teaming Behavioral Prompts (seeds) through an expert-defined schema. Rigorous expert validation confirms seed plausibility and realism for meaningful LLM safety evaluation. We also provide an expert-validated, finance-specific rubric that goes beyond disclaimer checks, aligns more closely...

论文介绍 现有安全基准未能覆盖金融领域特有的风险,如违反监管合规、助长欺诈及破坏系统性信任。本文与金融专家合作,提出了FinRED,一个专家引导的金融LLM红队测试框架。它构建了一个将全球标准映射到具体威胁的双层分类法,并设计了可扩展的流程,将真实金融文档转化为上下文丰富的测试种子,以评估LLM在复杂金融场景下的安全性。

Low-Cost Multi-Precision Systolic Arrays for Accelerating FHE NTTs on AI ASICs

第一作者: George Alexakis · 方向: 密码学协议

Abstract:Fully Homomorphic Encryption (FHE) ensures robust data privacy but suffers from prohibitive computational overhead. Accelerating FHE on AI hardware like Tensor Processing Units (TPUs) is promising, yet fundamentally limited by a precision mismatch: TPUs are optimized for 8-bit arithmetic, whereas FHE and its critical parts such as the Number Theoretic Transform (NTT), demand high precision. Current approaches bridge this gap using matrix decomposition to execute NTT computations on low-precision matrix engines. However, reconstructing the full-precision results requires shift-and-add accumulation that does not match the dataflow of matrix multiplication. This forces offloading full-precision reconstruction from matrix engines to vector processors that disrupts the matrix multiplication dataflow, creating significant performance bottleneck. To resolve this limitation, we...

论文介绍 全同态加密(FHE)计算开销巨大,其关键操作数论变换(NTT)所需的高精度与AI硬件(如TPU)优化的低精度不匹配。现有矩阵分解方法在重构全精度结果时会破坏数据流,造成性能瓶颈。本文提出一种低成本的多精度脉动阵列架构,旨在直接在矩阵引擎内完成高精度NTT计算,以解决这一根本性限制,提升FHE在AI ASIC上的加速效率。

Heterogeneous LLM Debate Under Adversarial Peers: Honest Gains, Replacement Costs, and Resilience

第一作者: Prashanti Nilayam · 方向: AI 安全

Abstract:Heterogeneous LLM debate is motivated by the promise that diverse peers correct one another, but the same exchange that carries correction also carries adversarial influence. We measure which dominates by tracking how a heterogeneous peer changes the honest agents' revision behavior: how often they change their answer, and whether the change is corrective or harmful. We compare matched panels (homogeneous baseline, honest-mixed, and adversarial-mixed) and contaminated panels in which a malicious same-family peer is already present, spanning four model families and three reasoning benchmarks. An honest heterogeneous peer sharply lowers harmful revision, and an adversarial one reverses it. For Llama-3.1-70B defenders on MATH-hard, the honest-slot harmful-revision rate falls from 89% in the homogeneous panel to 35% with an honest peer, and an adversarial peer returns it to 90%...

论文介绍 异质大型语言模型(LLM)辩论中,多样化的同伴既可能纠正错误也可能施加对抗性影响。本文通过实证研究,测量了异质同伴如何改变诚实代理的修正行为,包括修正频率及其效果(纠错或有害)。研究对比了同质、诚实混合及对抗混合的面板,并跨越多个模型家族和推理基准。结果表明,诚实的异质同伴能显著降低有害修正,而对抗性同伴则会使有害修正率升高。

DISARM: Target Electronic Device Informed Mitigation of Software Runtime Side-Channel Vulnerabilities

第一作者: Tasneem Suha · 方向: 密码学协议

Abstract:Program runtime or timing attacks exploit variations in a program's execution times to extract sensitive information from the program (e.g. encryption keys, sensitive variable data, intellectual property). State-of-the-art solutions to runtime side-channel attacks attempt to balance the execution time of the sensitive code for different control flow paths to eliminate the timing leakage. However, during the mitigation process, most techniques do not consider the underlying hardware or device on which the target program is supposed to run on. This can lead to over-fixing (unnecessary extra operations), under-fixing (not solving the imbalance properly), and even failures. We propose DISARM, a joint hardware-software methodology (unlike any existing solution) for mitigating runtime side-channel vulnerabilities that utilizes timing values from real embedded devices to generate...

论文介绍 程序运行时或时序攻击利用执行时间变化泄露敏感信息。现有缓解技术旨在平衡不同控制流路径的执行时间,但大多未考虑程序实际运行的底层硬件,可能导致过度修复、修复不足或失败。本文提出DISARM,一种利用真实嵌入式设备时序值的软硬件协同方法,以生成更精确的缓解方案,有效应对运行时侧信道漏洞。

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

第一作者: Haotian Xu · 方向: AI 安全

Abstract:Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses are largely incompatible with speculative inference: they either introduce additional computation or disrupt the draft-verify mechanism, negating acceleration benefits. This reveals a fundamental incompatibility between current safety methods and speculative decoding. We propose SafeSpec, a safety-aware speculative inference framework that integrates risk estimation directly into the verification process. SafeSpec attaches a lightweight latent safety head to the target model to jointly evaluate semantic validity and safety in a single forward pass. When unsafe generations are detected, SafeSpec applies rollback and safety-guided reflective multi-sampling to recover safe continuations rather than terminating generation. We model jailbreak...

论文介绍 推测推理能加速大语言模型(LLM)解码,但缺乏内生的安全性保证,且现有的安全防御措施大多与推测推理不兼容。本文提出SafeSpec,一个安全感知的推测推理框架。它通过在目标模型上附加一个轻量级安全头,在单次前向传播中联合评估语义有效性和安全性。当检测到不安全生成时,SafeSpec会执行回滚和安全引导的反射性多采样,以恢复安全延续而非终止生成。

When Global Gating Is Enough: Admission-Time Hubness Control in Anisotropic Vector Retrieval Systems

第一作者: Prashant Kumar Pathak · 方向: 安全研究

Vector hubness, where a few points become nearest neighbors of many queries, creates a poisoning risk in retrieval-augmented generation (RAG): one injected document can influence unrelated requests. Existing defenses use periodic reverse-kNN scans, leaving an exposure window and repeated corpus-wide work. We study admission-time control, scoring each candidate against sentinel queries and quarantining hub-like documents before insertion. Across two 100,000-document corpora, five encoders, and disjoint attacker and defender query sets, a global gate achieves recall 1.0 at the decisive embedding-space point (>=0.92 across the effective range) and 0.91 +/- 0.07 on HotFlip attacks, with 1% false positives on general documents. A per-topic gate provides no reliable benefit, consistent with anisotropy coupling local and global visibility. Thresholds are maintained incrementally, with...

论文介绍 向量检索中的「hubness」现象(少数点成为众多查询的近邻)为检索增强生成(RAG)带来投毒风险:一个注入的文档可能影响不相关的请求。现有防御依赖周期性反向kNN扫描,存在暴露窗口且计算量大。本文研究了准入时控制方法,在文档入库前利用哨兵查询进行评估并隔离类hub文档。实验表明,全局门控能在关键嵌入空间点实现高召回率,并有效防御HotFlip攻击。

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots

第一作者: Gulshan Saleem · 方向: AI 安全

Abstract:Prompt injection is ranked as the most critical vulnerability in large language model (LLM) deployments by the OWASP Top 10 for LLM Applications, yet existing defenses operate at isolated pipeline stages and remain incomplete. Input filters cannot inspect retrieved documents, while output monitors cannot prevent malicious payloads from reaching the model. Consequently, retrieval-augmented generation (RAG) chatbots remain vulnerable to indirect injection, where a poisoned knowledge-base document compromises every user whose query retrieves it. We present a three-layer framework that intercepts both direct and indirect prompt injection throughout the inference pipeline. Layer 1 screens user input using a rule-based pattern library and a fine-tuned semantic anomaly classifier. Layer 2 enforces a provenance-based instruction hierarchy during context assembly, preventing retrieved...

论文介绍 提示注入是LLM部署中的关键漏洞,现有防御措施通常作用于管道的孤立阶段且不完整。例如,输入过滤器无法检查检索到的文档,输出监控器无法阻止恶意载荷到达模型。本文提出一个三层框架,以拦截整个推理管道中的直接和间接提示注入。该框架包含用户输入筛查、基于来源的指令层级强制执行以及输出检测,旨在为RAG聊天机器人提供全面保护。

PUFFERDOS: Efficient and Effective Attack String Generation for Regular Expression Denial of Service Vulnerabilities

第一作者: Shangzhi Xu · 方向: 系统安全

Abstract:ReDoS attacks constitute a critical class of resource-exhaustion vulnerabilities. In such attacks, adversaries exploit the pathological worst-case execution behavior of regular expression (regex) engines to induce highly asymmetric computational workloads, ultimately exhausting system resources and degrading service availability. To protect systems against ReDoS attacks, numerous detection techniques have been proposed that simulate the attack process by generating attack strings to proactively exploit ReDoS vulnerabilities at the early development stage and facilitate remediation. Existing techniques broadly fall into two classes: static analyses that search for pathological regex structures, and dynamic exploration methods that synthesize candidate attack strings. However, the generated attack strings are often impractical for real-world exploitation because they usually...

论文介绍 正则表达式拒绝服务(ReDoS)攻击利用正则引擎的病态最坏情况行为来耗尽系统资源。为在开发早期检测此类漏洞,需要生成用于主动探测的攻击字符串。现有技术生成的攻击字符串往往不切实际。本文提出PUFFERDOS,一种高效有效的攻击字符串生成方法,旨在生成更贴近真实利用场景的攻击字符串,以提升ReDoS漏洞检测的实用性和有效性。

G-Lox: Group-Adaptive, Privacy-Preserving Bridge Distribution with Two-Party Computation

第一作者: Baigang Chen · 方向: 系统安全

Abstract:We present G-Lox (group-adaptive Lox), a bridge-distribution system that preserves Lox-style distributor blindness while enabling hidden, stateful group-level adaptation. G-Lox places adaptive assignment logic behind a two-server privacy wall, so no single server learns group identifiers or group-to-bridge assignments. Private state access and state-dependent updates use two-server DPF/FSS protocols and secure two-party computation, supporting blockage reporting, transport-aware reassignment, and privacy-preserving group splitting. We evaluate G-Lox through system measurements and policy simulation. In our C++/EMP implementation over real TCP sockets, private state access has low client-visible overhead: across state sizes up to 2^16, communication remains in the low-KiB range per iteration. At M=1024, the client sends 1,968 bytes, receives 1,280 bytes, and completes an...

论文介绍 现有桥接分配系统在实现自适应分配时可能损害分配器的隐私性(即不知晓用户与桥接的映射关系)。本文提出G-Lox,一种在保持分配器盲性的同时支持隐藏的、有状态的组级自适应分配的系统。G-Lox将自适应逻辑置于由两台服务器构成的隐私墙之后,通过两方计算协议进行私有状态访问和更新,从而在保护隐私的前提下支持封锁报告和传输感知的重分配。

FloatDoor: Platform-Triggered Backdoors in LLMs

第一作者: Nils Loose · 方向: 软件安全

Abstract:Large language models (LLMs) are increasingly deployed in sensitive settings such as software engineering, where their outputs directly shape downstream artifacts. Recent work has shown that an identical model can produce measurably different outputs depending on the deployment platform, a consequence of non-associative floating-point arithmetic and divergent kernel implementations. We study the security implications of this platform-dependent variability and uncover a novel attack surface on LLM deployments. We introduce FloatDoor, the first input-independent, platform-triggered backdoor attack against generative LLMs. The compromised model exhibits adversary-chosen behavior when served on a target platform and is otherwise benign. FloatDoor is realized through two lightweight LoRA adapters, one that amplifies inter-platform numerical divergence and one that binds the...

论文介绍 该研究发现,大语言模型因不同硬件平台在浮点运算上的细微差异,会产生可测量的不同输出。作者据此提出首个输入无关、平台触发的后门攻击方法「FloatDoor」。该方法通过轻量级适配器放大平台间的数值分歧,并将特定行为绑定到目标平台,从而在模型被部署到特定平台时激活恶意行为,而在其他平台则表现正常。此工作揭示了LLM部署中一个由计算特性引发的新型安全风险。

Secure Coding Drift in LLM-Assisted Post-Quantum Cryptography Development: A Gamified Fix

第一作者: R.D.N. Shakya · 方向: 密码学协议

Abstract:The transition to Post Quantum Cryptography (PQC) introduces considerable implementation complexity, requiring strict adherence to constant-time execution, side channel resistance, and precise parametrisation. Simultaneously, large language models (LLMs) are heavily embedded in software development workflows, including cryptographic engineering. While LLMs improve productivity, evidence shows that they frequently generate insecure or suboptimal code, particularly in security critical domains. This paper introduces Secure Coding Drift in PQC, a novel socio technical vulnerability model capturing the gradual degradation of secure coding practices due to sustained reliance on LLM-generated code. Unlike prior work that focuses on static vulnerabilities, we conceptualise security risk as a longitudinal behavioural phenomenon rising from human AI interaction. To mitigate this, we...

论文介绍 本文聚焦于向后量子密码学过渡过程中,因过度依赖大语言模型辅助编程而导致安全编码实践逐渐退化的现象,提出了「安全编码漂移」这一社会技术漏洞模型。该研究认为,LLM虽然提升效率,但在安全关键领域常生成非最优或不安全的代码。为应对这一长期行为风险,作者引入了一种游戏化的干预方案,旨在通过主动学习来抵消LLM引入的安全漂移,增强密码工程的健壮性。

bioETH-Beacon: A Confidential On-Chain Genomic Beacon with Encrypted Counts, Filters, and Bounded Noise over a Fully Homomorphic EVM

第一作者: Christos Galanopoulos · 方向: 密码学协议

The Global Alliance for Genomics and Health (GA4GH) Beacon protocol lets researchers ask whether a genomic variant has been observed in a participating cohort and receive aggregate variant-level counts. As Beacon networks grow, two privacy risks remain: host institutions can see plaintext queries, and repeated rare-variant queries can support membership-inference attacks. We present bioETH-Beacon, a smart-contract prototype that runs the Beacon "aggregate count" query over encrypted data on a fully homomorphic Ethereum Virtual Machine (fhEVM). Hospitals upload encrypted marker-count entries, authorized researchers submit encrypted marker queries, and the contract returns an encrypted answer that is released, via an off-chain key-management service, only to the requester named in the contract's on-chain ACL. The design is organized as a 3x4 tier-by-query-family grid spanning genotype...

论文介绍 针对基因组信标协议中医院可查看明文查询及罕见变体查询可能泄露参与者身份的隐私风险,本文提出了「bioETH-Beacon」原型系统。该系统在支持全同态加密的以太坊虚拟机上运行智能合约,实现对加密数据的聚合计数查询。医院上传加密的基因标记数据,研究人员提交加密查询,合约返回加密结果。通过链上访问控制列表,确保只有授权请求者能解密答案,从而在保护数据隐私的同时实现基因组数据的共享研究。

Artificial Intelligence as Game Changer in Cybersecurity: What We Learned in 2025-2026, and how this is relevant for Africa

第一作者: Mikael Alemu Gorsky · 方向: 软件安全

Abstract:In 2025 and 2026, two events settled questions that had until then been speculative. In the first, a large language model executed the great majority of a state-aligned cyber-espionage campaign on its own, with human operators intervening at only a few decision points. In the second, the most capable cyber-relevant model was placed under a controlled-access program limited to a vetted set of United States technology firms, allied governments, and European standards bodies; that perimeter included no African government, operator, or university. Together the two events establish the argument of this paper: frontier language models have become a decisive instrument of cyber operations, and that instrument is built, owned, and rationed within a small circle from which Africa is absent. The paper documents Africa's exclusion on every count. The continent does not build frontier...

论文介绍 本文基于2025-2026年两个关键事件进行论述:一是大型语言模型主导了国家级网络间谍活动;二是最先进的网络能力模型被纳入由美国技术公司、盟国政府等组成的受限访问计划,而将非洲完全排除在外。文章据此论证,前沿语言模型已成为网络空间行动的决定性工具,但其构建、所有权和使用权集中于一个将非洲排除在外的封闭圈子。论文详细记录了非洲在技术构建、资源获取和标准制定等多方面的被边缘化状态。

Analyzing the Narration Gap in LLM-Solver Loops

第一作者: Zunchen Huang · 方向: AI 安全

Formal tools such as SAT and SMT solvers are increasingly embedded in language model reasoning pipelines when a safety or security critical question can be formulated in logic. Unlike chain of thought whose steps are sampled from the model distribution without formal guarantee, a solver produces a sound and independently verifiable answer. However, the soundness guarantee can be lost in the interaction between the solver and the model. The hybrid pipeline has three components: formalizing the question, deciding it, and narrating the result. Prior work has studied the formalization and decision, but not narration, which is the step that turns a formal tool's output into the user answer. To fill the narration gap, we first model the LLM-solver loop as a verified decision procedure. We further evaluate five open-sourced models under prompt injection, and we find certificate gating makes...

论文介绍 当安全或安全关键问题可被逻辑形式化时,形式化工具(如SAT/SMT求解器)正越来越多地被嵌入语言模型的推理流程中,以提供可验证的可靠答案。然而,在LLM与求解器交互的混合流程中,形式化保证可能在「叙述」步骤——即将求解器输出转化为用户答案的过程中——丢失。本文首次将LLM-求解器循环建模为一个经过验证的决策过程,并通过实验发现,采用「证书门控」机制能有效提升流程在提示注入攻击下的安全性,从而填补了该混合流程在叙述可靠性上的研究空白。

Passive-User Bell-State Loop-Back Key Establishment without Quantum Detectors at the User Nodes

第一作者: Luis Adrián Lizama-Pérez · 方向: 安全研究

Abstract:We propose and analyze a Bell-state extension of the Loop-Back quantum key distribution architecture for secret-key establishment between two passive users that do not require quantum transmitters or quantum detectors. In the proposed setting, a single active station, Alice, provides the entangled-state infrastructure, retains one qubit of an initially prepared Bell pair, and sends the traveling subsystem through two passive users, denoted by $B_1$ and $B_2$. Each passive user applies a local Pauli operation to the same traveling subsystem, so that the operation observed by Alice is only the effective composition $U_{\mathrm{eff}}=U_2U_1$. After the subsystem returns, Alice performs a Bell-state measurement and, using her private knowledge of the initial Bell state, deterministically identifies the effective Pauli operation. However, the individual factors $U_1$ and $U_2$...

论文介绍 本文提出一种扩展的Bell态环回量子密钥分发架构,允许两个不拥有量子发射器或探测器的被动用户之间建立安全密钥。该方案中,一个主动站点(Alice)提供纠缠态基础设施,保留一对Bell态量子比特中的一个,并将另一个发送给两个被动用户。每个用户对行经的量子比特施加一个本地Pauli操作。当该子系统返回时,Alice通过Bell态测量和对初始Bell态的私有知识,可以确定性地识别出用户施加的有效复合操作,进而实现密钥协商。

MemoryWAM: Efficient World Action Modeling with Persistent Memory

第一作者: Sizhe Yang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robust robotic manipulation in the real world requires not only an understanding of the current observation, but also memory and dynamics modeling. World action models (WAMs) possess these capabilities by jointly modeling visual foresight and actions conditioned on both current and historical observations, making them a promising paradigm for robotic manipulation. However, existing WAMs face a fundamental trade-off: methods with efficient inference typically condition only on a bounded window of recent observations and therefore struggle in non-Markovian environments, whereas methods that preserve long histories incur time and space costs that grow substantially with sequence length. To address this challenge, we introduce MemoryWAM, a world action model with efficient persistent memory. MemoryWAM uses a hybrid memory design that combines recent frames, event-boundary anchor...

论文介绍 当前的世界动作模型在机器人操作中面临一个根本权衡:高效推理的方法通常只依赖近期有限观测窗口,难以应对非马尔可夫环境;而保存长期历史的方法则随序列长度增加导致时间和空间成本剧增。为此,本文提出了「MemoryWAM」,一种具有高效持久记忆的世界动作模型。它采用混合记忆设计,结合近期帧和事件边界锚点,在保持长期依赖信息的同时实现高效推理,旨在提升机器人在复杂、需要长时记忆的真实世界任务中的操作鲁棒性。

Generating Robot Hands from Human Demonstrations

第一作者: Sha Yi · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robot learning has advanced rapidly in learning control, but learning the physical body of a robot remains much more difficult because jointly searching over design and control creates a very large combinatorial problem. Here, we present a data-driven framework for generating robot hands from human demonstrations. Instead of learning a complex controller together with each candidate design, we generate robot hand designs using the same simple control policy used after fabrication: matching fingertip positions through inverse kinematics. Using more than 4 million frames of human fingertip motion from everyday manipulation, our algorithm optimizes tree-structured robot hands to reproduce desired target motions. The framework produced both a 6-degree-of-freedom (DoF) general-purpose hand and lower-DoF task-specific hands with spatial four-bar mimic joints. To accelerate the...

论文介绍 为解决同时优化机器人设计与控制的巨大组合难题,本文提出一种数据驱动的框架,可从人类演示直接生成机器人手的设计。该框架的核心思路是,在设计阶段就使用与制造后相同的简单控制策略——即通过逆向运动学匹配指尖位置。基于超过400万帧的人类日常操作指尖运动数据,算法优化了树状结构的机器人手,以复现目标运动。最终成功生成了一个6自由度通用手和若干低自由度任务专用手,为机器人手的设计提供了一种数据驱动的自动化方法。

Increasing Resilience of Continuum Robots via Motion Planning Algorithms

第一作者: Oxana Shamilyan · 方向: 导航与运动 · 来源: cs.RO

Abstract:This paper presents an experimental study of motion planning for resilient continuum robots. In this study we mainly focused on multi-criteria decision-making, its application for path-planning algorithms, impact on the generated path and execution time. To do this, we used two well-known algorithms for path planning, namely Genetic algorithm and A star algorithm, and modified them by adding the Analytical Hierarchy Process algorithm to evaluate the quality of the paths generated. In our experiment the Analytical Hierarchy Process considers four different criteria, i.e. distance, motors damage, mechanical damage of the robot's arm and accuracy, each considered to contribute to the resilience of a continuum robot. The use of different criteria is necessary to increase the time to maintenance operations of the continuum robot. We conducted the experiments using two different...

论文介绍 本文研究连续体机器人的运动规划以提高弹性。核心方法是将遗传算法和A*算法与层次分析法结合,后者评估路径在多标准下的质量,包括距离、电机损坏、机械损坏和准确性。实验验证该方法可延长机器人维护间隔,适用于高可靠性应用场景。

Fast Human Attention Prediction for Fixation-guided Active Perception in Autonomous Navigation

第一作者: Fatma Youssef Mohammed · 方向: 导航与运动 · 来源: cs.RO

Abstract:Human visual attention relies on structured scanpaths to efficiently process scenes, yet instilling this behavior into robot autonomy is in its infancy and hindered by the high,computational costs of existing predictive models. To address this, we introduce GazeLNN, a computationally lightweight,scanpath prediction model that leverages Liquid Neural Networks as its recurrent engine and employs MobileNetV3 for feature extraction. Operating auto-regressively, the architecture predicts sequential fixation heatmaps conditioned on the current visual stimulus and fixation history. Despite requiring only 0.61 GFLOPs, GazeLNN achieves state-of-the-art performance on the MIT Low Resolution dataset achieving 0.47 ScanMatch score. It outperforms existing recurrent baselines across diverse evaluation metrics, while reducing computational costs by 99.40% and accelerating inference by up to...

论文介绍 针对自主导航中人类视觉注意力集成的高计算成本问题,本文提出GazeLNN模型。该模型基于液态神经网络和MobileNetV3,自回归预测注视热图序列,计算高效仅需0.61 GFLOPs,在MIT数据集上达到先进性能,可用于机器人主动感知引导。

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation

第一作者: Zhenghao "Mark'' Peng · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Learning-based planners for sidewalk navigation can generate diverse candidate trajectories in real time, yet their scoring functions often fail to select the best trajectory in challenging situations, outputting trajectories that make the mobile robot drive onto grass, toward pedestrians, or in the wrong direction, even when better candidates exist in the same set. We call this the trajectory scoring gap: in real-world sidewalk navigation, the gap between an anchor-based planner's top choice and the best possible candidate is substantial, likely due to limited high-level scene understanding capability of the planner. Rather than replacing the planner with an end-to-end Vision-Language-Action model, we propose a VLM-Planner interface that uses a VLM to select a candidate index from the planner's proposal set and then fuse it with the planner's initial output. However, VLMs...

论文介绍 本文解决移动机器人人行道导航中学习规划器轨迹评分不佳的问题。提出VLM-Planner接口,使用视觉语言模型从规划器候选轨迹中选择最佳索引并与初始输出融合,增强场景理解能力,以应对复杂城市环境中的导航挑战。

TaCauchy: An Extensible FEM Framework for Vision-Based Tactile Simulation

第一作者: Hengfei Zhao · 方向: 策略学习 · 来源: cs.RO

Abstract:Vision-based tactile sensors require high-fidelity simulation for reinforcement learning, yet existing approaches struggle to provide accurate mechanical stress fields within GPU-accelerated robotics platforms. We present TaCauchy, an extensible Finite Element Method (FEM) framework that integrates rigorous physics-based force computation into Isaac Sim. Built on the Unified Incremental Potential Contact (UIPC) solver, TaCauchy directly computes Cauchy stress tensors from hyperelastic constitutive laws and projects them onto contact surfaces to obtain traction forces and pressure distributions, providing mechanical ground truth from first principles rather than empirical estimation. Our framework features automatic mesh generation with geometry-aware adaptive refinement and a modular sensor interface enabling rapid integration of diverse sensors (GelSight Mini, DIGIT, 9DTact)...

论文介绍 为视觉触觉传感器的高保真模拟,本文提出TaCauchy框架。基于有限元方法集成到Isaac Sim中,直接计算柯西应力张量并投影到接触面,提供物理准确的力反馈。框架支持自动网格生成和多种传感器接口,适用于机器人触觉研究。

CoLI: A Reproducible Platform for Continuum Robot Learning via Monolithic 3D Printing and Isomorphic Teleoperation

第一作者: Ziyuan Tang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Continuum robots offer strong potential for manipulation tasks due to their high degrees of freedom, compliant structures, and operational safety. However, their adoption in both research and practical applications has been hindered by reproducibility issues arising from complex fabrication and assembly processes, challenging kinematic modeling, and a lack of intuitive control interfaces. To address these challenges, we present a novel open-source continuum robot design. The platform features a simplified fabrication pipeline enabled by multi-material 3D printing, allowing the arm to be fabricated as a monolithic compliant structure with minimal assembly. Control is achieved through an isomorphic teleoperation interface that establishes a direct actuator-level mapping, eliminating the need for explicit kinematic modeling and providing a singularity-free mapping. Building on...

论文介绍 本文针对连续体机器人可重复性挑战,提出CoLI平台。使用多材料3D打印制造单体柔性结构简化组装,通过同构遥操作接口实现直接执行器映射,无需复杂运动学建模,便于控制和研究,促进机器人操作学习应用。

An Infrastructure-less, Control-Independent Solution to Relative Localisation of a Team of Mobile Robots using Ranging Measurements

第一作者: Paolo Golinelli · 方向: 导航与运动 · 来源: cs.RO

Abstract:The ability to localise teams of robots is essential for applications ranging from robotic fleets in unstructured environments to cooperative control and navigation tasks. In such contexts, fixed infrastructure is often unavailable, deployments must be fast and flexible, and system requirements must be minimal. We present a decentralised cooperative localisation algorithm that addresses all these challenges at once. The method is anchor-less, fully decentralised, and, unlike most existing approaches, does not require controlling the robots motion to ensure team observability. It relies only on local odometry, sparse inter-agent ranging measurements, and short-range communication, all of which are widely available in practice. The algorithm adopts a multi-hypothesis Bayesian framework that maintains the entire set of feasible solutions, ensuring robustness under transient...

论文介绍 本文提出一种去中心化协作定位算法,用于无固定基础设施的移动机器人团队。算法仅依赖局部里程计和稀疏测距测量,采用多假设贝叶斯框架维持可行解,确保鲁棒性,适用于快速部署和协调任务,如机器人车队管理。

Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm Vision-Language-Action Systems

第一作者: Yandong Wang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) models show strong capabilities in single and dual-arm robotic manipulation. Prior works show coordinated bimanual behaviors can emerge from end-to-end learning, leveraging large vision-language backbones with continuous action prediction. However, as bimanual tasks become tightly coupled and execution constraints become critical, implicit coordination alone is insufficient to ensure reliable, interpretable, and stable behavior. In this work, we propose Co-VLA, a coordination-aware bimanual manipulation framework introducing explicit structural priors into VLA models. We instantiate our method on a state-of-the-art vision-language backbone by replacing its monolithic action head with a Structured Action Expert (SAE) designed for bimanual coordination. Specifically, we introduce explicit structure at the action generation level with a modular...

论文介绍 本文提出Co-VLA框架用于双臂机器人协调操作。在视觉语言动作模型中引入结构动作专家,显式建模双臂协调,替代单一动作头以提升行为可靠性和可解释性,适用于复杂操作任务,基于先进视觉语言骨干实现。

Efficiently Linking Real Scenes with Synthetic Data Generation for AI-based Cognitive Robotics and Computer Vision Applications

第一作者: Paul Koch · 方向: 具身智能 · 来源: cs.RO

Abstract:AI vision models are a driving factor for the potential use case scenarios of cognitive robotics within in the industry and household applications. A large array of methods from semantic environment analysis towards 6D and grasping pose estimation have been proposed based on the latest AI achievements. However, such advancements require further strong and efficient methods w.r.t. training data and AI-architectures, which are capable in synergy to tackle current challenges, precision limits, and scalability beyond domain gaps. In this paper, we discuss these current limits and trends in the related state-of-the-art which are challenging those. Further we discuss our current work in progress on bridging the domain gap between simulations and real world applications by linking those in the training data generation.

论文介绍 本文探讨AI视觉模型在认知机器人应用中的训练数据挑战。讨论通过链接真实场景和模拟生成合成数据,以弥合域差距,提高模型精度和可扩展性,适用于工业和家庭环境中的机器人视觉任务。

Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

第一作者: Gia-Binh Nguyen · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion parameter architectures impose prohibitive computational burdens during downstream fine-tuning and real-time inference. In this work, we reveal a highly non-trivial architectural characteristic of these continuous control foundation policies (e.g., pi_0, GR00T-N1.5): despite being trained on diverse physical trajectories, they exhibit severe layer-wise representational redundancy. To exploit this, we introduce a structural compression pipeline that is entirely training-free, bypassing the need of existing methods to load full-scale models to learn optimized token reductions or dynamic layer selectors. Instead, using only a single forward pass via Centered Kernel Alignment to identify redundant layer features, we remove twin layers to...

论文介绍 本文研究了大规模视觉-语言-动作(VLA)模型在微调时面临的计算负担问题。作者揭示出这类连续控制基础策略存在显著的层间表示冗余。为此,提出了一种无需训练的结构压缩流程,仅需通过单次前向传播的中心核对齐来识别冗余层特征,并移除冗余层对,从而在保持性能的同时显著降低计算成本,提升了VLA模型下游应用的可行性。

FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching

第一作者: Francesco Argenziano · 方向: 导航与运动 · 来源: cs.RO

Abstract:Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments. Such agents must not only comprehend and navigate spatial layouts, but also reason about how these spaces evolve over time. In particular, humans interact with objects daily, causing them to change position throughout the environment and making it difficult for robots to reliably associate current observations with previously seen objects. However, these interactions are not random: human habits and routines induce spatio-temporally consistent patterns in object locations, which robotic agents can potentially learn and then exploit for downstream tasks such as navigation. To this end, we introduce FlowMaps, a latent flow matching model for estimating multimodal distributions over the future locations of dynamic objects in a continuous 3D space...

论文介绍 为使机器人理解动态环境中物体随时间的位置变化,本文提出了FlowMaps,一种基于潜在流匹配的模型。该模型能够估计未来物体在连续三维空间内位置的多模态概率分布。其核心思想是利用人类习惯与日常生活所导致的物体位置时空一致性模式,为下游如导航等任务提供先验知识。该工作旨在增强机器人对场景长期动态的理解能力。

Belt-Finger: An Affordable Soft Belt-Driven Gripper for Dexterous In-Hand Manipulation

第一作者: Boya Zhang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Parallel-jaw grippers are the default manipulator choice in robotics because they are simple, robust, and inexpensive. Their limited in-hand mobility, however, often forces large arm motions and restricts dexterous manipulation in confined workspaces. We present a parallel-gripper upgrade: a double-soft-belt-based finger module that preserves standard opening/closing while adding three in-hand degrees of freedom (DoF): translation, pitch, and roll. The mechanism is deliberately kept simple and engineered for inexpensive manufacturing and straightforward integration, preserving the reliability and precise control of traditional parallel grippers while greatly broadening the range of manipulation capabilities. To demonstrate the utility of the added DoFs, we integrate the gripper in two control pipelines. First, we adapt a model predictive controller for in-hand manipulation of...

论文介绍 针对传统平行夹持器手内操作自由度有限的问题,本文介绍了一种名为Belt-Finger的升级方案。该方案通过在标准平行夹持器上集成一个双软带驱动的指模块,在保留原有开合功能的同时,增加了平移、俯仰和滚转三个手内操作自由度。其设计强调结构简单、制造经济且易于集成,旨在以低成本显著扩展机器人末端执行器的操作能力范围。

Robust Assembly State Reasoning from Action Recognition for Human-Robot Collaboration

第一作者: James Fant-Male · 方向: 具身智能 · 来源: cs.RO

Abstract:Human Action Recognition (HAR) is frequently investigated in Human-Robot Collaboration (HRC) research to understand what actions have been performed and hence the state of a collaborative task. Accurately tracking an assembly state from HAR is however not fully investigated, and in realistic scenarios is not a trivial task. This research systematically investigates and compares methods for tracking assembly state using action recognition inputs. Investigations using two diverse datasets and five state tracking approaches, including logic-based, Hidden Markov Model (HMM), and neural network (NN) methods, show that optimal approaches are not uniform across different tasks and that different methods fail under different circumstances. Testing is performed using both simulated inputs with varying noise levels and realistic inputs from a HAR model. Results show NN and HMM methods...

论文介绍 在人类-机器人协作中,通过人类动作识别来推断任务状态至关重要。本文系统地研究和比较了利用动作识别输入来跟踪装配状态的不同方法。通过在两个数据集上测试基于逻辑、隐马尔可夫模型和神经网络等五种方法,发现在不同任务和噪声条件下最优方法并非固定。研究揭示了不同方法的失败场景,为实际应用中选择合适的跟踪方法提供了参考。

Frequency-Aware Flow Matching for Continuous and Consistent Robotic Action Generation

第一作者: Jianing Guo · 方向: 机器人操作 · 来源: cs.RO

Abstract:Flow matching has emerged as a standard paradigm for robotic manipulation owing to its strong expressive power for modelling complex, multimodal action distributions, alongside similar approaches like diffusion policy. However, existing methods rely on discretized action chunks, making them brittle to demonstrations collected at heterogeneous control frequencies and prone to temporally inconsistent actions that degrade control stability. In this paper, we propose Frequency-Aware Flow Matching (FAFM), which outputs continuous, temporally consistent actions. To handle heterogeneous frequency input, we transform discrete action sequences into the frequency domain with the discrete cosine transform (DCT), perform flow matching over the resulting coefficients, and reconstruct continuous actions via cosine basis expansion. To generate temporally consistent actions, we regularize the...

论文介绍 现有基于流匹配的机器人操作方法依赖于离散动作块,难以处理异构控制频率的数据,且易产生时序不一致的动作。本文提出频率感知流匹配(FAFM)来解决此问题。核心是将离散动作序列通过离散余弦变换转换到频域,在频域系数上执行流匹配,并通过余弦基展开重构出连续、时序一致的动作,从而提升控制的稳定性和对不同演示数据的兼容性。

Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform

第一作者: Hyeonna Choi · 方向: 具身智能 · 来源: cs.RO

Abstract:Biological experiment protocols are written in natural language, whereas automation systems rely on predefined control commands, creating a semantic gap that limits autonomous execution. Microplate-based automatic experiments are particularly challenging due to the need to simultaneously control well mapping, sample-reagent combinations, replicate placement, and parallel dispensing. This study proposes an agent-based protocol translation framework that converts natural-language microplate-based protocols into executable control commands for a robotic laboratory platform. A Parser Agent formalizes the natural-language protocol into a structured representation, and a rule-based mapping engine deterministically incorporates the operational constraints of the robotic laboratory platform to generate device-level control commands. A heterogeneous LLM Validation Agent verifies...

论文介绍 生物实验的自然语言协议与机器人平台的控制命令之间存在语义鸿沟。本文提出一个双智能体框架来桥接这一差距。解析智能体将自然语言协议形式化为结构化表示,规则映射引擎结合平台约束生成设备控制命令,而验证智能体则对生成命令的物理合理性进行检查。该框架旨在实现微孔板实验的自动化,将文字描述直接转化为可执行的机器人操作指令。

Pose6DAug: Physically Plausible Multi-view Object Swapping for Robot Data Augmentation

第一作者: Jonghoon Lee · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution. The standard remedy is to collect multi-view teleoperation data for every failure case, but this scales poorly in both cost and time. We introduce Pose6DAug, a failure-driven data augmentation framework that turns a policy's own successful episodes into targeted demonstrations for its failure modes, without any new data collection. Our key insight is that each successful episode already encodes a physically valid action trajectory together with calibrated multi-view observations. By swapping only the manipulated object while preserving this trajectory, we obtain new and physically grounded demonstrations. However, naive 2D video editing breaks...

论文介绍 视觉-语言-动作策略在面对训练分布外的物体时容易失败。本文提出Pose6DAug,一种失败驱动的数据增强框架。它能将策略自身的成功轨迹转化为针对其失败模式的新演示,无需新数据采集。关键在于利用成功轨迹中已有的物理合理动作轨迹和多视角观察,仅通过交换被操作物体来生成新演示。该工作强调在多视角下进行物理合理的物体替换以保证数据有效性。

VFILC: Accurate Frequency Extrapolations in Imitation Learning via Sampling Frequency ILC

第一作者: Nozomu Masuya · 方向: 模仿学习 · 来源: cs.RO

Abstract:Conventional neural network (NN)-based imitation learning methods for variable-speed motion either restricted their scope to interpolated speeds, or generated unpredictable motions when extrapolating beyond trained velocity ranges. Variable-frequency imitation learning (VFIL) enabled extrapolations of speeds by linking the NN model's sampling frequency to the motion frequency, whereas its open-loop configuration caused frequency errors, especially in the extrapolated high-frequency settings. This study proposes variable-frequency imitation learning with iterative learning control (VFILC) based on a combination of VFIL and iterative learning control (ILC) with both feedforward and feedback parts, the former taking advantage of VFIL and the latter adjusting the frequency errors. The experimental results showed that the proposed method successfully and accurately extrapolated...

论文介绍 传统的模仿学习在处理变速度运动时,外推至训练速度范围之外会产生不可预测的运动。本文提出VFILC方法,将可变频率模仿学习(VFIL)与迭代学习控制(ILC)相结合。VFIL通过将模型采样频率与运动频率关联实现速度外推,而ILC的反馈部分用于校正频率误差,尤其是高频下的误差。实验表明,该方法能成功且准确地实现超出训练范围的速度外推。

MirrorDuo: Reflection-Consistent Visuomotor Learning from Mirrored Demonstration Pairs

第一作者: Zheyu Zhuang · 方向: 模仿学习 · 来源: cs.RO

Abstract:Image-based behaviour cloning leverages demonstrations captured from ubiquitous RGB cameras. However, it remains constrained by the cost of collecting diverse demos, especially for generalizing across workspace variations. We propose MirrorDuo, a reflection-based formulation that operates on image, proprioception, and full 6-DoF end-effector action tuples, generating a mirrored counterpart for each original demonstration, effectively achieving "collect one, get one for free". It can be applied as a data augmentation strategy for existing learning pipelines, such as standard behaviour cloning or diffusion policy, or as a structural prior for reflection-equivariant policy networks. By leveraging the overlap between the original and mirrored domains, MirrorDuo achieves significantly improved performance under the same data budget when demonstrations are evenly distributed across...

论文介绍 本文针对基于图像的行为克隆在收集多样化演示时成本高昂的问题,提出了MirrorDuo方法。该方法通过对每条原始演示生成其镜像对应体,实现「收一得二」。它可作为现有学习流程(如标准行为克隆或扩散策略)的数据增强策略,或作为反射等变策略网络的结构先验。通过利用原始域与镜像域的重叠,在演示均匀分布时,能在相同数据预算下显著提升性能。

A Neuromorphic Reinforcement Learning Framework for Efficient Pathfinding in Robotic Mobile Fulfillment Systems

第一作者: Junzhe Xu · 方向: 策略学习 · 来源: cs.RO

Abstract:Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Systems (RMFS) a challenging problem for conventional search- and rule-based methods, which typically suffer from high computational complexity and long decision latency. While reinforcement learning (RL) has emerged as a powerful alternative, deploying learned policies with extreme energy efficiency on resource-constrained hardware remains an open challenge. We present SDQN-RMFS, an end-to-end framework that achieves high-fidelity deployment of an RL-trained policy from a full-precision artificial neural network (ANN) through to a neuromorphic chip. By computing only when triggered by sparse events, this framework unlocks ultra-low-power RMFS pathfinding. Our full-stack pipeline operates as follows: an ANN policy is first efficiently trained...

论文介绍 动态环境、有限空间和实时约束使机器人移动履行系统(RMFS)的路径规划面临挑战。本文提出了SDQN-RMFS框架,实现强化学习训练策略从全精度人工神经网络到神经形态芯片的端到端高保真部署。该框架通过仅在事件触发时计算,实现了超低功耗的RMFS路径规划。

Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

第一作者: Jinghan Yang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions can cause irreversible harm, making generalizable and interpretable failure detection essential. We observe that successful and failed rollouts carry systematically different information-theoretic signatures. Building on this, we formalize VLA control as a closed-loop information pipeline and derive the Triple Information-theoretic (Tri-Info) signals that capture whether actions remain diverse, temporally consistent, and coupled to state transitions. Across six VLA models and three benchmark environments, Tri-Info matches the strongest baselines in-domain. Moreover, Tri-Info transfers across architectures, environments, and the sim-to-real gap without retraining, reaching 83\% accuracy on real-world tasks where prior detectors collapse to...

论文介绍 视觉语言动作(VLA)模型日益被用于多种任务,但其黑盒性质使通用且可解释的失败检测至关重要。本文观察到成功与失败的运行轨迹携带系统性不同的信息论特征。基于此,将VLA控制形式化为闭环信息管道,并推导出三重信息论信号,用以捕捉动作是否保持多样性、时间一致性及与状态转换的耦合。该方法在跨架构、环境及仿真到现实场景时无需重新训练。

Evaluation of Augmented Reality-based Intuitive Interface for Robot-Assisted Transesophageal Echocardiography: A User Study

第一作者: Xiu Zhang* · 方向: 机器人操作 · 来源: cs.RO

Abstract:TransEsophageal Echocardiography (TEE) is essential for diagnosing and guiding Structural Heart Disease (SHD) interventions. However, manual TEE manipulation demands significant operator expertise, is physically demanding, and exposes clinicians to radiation when performed alongside fluoroscopy. Robotic-assisted TEE systems have been introduced to improve probe handling and reduce operator fatigue, yet the design of intuitive and effective user interfaces remains an open challenge. This study presents and evaluates a model-enhanced, Augmented Reality (AR)-based intuitive interface for robot-assisted TEE, designed to improve spatial awareness and control intuitiveness. A robotic TEE platform integrated with electromagnetic tracking and a virtual simulator was used to compare three user interfaces differing in visualization and interaction modalities: 2D jointlevel (2D-JI), 3D...

论文介绍 手动操作经食道超声心动图(TEE)对操作者要求高、体力消耗大。本研究设计并评估了一种基于模型的增强现实(AR)直观界面,用于机器人辅助TEE。该界面旨在改善空间感知和控制直觉性。通过集成电磁跟踪和虚拟仿真器的机器人TEE平台,比较了三种不同可视化和交互模式的用户界面。

SWAP: Symmetric Equivariant World-Model for Agile Robot Parkour

第一作者: Kaixin Lan · 方向: 具身智能 · 来源: cs.RO

Abstract:While latent world models enable the proactive predictions required for extreme parkour, their purely data-driven nature forces them to redundantly encode left-right symmetric interactions as independent patterns. This inflates the learning burden and hinders the capture of geometric regularities, restricting the latent space's efficiency for downstream policies. To address this, we propose SWAP, an end-to-end equivariant symmetric world model. This framework embeds symmetry directly into both the world model and the actor-critic networks. In real-world tests, the robot leaps across a 2.13 m gap and climbs a 1.63 m platform, breaking records for quadruped parkour. Furthermore, the framework exhibits robust geometric generalization to unseen mirrored terrains and exceptional zero-shot transferability across diverse outdoor environments. These results demonstrate that symmetry...

论文介绍 潜在世界模型可用于极端跑酷所需的主动预测,但其纯数据驱动特性会将左右对称交互冗余编码为独立模式。本文提出SWAP,一种端到端的等变对称世界模型框架。该框架将对称性直接嵌入世界模型和Actor-Critic网络中,显著提升了学习效率和潜在空间的利用效率,并在真实世界测试中实现了跨环境的强鲁棒性和零样本迁移能力。

Deep-Unfolded Coordination

第一作者: Hunter Kuperman · 方向: 具身智能 · 来源: cs.RO

Abstract:Distributed optimization is a highly scalable and structurally transparent technique to solve multi-agent robotics problems; however, such methods often suffer from the need for highly-specialized, problem-specific hyperparameter tunings. In this work, we propose Deep Coordinator, a deep-unfolding framework that learns to dynamically adjust the hyperparameters of ADMM-DDP, a popular distributed solver for robotics tasks, at solve-time in response to optimizer performance. Our architecture consists of unrolling a fixed number of ADMM-DDP iterations into a neural network with learnable functions between layers mapping the optimizer state to the next hyperparameters. To the best of our knowledge, Deep Coordinator is the first deep-unfolding framework to adapt the penalty parameters of a non-convex optimizer at solve-time; we show that the mainstream supervised approach can yield...

论文介绍 分布式优化是解决多智能体机器人问题的可扩展方法,但其常需针对特定问题进行繁琐的超参数调优。本文提出Deep Coordinator,一种深度展开框架,能在求解时学习动态调整ADMM-DDP优化器的超参数。该框架将固定次数的ADMM-DDP迭代展开为神经网络,层间嵌入可学习函数以根据优化器状态映射至下一组超参数。

Co-policy: Responsive Human-Robot Co-Creation for Musical Performances

第一作者: Xuetao Li · 方向: 多模态具身 · 来源: cs.RO

Abstract:Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route for generative models to participate in that creativity through physical action rather than disembodied digital content. In robotic music co-creation, it is challenging to connect semantic musical understanding with real-time and physically executable performance. We present Co-policy, a framework for human-robot musical co-creation that separates semantic intent grounding, constrained musical variation, and visuomotor execution. To ground musical semantics, Co-policy uses pre-inference semantic anchors and a fine-tuned Qwen-vl planner (F-Qwen) to transform speech, live musical seeds, and visual observations into structured co-creation plans. To support low-latency execution, Co-policy introduces a Gaussian-Mixture Visuomotor Policy (GMP), implemented as a...

论文介绍 本文提出Co-policy框架,用于实现人机音乐共同创作。该框架将语义意图接地、约束性音乐变化和视觉运动执行分离。它使用预推理语义锚点和微调的Qwen-vl规划器将语音、实时音乐种子和视觉观察转化为结构化共创计划,并引入高斯混合视觉运动策略以实现低延迟执行。

One-to-Two Acting: A Novel Framework for Single-arm Agent Action Expansion to Dual Arms

第一作者: Youbin Yao · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dual-arm manipulation can improve throughput via parallel execution, but collecting bimanual demonstrations for training is costly and difficult. We present ExS2D, a hierarchical action expansion framework that enables dual-arm manipulation from single-arm supervision. ExS2D first generates structured subtasks from textual instructions while explicitly capturing temporal precedence. It then grounds each subtask into executable actions through subtask-guided action mapping in observation. Finally, precedence-aware action allocation and synchronized planning are performed by a multimodal large language model driven coordinator to select collision-free dual-arm executions. Simulation experiments demonstrate that ExS2D reduces the average execution steps by 54.4% while maintaining a comparable success rate to a single-arm baseline. Real-robot experiments on four tasks further...

论文介绍 双臂操作可通过并行执行提高效率,但收集双臂演示成本高昂。本文提出ExS2D,一种层次化动作扩展框架,能从单臂监督中实现双臂操作。该框架首先从文本指令生成结构化子任务,然后将每个子任务通过子任务引导的动作映射落地为可执行动作,最终由多模态大语言模型驱动的协调器进行优先级感知的动作分配和同步规划,以实现无碰撞的双臂执行。

TIDY: Thermal Infrared Image Denoising via Wavelet Domain Entropy and Directional Stripe Index

第一作者: Tai Hyoung Rhee · 方向: 具身智能 · 来源: cs.RO

Abstract:Thermal infrared (TIR) imaging has been a popular choice for field robotics due to its robust perception capability under low light visual degradation, but it suffers from severe stochastic and fixed-pattern noise that breaks downstream estimation. This noise is intensified indoors due to low thermal contrast and uniform temperature distributions, contributing to the relative lack of indoor TIR deployments. Existing TIR denoising methods exhibit a poor accuracy-efficiency tradeoff, either too slow for online deployment required in robotics or insufficiently robust to severe degradation, while typically being trained on synthetic noise. Addressing these problems, we propose TIDY, a lightweight wavelet-domain denoiser trained on real clean-noisy TIR data. By reformulating TIR denoising in the wavelet domain, TIDY explicitly disentangles noise from structural content, enabling...

论文介绍 该研究针对热红外图像在室内机器人部署中因低热对比度和均匀温度分布导致严重噪声的问题,现有去噪方法在准确率与效率上权衡不佳。提出TIDY,一种基于小波域熵和方向条纹指数的轻量级去噪器,通过重构到小波域分离噪声与结构内容,训练于真实数据,以提升机器人视觉感知的鲁棒性。

EquiVLA: A General Framework for Rotationally Equivariant Vision-Language-Action Models

第一作者: Thien-Loc Ha · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have emerged as a powerful paradigm for generalist robot manipulation, yet they lack geometric inductive biases: policies trained at specific orientations require substantially more data to generalize across rotational configurations. We present \textsc{EquiVLA}, the first general framework for end-to-end $\mathrm{SO}(2)$-equivariant VLA models, applicable to any architecture coupling a frozen vision-language backbone with a flow-matching Diffusion Transformer action head. \textsc{EquiVLA} introduces \textsc{EquiPerceptor}, which produces approximately $\mathrm{SO}(2)$-equivariant visual representations from frozen ViT features; and \textsc{EquiActor}, an exactly $\mathrm{SO}(2)$-equivariant flow-matching Diffusion Transformer action head. Together, they establish an approximate $\mathrm{SO}(2)$ equivariance chain from camera observations to...

论文介绍 该研究指出视觉-语言-动作(VLA)模型缺乏几何归纳偏置,导致在旋转配置间泛化需大量数据。提出EquiVLA,首个端到端SO(2)-等变VLA框架,适用于耦合冻结视觉-语言骨干与流匹配扩散Transformer动作头的架构。引入EquiPerceptor和EquiActor,从观测到动作建立等变链,提升机器人操作的泛化能力。

Start Right, Arrive Right: Asynchronous Execution via Initial Noise Selection

第一作者: Trong-Bao Ho · 方向: 导航与运动 · 来源: cs.RO

Abstract:Action chunking enables robot policies to produce temporally coherent behavior, but generating multi-step action sequences with flow-based policies incurs latency that is incompatible with real-time control. Under asynchronous execution, the robot continues executing the current chunk while the next one is generated, causing even minor delays to create inconsistencies at chunk boundaries. Existing methods address this problem by steering generation toward the already executed action prefix. We instead show that prefix consistency can be achieved by selecting an appropriate initial noise before generation begins, allowing the unmodified flow ODE to produce a coherent next chunk. This reframes asynchronous inference as a noise selection problem rather than a trajectory steering problem. We introduce \textbf{PAINT}, a training-free method that finds this noise via backward Euler...

论文介绍 该研究针对动作分块策略在异步执行中因延迟导致块边界不一致的问题,现有方法通过引导生成来保持前缀一致性。提出PAINT方法,将异步推理重构为噪声选择问题,在生成前通过反向欧拉法选择初始噪声,使未修改的流ODE产生连贯的下一块,无需训练,提升实时控制性能。

Data Standards for Humanoid Robotics: The Missing Infrastructure for Physical AI

第一作者: Shaoshan Liu · 方向: 多模态具身 · 来源: cs.RO

Abstract:The scalability of humanoid robots will depend not only on models and hardware, but also on whether physical experience can accumulate across robots, tasks, organizations, and time. Drawing on the authors' work in developing ISO/WD 26264-1, Humanoid robot datasets -- Part 1: General requirements, within ISO/TC 299/WG 16, this article argues that data standards are becoming foundational infrastructure for Physical AI. We develop three insights. First, humanoid robot data is embodied interaction data, not a collection of isolated digital samples; a useful dataset must preserve the relationship among robot body, action, task, scene, execution trace, and outcome. Second, its value depends on physical coherence: multimodal streams are reusable only when timing, coordinate frames, calibration, kinematics, units, and synchronization assumptions remain inspectable. Third, the main...

论文介绍 该研究认为人形机器人的可扩展性依赖于数据标准的建立,以支持跨机器人、任务和时间的物理经验积累。基于ISO标准开发工作,提出数据标准作为物理AI的基础基础设施,强调人形机器人数据需保留身体、动作、任务等关系,并保证多模态流的物理一致性,以促进数据复用和共享。

Temporal Self-Imitation Learning

第一作者: Yinsen Jia · 方向: 机器人操作 · 来源: cs.RO

Abstract:Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while rare efficient behaviors may be forgotten during training. We argue that temporal efficiency itself provides a powerful and underutilized source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 15 distinct long-horizon manipulation tasks, TSIL consistently...

论文介绍 该研究针对长时程机器人操作中奖励塑形仍可能导致低效交互和高效行为遗忘的问题,提出时间自模仿学习(TSIL)框架。TSIL挖掘学习中生成的时序高效成功轨迹,将其转化为监督信号,通过配置自适应时间目标和效率加权自模仿进行策略改进,在多个任务中提升学习效率。

VOiLA: Vectorized Online Planning with Learned Diffusion Model for POMDP Agents

第一作者: Marcus Hoerger · 方向: 具身智能 · 来源: cs.RO

Abstract:Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Decision Process (POMDP) provides a powerful framework for such a capability. Although POMDP-based planning has advanced significantly, its application to real-world problems is often limited by the difficulty of obtaining faithful POMDP models. We present Vectorized Online planning wIth Learned diffusion model for POMDP Agents (VOiLA), a framework that learns task-agnostic POMDP models for online planning under uncertainty. VOiLA learns transition and observation samplers using conditional diffusion models and learns observation-likelihood models for particle-based belief updates. To enable efficient online planning, the diffusion samplers are distilled into compact feedforward generators and integrated with Vectorized Online POMDP Planner (VOPP), an online POMDP...

论文介绍 该研究针对不确定性下机器人规划中POMDP模型难以获取的限制,提出VOiLA框架,学习任务无关的POMDP模型用于在线规划。使用条件扩散模型学习转移和观测采样器,并蒸馏为紧凑前馈生成器,与向量化在线POMDP规划器集成,以实现高效在线决策。

Bidirectional Tutoring for Developmental Motor Learning in Robots: Co-Developed Interaction Dynamics Support Stable Learning

第一作者: Rui Fukushima · 方向: 具身智能 · 来源: cs.RO

Abstract:Infants are well known to develop their motor skills through dense interaction with caregivers. Although such social interaction is crucial for human development, motor-skill learning in robots is often treated as a unidirectional process in which robots passively receive demonstrations from tutors. This overlooks a key property of social interaction: it is inherently bidirectional, with tutor and learner dynamically adapting to each other. In such interactions, the robot's past experiences may function as prior constraints that shape the dynamics of their co-developed trajectories. We hypothesize that bidirectional tutoring allows such constraints to guide the formation of consistent behavioral patterns that preserve behavioral coherence and support generalization, whereas unidirectional interaction lacks such constraints and leads to broader, less consistent behavioral...

论文介绍 该研究指出机器人运动学习常被视为单向过程,忽略了交互的双向性。提出双向辅导框架,强调辅导者和学习者动态适应,利用机器人的先验经验约束共同开发轨迹,以形成一致行为模式,支持行为连贯性和泛化,不同于单向学习导致的行为不一致。

Comparative Study on Agility, Efficiency, and Impact Absorption of Bipedal Robots with Active Toes

第一作者: Joong-Gil Kim · 方向: 具身智能 · 来源: cs.RO

Abstract:Human legs exhibit high efficiency, agility, and impact absorption, with toes playing a crucial role in these capabilities. While many attempts have been made to implement human-like toes in robots, they have not fully replicated human characteristics nor rigorously validated their benefits. We propose a 14-DOF biped robot emulating human toes' lightweight, high-torque, robust nature. To quantitatively analyze the effectiveness of the active toes in terms of agility, efficiency, and impact absorption, we developed a high-fidelity simulation training environment that reflects actual actuators with coupled transmissions and accurate power consumption. To ensure a fair comparison between configurations with and without active toes, we designed a minimal RL reward function and applied an identical training procedure to both. The simulation results indicate that, at 1.33 m/s...

论文介绍 该研究针对仿生脚趾在机器人中效益验证不足的问题,提出一个14自由度双足机器人,模拟人类脚趾的轻量高扭矩特性。通过高保真仿真环境和统一强化学习训练,定量比较有无主动脚趾在敏捷性、效率和冲击吸收上的性能,为仿生机器人设计提供依据。

ForEnt: A Multi-Modal Dataset for Characterizing Quadruped Robot Entrapments in Forest Environments

第一作者: Natapat Kirdwichai · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Legged robots are increasingly deployed in forests for ecological surveying and monitoring, yet their autonomy is often interrupted consequent to the challenges posed in traversing forest environments. Forest entrapments, for example, when a robot's legs are ensnared in vines or other vegetation, result in loss of stability and toppling. Such events not only disrupt the mission and require manual intervention, but also risk damage to the robot hardware. To address the absence of a dedicated dataset to investigate these failure modes in forest environments, we present ForEnt, a multi-modal dataset collected with the low-cost Unitree Go2 quadruped across eight forest sites in the Southampton Common Woodlands, UK. For our dataset, over approximately 1.7 km of traversals in 11 sequences were conducted, yielding 69 recorded entrapment events. ForEnt includes time-synchronized RGB-D...

论文介绍 针对四足机器人在森林环境中易被藤蔓等植被缠绕、导致任务中断和硬件损坏的问题,本研究构建了首个专门的多模态数据集「ForEnt」。该数据集基于低成本 Unitree Go2 机器人,在英国南安普敦公共林地收集了超过 1.7 公里的轨迹数据,记录了 69 次被困事件,包含时间同步的 RGB-D 等信息。此数据集旨在为研究森林环境下的机器人失败模式提供基准,以推动机器人自主性和鲁棒性的提升。

Safe Local Navigation for Ackermann-Steered Robots in Unmapped Environments

第一作者: Christian Schaible · 方向: 导航与运动 · 来源: cs.RO

Abstract:A control framework is proposed for safe local navigation of mobile robots equipped with Ackermann steering in unmapped environments where a global goal is absent. Based on local obstacle detections, the safest heading angle is determined along the direction of the largest open space ahead of the vehicle. Guided by this direction, bounding lines are constructed on the left and right sides of the vehicle to achieve obstacle separation. These bounding lines are obtained by solving a convex quadratic optimization that maximizes vehicle-to-obstacle clearance. Optionally, conditions are imposed on the bounding lines to preserve parallelism and smooth abrupt changes from prior control steps. A feedback-linearizing controller is then used to regulate the vehicle's distance from one or both bounding lines, effectively enabling tracking of a local reference path that preserves safety...

论文介绍 本文提出一种面向配备阿克曼转向机构的移动机器人的安全局部导航框架,适用于无全局目标且环境未知的场景。该框架基于局部障碍物检测,确定车辆前方最大开阔空间方向作为最安全航向,并通过求解凸二次优化问题构建左右边界线,以最大化车辆与障碍物的距离。控制器通过调节车辆与这些边界线的距离,实现对局部参考路径的安全跟踪,从而在没有先验地图的情况下实现可靠避障。

DF-ExpEnse: Diffusion Filtered Exploration for Sample Efficient Finetuning

第一作者: Calvin Luo · 方向: 多模态具身 · 来源: cs.RO

Abstract:A natural recipe for intelligent robotic decision-making is initializing from pretrained generative control policies, which have summarized offline experience, and adapting them to self-collected online experience. We present DF-ExpEnse, an exploration technique that improves the quality of online experience collection, thus increasing finetuning sample-efficiency. DF-ExpEnse leverages the multimodal modeling capabilities of the generative control policy to create an expressive and tractably evaluatable candidate set. It then utilizes an ensemble of critics to identify the action that best balances quality with high exploration interest. In fleet settings, DF-ExpEnse further enables cross-agent communication to facilitate collaborative exploration as a group. DF-ExpEnse can be seamlessly integrated with existing strategies that finetune pretrained generative control policies...

论文介绍 在利用预训练生成式控制策略进行机器人决策时,高效地收集在线经验以微调策略是关键挑战。本文提出「DF-ExpEnse」探索技术,它利用生成策略的多模态建模能力,构建一个可表达且易于评估的候选动作集。随后,通过一个评论家集成来识别兼顾质量与探索兴趣的最佳动作。在多机器人场景下,该方法还能促进智能体间的协作探索。该技术可与现有策略微调方法无缝集成,旨在提高在线学习效率。

Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation

第一作者: Ahmad Farooq · 方向: 策略学习 · 来源: cs.RO

Abstract:Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets. We present the first end-to-end framework for safety verification of learned multi-agent communication policies through policy abstraction: neural policies are distilled into interpretable decision trees, then formally verified, with empirical validation confirming that verified safety properties transfer to original networks. Our four-stage pipeline consists of domain-specific feature extraction from agent observations, decision tree distillation achieving 97.9% +/- 1.2% fidelity to neural policies, automated translation to PRISM probabilistic model checker specifications with complete...

论文介绍 多智能体强化学习中涌现的通信策略缺乏形式化安全保证,这限制了其在无人机群等安全关键系统中的应用。本文提出首个端到端框架,用于验证学习到的多智能体通信策略的安全性。该框架通过策略抽象,将神经网络策略蒸馏为可解释的决策树,并进行形式化验证。实验表明,经验证的安全属性能够转移到原始神经网络。此方法为复杂多智能体系统的安全部署提供了新的思路。

Fail-RAG : A Retrieval Augmented Generation Informed Framework for Robot Failure Identification

第一作者: Ameya Salvi · 方向: 具身智能 · 来源: cs.RO

Abstract:Industry automation is witnessing an evolution in robotics driven by both technological breakthroughs and societal changes: progress towards generalist robots, embodied and physical artificial intelligence (AI), and increasing labor shortage in this http URL intelligent autonomous robot needs to not only act according to planned motions but also react to any unexpected events. In this study, we focus on such unexpected events in warehouses where robots are used for material handling. Specifically, we refer to any unexpected events as failures and develop methods to detect robot operations related failures. Rule-based detection methods may break since the form of failures could change due to the dynamic nature of both environments and tasks. We propose 'Fail-RAG', a Retrieval Augmented Generation (RAG)-based failure detection framework where failure images and context...

论文介绍 在仓储自动化中,机器人需能应对任务执行中的意外失败事件。本文针对基于规则的检测方法因环境动态变化而可能失效的问题,提出了「Fail-RAG」框架。该框架利用检索增强生成技术,通过检索相关的失败图像和上下文信息来增强对当前运行状态的判断,从而更灵活、准确地识别机器人操作故障。这为提升工业机器人的自主性与可靠性提供了一种新的数据驱动方法。

One Demo is Worth a Thousand Trajectories: Action-View Augmentation for Visuomotor Policies

第一作者: Chuer Pan · 方向: 机器人操作 · 来源: cs.RO

Abstract:Visuomotor policies for manipulation have demonstrated remarkable potential in modeling complex robotic behaviors, yet minor alterations in the robot's initial configuration and unseen obstacles easily lead to out-of-distribution observations. Without extensive data collection effort, these result in catastrophic execution failures. In this work, we introduce an effective data augmentation framework that generates visually realistic fisheye image sequences and corresponding physically feasible action trajectories from real-world eye-in-hand demonstrations, captured with a portable parallel gripper with a single fisheye camera. We introduce a novel Gaussian Splatting formulation, adapted to wide FoV fisheye cameras, to reconstruct and edit the 3D scene with unseen objects. We utilize trajectory optimization to generate smooth, collision-free, view-rendering-friendly action...

论文介绍 视觉运动策略容易因初始条件或障碍物变化而导致性能下降,而大规模数据收集成本高昂。本文提出一个有效的数据增强框架,能从真实的手持鱼眼演示中,生成视觉上逼真的鱼眼图像序列及对应物理可行的动作轨迹。该方法引入了适配广角鱼眼相机的高斯泼溅公式,以重建和编辑包含未见物体的3D场景,并利用轨迹优化生成平滑无碰撞的动作序列,从而用一个演示生成大量训练数据。

pdSTL: Probabilistic Differentiable Signal Temporal Logic for Stochastic Systems

第一作者: Bennett Dogbey · 方向: 导航与运动 · 来源: cs.RO

Abstract:Autonomous robots operating in uncertain environments must satisfy complex temporal and safety specifications despite stochastic dynamics and sensing noise. While Signal Temporal Logic (STL) offers robustness measures for gradient-based optimization, existing extensions either lack differentiability or ignore belief-space uncertainty. We introduce pdSTL (probabilistic differentiable Signal Temporal Logic), a framework that unifies probabilistic semantics with differentiable robustness over belief trajectories. pdSTL employs interval-valued probabilistic semantics to compute conservative satisfaction bounds, propagated compositionally through the STL syntax tree. We formulate the temporal robustness evaluation as a recurrent, LSTM-style unfolding of STL operators, enabling linear-time, differentiable monitoring suitable for end-to-end trajectory optimization. We validate pdSTL...

论文介绍 在存在随机动力学和感知噪声的不确定环境中,自主机器人需要满足复杂的时序与安全规范。信号时序逻辑提供了基于梯度优化的鲁棒性度量,但现有扩展在可微性或对信念空间不确定性的处理上存在不足。本文引入 pdSTL 框架,它统一了概率语义与可微分鲁棒性,通过区间值概率语义计算保守的满足度边界,并采用类似 LSTM 的递归展开实现线性时间、可微分的时序监控,适用于端到端轨迹优化。

SCAN-Planner: Spatial Collision-Aware Local Planning for Route-Guided Long-Range Quadruped Navigation

第一作者: Han Zheng · 方向: 导航与运动 · 来源: cs.RO

Abstract:Quadruped robots are increasingly expected to navigate through narrow passages, cluttered indoor scenes, and large-scale 3D unstructured environments. Existing local planners commonly approximate the robot using isotropic geometric inflation or rely on planar and elevation-map representations, leading to conservative motion in tight spaces and limited reasoning about overhanging structures. This letter presents SCAN-Planner, a spatial collision-aware local planning framework for long-range quadruped navigation. A yaw-aware twin-cylinder footprint is used to model the elongated robot body, enabling whole-body collision evaluation through sparse queries in an inflated 3D occupancy map. We further introduce a projected A* search that generates collision-free guidance on an interpolated ground-following surface, with z-gradient suppression to avoid obstacles horizontally while...

论文介绍 四足机器人需在狭窄通道、杂乱室内及大规模非结构化3D环境中导航。现有局部规划器常因简化机器人模型或使用平面地图而导致动作保守且忽略头顶障碍。本文提出「SCAN-Planner」,一个空间碰撞感知的局部规划框架。它采用偏航感知的双圆柱足迹模型来表征机器人身体,通过在膨胀的3D占据图中进行稀疏查询实现全身碰撞评估。进一步引入基于插值地表的投影A*搜索,生成无碰撞引导路径,并通过梯度抑制实现水平避障,从而支持长距离导航。

A Categorial and Sheaf-Theoretic Semantics for Autonomic Component Ensembles

第一作者: Manuel Hernández · 方向: 具身智能 · 来源: cs.RO

Abstract:The proliferation of large-scale, decentralized systems of autonomous agents, such as swarms of robots and networked cyber-physical systems, presents a formidable challenge to traditional formal methods. The Software Component Ensemble Language (SCEL) offers a formal model for such systems, but its operational semantics is not ideal for reasoning about global, structural, and emergent properties. This report proposes a new, multi-layered mathematical model for SCEL using category theory and sheaf theory. We argue that a society of robots described in SCEL can be formally modeled as a sheaf on a topological space, where components are points, ensembles are open sets, and distributed knowledge forms the sheaf's data. In this framework, computational processes like information sharing become equivalent to the sheaf-theoretic operation of "gluing" local data. System failures can...

论文介绍 本文针对大规模分散自主系统(如机器人蜂群)的形式化推理挑战,提出基于范畴论和束论的多层数学模型。该模型将机器人社会形式化为拓扑空间上的束,其中组件对应点、集合对应开集、分布式知识构成数据,并通过“粘合”操作描述信息共享过程,从而支持对系统全局、结构和涌现属性的分析。

Proprioceptive Invariant State Estimation for Humanoid Robots on Non-Inertial Ground

第一作者: Falak Mandali · 方向: 具身智能 · 来源: cs.RO

Abstract:This paper presents an invariant extended Kalman filtering (InEKF) approach for real-time state estimation of humanoid robots operating on non-inertial ground using only onboard proprioceptive sensing. The proposed approach estimates the robot's base position and velocity relative to the moving ground frame without requiring direct measurements of ground motion or externally mounted sensors. By exploiting kinematic constraints at the stance foot through foot-mounted IMUs, the filter accounts for ground-induced nonlinearities in the process and measurement models while remaining fully proprioceptive. The estimator is formulated to admit a right-invariant measurement model, enabling favorable error dynamics under large initial uncertainties. Observability analysis establishes conditions under which the robot's relative base position and velocity are observable with respect to...

论文介绍 本文研究人形机器人在非惯性地面上仅使用本体感觉传感的实时状态估计问题。核心方法是提出一种不变扩展卡尔曼滤波器,利用足部IMU的运动学约束处理地面引起的非线性,从而估计机器人基座相对于移动地面帧的位置和速度,无需外部传感器,增强了在复杂环境下的适用性。

Simulating Robotic Locomotion in Sand: Resistive Force Theory in an Open-Source Physics Engine

第一作者: Ryan Walker Brown · 方向: 导航与运动 · 来源: cs.RO

Abstract:Recent advancements in Resistive Force Theory (RFT) enable approximation of ground reaction forces for locomotion in sand without the computational expense of modeling interactions with individual grains. However, these tools have been absent in 3D physics engines commonly used for robot simulation. We explore if resistive force approximations are sufficient, when integrated with standard dynamics calculations, to provide a stable substrate for a freely walking robot. To determine this, we implement 3D Granular Resistive Force Theory (3D RFT) in a physics simulation engine, MuJoCo. We verify simulations in multiple scenarios to demonstrate that key trends due to end effector shape, speed, and loading are preserved. Our implementation predicts walking distance and foot sinkage of a 12-Degree of Freedom hexapod robot within 20\% of experiments in sand. While RFT has inherent...

论文介绍 本文探索将3D颗粒阻力理论集成到开源物理引擎MuJoCo中,以模拟机器人在沙中的运动。通过实现3D RFT来近似地面反应力,并验证模拟在多个场景下能保持关键趋势(如末端执行器形状和速度的影响),预测六足机器人的行走距离与实验相符,从而支持低成本、高效的颗粒介质运动仿真。

Playful Agentic Robot Learning

第一作者: Junyi Zhang · 方向: 具身智能 · 来源: cs.RO

Abstract:Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions. We study Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play as a continual skill-learning stage before downstream tasks arrive. We introduce RATs, Robotics Agent Teams designed for play-time skill acquisition. During play, RATs proposes novel yet learnable exploratory tasks, plans and executes robot-code policies, verifies intermediate progress, diagnoses failures, retries with dense, step-level feedback, and distills successful executions into a persistent code skill library. At test time, the agent reuses relevant skills from this frozen library to help solve new tasks. Experiments in LIBERO-PRO and...

论文介绍 本文研究通过自导向游戏实现机器人持续技能学习的问题。核心方法是引入RATs框架,其中具身编码代理在游戏阶段提出探索性任务、执行代码策略、诊断失败并将成功执行蒸馏到技能库中。这允许代理在下游任务前积累可重用技能,提升任务泛化能力和自主性。

DiffusionVS: A Generative Framework for Robust Visual Servoing Based on Diffusion Policy

第一作者: Hongkang Cui · 方向: 机器人操作 · 来源: cs.RO

Abstract:Visual servoing is a fundamental technique in robotic manipulation and navigation. Regression-based visual servoing frequently experiences trajectory jitter as a result of noise-sensitive single-step mappings and the accumulation of errors during distribution shifts. In contrast, Diffusion Policy maintains temporal consistency by predicting action sequences and improves robustness through implicit data augmentation. This paper presents a novel diffusion-based servoing method. Based on Diffusion Policy, the proposed approach uses normalized image coordinates of observed tag corners as input and generates camera velocity through conditional denoising. To overcome the generalization limitations of models trained on static datasets, an online training paradigm is adopted, continuously expanding the diversity of training data through interactive experience collection. This strategy...

论文介绍 本文针对视觉伺服中的轨迹抖动和鲁棒性问题,提出基于扩散策略的生成框架。方法使用归一化图像坐标作为输入,通过条件去噪生成相机速度序列以保持时间一致性,并采用在线训练范式收集交互经验扩展数据多样性,从而改善对噪声和分布偏移的适应性。

3D Scene Graphs: Open Challenges and Future Directions

第一作者: Dennis Rotondi · 方向: 机器人操作 · 来源: cs.RO

Abstract:3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including manipulation, navigation, task planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real-world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize...

论文介绍 本文综述3D场景图(3DSGs)作为时空AI的表示方法,结合几何、语义和关系抽象。通过统一定义和分析建模选择,指出领域在构建管道和评估协议上的碎片化问题,并讨论开放挑战(如鲁棒部署),旨在促进方法比较和未来研究方向,适用于机器人操作和场景理解等任务。

WorkBenchMark: A LEGO-Based Assembly Benchmark with an Assembly-by-Disassembly Baseline for the Smart Manufacturing League

第一作者: Wenbo Ma · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:We introduceWorkBenchMark, a LEGO Duplo-based robotic assembly benchmark motivated by the RoboCup Smart Manufacturing League. Robotic assembly couples low-level manipulation with task-level symbolic reasoning under physical constraints, a combination that current end-to-end learning methods do not yet solve reliably. The benchmark provides 400 tasks across four complexity tiers. We provide an open-vocabulary perception, Assembly-by-Disassembly baseline solution. Our planning-based pipeline outperforms a modern vision-language-action approach across all tiers. The benchmark, simulation environment, and baseline implementation will be released openly to support the broader robotic assembly community.

论文介绍 本文提出WorkBenchMark,一个基于LEGO Duplo的机器人装配基准,包含400个跨四个复杂度层级的任务。核心是提供开放词汇感知和基于规划的基线解决方案(装配-拆卸方法),该方案在基准上优于视觉-语言-动作方法,旨在推动机器人装配中低级操作与高级推理结合的研究。

Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots

第一作者: Khurram Javed · 方向: 策略学习 · 来源: cs.RO

Abstract:We built a robot called the Robotroller that actuates an Atari CX40+ controller and a device called the Atari Devbox that renders the game frame and the reward signal from the Arcade Learning Environment on a screen. The Robotroller and the Atari Devbox, together with an off-the-shelf camera and a desktop computer, constitute a system that can be used to study reinforcement learning algorithms in the physical world. We call the full system Physical Atari. In this paper, we detail the key decisions that make Physical Atari a robust and accessible platform. To make the system robust, we designed the Robotroller so that all movement is done through bearings, which reduces wear. Additionally, we wrote software that monitors the state of the servos at a high frequency and intervenes to limit stress. To make the system accessible, we used affordable off-the-shelf components and...

论文介绍 本文构建Physical Atari系统,包括Robotroller和Atari Devbox,用于在物理机器人上进行实时强化学习研究。通过定制设计(如使用轴承减少磨损)和廉价组件,确保系统鲁棒且可访问,从而支持强化学习算法在真实环境中的实验验证和推广。

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

第一作者: Yuyang Zhang · 方向: 具身智能 · 来源: cs.RO

Abstract:World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit...

论文介绍 研究问题是世界动作模型是否真的需要视频生成来连接视觉建模和机器人控制。核心方法是提出ImageWAM框架,利用预训练图像编辑模型进行动作预测,专注于目标帧转换和动作相关视觉差异。这可能提高机器人控制效率,减少推理成本和长时预测误差。

3D-DLP: Self-Supervised 3D Object-Centric Scene Representation Learning

第一作者: Ellina Zhang · 方向: 机器人操作 · 来源: cs.RO

Abstract:We introduce 3D-DLP, a self-supervised object-centric representation learning model that decomposes scene-level RGB-D or voxel observations into a set of 3D latent particles. Building on the Deep Latent Particles (DLP) framework, each particle encodes disentangled attributes, including 3D keypoint position, bounding box dimensions, and appearance features, and represents a distinct entity in the scene. The model learns interpretable per-particle segmentation maps through an end-to-end self-supervised reconstruction objective. We demonstrate on both simulated and real-world datasets that the learned latent space is interpretable and controllable: by manipulating particle positions and decoding, we can generate novel scene configurations. Furthermore, we show that leveraging these compact 3D latent particles for downstream robotic manipulation improves performance over baselines...

论文介绍 研究问题是学习可解释的3D场景表示以支持机器人操作。核心方法是提出3D-DLP模型,将RGB-D或体素观察分解为潜在粒子,编码解耦属性如关键点位置。通过自监督重建学习分割图,可用于生成新场景配置和改进机器人操作性能。

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

第一作者: Juncheng Ma · 方向: 具身智能 · 来源: cs.CV

Abstract:Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity. These limitations have sparked interest in egocentric human video as a scalable, substantially lower-cost, and more diverse alternative for embodied model pretraining. However, its effectiveness compared to teleoperated real-robot data remains underexplored. To address this question, we conduct a systematic study comparing egocentric human video and teleoperated real-robot trajectories as pretraining data sources for embodied foundation models, under fixed post-training and...

论文介绍 研究问题是自我中心人类视频能否作为比真实机器人数据更优的预训练来源。核心方法是进行系统比较,固定后训练设置,评估不同数据源对具身基础模型的影响。意义在于为数据瓶颈提供可扩展、低成本的替代方案。

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

第一作者: Ganlin Yang · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historical context, they either suffer from severe information bottlenecks, incur high latency via decoupled dual systems, or rely on unselective buffers that accumulate massive visual redundancies. To address these limitations, we introduce EventVLA, an end-to-end framework founded on the concept of sparse visual evidence memory that comprises two core components: foundational visual anchors to retain initial and short-term contexts, and a dynamic Keyframe Evidence Memory (KEM) module. Specifically, KEM directly predicts future keyframe probabilities from the VLA's latent embeddings to autonomously capture and store sparse...

论文介绍 研究问题是长时机器人操作中的记忆瓶颈,导致任务相关线索被遮挡时策略失效。核心方法是提出EventVLA框架,基于稀疏视觉证据记忆,动态捕获关键帧。这能增强VLA策略在复杂场景下的鲁棒性和性能。

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding

第一作者: Jianing Li · 方向: 多模态具身 · 来源: cs.CV

Abstract:Recently, vision-language models (VLMs) have made significant progress in 3D scene understanding, driving advances in applications such as embodied intelligence and robotic vision. However, existing approaches typically either rely directly on explicit 3D inputs (e.g., point clouds or RGB-D sequences), or introduce an additional 3D geometry encoder to derive 3D-aware visual tokens from 2D images. Such designs structurally decouple 3D geometric perception from the rich 2D semantics learned via vision-language pre-training, hindering the development of a unified 3D vision-language representation. In this work, we propose Occ-VLM, a novel framework for 3D scene understanding that operates purely on posed RGB images and employs a single 2D vision encoder. Specifically, Occ-VLM reconstructs 3D scene occupancy as an auxiliary geometric prior, which is utilized to spatially associate...

论文介绍 研究问题是统一3D几何感知与2D语义表示以改进室内场景理解。核心方法是提出Occ-VLM,仅使用姿态RGB图像,通过重建3D占用作为几何先验来空间关联视觉标记。可能应用于具身智能和机器人视觉任务。

Mix-QVLA: Task-Evidence-Aware Mixed-Precision Quantization of Vision-Language-Action Models

第一作者: Navin Ranjan · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:We propose Mix-QVLA, a task-evidence-aware mixed-precision PTQ framework for VLA models. Mix-QVLA anchors each quantized variant to the full-precision action-token reference decision and evaluates whether quantization preserves task-relevant evidence across key VLA functional boundaries. It computes normalized gradient-weighted task-evidence maps from boundary activations and compares full-precision and quantized maps using evidence-mass and attribution-distribution distortion, capturing changes in both the strength and allocation of decision-supporting evidence. A soft-bottleneck objective aggregates boundary-level degradation into layer-wise sensitivity scores. Mix-QVLA further models sensitivity throughout task execution, capturing phase-dependent shifts in layer importance rather than assuming a fixed sensitivity profile. The resulting evidence- and time-aware scores guide...

论文介绍 研究问题是VLA模型的高效部署,面临量化导致任务证据丢失的挑战。核心方法是提出Mix-QVLA框架,进行任务证据感知的混合精度后训练量化,评估并保持决策支持证据。这有助于在资源受限设备上部署VLA模型。

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

第一作者: Wenli Xiao · 方向: 机器人操作 · 来源: cs.AI

Abstract:Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to...

论文介绍 研究问题是真实世界机器人策略开发中依赖人工监督的瓶颈。核心方法是提出ENPIRE框架,通过物理反馈循环自动化策略改进,包括环境重置、执行、验证和迭代。这能减少人工干预,加速机器人操作算法研究。

市场总览

美股市场技术面呈现分化,标普500 ETF(SPY)和纳斯达克100 ETF(QQQ)均保持多头排列,RSI分别为54.1和59.1,接近52周高点,技术动量偏上行;但个股如微软(MSFT)和特斯拉(TSLA)显示空头排列,RSI低于50,板块内部走势不一。加密市场整体极度恐慌,恐慌贪婪指数为20,比特币(BTC-USD)价格63,876.33,RSI14为39.5,趋势熊市,空头排列,以太坊(ETH-USD)和Solana(SOL-USD)类似,技术面偏下行。中概股板块普遍承压,阿里巴巴(BABA)RSI24.7超卖,空头排列,拼多多(PDD)和京东(JD)也处于下行趋势。商品外汇中,黄金期货(GC=F)触发MACD金叉但趋势中性,RSI38.8;原油期货(CL=F)RSI31超卖但趋势中性;美元指数(DXY)呈现多头排列,MACD金叉,RSI69接近超买,显示短期强势。整体市场情绪谨慎,技术信号显示资产间走势分歧显著。

今日关注

DX-Y.NYB 美元指数 DXY
偏上行

当前价格100.76,RSI14为69接近超买区域,MACD金叉且值为0.3927高于信号线0.3043,趋势为多头排列,技术信号包括接近52周高和MACD金叉,显示短期动量偏强。

BABA 阿里巴巴 (BABA)
偏下行

当前价格107.1,RSI14为24.7处于超卖状态,MACD为-6.2971低于信号线-4.8371,趋势为空头排列,技术信号包括RSI超卖,表明下行压力持续。

QQQ Nasdaq 100 ETF
偏上行

当前价格740.62,单日涨幅2.51%显著,RSI14为59.1处于正常偏强区间,趋势为多头排列,技术信号包括接近52周高,显示技术动量向上。

^VIX VIX 恐慌指数
偏下行

当前价格16.78,单日跌幅9%,RSI14为46,趋势为空头排列,技术信号包括死叉(SMA50↓SMA200)和MACD死叉,表明市场恐慌情绪缓解。

CL=F WTI 原油期货
中性

当前价格76.3,RSI14为31接近超卖,但趋势为中性,信号列表为空,MACD为-5.1903低于信号线-3.706,技术面显示下行但趋势未确认反转。

全部资产

^VIX

VIX 恐慌指数

$16.78 -9.00%
5 日
-13.68%
距 52w 高
-52.5%
RSI(14)
46.0
趋势
空头
SMA 20 / 50 / 200
17.40 / 17.79 / 18.56
MACD / 信号
-0.108 / -0.053
死叉(SMA50↓SMA200) (6 天前)MACD 死叉 (今天)空头排列

^TNX

10Y 美债收益率 (%)

$4.45 -0.27%
5 日
-0.27%
距 52w 高
-10.9%
RSI(14)
47.4
趋势
多头
SMA 20 / 50 / 200
4.49 / 4.43 / 4.21
MACD / 信号
0.004 / 0.017
接近 52 周低多头排列

DX-Y.NYB

美元指数 DXY

$100.76 -0.09%
5 日
+1.02%
距 52w 高
-0.2%
RSI(14)
69.0
趋势
多头
SMA 20 / 50 / 200
99.67 / 98.96 / 98.70
MACD / 信号
0.393 / 0.304
MACD 金叉 (2 天前)接近 52 周高多头排列

SPY

S&P 500 ETF

$746.74 +0.78%
5 日
+1.22%
距 52w 高
-1.8%
RSI(14)
54.1
趋势
多头
SMA 20 / 50 / 200
747.08 / 729.66 / 688.36
MACD / 信号
3.930 / 5.524
接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$740.62 +2.51%
5 日
+3.28%
距 52w 高
-1.1%
RSI(14)
59.1
趋势
多头
SMA 20 / 50 / 200
726.88 / 693.18 / 628.63
MACD / 信号
9.523 / 11.111
接近 52 周高多头排列

AAPL

Apple

$298.01 +0.70%
5 日
+0.81%
距 52w 高
-6.1%
RSI(14)
50.9
趋势
多头
SMA 20 / 50 / 200
303.40 / 288.74 / 268.19
MACD / 信号
1.180 / 3.216
多头排列

MSFT

Microsoft

$379.40 +0.13%
5 日
-2.80%
距 52w 高
-31.7%
RSI(14)
34.9
趋势
空头
SMA 20 / 50 / 200
413.15 / 412.98 / 451.35
MACD / 信号
-8.571 / -3.544
空头排列

NVDA

Nvidia

$210.69 +2.95%
5 日
+2.84%
距 52w 高
-10.9%
RSI(14)
50.4
趋势
多头
SMA 20 / 50 / 200
211.79 / 209.31 / 189.90
MACD / 信号
-1.067 / -0.113
多头排列

GOOGL

Alphabet

$368.03 +1.17%
5 日
+2.87%
距 52w 高
-9.9%
RSI(14)
49.1
趋势
多头
SMA 20 / 50 / 200
371.63 / 367.37 / 311.10
MACD / 信号
-1.850 / -0.734
多头排列

TSLA

Tesla

$400.49 +1.04%
5 日
+0.34%
距 52w 高
-19.7%
RSI(14)
47.0
趋势
空头
SMA 20 / 50 / 200
413.70 / 402.49 / 416.96
MACD / 信号
-2.850 / -0.505
空头排列

META

Meta

$577.22 +1.70%
5 日
+1.55%
距 52w 高
-27.5%
RSI(14)
42.8
趋势
空头
SMA 20 / 50 / 200
599.48 / 621.90 / 654.92
MACD / 信号
-11.455 / -9.861
空头排列
加密恐慌贪婪
20
极度恐慌
加密总市值
$2.28 T
-0.42% / 24h
BTC 主导率
56.2%
ETH 9.1%
24h 成交量
$50.3 B
活跃币 17,413

BTC-USD

Bitcoin

$63,876.33 -0.57%
5 日
-2.63%
距 52w 高
-49.4%
RSI(14)
39.5
趋势
空头
SMA 20 / 50 / 200
63,691.54 / 72,111.41 / 76,726.91
MACD / 信号
-2,081.029 / -2,568.753
空头排列

ETH-USD

Ethereum

$1,726.60 -0.73%
5 日
-3.56%
距 52w 高
-65.1%
RSI(14)
41.1
趋势
空头
SMA 20 / 50 / 200
1,709.25 / 1,991.37 / 2,363.89
MACD / 信号
-68.409 / -90.192
空头排列

SOL-USD

Solana

$73.34 +0.23%
5 日
-0.10%
距 52w 高
-71.0%
RSI(14)
50.1
趋势
空头
SMA 20 / 50 / 200
69.01 / 79.76 / 97.41
MACD / 信号
-1.866 / -3.043
空头排列

BABA

阿里巴巴 (BABA)

$107.10 -0.32%
5 日
-4.96%
距 52w 高
-44.4%
RSI(14)
24.7
趋势
空头
SMA 20 / 50 / 200
120.91 / 129.30 / 149.19
MACD / 信号
-6.297 / -4.837
RSI 超卖空头排列

PDD

拼多多 (PDD)

$79.56 -0.38%
5 日
-2.14%
距 52w 高
-42.9%
RSI(14)
32.0
趋势
空头
SMA 20 / 50 / 200
85.43 / 93.76 / 110.59
MACD / 信号
-4.046 / -3.933
接近 52 周低空头排列

JD

京东 (JD)

$27.57 -1.22%
5 日
-1.75%
距 52w 高
-25.2%
RSI(14)
35.6
趋势
空头
SMA 20 / 50 / 200
29.07 / 30.07 / 30.25
MACD / 信号
-0.650 / -0.511
空头排列

0700.HK

腾讯控股 (0700.HK)

HK$451.20 +2.50%
5 日
-2.67%
距 52w 高
-33.9%
RSI(14)
47.7
趋势
空头
SMA 20 / 50 / 200
449.65 / 468.21 / 565.21
MACD / 信号
-4.277 / -4.887
空头排列

GC=F

黄金期货

$4,225.30 +0.03%
5 日
+0.24%
距 52w 高
-24.4%
RSI(14)
38.8
趋势
中性
SMA 20 / 50 / 200
4,360.96 / 4,546.97 / 4,439.58
MACD / 信号
-90.777 / -91.906
MACD 金叉 (2 天前)

CL=F

WTI 原油期货

$76.30 -0.39%
5 日
-10.11%
距 52w 高
-36.1%
RSI(14)
31.0
趋势
中性
SMA 20 / 50 / 200
87.48 / 93.89 / 73.73
MACD / 信号
-5.190 / -3.706

USDCNY=X

美元 / 人民币

¥6.77 +0.00%
5 日
+0.04%
距 52w 高
-6.1%
RSI(14)
42.8
趋势
空头
SMA 20 / 50 / 200
6.77 / 6.80 / 6.96
MACD / 信号
-0.010 / -0.012
接近 52 周低空头排列
风险提示

本报告基于公开行情数据计算的技术指标,仅供技术指标解读参考。过去走势不代表未来表现,技术分析具有局限性,投资决策需综合考虑多种因素。

Mideast Live Updates: Strains Emerge on First Day of U.S.-Iran Talks

Negotiations were expected to continue through the night, a U.S. official said. Iranian negotiators insisted on an end to the war in Lebanon as a condition for further talks, state media reported, as President Trump renewed threats against Iran.

中文摘要 美国与伊朗在瑞士举行首轮直接谈判,预计持续通宵。伊朗要求以结束黎巴嫩战争为继续对话的条件,同时特朗普总统对伊朗发出新威胁。

Here’s the latest.

中文摘要 此条目为“最新消息”更新,但未提供具体内容,信息不足。

Australia politics live: Tim Wilson says it’s ‘silly’ to ask if house prices should go up or down amid debate over property tax reform

Follow today’s news live Get our breaking news email, free app or daily news podcast Hume has no problem with more debate on NDIS changes, but criticises ‘horse trading’ between Labor and the Greens Jane Hume is speaking to RN Breakfast now, saying the Coalition has no problem with increased scrutin

中文摘要 澳大利亚政治实时更新中,议员蒂姆·威尔逊称在财产税改革辩论中询问房价是否应涨跌是“愚蠢的”。简·休姆表示联盟党不反对就NDIS变化进一步讨论,但批评工党与绿党之间的“交易”。

Abelardo De La Espriella, Trump-Backed Rightist, Headed for Win in Colombia

A victory for Abelardo De La Espriella, a lawyer with no previous political experience, would be a rebuke to the left and another win for the right in Latin America.

中文摘要 特朗普支持的右翼候选人阿贝拉多·德·拉·埃斯普列利亚在哥伦比亚总统选举中可能获胜。他是一名无政治经验的律师,其胜利将是对左翼的打击,并加强拉丁美洲右翼力量。

Far-right millionaire Abelardo de la Espriella wins Colombia’s presidential runoff

Leftwing opponent alleges vote count irregularities after Trump-endorsed lawyer secures narrow majority The Trump-admiring far-right millionaire lawyer and self-styled “outsider” Abelardo de la Espriella has won Colombia’s presidential runoff, defeating the leftwing senator Iván Cepeda. With 99.98%

中文摘要 极右翼百万富翁律师阿贝拉多·德·拉·埃斯普列利亚在哥伦比亚总统决选中获胜,击败左翼参议员伊万。左翼对手指控计票存在违规行为。

US-Iran talks strained as Trump threats spark Iranian walkout

Negotiations expected to continue through the night despite disruption caused by US president’s threat to bomb Iran and kidnap negotiating team High-stakes talks between the US and Iran were expected to continue into the early hours of Monday in Switzerland, a US official said, after a tense start t

中文摘要 美伊谈判因特朗普总统威胁轰炸伊朗并绑架谈判团队而引发紧张,伊朗一度退场。但谈判预计在瑞士继续进行至周一凌晨。

Iran war live: First day of US talks covers Lebanon, Hormuz, frozen assets

Trump threatens to hit Iran 'very hard' amid talks, prompting Ghalibaf to warn the US to take care with its rhetoric.

中文摘要 美伊谈判首日议题包括黎巴嫩冲突、霍尔木兹海峡和冻结资产。特朗普在谈判中威胁“严厉打击”伊朗,伊朗谈判代表加利巴夫警告美国谨慎言辞。

First round of direct US-Iran talks since deal expected to continue through the night

The US president, who is not at the talks, had earlier exchanged warnings with Iran's negotiator over clashes between Israel and Hezbollah in Lebanon.

中文摘要 美伊自协议以来的首轮直接谈判预计在瑞士持续通宵。特朗普总统未出席谈判,但早前就黎巴嫩境内以色列与真主党冲突与伊朗代表交换警告。

U.S. and Iranian Officials to Meet for Peace Talks in Switzerland

Amid a volley of threats, the brief talks largely focused on the war in Lebanon, according to Iranian state media.

中文摘要 美国与伊朗官员在瑞士举行和平谈判。据伊朗国家媒体报道,在双方互发威胁的背景下,短暂会谈主要聚焦黎巴嫩战争。

Almost three tonnes of cocaine found buried under Sydney property in Australia’s biggest ever seizure, police say

Australian federal police arrested and charged two men after allegedly finding 2.7 tonnes of cocaine in ‘bunkers’ under shipping containers Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Police have made what they say is Australia’s

中文摘要 澳大利亚联邦警方在悉尼一处房产下发现2.7吨可卡因,创下该国最大缉毒记录。两名男子已被逮捕并起诉,毒品藏匿在集装箱下的“掩体”中。

Trump faces fresh bipartisan criticism on Iran deal as Vance hails peace talks

Objections comes as Trump threatens to renew attacks on Iran if it doesn’t rein in its proxy in Lebanon US political figures from left and right voiced fresh objections on Sunday to Donald Trump’s provisional deal with Iran – even as the US president made new threats while Vice-President JD Vance ha

中文摘要 特朗普的临时伊朗协议遭到美国两党政治人物的新批评,尽管副总统万斯称赞和平谈判。特朗普威胁称,若伊朗不控制其在黎巴嫩的代理人,将重新攻击伊朗。

Leading Lebanese conservationist dies after Israeli airstrike on her home

Mona Khalil died Friday after an Israeli airstrike hit her beachside home two weeks ago. She's credited with creating a conservation movement in southern Lebanon to protect sea turtle nesting grounds.

中文摘要 黎巴嫩著名环保主义者莫娜·哈利勒在以色列对其海滨住宅的空袭中身亡。她因在黎巴嫩南部开创保护海龟筑巢地的环保运动而闻名。

Colombia weighs peace talks against a tougher approach

Colombia's government is touting a rare peace deal with a rebel group. But the front-runner in today's presidential election says he'll abandon negotiations. NPR's John Otis reports.

中文摘要 哥伦比亚政府正宣扬与一个叛乱团体达成的罕见和平协议。然而,在今日总统选举的领跑者表示将放弃谈判。

Inghams Shares Tumble as Farms Locked Down in Fight Against H5N1

Shares in Australia-listed chicken producer Inghams Group Ltd. dropped as much as 14% in early Monday trading after the company said it had locked down its Western Australia operations following the detection of the H5N1 avian influenza in the state.

中文摘要 澳大利亚鸡肉生产商Inghams Group股价周一早盘最多下跌14%,因西澳大利亚州检测到H5N1禽流感,公司封锁了当地农场。

Abaxx Calls for Probe, Hires Paul Weiss After Short Seller’s Attack

Abaxx Technologies Inc. is asking Canadian market regulators to investigate whether manipulative trading occurred in its shares, and has hired law firm Paul Weiss for help in response to a short-selling campaign by Viceroy Research.

中文摘要 Abaxx Technologies请求加拿大市场监管机构调查其股票是否被操纵交易,并聘请Paul Weiss律师事务所,以应对做空机构Viceroy Research的做空活动。

US, Iran Meet in Switzerland as Trump Threatens Iran

The US and Iran began talks in Switzerland on a peace deal to settle the issue of the Islamic Republic’s nuclear program as President Donald Trump once again threatened strikes if Hezbollah keeps attacking Israel. Bloomberg's Wendy Benjaminson has the latest. (Source: Bloomberg)

中文摘要 美国与伊朗在瑞士展开和平谈判,讨论伊朗核计划,但特朗普总统威胁若真主党继续攻击以色列将采取打击行动。

Gold Holds Decline After Trump Warns Iran During Peace Talks

Gold held a decline after US President Donald Trump issued a fresh threat to strike Iran, raising tensions during high-level talks to find a permanent resolution to the war that’s roiled global markets.

中文摘要 美国总统特朗普在高级别和平谈判期间威胁打击伊朗,导致黄金价格持续下跌,市场紧张情绪加剧。

Is Germany looking again at coal-powered electricity?

It had planned to abandon the fuel, but the higher cost of natural gas may make it think again.

中文摘要 德国原计划放弃燃煤发电,但因天然气成本上升,可能重新考虑使用煤电。

Eileen Gu on Values That Define Success

Skiing synthetizes all of the values and reflects back to me those values I judge myself by,” Olympic gold medalist Eileen Gu tells Bloomberg in an exclusive interview in Hong Kong. The freestyle skier also describes what a week trailing her would look like, if someone can keep pace. (Source: Bloomb

中文摘要 奥运金牌得主谷爱凌在香港接受专访,分享她定义成功的价值观。

Oil Rises, US Futures Slide on Tense US-Iran Talks: Markets Wrap

US stock futures dropped while oil climbed for a fourth day as talks between Washington and Tehran over a peace deal were clouded by a renewed threat from President Donald Trump to strike Iran.

中文摘要 美国股市期货下跌,油价连续第四天上涨,因美伊和平谈判受特朗普威胁打击伊朗的影响而蒙上阴影。

Oil Climbs After Fresh Trump Threat as US-Iran Peace Talks Begin

Oil gained after President Donald Trump threatened strikes on Iran if Hezbollah keeps attacking Israel, raising concerns about progress for peace talks between Washington and Tehran.

中文摘要 油价在美国总统特朗普威胁若真主党继续攻击以色列将打击伊朗后上涨,引发对美伊和平谈判进展的担忧。

Singapore Dollar Set to Gain Despite Hawkish Fed, Analysts Say

The Singapore dollar is poised to strengthen against the US dollar in the second half of the year despite a hawkish Federal Reserve boosting sentiment toward the greenback, according to strategists.

中文摘要 分析师预测,尽管美联储持鹰派立场支撑美元,新加坡元在下半年仍可能对美元升值。

Trump Ally Wins Initial Count in Colombian Presidential Vote

Conservative lawyer Abelardo de la Espriella narrowly won the preliminary vote count in Colombia’s presidential election, likely sweeping aside Gustavo Petro’s leftist movement and realigning Bogota with the US.

中文摘要 保守派律师Abelardo de la Espriella在哥伦比亚总统选举初步计票中险胜,可能击败Gustavo Petro的左翼运动,使波哥大与美国关系重新结盟。

【今朝の5本】仕事を始める前に読んでおきたい厳選ニュース

米イラン和平協議、ホルムズ原油輸送続く、英首相が近く退任表明か、重要鉱物、日産取締役案に不支持

中文摘要 内容涵盖美伊和平协议、霍尔木兹海峡原油运输持续、英国首相可能即将宣布辞职、重要矿物供应、日产汽车董事方案不被支持。

(cs2)就在今天!

猎鹰夺得科隆major冠军,niko等了10年终于结束了 16 个帖子 - 16 位参与者 阅读完整话题

美团把GPT-5.5、Claude Opus 4.8免费使用

tabbit是美团近期上线的AI应用,分国际版和国内版双轨。 国际版:免费接入GPT-5.5、Claude Opus 4.8、Gemini 3.5 Flash,外加Kimi-2.6、GLM-5.1、MiniMax-M3。 国内版:仅提供国内模型(Kimi、GLM、MiniMax等),无海外旗舰。 国际版链接:https://www.tabbit.ai/ 36 个帖子 - 31 位参与者 阅读完整话题

又有人向openai举报team、优惠码、K12

OpenAI Developer Community – 21 Jun 26 Report on Multiple Abuse Methods Involving ChatGPT Business / Team, Codex,... ChatGPT Bugs chatgpt Hello OpenAI Team, I would like to report several abuse methods that appear to be occurring. These issues involve ChatGPT Business / Team, Business Codex, enterpr

【CHY公益站】站长即将失踪通知

马上就要期末考试了,大概是7月6号考,在这段时间我如果不是啥都干完了没事干是不会动电脑了 、带-free后缀的模型 、前面有 开发公司/ 的模型 7 个帖子 - 7 位参与者 阅读完整话题

【九幺】继续放送公益站余额

本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已

知道越多,越害怕发言

即使有些知识我认为已经深耕过了我还是不敢发表见解,老是认为自己学的还不够,怕有人跳出来打我脸 但是有些人,其他平台等一些人明明一知半解还敢大大大的发声,搞得我一直很郁郁… 为什么我敢在linuxdo发?那当然是因为linuxdo环境很好啊 ( 26 个帖子 - 23 位参与者 阅读完整话题

各位佬,最近我打算写一本小说<<我以残剑守蓝星>>,核心设定已经完成,前三卷人物设定和故事线也有。

我不会多说,避免破坏大家的阅读体验,下面放一张ai生成的核心设定图谱: 实在没人一起讨论,自己一个人实在太那啥了;一个人的思想太单一了,大家一起集思广益。 下面放一些,设定的文档,只有文件名,没有内容,可以大概看看: 这本书,我设定并没有只打算,只作为自己去写的书籍,这是一本可扩展,多个世界宇宙的小说;后面,这本我完本后,会开源出来,为什么现在没有开源,一是我怕被盗取(这些设定花了我不少脑细胞,真的很头疼),而是,设定需要以当前小说去验证。再加一个核心目录的截图。 后面,完本后,我也会出一份,扩展续作的文档。 对了,现在本人,没女朋友,工作也不稳定,上班也不能手机,各位佬不要催更,如果消息没有