每日简报

2026-06-18

← 历史归档

DeusData/codebase-memory-mcp

C · ★ 5,294 · 🍴 489 · 📈 371 stars today

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

中文介绍 一款高性能的代码智能服务器,采用MCP协议。它能将代码库快速索引成持久化的知识图谱,实现亚毫秒级查询,支持158种语言,并能大幅减少AI对话所需的Token消耗。适用于需要深度理解、检索和交互大型代码库的开发者和工具链。

n0-computer/iroh

Rust · ★ 9,648 · 🍴 451 · 📈 421 stars today

IP addresses break, dial keys instead. Modular networking stack in Rust.

中文介绍 一个基于Rust语言的模块化网络协议栈。其核心创新在于使用“密钥”而非易变的IP地址作为网络标识,从而建立更稳定、去中心化的网络连接。适用于构建对网络身份有更高要求、避免地址漂移的点对点应用或分布式系统。

Panniantong/Agent-Reach

Python · ★ 33,181 · 🍴 2,669 · 📈 1,161 stars today

Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.

中文介绍 一个为AI代理提供全网信息访问能力的工具。它通过单一命令行接口,免API费用地读取和搜索Twitter、Reddit、YouTube、GitHub、B站、小红书等多个平台的内容。旨在让AI代理能够获取和理解广阔的互联网信息。

meshery/meshery

TypeScript · ★ 11,019 · 🍴 3,458 · 📈 196 stars today

Meshery, the cloud native manager

中文介绍 一款云原生管理工具,用于管理Kubernetes和各种云原生基础设施。它提供统一的控制平面,帮助用户部署、配置和运维云原生应用和服务网格,简化多集群环境的复杂管理操作。

obra/superpowers

Shell · ★ 231,069 · 🍴 20,541 · 📈 1,129 stars today

An agentic skills framework & software development methodology that works.

中文介绍 一个旨在提升AI代理能力的技能框架与软件开发方法论。它提供了一套可行的实践指南,帮助开发者构建和集成有效的代理技能,将AI代理应用于实际的软件开发流程中。

google-research/timesfm

Python · ★ 21,897 · 🍴 2,130 · 📈 606 stars today

TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.

中文介绍 由谷歌研究团队开发的时间序列基础模型。它是一个经过大规模预训练的模型,专门用于进行时间序列数据的预测,可作为通用基底模型应用于各种时间序列分析场景。

RocketChat/Rocket.Chat

TypeScript · ★ 45,582 · 🍴 13,656 · 📈 22 stars today

The Secure CommsOS™ for mission-critical operations

中文介绍 一款安全、开源的全渠道通讯平台,定位为“安全通讯操作系统”。它专为对稳定性和安全性要求高的关键任务场景设计,提供实时聊天、邮件支持等企业级通讯功能,是Intercom、Zendesk等的开源替代方案。

continuedev/continue

TypeScript · ★ 33,901 · 🍴 4,699 · 📈 49 stars today

open-source coding agent

中文介绍 一款开源的AI编码代理。它能够集成到开发环境中,通过理解代码上下文,辅助开发者进行代码生成、补全、调试和问答,旨在提升编程效率。

penpot/penpot

Clojure · ★ 50,086 · 🍴 3,234 · 📈 70 stars today

Penpot: The open-source design tool for design and code collaboration

中文介绍 一款开源的设计与协作平台。它支持矢量设计、原型制作,并强调设计与代码之间的无缝协作,允许设计师和开发者在同一平台上高效工作,是Figma等商业工具的开源替代品。

krahets/hello-algo

Java · ★ 127,446 · 🍴 15,175 · 📈 96 stars today

《Hello 算法》:动画图解、一键运行的数据结构与算法教程。支持简中、繁中、English、日本語,提供 Python, Java, C++, C, C#, JS, Go, Swift, Rust, Ruby, Kotlin, TS, Dart 等代码实现

中文介绍 《Hello 算法》是一本数据结构与算法教程,以动画图解和一键运行的代码示例为特色。它支持Python、Java、C++等多种编程语言,适合希望通过直观方式学习编程基础的初学者和开发者。

Universal-Debloater-Alliance/universal-android-debloater-next-generation

Rust · ★ 7,643 · 🍴 328 · 📈 457 stars today

Cross-platform GUI written in Rust using ADB to debloat non-rooted Android devices. Improve your privacy, the security and battery life of your device.

中文介绍 一款用Rust编写的跨平台GUI工具,通过ADB命令为未Root的安卓设备卸载预装臃肿软件。它可以帮助用户提升设备的隐私安全性和电池续航能力,界面直观易用。

mattpocock/skills

Shell · ★ 133,590 · 🍴 11,603 · 📈 1,523 stars today

Skills for Real Engineers. Straight from my .claude directory.

中文介绍 一套面向“真正工程师”的技能集合,来源于作者的 `.claude` 目录配置。它本质上是一组用于AI助手(如Claude)的提示词和技能模板,旨在引导AI更有效地辅助软件工程任务。

yairm210/Unciv

Kotlin · ★ 10,656 · 🍴 1,844 · 📈 24 stars today

Open-source Android/Desktop remake of Civ V

中文介绍 一款开源的《文明5》游戏重制版,支持Android和PC桌面平台。它忠实再现了原版策略游戏的核心玩法,让玩家可以在移动端和电脑上免费体验经典的回合制策略。

freeCodeCamp/freeCodeCamp

TypeScript · ★ 449,140 · 🍴 45,089 · 📈 757 stars today

freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.

中文介绍 freeCodeCamp.org 的开源代码库与课程体系。它提供完全免费的编程课程,涵盖数学、编程和计算机科学等领域,通过实践项目帮助学习者掌握全栈开发等技能,是自学编程的知名平台。

bytedance/UI-TARS-desktop

TypeScript · ★ 36,697 · 🍴 3,701 · 📈 150 stars today

The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra

中文介绍 字节跳动推出的一套开源多模态AI代理技术栈。它旨在连接前沿的AI视觉/语言模型与代理基础设施,为构建能够理解和操作用户界面的智能代理提供框架和工具。

nautechsystems/nautilus_trader

Rust · ★ 23,809 · 🍴 3,017 · 📈 98 stars today

Production-grade Rust-native trading engine with deterministic event-driven architecture

中文介绍 一个生产级别的、原生使用Rust编写的量化交易引擎。它采用确定性的事件驱动架构,提供高性能和低延迟的交易执行能力,适合专业交易员和机构构建复杂的交易策略。

chatwoot/chatwoot

Ruby · ★ 32,372 · 🍴 7,713 · 📈 264 stars today

Open-source live-chat, email support, omni-channel desk. An alternative to Intercom, Zendesk, Salesforce Service Cloud etc. 🔥💬

中文介绍 一款开源的全渠道客户支持平台。它集成了实时聊天、邮件等多种沟通渠道,提供统一的客服工作台,是Intercom、Zendesk等商业客服软件的强大开源替代方案。

calesthio/OpenMontage

Python · ★ 5,328 · 🍴 1,007 · 📈 98 stars today

World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.

中文介绍 全球首个开源的、代理驱动的视频制作系统。它包含多个处理管道和大量工具,能将AI编码助手转变为完整的视频制作工作室,支持从创意到成片的自动化流程。

alexzhang13/rlm

Python · ★ 4,924 · 🍴 822 · 📈 43 stars today

General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.

中文介绍 一个通用的、即插即用的递归语言模型推理库。它支持在多种沙盒环境中运行递归语言模型,提供了灵活的框架,方便研究者实验和部署此类新型语言模型架构。

makeplane/plane

TypeScript · ★ 51,294 · 🍴 4,563 · 📈 89 stars today

🔥🔥🔥 Open-source Jira, Linear, Monday, and ClickUp alternative. Plane is a modern project management platform to manage tasks, sprints, docs, and triage.

中文介绍 一款现代的开源项目管理平台,定位为Jira、Linear、Monday.com的替代品。它提供任务、冲刺、文档管理和问题分流等功能,帮助研发团队更高效地进行项目规划和协作。

How to Build a Claude Code Agent Team That Runs in Loops (Exact Setup Inside)

@zodchiii · 22.7K 粉丝 · 1.2M 阅 · 500 赞 · 71 转

Most setups run agents once and hand you whatever comes out. A team that runs in loops keeps going until the work actually passes. Below is the setup in 3 files: the agents, the loop that drives

中文介绍 分享构建一个可在循环中运行的 Claude 代码代理团队的具体设置。通过三个文件定义代理和驱动循环,使团队持续工作直到任务通过验收,而非一次执行即止。

How To Build a Second Brain That Runs Itself With Obsidian (Full Course)

@eng_khairallah1 · 67.2K 粉丝 · 1.2M 阅 · 510 赞 · 73 转

You read maybe two hundred articles this year. A few dozen papers. Hundreds of threads. Save this Every second-brain method ever sold to you, Zettelkasten, PARA, the graph view, the daily note,

中文介绍 对比 Zettelkasten、PARA 等流行的第二大脑方法,介绍如何用 Obsidian 构建一个能自动整理所读文章、论文和线程的自运行系统,属于知识管理课程分享。

Owning vs. Renting Intelligence

@lqiao · 95.4K 粉丝 · 501.6K 阅 · 509 赞 · 86 转

Mythos got shut down this week. Whether you agreed with the decision or not is almost beside the point. A company built on top of intelligence it didn't control suddenly found itself exposed to

中文介绍 以 Mythos 平台关闭为例,探讨构建于非自有智能(如依赖特定大模型)之上的公司的脆弱性,强调企业应追求对核心智能的「拥有」而非「租用」。

Using Claude to go Viral on X… (Mr. Beasts Framework)

@mattepstein · 35.6K 粉丝 · 393.3K 阅 · 504 赞 · 26 转

Have you seen any of the launches below on your timeline? (you probably have).. What if I told you they all followed a repeatable viral science that can be 95% automatable with claude. In this

中文介绍 介绍一套可重复的爆款公式,并声称其中 95% 的环节可通过 Claude 实现自动化。内容侧重于如何利用 Claude 执行特定框架,以实现内容在 X 平台上的病毒式传播。

How to Create Loops with Claude

@mikenevermiss · 10.8K 粉丝 · 261.4K 阅 · 568 赞 · 67 转

stop making prompts. start designing loops. a prompt gets you one response. a loop gets you a system that keeps working after you close the laptop. Boris Cherny, who runs Claude Code at Anthropic, put

中文介绍 引用 Anthropic 专家的观点,强调从设计「提示」转向设计「循环」。提示只获得一次响应,而循环能创建一个在用户离开后仍持续工作的系统。

Lazymaxxing TikTok Slideshows: 600/month for $2

@athcanft · 19.1K 粉丝 · 205.6K 阅 · 514 赞 · 23 转

I've been mass-producing TikTok slideshows with AI and scheduling them weeks in advance. Zero filming. Zero editing. Zero daily posting grind. This article breaks down the exact system, step-by-step,

中文介绍 分享一套使用 AI 批量生产 TikTok 幻灯片视频并提前数周排期的系统。该方法无需拍摄、剪辑或每日发布,成本仅为每月 2 美元,月产可达 600 条。

Three Ways Codex Can Use a Computer

@jxnlco · 105.9K 粉丝 · 204.4K 阅 · 504 赞 · 47 转

Update: Computer Use is now Available in the EU/UK ;) Enjoy! There are three ways for Codex to use a computer: Computer Use, the Chrome extension, and the in-app browser. They overlap just enough to

中文介绍 介绍 Codex 使用计算机的三种方式:Computer Use 功能、Chrome 扩展和应用内浏览器。文章详细比较了这三种方法的异同与适用场景。

Factory 2.0: From coding agents to software factories

@matanSF · 20.2K 粉丝 · 123.1K 阅 · 529 赞 · 60 转

In 2023, we launched Factory with the mission to bring autonomy to software engineering. While others were using models to speed up coding, we set out to deploy autonomous Droids across the

中文介绍 阐述 Factory 2.0 的愿景:从编码代理升级为软件工厂。旨在超越使用模型加速编码的层面,目标是部署能够自主执行跨团队复杂工程任务的智能体。

Zen and the Art of AI Research

@jxmnop · 50.7K 粉丝 · 114.1K 阅 · 504 赞 · 57 转

So you want to do AI research? It's true that no one really teaches you how. Not directly, anyway. But it turns out that the way to get started is pretty simple: some combination of (i) reading and

中文介绍 探讨如何进行 AI 研究,指出入门路径是阅读与动手实践的结合。文章旨在分享一种平衡的研究心态与方法论,而非具体技术。

ORACLE: Official AI Agents Trade on Polymarket

@OracleMindAI · 21.0K 粉丝 · 105.0K 阅 · 2.8K 赞 · 582 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 介绍在预测市场 Polymarket 上运行的官方 AI 代理 ORACLE。文章指出,到 2026 年,自主 AI 代理已成为该市场最有效的策略之一,交易活动占比超 30%。

The Art of Loop Engineering

@sydneyrunkle · 7.9K 粉丝 · 74.7K 阅 · 565 赞 · 87 转

Agents are useful because they help us automate work by taking actions in the real world. But getting agents to do valuable work reliably takes more than just a good model: it requires a carefully

中文介绍 探讨「循环工程」的艺术。指出要让代理可靠地完成有价值的工作,除了好模型,更需要精心设计的系统架构来支撑其在现实世界中采取行动。

Agent harness engineering with Claude: 14-step roadmap from one agent to a self-improving system.

@0xCodez · 8.2K 粉丝 · 72.8K 阅 · 508 赞 · 79 转

Everyone’s talking about loops. Almost no one is talking about what the loop runs on. 9 out of 10 builders run Claude Code on the default harness - no rules, no subagents, no hooks, no memory. Then

中文介绍 提供一份 14 步路线图,指导如何为 Claude 代码构建从单一代理到自改进系统的代理执行框架(harness)。重点在于规则、子代理、钩子和记忆系统的配置。

The Window Has Closed

@AndrewCurran_ · 53.9K 粉丝 · 62.8K 阅 · 569 赞 · 61 转

If you used Fable while it was available, you know it is special in ways that will not show up on benchmarks. I post benchmarks all the time because they matter to many people, but for a long time

中文介绍 评价已下线的模型 Fable,认为其特殊优势无法在基准测试中体现。作者虽常发基准,但强调某些模型的实际体验(如创作流畅度)远胜于测试分数。

Agentic Code Review

@addyosmani · 400.1K 粉丝 · 52.9K 阅 · 522 赞 · 41 转

Coding agents are extraordinarily good now and getting better fast. The interesting consequence is that the hard part of engineering moved from writing code to deciding whether to trust it, which

中文介绍 探讨「代理式代码审查」。认为随着编码代理能力增强,工程难点已从编写代码转向决定是否信任这些代码,因此需要新的审查策略。

The Stanford STORM Method: How to Make Claude Research Like a PhD in Minutes

@heynavtoor · 143.5K 粉丝 · 47.7K 阅 · 538 赞 · 70 转

Most people use Claude like a search box. Ask, answer, close tab. They are leaving the best feature locked. Save this :) Stanford built a research system called STORM. In peer reviewed testing it

中文介绍 介绍斯坦福大学的 STORM 研究系统,并指导如何利用其方法论让 Claude 在几分钟内完成博士生级别的主题研究。旨在挖掘 Claude 深度研究潜力。

Claude's June 15 Agent SDK billing split: what it means for Agent builders

@unicity_labs · 125.6K 粉丝 · 47.6K 阅 · 534 赞 · 659 转

Programmatic agent usage moves onto a separate metered credit today. Here is what changes, and the controls we are building into AstridOS to manage it. What changed today Today, 15 June 2026,

中文介绍 解读 Anthropic 于 2026 年 6 月 15 日生效的代理 SDK 计费变更。说明编程式代理用量现在将使用单独的计量信用,并介绍其平台 AstridOS 的相关管理控制功能。

How to Start Your AI Content Journey: A Complete Step-by-Step Guide

@israfill · 8.8K 粉丝 · 46.5K 阅 · 508 赞 · 59 转

This is the guide I wish existed when I started. Follow it in order. Do not skip steps. Every command, every file, every decision is documented below. Who This Guide Is For You want to start

中文介绍 发布一份完整的分步指南,旨在帮助新手从零开始 AI 内容创作之旅。指南按顺序记录了所有必要的命令、文件和决策,属于保姆级教程。

How to build a self-improvement loop for your Skills

@zachlloydtweets · 10.6K 粉丝 · 43.5K 阅 · 512 赞 · 53 转

There’s been a lot of chatter about using “loops” lately to drive agents, and I think this has been accompanied by a bit of “what actually is a loop”? I can’t speak for everyone else using the term,

中文介绍 解析「循环」在代理开发中的实际含义,并澄清相关概念。旨在区分当前流行的术语,并分享作者在使用循环驱动代理(如用于技能提升)方面的实践见解。

My Thoughts on Loop Engineering

@samueljmcd · 857 粉丝 · 42.2K 阅 · 511 赞 · 52 转

Loop engineering is the new label. The hard part is the one it has always been. Verification. Tip: You can copy and paste this article into Claude and ask for the best insights if you don't want to

中文介绍 评论「循环工程」这一新标签,认为其核心难点始终是「验证」。文章可直接喂给 Claude 提取关键见解,属于对当前代理工程趋势的简短反思。

The 10 rules to ship truly polished UI with Claude

@kvnkld · 7.2K 粉丝 · 42.0K 阅 · 531 赞 · 25 转

People keep asking how the UI components I post end up looking so polished, or what prompts I use. So here’s a breakdown of the most important things: Polish is not a feature you prompt for. You can't

中文介绍 总结使用 Claude 打磨出精致 UI 组件的 10 条规则。强调「精致」并非直接提示可得,而是需要一套系统性的提示策略和设计原则。

🔬 The Self-Driving Lab — Joseph Krause, Radical AI

Radical AI's Joseph Krause on why the moat in materials is the lab, not the model

中文介绍 Radical AI的Joseph Krause表示,材料科学领域的竞争壁垒在于实验室,而非AI模型本身。

Introducing LifeSciBench

Introducing LifeSciBench, an expert-authored, expert-reviewed benchmark for evaluating how AI systems handle real-world life science research tasks and decisions.

中文介绍 OpenAI推出LifeSciBench基准测试,旨在评估AI系统处理真实世界生命科学研究任务与决策的能力。

Unlocking UK house-building with AI-accelerated planning

UK government partners with Google DeepMind to build a new AI-powered prototype aimed at faster housing decisions.

中文介绍 英国政府与谷歌DeepMind合作,旨在利用AI加速住房规划审批流程,以促进房屋建设。

Want to get a data center online quickly? Give it some flex.

At the end of a tense and scoreless first half of a soccer match between the English men’s team and rival Germany, millions of Brits let out a collective sigh and did what they so often do in moments of stress: They made tea. That wave of electric kettles clicking on, however, caused a different…

中文介绍 为快速让数据中心上线运营,需要赋予电网更多灵活性,以应对瞬间的高电力需求。

Predicting model behavior before release by simulating deployment

OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.

中文介绍 OpenAI推出“部署模拟”方法,通过模拟真实对话数据来预测AI模型部署后的行为,以提升安全性评估。

Why do South Koreans love AI so much?

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. When I landed in Seoul after a grueling 12-hour flight from San Francisco, I walked through an unmanned immigration checkpoint, where a machine scanned my face an

中文介绍 分析探讨韩国社会对人工智能接受度高的原因,体现在无人值守设施等日常应用中。

Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents

第一作者: Ankita Samaddar · 方向: AI 安全

Abstract:With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents trained via Reinforcement Learning (RL). These agents employ neurosymbolic approaches such as behavior trees with learning-enabled components (LECs) to learn, reason, adapt, and implement security rules while maintaining critical operations. However, these autonomous networks are partially observable systems, i.e., the cyber-attacker's (red agent's) actions are not observable, making it difficult for the defender to predict red actions, learn red policies, or assess the attacker's intrusion levels. To address this, we propose a Policy Learning Technique using imitation learning to learn policies for partially observable RL agents with discrete states and discrete actions. We apply this technique in an autonomous cyber environment to predict red...

论文介绍 研究部分可观测网络环境中红队行为不可预测导致的防御挑战。提出一种基于模仿学习的策略学习技术,用于神经符号自主网络代理,通过从观测数据学习红队策略来增强防御。该方法可应用于自动化网络安全系统,提高对未知攻击的适应能力。

Gatling: Rapid-Fire Consensus from Parallel Composition

第一作者: Giulia Scaffino · 方向: 密码学协议

Abstract:Consensus protocols form the core of blockchains and other replicated state machines, ensuring that all correct nodes process the same totally ordered log of input transactions. In fault-free executions, performance is driven by the good-case transaction latency -- the time between a transaction becoming known to all nodes and its confirmation by the consensus protocol -- which depends on both how frequently proposals are made and, once made, how quickly they are confirmed. While prior work has established tight lower bounds on confirmation latency that modern protocols already achieve, it remains open whether the inter-proposal time can be further reduced below the state-of-the-art of one network delay. We introduce Gatling, an atomic broadcast protocol that achieves arbitrarily small inter-proposal times under rotating leader schedules; in particular, smaller than the...

论文介绍 探讨区块链共识协议中交易延迟的优化问题。提出Gatling原子广播协议,通过并行组合和旋转领导调度实现极小的提案间隔时间,突破现有网络延迟下限。该设计可能提升分布式系统的吞吐量和响应速度。

Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners

第一作者: Xiaojun Jia · 方向: 软件安全

Abstract:Agent skills are emerging as an important attack surface in LLM-based systems. Through an empirical study of existing skill scanners, we find that current defenses primarily rely on textual descriptions, manifests, and source code as the main signals for security analysis, which can leave visually conveyed malicious intent insufficiently examined. This creates a practical blind spot: harmful operational instructions hidden in images may bypass scanning while still being recoverable by multimodal agents during deployment. To systematically investigate this threat, we propose SkillCamo, a document-mediated multimodal instruction attack that conceals malicious instructions within images bundled with a skill while rewriting the surrounding documentation to naturally reference those images as part of the normal workflow. Thus, the attack does not rely on the image alone, but on the...

论文介绍 分析多模态LLM系统中技能扫描器的安全缺陷,指出其主要依赖文本分析而忽略图像中的恶意指令。提出SkillCamo攻击,通过在图像中隐藏恶意指令并重写文档来绕过扫描。该研究揭示了多模态安全的新威胁,为扫描器改进提供方向。

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models

第一作者: Nicola Franco · 方向: AI 安全

Abstract:We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundreds of thousands of adversarial attempts were generated and every apparent success was independently re-adjudicated by a panel of three judge models (majority vote). Both models resist the majority of attacks, but the residual surface is larger than aggregate framing suggests: it is dominated by adaptive iterative attacks, while static obfuscation is near-fully neutralised. The strongest adaptive search (tree-of-attacks) breaks Opus 4.8 on 11.5% of intents overall, whereas Fable 5 stays in the single digits (6.1% worst-case). Aggregate rates therefore should not be read as...

论文介绍 评估Anthropic两个前沿LLM对自动越狱攻击的鲁棒性。使用HackAgent框架进行大规模红队测试,发现自适应迭代攻击更易成功。结果表明,静态防御有效但自适应攻击构成显著威胁,需加强LLM安全防护。

Multi-Source Cybersecurity Logs: An ATT&CK-Labeled Dataset and SLM Evaluation

第一作者: Abir Ashab Niloy · 方向: AI 安全

Abstract:Multi-stage cyberattacks span system, network, and browser logs. Detecting them requires correlating events across all three sources. Machine learning methods can learn these cross-source patterns, but they need labeled multi-source data. Existing public datasets fall short. Network-only datasets such as CICIDS and UNSW-NB15 miss host and browser activity. Host-focused datasets such as LMDG and CICAPT-IIoT lack browser telemetry. ATLAS includes all three sources but labels events only as malicious or benign, without MITRE Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK) technique granularity. No public dataset combines all three sources with per-entry ATT&CK technique labels. We close the gap by building a multi-source log dataset of 870 sessions (70 attack, 800 benign) and approximately 2.3 million events. We captured system, network, and browser activity...

论文介绍 针对多阶段网络攻击检测需要多源数据的问题,构建了一个包含系统、网络和浏览器日志的ATT&CK标签数据集。评估小型语言模型在该数据集上的性能,为跨源攻击模式学习提供基础数据支持。

Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports

第一作者: Ahmed Ryan · 方向: AI 安全

Abstract:Classifying Cyber Threat Intelligence (CTI) using MITRE Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK) is essential for proactive defense, but historically required extensive human effort. Pre-Large Language Model (LLM) automation sped up this process, but could not resolve the complex language and multi-step attack patterns found in unstructured CTI reports. LLMs addressed previous limitations by using contextual reasoning to understand unstructured text. However, current evaluations rely on simplified, single-technique sentences that ignore the complexity of real-world CTI reports, which often leads to inflated performance results. Consequently, the baseline performance of open-source LLMs on complex unstructured CTI reports remains unevaluated. To address this gap, we constructed a ground-truth dataset of 2,076 human-annotated sentences (1,281...

论文介绍 评估开源LLM在无结构CTI报告中进行多标签ATT&CK技术分类的能力。构建了真实数据集进行基准测试,指出现有评估过于简化,揭示了LLM在复杂报告上的性能基线,助力自动化威胁情报分析。

Structural Role Injection in Handlebars-Templated LLM Prompts: Triple-Brace Interpolation, Delimiter Family, and the Limits of HTML Auto-Escaping

第一作者: Mohammadreza Rashidi · 方向: AI 安全

Abstract:Large language model applications build prompts from templates, and Handlebars is a widely used templating engine and the default prompt-template format in Microsoft Semantic Kernel. Its double-brace {x} expression HTML-escapes the interpolated value and is documented as the safe default; its triple-brace {x} expression inserts the value raw. We show that this choice silently governs an application's exposure to structural role injection, where attacker-controlled data carries chat role delimiters that forge a higher-privilege turn. A model-free analysis establishes the mechanism: Handlebars escaping rewrites angle brackets but not square brackets, colons, or Markdown hashes, so it neutralises ChatML, Llama-3, and XML role delimiters (survival rate 0.00) while leaving Llama-2 [INST], legacy Human:/Assistant:, and Markdown ### delimiters intact (survival rate 1.00 for the last...

论文介绍 研究Handlebars模板引擎中HTML转义对LLM提示结构化角色注入攻击的影响。分析显示,双花括号转义可防御某些分隔符,但三花括号存在风险。该工作指导安全模板设计,防范提示注入漏洞。

An Empirical Analysis of AI Slop in Music Streaming

第一作者: Stanley Wu · 方向: AI 安全

Abstract:Generative AI models lower the bar for content creation, making it easy for any user to create professional-looking images, text and music with minimal effort. This has enabled a new cottage industry around creation of "AI slop" mass quantities of mediocre content produced to generate revenue, often through misrepresentation as human-authored content, or scams involving automated scripts and fake consumption. While there are obvious parallels between the AI-slop industry and "traditional" email spam networks, it might be too early to determine if AI slop generation can grow into a similar self-sustaining industry. In this paper, we look specifically at the music industry, and explore the question: Can we prevent AI music slop from growing into a self-sustaining shadow industry? To answer this question, we characterize the current state of AI slop in music, and its pipeline...

论文介绍 实证分析音乐流媒体中AI生成劣质内容(AI slop)的现象,探讨其可否发展为自给自足的阴影产业。通过特征化当前状态和生产流程,研究防止AI音乐内容恶化的策略,应对生成式AI带来的内容质量挑战。

ShellGames: Speculative LLM-Driven SSH Deception

第一作者: Umberto Salviati · 方向: AI 安全

Abstract:Cyber deception and Moving Target Defense are promising strategies that aim to disrupt adversaries by increasing uncertainty. However, sustaining long-lived, credible interactive sessions with adversaries remains an open challenge. Large Language Models (LLMs) offer a promising path toward more dynamic deception systems, but suffer from key limitations that fundamentally limit their applicability, including: lack of persistent state, output inconsistencies, hallucinations, latency, and susceptibility to behavioral subversion that may reveal the deception. We propose ShellGames, an SSH shell simulator based on LLM designed to address these limitations. ShellGames combines five complementary techniques: (i) Automatic Chain-of-Thought and few-shot learning to improve correctness; (ii) memory management to maintain system state coherency; (iii) speculative command execution to...

论文介绍 本文针对基于大型语言模型的网络欺骗系统缺乏持久状态、输出不一致等问题,提出了ShellGames。该系统是一个SSH shell模拟器,结合了思维链提示、记忆管理、推测性命令执行等技术,旨在生成更具可信度和持续性的欺骗会话,从而在对抗中增加不确定性,提升网络主动防御能力。

Children Are Not the Enemy: Child-Fit Security as an Alternative to Bans and Surveillance

第一作者: Kopo M. Ramokapane · 方向: 隐私保护

Abstract:Digital technologies are now central to children's learning, play, communication, identity formation, and social participation. Yet dominant approaches to children's online safety often rely on containment mechanisms, including bans, age gates, parental controls, monitoring, and screen-time restrictions. These approaches can be useful in specific contexts, but they often frame child protection primarily as a problem of restricting access to systems designed for adults. In this paper, we argue that this framing is inadequate for children's digital lives and insufficient as a security paradigm. We propose Child-fit security, a design paradigm in which technologies likely to be used by children treat a child as legitimate users, not attackers to be excluded, vulnerabilities to be patched, or risks to be managed. In this paradigm, children's wellbeing, development, privacy...

论文介绍 当前保护儿童在线安全的主流方法,如禁令和监控,常将儿童视为需被限制的潜在风险方。本文提出「Child-fit安全」范式,主张将儿童视为合法用户而非攻击者,在技术设计阶段就纳入其福祉、发展和隐私需求。该范式旨在构建更符合儿童使用场景、更具包容性的数字环境,而非简单地实施访问限制。

Beyond Native Success: Auditing Deployment-Interface Exposure of CLIP Backdoors

第一作者: Kunlan Xiang · 方向: 安全研究

Abstract:Contrastive Language-Image Pre-training models are widely reused across downstream interfaces, including feature extraction, retrieval, reranking, and selection. Existing CLIP backdoor, however, usually validate attacks on a small attack-native task, leaving unclear whether the same poisoned checkpoint remains exposed, weakens, or becomes not applicable when reused through other interfaces. We introduce DIFE, a Deployment-Interface Footprint Evaluation framework that audits backdoored CLIP checkpoints across deployment interfaces. DIFE makes various evaluations comparable by specifying each interface's component readout, trigger channel, target event, reference condition, and metric. DIFE also introduces effective-footprint diagnosis to identify the reusable CLIP component or component combination that carries exposure and explains where risk transfers. Auditing reproduced...

论文介绍 现有针对CLIP模型的后门攻击研究多在单一攻击任务上验证,但模型常通过特征提取、检索等不同接口被复用。本文引入DIFE框架,用于系统性审计受污染的CLIP检查点在不同部署接口下的暴露风险。该框架通过定义统一评估组件,使跨接口风险比较成为可能,并能诊断风险转移的具体路径。

Anywhere, Any-Stymie: Remote Activation of Trojan Malware on LiDAR with Modulated Signals

第一作者: R. Spencer Hallyburton · 方向: 系统安全

Abstract:LiDAR sensors are widely deployed in autonomous systems for 3D perception and safety-critical decision-making. We identify a previously unexplored attack surface in which dormant malware embedded in the LiDAR sensing pipeline remains inactive during normal operation and can be externally triggered after deployment, without requiring access to sensor hardware or networking at attack time. To operationalize this threat, we design malware capable of low-level point-cloud manipulation and embed it into LiDAR firmware. This malware was developed in a closed research test environment with vendor technical support, rather than by exploiting an inherent production supply-chain vulnerability. To selectively trigger attack activation, we design and implement an optical trigger that remotely activates the malware by delivering a modulated signal into the sensing environment. Once...

论文介绍 本文揭示了一种新的LiDAR传感器攻击面:攻击者可在设备中预埋休眠恶意软件,并在部署后通过调制光信号远程触发。该恶意软件能篡改点云数据,影响自动驾驶系统的3D感知与决策。此攻击无需在触发时访问硬件或网络,对依赖LiDAR的自主系统构成潜在的隐蔽威胁。

An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

第一作者: Joseph Walusimbi · 方向: 密码学协议

Abstract:Banks simultaneously face signature-based fraud (card-not-present attacks, account takeover, ATM cloning) and behavioural financial crime (structuring, layering, mule networks, business email compromise) -- two threat families with fundamentally different detection requirements. Static rule engines that reliably catch brute-force and high-velocity events are structurally blind to business-email-compromise (BEC) payment redirection, session hijacking, and money-laundering layering, which are engineered to appear indistinguishable from legitimate activity at the individual transaction or session level. This paper presents an AI security agent for retail and corporate banking that addresses this gap through a three-component fusion architecture operating on two parallel event streams: a transaction stream (card fraud, ACH/wire fraud, AML categories) and a session stream (account...

论文介绍 银行面临基于规则的交易欺诈和行为性的金融犯罪两类不同威胁。本文提出一个面向零售与企业银行业务的AI安全代理,它采用三组件融合架构,同时分析交易流和会话流数据。该系统旨在克服静态规则引擎的局限,检测如商务电邮诈骗、会话劫持和洗钱分层等复杂威胁。

SNAS: A Multi-Layer Defense-in-Depth Architecture for Secure Egress in Sandboxed Workloads

第一作者: Niranjan Kumar Sharma · 方向: 系统安全

Abstract:Snowpark enables data engineering and AI/ML workloads in Snowflake by executing user-defined functions in secure sandboxes. Many of these workloads require external connectivity to access cloud APIs, external databases, or feature stores, creating a dependability challenge: how to provide transparent network access while preserving strict multi-tenant isolation and resource fairness. This paper presents Secure Network Access in Snowpark (SNAS), a production architecture for secure external communication from sandboxed workloads. SNAS combines Extended Berkeley Packet Filter (eBPF) packet filtering, Generic Network Virtualization Encapsulation (GENEVE) overlay networks, and distributed egress proxies for policy-driven egress control with low overhead. We describe the design, deployment, and measured production behavior of SNAS, including an eBPF-based bandwidth limiter using...

论文介绍 在云平台沙盒中运行用户任务时,如何安全地提供外部网络连接是一个挑战。本文介绍了SNAS架构,用于Snowpark沙盒工作负载的安全出口通信。该架构结合了eBPF数据包过滤、GENEVE网络封装与分布式出口代理,实现了策略驱动的、低开销的出站流量控制,在保持多租户隔离的同时满足外部访问需求。

PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents

第一作者: Aaditya Pai · 方向: AI 安全

Abstract:Prompt injection defenses evaluated on synthetic benchmarks do not generalize to real enterprise documents, which are longer, denser, and interleave legitimate authority language with factual content. We demonstrate this gap with a real-document benchmark of 122 tasks across five professional domains (financial, legal, medical, scientific, DevOps) using actual SEC filings, Federal Register rules, PubMed abstracts, arXiv papers, and GitHub postmortems. Paraphrasing, the strongest defense on synthetic benchmarks, shows no statistically significant attack success rate reduction on real documents (p=0.500) while degrading utility from 91.8% to 82.8%. We introduce PARSE (Provenance-Aware Retrieval Sanitization), a domain-aware, fact-preserving sanitization pipeline that classifies each sentence by injection likelihood, extracts structured facts before rewriting, and verifies fact...

论文介绍 现有基于合成数据评估的提示注入防御方法,在面对真实、冗长的专业文档时效果不佳。本文通过构建涵盖金融、法律等五个领域的基准,验证了这一缺陷,并提出了PARSE防御管线。该方法通过识别句子注入可能性、提取结构化事实并改写,在净化注入指令的同时,尽可能保留文档中的有用事实信息。

Bifrost: Hybrid TEE-FHE Inference for Privacy-Preserving Transformer and LLM Serving

第一作者: Chenghao Chen · 方向: 密码学协议

Abstract:Cloud-hosted transformer and large language model (LLM) inference creates a direct confidentiality problem: user prompts may contain sensitive code, business data, personal information, or regulated documents, yet remote serving exposes intermediate state to the cloud software stack and accelerator runtime. Fully homomorphic encryption (FHE) keeps accelerator-side execution ciphertext-only, but end-to-end LLM inference remains expensive because linear layers are interleaved with non-linear, cache-state, and refresh-sensitive operators. CPU trusted execution environments (TEEs) can execute those operators natively, but a CPU TEE alone does not define how an untrusted accelerator should participate. We present Bifrost, a hybrid TEE-FHE serving architecture in which secrets are provisioned only to an attested CPU TEE, while the accelerator, device memory, driver/runtime stack...

论文介绍 云端大语言模型服务存在用户提示泄露风险。全同态加密能保护计算隐私但效率低,而CPU可信执行环境无法直接保护加速器。本文提出Bifrost混合架构,将需要隐私保护的算子在CPU TEE中执行,将计算密集的线性层交由加速器使用同态加密处理。该架构旨在为transformer和LLM推理提供实用的隐私保护方案。

SoK: AI-Augmented Binary Reversing

第一作者: Yujeong Kwon · 方向: 软件安全

Abstract:Binary reversing is fundamental to software understanding, vulnerability discovery, malware investigation, and firmware auditing. However, it remains inherently challenging due to the irreversible loss of semantic information during compilation. Recent advances in machine learning, large language models (LLMs), and agentic AI systems have accelerated the adoption of AI-augmented binary reversing. Yet, the resulting body of work has become increasingly fragmented across reversing domains, artifact representations, learning approaches, and evaluation practices. This paper presents the first comprehensive systematization of knowledge on AI-augmented binary reversing. We analyze 144 research papers published since 2015, and organize them into 22 binary reversing domains according to the inference tasks. We further introduce a unified taxonomy spanning conventional and AI-augmented...

论文介绍 该论文系统化地总结了人工智能增强二进制逆向工程领域的研究。作者分析了自2015年以来的144篇相关论文,将逆向任务归纳为22个领域,并建立了一个统一的分类法。这项研究梳理了传统方法与AI增强方法的现状与挑战,为该领域的后续发展提供了全面的知识框架。

OTRO: Oblivious Tokenization Path with Square-Root ORAM

第一作者: Jonghyun Lee · 方向: AI 安全

Abstract:The CPU-side large language model (LLM) tokenizer is a critical security gap in LLM serving through a confidential computing stack with CPU and GPU trusted execution environments (TEEs). Tokenizers converts the prompts through table-driven lookups, and the resulting memory access patterns are a powerful source of side-channel leakage. Recent work demonstrates end-to-end recovery of user prompts from tokenizer access pattern on production Intel TDX. However, a drop-in use of the popular tree-based Oblivious RAMs (e.g., PathORAM) to prevent access-pattern leakage introduces $\sim$13$\times$ tokenizer slowdown, resulting in 10-58% higher time-to-first-token (TTFT). In this paper, we present OTRO, an efficient, oblivious tokenization path tailored to latency-critical LLM serving. OTRO relies on square-root ORAM for fast single-access lookups, but avoids its prohibitive...

论文介绍 本文针对大模型服务中CPU端分词器的侧信道泄露问题,提出了OTRO方案。该方案基于Square-Root ORAM设计了高效的遗忘分词路径,旨在防止攻击者通过内存访问模式窃取用户提示词。实验表明,它能有效保护隐私并将性能开销控制在较低水平。

ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software

第一作者: Xiang Mei · 方向: 软件安全

Abstract:Achieving reproducibility, quantity, and diversity in vulnerability datasets has long been viewed as an inherent three-way trade-off, where improving one dimension often comes at the cost of the others. In practice, reproducibility has been the dimension most often neglected. This has limited what can be automatically extracted from historical bug datasets, and has reduced their utility for downstream security research. In this work, we propose a method to produce a new security dataset which ensures reproducibility for diverse vulnerabilities at scale by identifying the key obstacles to large-scale bug reproduction and addressing them with general solutions. Using this method, we introduce full reproducibility to the largest open source software vulnerability dataset (OSS-Fuzz) and construct the ARVO dataset (an Atlas of Reproducible Vulnerabilities in Open-source software)...

论文介绍 本研究解决了安全漏洞数据集在可复现性、数量和多样性上难以兼顾的挑战。作者提出了一套方法,能大规模地构建可复现的漏洞样本。基于此方法,他们对OSS-Fuzz数据集进行增强,建立了名为ARVO的全新可复现漏洞图集,以支持下游安全研究。

Cache to the Future: A Distributed Webpage Archive for Internet Blackouts

第一作者: Ross Evans · 方向: AI 安全

Abstract:Internet blackouts, occurring due to technological mishaps or intentional governmental action, prevent citizens from accessing the internet. Citizens in regions where internet blackouts are common have utilized blackout-resistant technologies to maintain communication. Such technologies often rely on mobile mesh networks to provide limited messaging services. However, no technology currently exists which can provide continued access to knowledge sources on the web during a blackout. We present Cache to the Future (CttF): a system to cache and deliver static content hosted on the web during a blackout. CttF's distributed community ratings crowdsources caching at scale while cryptographic constructs (digital signatures, proofs-of-work) mitigate adversarial interference. Our realistic simulations demonstrate CttF delivering content at city-scale across a wide range of benign and...

论文介绍 为了应对互联网中断导致的知识访问中断,本文提出了Cache to the Future系统。该系统能缓存并分发网页上的静态内容,其分布式社区评级机制可实现大规模众包缓存,而数字签名和工作量证明等密码学手段则用于防御恶意干扰。模拟实验验证了其在城市规模下的有效性。

Safety, Security, and Cognitive Risks in Neuro-Symbolic AI

第一作者: Manoj Parmar · 方向: 软件安全

Abstract:Neuro-symbolic AI (NeSy) pairs neural perception with symbolic reasoning, making it attractive for high-stakes domains where explainability and structured inference are required. However, this hybrid architecture introduces an enlarged attack surface spanning five layers: neural perception, symbolic knowledge bases, reasoning engines, agentic orchestration, and data stores -- each exploitable in ways absent from purely neural systems. This paper makes six contributions: (1) formal definitions of NeSy Attack Surface, Symbolic Integrity Violation (SIV), and Cross-Layer Amplification Ratio $\mathcal{X}$, decomposed into neural-caused and autonomous symbolic sensitivity components; (2) a unified threat model extending MITRE ATLAS with 11 NeSy-specific tactic extensions and a five-profile attacker taxonomy; (3) a symbolic-layer threat catalogue covering knowledge graph (KG)...

论文介绍 本文系统分析了神经符号AI的安全风险。作者指出,这种结合神经网络与符号推理的混合架构,其攻击面横跨神经感知、知识库、推理引擎、智能体编排和数据存储五个层次。论文形式化定义了相关概念,并扩展了威胁模型,以识别与纯神经网络不同的新型攻击向量。

LineageMark: Multi-user White-box Watermarking for Contribution Tracing in Model Derivation Chains

第一作者: Bingxue Zhang · 方向: AI 安全

Abstract:In open large language model (LLM) ecosystems, models are frequently adapted across multiple domains and applications, forming multi-stage derivation chains. Consequently, tracking and verifying historical contributions is essential for model provenance and intellectual property protection. However, existing watermarking methods are mainly designed for single-user, one-time embeddings, often fail under repeated model derivation and incremental updates. To address this problem, we propose LineageMark, a multi-user white-box watermarking framework for model derivation chains. The framework encodes watermarks in model parameters using a projection-based approach. Stable carriers are first selected to reduce sensitivity to model changes, each watermark bit is then represented as a projection statistic over these carriers. Additional watermark insertions introduce only bounded...

论文介绍 针对开源大模型在多次衍生和更新后贡献难以追溯的问题,本文提出了LineageMark框架。该框架采用基于投影的白盒水印技术,将水印信息嵌入模型参数的稳定载体中。多个用户可在模型衍生链的不同阶段插入水印,且对模型性能影响有限,从而实现知识产权保护和溯源。

TrustErase: Auditable Instant Machine Unlearning with Passport-Embedded Representations

第一作者: Rutger Hendrix · 方向: 隐私保护

Abstract:The demand for privacy-compliant AI has amplified the need for machine unlearning; yet, existing retraining or distillation-based methods remain unverifiable and computationally costly. We introduce TrustErase, a verifiable, data-free unlearning framework leveraging passport-embedded representations for instant, modular, and auditable forgetting. By treating passports as cryptographic keys within parameter-efficient adaptation layers, TrustErase enables the removal of specific classes or datasets through simple deactivation, without retraining, fine-tuning, or access to the original data. A singular value based decomposition conceals passports within model weights, ensuring that unlearning actions remain transparent and provably compliant. Evaluations on MNIST, CIFAR10 and CIFAR100 show that TrustErase matches or exceeds state-of-the-art benchmarks such as DELETE, L2UL, and...

论文介绍 本文提出TrustErase框架,用于实现可验证、即时且无需原始数据的机器遗忘。其核心是将“护照”视为密钥嵌入模型的参数高效适配层中。通过特定的分解技术,移除某个类别或数据集只需简单地禁用对应的护照,无需重新训练,且删除操作透明可审计。

Graph neural networks at war: integrating cybersecurity and drone intelligence in the Israeli-Iranian conflict

第一作者: Sozan Sulaiman Maghdid · 方向: AI 安全

Abstract:Physical cyber systems have brought about new threats and challenges in detection and immediate response. This study examines how Graph Neural Networks (GNNs) can be used to aid cybersecurity and drone management in a physical cyber system comprising of cyber intrusions and unmanned aerial vehicles (UAVs). By providing a bridge between structural understanding of graphical neural networks, this work has provided an integrated procedure that allows intrusion detection systems to educate on underlying network structures, identify malicious activity, and facilitates drone response measures. Based on an emulation-based case study, cyberattacks models were created to provoke the responses of the drones, which proved that graph-based learning can assist with the situational awareness, swarm coordination, and adaptive maneuver. According to the performance valuation, this method has...

论文介绍 该研究探讨了如何利用图神经网络来整合物理网络系统中的网络安全与无人机管理。作者提出了一种集成方案,通过GNN理解底层网络结构以检测入侵,并指导无人机的协调响应。基于仿真的案例研究表明,这种方法有助于提升态势感知和自适应机动能力。

Quantifying quantum risk: a measure of crypto agility

第一作者: Coryan Wilson-Shah · 方向: 安全研究

Abstract:Because of their ability to enable new forms of cryptanalysis, quantum computers pose a threat to the cryptographic algorithms that are widely used to secure contemporary computer systems. A practical quantum computer may emerge within the next ten years or so, but due to theorised "harvest now, decrypt later" style attacker behaviour, mitigations are necessary today. Recent advances in cryptography and security architecture show promise in supporting the design of systems that exhibit resilience against quantum-enabled cryptanalysis, however there is a key gap in the literature around the subject of deriving tolerances for such systems. In this paper, we introduce the concept of rotation time as a measure of crypto agility, and derive an approximation that links rotation time tolerance to security risk tolerance. Historical CVE data is used to calculate illustrative values...

论文介绍 量子计算机对广泛使用的加密算法构成威胁,可能在未来十年内出现,且存在“现在窃取,以后解密”的攻击风险。本文引入旋转时间作为加密敏捷性的度量,并推导出旋转时间容差与安全风险容差之间的近似关系,使用历史CVE数据计算示例值。该研究填补了系统抗量子攻击能力容差推导的文献空白,有助于设计量子弹性系统。

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios

第一作者: Hankyul Baek · 方向: AI 安全

Abstract:AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they can read, update, and disseminate sensitive information. Much of prior research on data leakage risks in agents has focused on adversarial data exfiltration through prompt injections and jailbreaks. However, sensitive information may also be exposed during non-adversarial use, creating leakage risks even when users issue benign requests. We report a joint evaluation by the Singapore AI Safety Institute and the Korea AI Safety Institute examining agent data leakage in 12 realistic, non-adversarial tasks spanning customer support, DevOps, web automation, and enterprise and personal productivity. The evaluation covers five risk types: lack of data awareness, audience awareness, policy compliance, data minimization, and...

论文介绍 AI代理在访问电子邮件、数据库等工具时可能暴露敏感信息,现有研究多关注对抗性攻击,但非对抗性使用也存在泄漏风险。本文由新加坡和韩国AI安全研究所联合评估,覆盖客户服务、DevOps等12个现实任务,考察数据意识、受众意识等五种风险类型。该评估为代理安全设计提供实践参考,强调在日常使用中加强数据保护。

Fractional Verkle Trees: A Hypertree Decomposition and Verified Proof Serialization Architecture for High-Performance Blockchain State Accumulators

第一作者: Ekleen Kaur · 方向: 密码学协议

Abstract:Modern blockchain state management faces a critical scalability bottleneck: maintaining cryptographic commitments over hundreds of millions of entries becomes computationally prohibitive. Ethereum's transition to Verkle Trees: polynomial commitment accumulators reducing proof sizes from O(width * depth) to O(depth) via constant-size IPA vector commitments, is a critical step toward stateless operation. Yet, current implementations exhibit pathological characteristics that burden home validators. We identify four inefficiencies in the reference go-verkle implementation \cite{kaur2025goverkle, kaur2025goethereum}: (1) phantom node creation during non-existent account deletion; (2) 64-byte database keys triggering excessive LSM-tree compaction; (3) redundant memory copying in proof deserialization; (4) a Proof of Absence wire format incompatibility causing non-deterministic...

论文介绍 区块链状态管理面临可扩展性瓶颈,Verkle树通过多项式承诺减少证明大小,但参考实现存在低效问题。本文提出分数Verkle树架构,采用超树分解和验证证明序列化,针对幽灵节点创建、数据库键冗余、内存复制等四类低效进行优化。该研究旨在提升区块链性能,支持无状态操作,减少家庭验证器的计算负担。

Loss Landscape Poisoning: Targeted Extraction of Unseen Training Data from LLMs

第一作者: Md Abdullah Al Mamun · 方向: 隐私保护

Abstract:Large Language Models are increasingly trained on proprietary or sensitive data, from private healthcare and financial records to user conversations containing secrets. Ensuring the privacy of such data against extraction attacks has become a central concern. In this paper, we ask whether an attacker who can poison a portion of the training data can facilitate the leakage of a separate target record they have no access to. We answer in the affirmative and show that such leakage can be induced by a poisoning mechanism that reshapes the model's local loss landscape around the target completion. Our key insight is that poisoning to create a sharp loss minimum at the target, surrounded by elevated loss on nearby alternatives, forces the model to memorize the target as the unique low-loss solution in its neighborhood. The attack requires no architectural changes, and generalizes...

论文介绍 大语言模型训练数据可能包含敏感信息,提取攻击威胁隐私。本文探讨攻击者通过中毒部分训练数据,诱导模型泄漏未访问的目标记录。核心机制是重塑模型在目标附近的损失景观,创建尖锐最小值迫使模型记忆目标。该攻击无需修改架构,泛化到多种模型,对隐私保护防御提出新挑战。

Timestamp-Aware Spatio-Temporal Graph Contrastive Learning for Network Intrusion Detection

第一作者: Jianli Dai · 方向: AI 安全

Abstract:Given their effectiveness in modeling the relational structure among network traffic flows, graph neural networks (GNNs) have been widely adopted in network intrusion detection systems (NIDSs). However, most existing GNN-based NIDS approaches focus on the relational structure of traffic flows, and treat them as temporally independent, which limits their ability to cope with evolving attack behaviors. Moreover, their reliance on supervised or semi-supervised learning often restricts generalization to unseen attacks. To address these limitations, we propose a novel self-supervised GNN-based framework. To the best of our knowledge, the proposed model is among the first self-supervised GNN-based NIDS models to explicitly leverage real timestamps, which provides faithful temporal dependencies for representation learning. We first construct a series of temporal graphs from network...

论文介绍 图神经网络在网络入侵检测中建模流量关系,但多数方法忽略时间依赖性,限制对演化攻击的应对。本文提出自监督GNN框架,首次显式利用真实时间戳构建时空图,进行对比学习以捕获时序依赖。该框架不依赖标签,提升对未见攻击的泛化能力,适用于动态网络环境。

Securing Multi-Agent GIS Systems: Risk Evaluation and Prompt Hardening Optimization

第一作者: Kyle Gao · 方向: 软件安全

Abstract:Agentic systems are increasingly integrated with geographic information systems (GIS), where multi-agent coordination enables complex conversational and spatial analysis but introduces security risks. This work presents a security-oriented framework for risk identification, evaluation, and mitigation in a multi-agent GIS system while maintaining adaptability to broader agentic architectures. We test the agentic system of a commercial geospatial partner while developing a modular state-machine-based orchestration framework that abstracts agent behavior into reusable components. We evaluate robustness using a red-teaming framework with an adaptive attacker LLM and a deterministic judge that produces binary outcomes with supporting rationales across multi-turn attacks. We further improve resilience with a prompt optimization framework that treats prompts as structured signatures...

论文介绍 多代理系统与地理信息系统集成带来安全风险,涉及协调复杂空间分析。本文提出安全导向框架,用于风险识别、评估和缓解,采用状态机编排抽象代理行为。通过红队测试评估鲁棒性,使用自适应攻击LLM和确定性法官,并应用提示优化增强弹性。该框架适用于更广泛的代理架构,提升系统安全性。

Security and Human-Centered Assessment of BACnet-Controlled DALI Infrastructure in an Educational Building Automation Testbed

第一作者: Ariton Verush · 方向: 网络安全

Abstract:Building automation and control systems integrate heating, ventilation, air conditioning, lighting, sensing, and management functions through specialized communication protocols. While this integration enables flexible building operation, it also creates complex cyber-physical environments that are difficult to inspect, secure, and explain to new analysts. This paper presents a practical security and human-centered case study of a BACnet/IP building automation testbed with DALI lighting infrastructure, investigated during a domotics-oriented cybersecurity hackathon in Thun, Switzerland in April 2026. The study combines network-oriented enumeration, object-level inspection, physical rack analysis, and reflective HCI analysis of tool-supported learning. Using Yabe and BACteria, the work documents observable BACnet services, reconstructs structured object hierarchies, identifies...

论文介绍 建筑自动化系统集成多种功能,但复杂网络物理环境难以检查和保障。本文对BACnet/IP测试床进行安全与以用户为中心的评估,结合网络枚举、对象检查和HCI分析,记录可观测服务并重构对象层次。研究基于瑞士图恩的网络安全黑客马拉松,为工具支持学习提供反思,有助于新分析师培训。

Verifiable computations for dynamic encrypted control

第一作者: Sebastian Schlor · 方向: 系统安全

Abstract:Encrypted control can preserve the privacy of data and parameters while the necessary computations can be outsourced to a cloud server. To ensure the integrity of the received values from the cloud, i.e., that they have not been changed, however, strong assumptions or verification algorithms are needed. Previous methods require computationally expensive cryptographic protocols or are only applicable to static computations. In this paper, we present a novel type of verification algorithm for linear dynamic encrypted control. We utilize system-theoretic input-output properties of the controller for artificial challenge signals, which are processed in the cloud in parallel with the requested control input, to check the correctness of the results at the plant. This results in almost no additional computational load, wrong computations are revealed with high probability, and no...

论文介绍 加密控制可外包计算到云端保护隐私,但需确保结果完整性。现有方法依赖昂贵协议或仅适用于静态计算。本文提出动态加密控制的验证算法,利用控制器输入输出属性生成挑战信号,在云端并行处理以检查正确性。该方法几乎无额外计算负荷,能高概率检测错误,无需强假设。

Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond

第一作者: Hobin Kim · 方向: AI 安全

Abstract:Large language models (LLMs) are widely used to fulfill users' information needs; users ask LLMs about the weather, pose educational questions, and consult them for legal assistance. One particularly understudied area is digital security and privacy (S&P), where users may seek LLMs' help on how to secure their online accounts or protect their computers from cyber attacks. To the best of our knowledge, no prior study has collected or analyzed the S&P questions users ask LLMs; prior research on LLM response quality relied on expert-authored S&P misconceptions or FAQs rather than user queries. Drawing from WildChat, a dataset of 3.2M user-LLM conversations collected in the wild, our study identifies 14,727 S&P prompts and categorizes them into nine categories covering a wide range of S&P topics. From the S&P prompts, we sampled 450 and performed a thematic analysis to...

论文介绍 该研究关注用户向大语言模型提出的数字安全与隐私问题。基于包含320万次真实对话的WildChat数据集,作者识别出14727条相关提示,并将其归类为九个类别。通过对450条提示的深入主题分析,研究旨在了解用户实际的安全关切以及模型的应对表现,填补了以往研究依赖专家构造问题而非真实用户查询的空白。

Differential Privacy of Gaussian Process Posterior Sampling

第一作者: Tomasz Maciazek · 方向: 隐私保护

Abstract:We study the privacy of releasing posterior sample paths from a Gaussian process (GP) when the entire training set including covariates and responses is private. Unlike standard differential-privacy (DP) mechanisms that add external noise, posterior sampling is random by construction. We show that this intrinsic randomness yields DP guarantees by deriving explicit Rényi-DP bounds for GP posterior sample-path release. The bounds separate posterior-mean leakage from data-dependent posterior-covariance leakage showing that meaningful privacy depends sharply on effective ridge regularisation. We apply membership-inference attacks to show that empirical leakage follows the predicted dependence on regularisation, posterior variance and the number of released posterior sample-paths. Utility experiments on downstream posterior-sampling tasks identify noisy-observation regimes where...

论文介绍 本文研究了在训练数据集完全私密的情况下,发布高斯过程后验样本路径的隐私性。不同于添加外部噪声的常规差分隐私机制,后验采样本身具有内在随机性。作者推导出了显式的Rényi-DP界限,证明这种内在随机性可以提供差分隐私保证,并指出隐私保护程度与有效的正则化参数密切相关。研究通过成员推理攻击和效用实验验证了理论分析。

Security-Induced Braess Paradoxes in Service Function Chain Orchestration

第一作者: Daniel Commey · 方向: AI 安全

Abstract:NFV/SDN orchestration lets operators instantiate and steer traffic through virtual firewalls, IDS/IPS replicas, WAF clusters, zero-trust gateways, backup inspection paths, and migration targets on demand. Operators often treat these options as monotone improvements: more inspection capacity, lower nominal latency, or broader placement flexibility should not degrade the service. That intuition can fail even when the new option is locally attractive. We study a security-induced Braess paradox in service function chain (SFC) orchestration, where adding a defensive option worsens the post-adaptation equilibrium by concentrating traffic and adversarial value on shared security resources. We define Braessian security-management actions, derive a sufficient condition for paradox emergence under affine load-dependent VNF delay, and give a pre-deployment orchestration screen that...

论文介绍 本文揭示了网络功能虚拟化与软件定义网络编排中可能出现的一个反直觉现象:添加新的防御性资源(如虚拟防火墙或入侵检测系统副本)反而可能恶化系统整体性能,即「安全诱导的Braess悖论」。研究定义了相应的安全操作,在仿射负载相关延迟模型下推导了悖论出现的充分条件,并提出了一个事前编排筛选方法,以避免此类资源分配导致的负面均衡。

Cordon: Semantic Transactions for Tool-Using LLM Agents

第一作者: Zheng Chen · 方向: AI 安全

Abstract:Tool-using LLM agents are shifting the unit of computation from explicit human-issued commands to model-driven tasks with stateful consequences. Yet today's agent runtimes still expose tools as isolated RPCs. This interface gives runtimes a convenient integration point, but it lacks a task-scoped execution boundary for commit, rollback, recovery, and audit across multi-step agent workflows. We argue that this mismatch calls for a runtime containment boundary rather than another per-call guardrail. This paper introduces Cordon, a transactional runtime system for staging and validating irreversible agent effects before commit. A semantic transaction is a task-level execution boundary that binds tool intents and runtime-tracked result lineage to reversible local state, staged external effects, delegated authority, and audit metadata. Cordon implements this abstraction with a...

论文介绍 本文从一个关于两种归纳理论(开放归纳和子句集循环)是否不可比的问题出发,用一个简洁的证明给出了否定答案。证明的核心在于,纯语法系统中的规则无法触及像「a+b」与「b+a」这类涉及常数顺序的项。由此,作者提炼出「语法不变性原则」,阐明了此类论证的普遍形态,并探讨了该形态在已知非形式推理中的类似表现。

Syntactic Systems Cannot See Semantic Invariants

第一作者: Fabio F.G. Buono · 方向: 安全研究

Abstract:We start from a small open question, where Hetzl and Vierling asked whether two theories of induction, open induction and clause set cycles, are incomparable. They proved one direction and left the other open. Here we close it, and the proof is almost embarrassingly short, because the rules for addition can only fire when the first argument is $0$ or a successor, a Skolem constant is neither, so the terms $a{+}b$ and $b{+}a$ can never be touched, and a machine that can never touch them can never prove they are equal. The thing that separates the two theories is the order of two constants, and that order is a fact about numbers, not about symbols. We extract from this proof a small general principle, the Syntactic Invariance Principle, that names the shape of such arguments. We then close with a few speculative remarks on how this same shape appears, informally, in the known...

论文介绍 本文从一个开放问题出发,即开放归纳和子句集循环两个归纳理论是否不可比。作者通过一个简短的证明关闭了其中一个方向,证明由于加法规则限制,Skolem 常数既不是 0 也不是后继,导致项 a+b 和 b+a 无法被处理,因此无法证明相等。核心贡献是从中提取了语法不变性原则,该原则命名了此类论证的结构,并在已知逻辑问题中非正式出现。

Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement

第一作者: Mingtong Zhang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism of practicing and learning from feedback. In this paper, we propose VERITAS, a generator-verifier framework for generalist robot policies for inference-time policy steering and self-improvement. We use a pre-trained generalist robot policy as a ``generator'' and pair it with a gradient-free ``visual verifier'' that evaluates actions at inference time. This framework enables inference-time steering that improves policy performance without additional training. We demonstrate that inference-time verification consistently outperforms vanilla generalists without training on additional demonstration data. Additionally, we demonstrate that the verified rollouts provide effective supervision for offline policy improvement: policies fine-tuned on verified self-generated...

论文介绍 本文介绍了EBench,一个用于诊断通用移动操作机器人策略的仿真基准。它包含26个任务,并沿能力与泛化两个维度进行标注。研究评估了当前先进的通用模型,发现即使整体成功率相近,不同模型的能力画像也差异显著。例如,某模型在操作任务上表现优异但在灵巧任务上崩溃。EBench超越了单一的成功率指标,能从多个代表性角度分析策略的泛化能力。

EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies

第一作者: Ning Gao · 方向: 机器人操作 · 来源: cs.RO

Abstract:We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single success-rate scalar. EBench comprises 26 diverse and challenging manipulation tasks annotated along 5 capability dimensions and 4 generalization dimensions. We evaluate state-of-the-art generalist manipulation models including $\pi_0$, $\pi_{0.5}$, XVLA, and InternVLA-A1, and reveal that models with near success rates exhibit strikingly different capability profiles: $\pi_{0.5}$ achieves the highest test success rate and the best train--test retention, whereas InternVLA-A1 dominates mobile manipulation but collapses on dexterous tasks, and XVLA exhibits strengths on a disjoint set of atomic skills compared to other policies. Beyond capability profiling, EBench analyzes the generalization ability from 4 representative perspectives, identifying the impact of different...

论文介绍 本文提出了EBench,一个用于全面诊断通用移动操作策略的仿真基准。该基准包含26个多样化挑战任务,并沿5个能力维度和4个泛化维度进行标注。通过评估π0、π0.5、XVLA和InternVLA-A1等先进模型,研究揭示了不同策略在能力分布上的显著差异,例如π0.5在成功率和保留性上表现突出,而InternVLA-A1在移动操作上占优但灵巧任务不足。EBench还从多个视角分析泛化能力,为理解和改进移动操作策略提供了系统化工具。

Beyond Failure Recovery: An Engagement-Aware Human-in-the-loop Framework for Robotic Systems

第一作者: Jiaying Fang · 方向: 具身智能 · 来源: cs.RO

Abstract:Conventional human-in-the-loop approaches typically involve users only when a robot encounters failure or uncertainty, treating humans primarily as tools for improving robot performance. However, in many human-centered robotics settings, interaction should support engagement by keeping users involved in decision-making rather than limiting them to failure-driven interventions. This is particularly compelling in physical caregiving, where mobility limitations can reduce users' ability to intervene or modulate the robot's behavior in the moment. As a result, failure-driven interaction policies may relegate users to passive observers for long stretches of the task. For example, a user with mobility limitations may feel less engaged when being continuously and passively fed by a robot. At the same time, overly frequent interaction can be tiring and increase the user's workload. To...

论文介绍 传统的人机在环方法主要在机器人失败或不确定时才引入人类,将人视为提升性能的工具。本文提出了一种参与感知的框架,旨在通过让人类持续参与决策来维持其参与感,而非仅在故障时干预。这对于身体护理等场景尤为重要,因为用户行动受限可能使其长时间处于被动状态。该框架旨在平衡交互频率,既避免用户沦为被动观察者,也防止过度交互增加其负担。

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

第一作者: Ralf Römer · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets. Despite their strong empirical performance in robotic manipulation, VLAs lack mechanisms to quantify confidence in their predictions and to detect when their actions may be unreliable. This presents a critical limitation for real-world deployment in non-stationary environments, where models inevitably encounter scenarios outside their pretraining distribution and may fail without warning. To address this, we derive an efficient method for quantifying epistemic uncertainty in flow-matching models by leveraging velocity-field disagreement (VFD) across a small ensemble. We successfully use this uncertainty estimate for failure detection during deployment and active fine-tuning of flow-based VLAs. To this end, we...

论文介绍 视觉-语言-动作模型在实际部署中缺乏预测置信度的量化机制,难以应对超出预训练分布的场景。本文提出一种高效方法,通过小规模模型集成中的速度场差异来量化流动匹配模型的认知不确定性,并成功将该估计用于运行时故障检测和模型主动微调,提升了模型在非平稳环境中的可靠性。

LAGO Policy: Latency-Aware Asynchronous Diffusion Policies with Goal-Directed Collision-Free Planning for Smooth Manipulation

第一作者: Guowei Shi · 方向: 机器人操作 · 来源: cs.RO

Abstract:Diffusion-based visuomotor policies deployed with asynchronous inference often exhibit inter-chunk discontinuities and lack explicit mechanisms for obstacle-aware execution, leading to jerky motions and collisions that hinder reliable manipulation in real-world scenes. To address these issues, we propose LAGO Policy, a unified asynchronous action-generation framework that integrates trajectory optimization with diffusion policy for smooth and safe execution. LAGO Policy improves inter-chunk consistency via latency-aware classifier-free guidance conditioning on future actions. It further enables goal-directed collision-free trajectory planning by predicting a task-relevant interaction goal from demonstrations. Finally, spatial-temporal trajectory optimization refines the actions to be executed for low-jerk and feasible motion. Extensive real-world experiments demonstrate that...

论文介绍 基于扩散策略的异步推理容易导致动作块间运动不连续和碰撞。本文提出LAGO策略,一个统一的异步动作生成框架,通过延迟感知的无分类器指导提高动作块间一致性,并利用演示预测任务相关的交互目标以实现无碰撞轨迹规划。最终通过时空轨迹优化生成平滑、可行的动作序列,提升真实场景操作的可靠性。

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation

第一作者: Tianyi Lu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiting their capacity for reasoning-intensive long-horizon tasks. To address this, existing approaches adopt Chain-of-Thought (CoT) reasoning to enable subgoal decomposition and spatial anticipation. However, those methods lack a unified architecture for effective cross-modal reasoning and fail to explicitly include inverse reasoning ability based on the target state. We argue that manipulation planning naturally decomposes into prediction, anticipating the next visual state, and inverse dynamics, inferring the actions to reach it. Bridging both requires a unified autoregressive architecture that interleaves textual and visual reasoning in a single generation process. We propose \textbf{ThinkingVLA}, a generative model that realizes this decomposition within a unified...

论文介绍 现有视觉-语言-动作模型缺乏显式推理,限制了其处理长时程任务的能力。本文提出ThinkingVLA,通过统一的自回归架构,在单一生成过程中交错进行文本和视觉推理。该模型将操作规划自然分解为预测下一视觉状态和推断达到该状态的动作,从而实现了更有效的跨模态推理与规划。

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space

第一作者: Bochen Yang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit deliberation. Directly decoding actions from vision-language backbone representations enables low-latency control, whereas explicit reasoning through textual chains, pixel-level subgoals, or action search can improve planning but incurs substantial latency and computational cost. We propose PearlVLA, a VLA framework that moves deliberation into the latent space of a vision-language model (VLM). PearlVLA separates VLM meta-query representations into a fixed visual grounding branch and an iterative latent plan branch. At each refinement round, a plan-conditioned world query probes a lightweight frozen latent world model for an action-free future observation latent, which is fed back to guide plan refinement. A future-guided RefineNet then applies scheduled residual...

论文介绍 当前VLA模型在高效动作生成与显式推理之间存在权衡。本文提出PearlVLA框架,将规划过程移入视觉-语言模型的潜在空间。它将模型表征分离为视觉定位分支和迭代潜在计划分支,通过计划条件化的世界查询与一个轻量级潜在世界模型交互,迭代获取未来观测以指导计划细化,从而平衡规划深度与计算效率。

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

第一作者: Haoqi Yuan · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether this scaling recipe can be applied to robotic manipulation to achieve genuine generalization. This is challenging because, unlike text, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity, making alignment and scale simultaneously difficult. We present Qwen-RobotManip, a generalizable Vision-Language-Action foundation model built on Qwen-VL. Qwen-RobotManip introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. This alignment capability in turn enables Qwen-RobotManip to absorb manipulation data at a scale...

论文介绍 本技术报告探讨了将语言和多模态基础模型的成功经验——即对齐异构数据并规模化训练——应用于机器人操作基础模型的可能性。报告介绍了Qwen-RobotManip,一个基于Qwen-VL的可泛化VLA基础模型。该模型引入了跨表征、运动和行为维度的统一对齐框架,使得大规模多源训练得以协调进行,从而增强了模型吸收多样化操作数据的能力。

HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning

第一作者: Taowen Wang · 方向: 策略学习 · 来源: cs.RO

Abstract:Humanoid robots promise whole-body interaction in human-centered environments, but scalable policy learning remains difficult because task-level decision-making and whole-body dynamic execution are tightly coupled. A practical solution is hierarchical control, where a high-level policy predicts intermediate whole-body actions and low-level general motion trackers (GMTs) execute them as stable humanoid motion. However, existing benchmarks rarely evaluate the policy-tracker interface itself, leaving open whether intermediate whole-body actions are executable, robust under task distribution shifts, and transferable across different GMT backends. We introduce HumanoidArena, a simulation-first benchmark for egocentric hierarchical whole-body learning. The benchmark formulates policy learning as a hierarchical decision making problem: a high-level policy converts egocentric vision...

论文介绍 人形机器人的分层控制策略中,高层策略与低层运动跟踪器之间的接口评估不足。本文提出HumanoidArena,一个以自我为中心视角、用于评估分层全身学习的模拟基准。该基准将策略学习形式化为一个分层决策问题,重点评估中间级全身动作的可执行性、鲁棒性以及跨不同运动跟踪器后端的可迁移性。

ED3R: Energy-Aware Distributed Disaster Detection Enabled by Cooperative Robotic Agents

第一作者: Lina Magoula · 方向: 具身智能 · 来源: cs.RO

Abstract:Robotics are expected to support environmental monitoring and natural disaster management, where decisions must be made under uncertainty, resource limitations, and strict operational constraints. In critical missions, such as wildfires, robotic agents must not only identify hazardous events with sufficient confidence, but also manage the energy cost and time until detection. This paper introduces ED3R, an energy-aware distributed framework for wildfire detection under uncertainty. ED3R enables hierarchical cooperative decision-making between a robot and a remote controller. The remote controller decides upon the robot's motion, while the robot senses the environment and decides where to execute the wildfire detection (onboard or remotely) and how. The common goal is to detect wildfires with a required confidence while minimizing the energy consumed by any robot operation...

论文介绍 在资源受限的灾害监测任务中,机器人需在不确定性下做出高置信度决策并管理能耗。本文提出ED3R,一个用于野火检测的能源感知分布式框架。该框架支持机器人与远程控制器之间的分层协同决策,控制器决定机器人的运动,机器人则决定检测任务的执行方式(本地或远程),共同目标是在满足置信度要求的同时最小化能源消耗。

ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI

第一作者: Hong Yang · 方向: 导航与运动 · 来源: cs.RO

Abstract:Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visual observations. Yet existing visual and embodied question answering benchmarks often provide limited control over the reasoning dependencies being tested, making it difficult to distinguish grounded embodied reasoning from shortcut-driven visual or linguistic pattern matching. We present ERQA-Plus, a diagnostic benchmark for reasoning in embodied AI. ERQA-Plus contains 1,766 question-answer instances grounded in 711 robot-centric images and organized according to a structured taxonomy spanning perceptual, action-centric, social-interaction, navigation-environmental, and contextual commonsense reasoning. The dataset is constructed using a multi-stage generation...

论文介绍 通用具身智能体需要超越物体识别的推理能力。现有基准对测试的推理依赖控制有限。本文提出ERQA-Plus,一个针对具身AI推理的诊断性基准。该基准包含1766个问答实例,基于711张机器人中心图像,按照涵盖感知、动作、社交互动、导航环境及上下文常识等结构化分类法组织,旨在区分真实的具身推理与基于捷径的模式匹配。

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation

第一作者: Xingyuming Liu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Language-Action (VLA) models for robotics rely solely on RGB observations. This limits their ability to perceive physical properties that are difficult or impossible to infer from RGB cameras, such as temperature, sound, or radar response. We present MuseVLA, an adaptive multimodal sensing VLA model that integrates novel sensors as on-demand tools for robotic manipulation. Given a task instruction and visual context, MuseVLA first generates a sensor token and target description that select the sensing modality to invoke and what to attend to, analogous to a tool call with arguments. It then converts the selected sensor measurement into a grounded sensor image, a unified intermediate representation that encodes heterogeneous readings for multimodal fusion and action...

论文介绍 针对现有视觉-语言-动作模型仅依赖RGB观测的局限,本文提出自适应多模态感知模型MuseVLA。该模型在给定任务指令后,能像工具调用一样选择并整合温度、声音等非视觉传感器数据。其核心是将异构传感器读数转换为统一的「接地传感器图像」中间表示,从而在单一模型内实现多模态融合与动作生成,以增强机器人对物理世界的感知与交互能力。

GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments

第一作者: Jiawei Zhang · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Training embodied agents in the real world requires skilled operators and expensive hardware. Simulation environments offer a compelling alternative by enabling large-scale, cost-effective data augmentation. Consequently, rapidly constructing high-fidelity simulation scenes with a minimal sim-to-real gap has become a critical objective in robot learning. While reconstruction-based methods provide superior visual quality, current workflows are hindered by inefficient data acquisition and subpar foreground object extraction. We thus propose GASE, a highly automated system for simulation scene construction. GASE leverages multi-view video streams from panoramic camera arrays to enable rapid environment scanning. To ensure high-quality asset generation, our pipeline introduces a camera-pose-based strategy that robustly extracts objects across frames in the 2D domain, followed by...

论文介绍 为快速构建高保真仿真场景以缩小虚实差距,本文提出高度自动化的GASE系统。该系统利用全景相机阵列采集多视角视频,实现环境的快速扫描。其核心创新在于采用基于相机位姿的策略,在2D域内鲁棒地提取跨帧前景物体,并通过基于高斯溅射的流程生成高质量资产,旨在为机器人学习提供一种高效、低成本的场景构建方法。

MagicSim: A Unified Infrastructure for Executable Embodied Interaction

第一作者: Haoran Lu · 方向: 具身智能 · 来源: cs.RO

Abstract:Robot learning and embodied agents now require simulation to serve as a shared execution substrate linking control, skills, and planning, not only as a renderer, controller testbed, or fixed task environment. Existing pipelines split these layers with "magic" actions, disconnected training environments, or forward-only renders that cannot reproduce, evaluate, and annotate the same episode. We present MagicSim, an embodied interaction infrastructure built around one deterministic batched runtime and a shared Markov decision process (MDP). From YAML-first specifications that decouple contents, placement, behavior, and agent exposure, MagicSim constructs diverse executable worlds spanning task families, interaction regimes, physics, layouts, sensors, avatars, and robot embodiments in one reset-and-step loop. A common execution interface grounds high-level commands through...

论文介绍 现有机器人学习仿真流程常将控制、技能和规划层割裂。本文提出MagicSim,一个构建在确定性批量运行时和共享马尔可夫决策过程之上的具身交互基础设施。它通过以YAML为先的规范解耦内容、行为和代理,能在单一的重置-步进循环中构建跨越多种任务家族、物理模拟和机器人形态的可执行世界,为机器人学习提供了一个统一、可复现的执行基础。

When Robots Sleep: Offline Skill Consolidation for Shared-Policy Robot Learning

第一作者: Nethmi Jayasinghe · 方向: 模仿学习 · 来源: cs.RO

Abstract:Robots that learn over long deployments must add new skills without losing the shared policy structure that makes earlier skills reusable. We study sequential robot skill learning, where previous trajectories and task losses may be unavailable, and the deployed policy must remain a single shared controller without task-specific heads, routing, or adapters. We identify skill-coupling collapse, a failure mode in which individual skill success remains non-trivial while reliability among related skills deteriorates. We propose Sleeping Robots, a wake-sleep framework that learns each new skill during wake and consolidates the shared policy offline during sleep using compact frozen skill memories: frozen critics with unordered state buffers for reinforcement learning and frozen actor snapshots with unordered observation buffers for imitation learning. During sleep, these memories...

论文介绍 在机器人长期部署中持续学习新技能,同时保持共享策略的通用性是一个挑战。本文研究顺序技能学习,识别了「技能耦合崩溃」现象,并提出了睡眠机器人框架。该框架在「清醒」期学习新技能,在「睡眠」期利用冻结的技能记忆体(如批评家/演员快照)离线整合共享策略,使得在无需历史数据和任务特定模块的情况下,巩固并维护一个统一的控制器。

AnnotateAnything: Automatic Annotation of 3D Assets for Robot Manipulation

第一作者: Haoran Lu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Simulation enables scalable robot data collection, but raw 3D assets provide only geometry, lacking the semantic, interactive, and physical knowledge needed to specify where and how robots should act. In this work, we present AnnotateAnything, a general automatic annotation framework that converts passive 3D assets into manipulation-ready assets with structured, diverse, and executable manipulation labels. AnnotateAnything is built around two complementary pipelines. First, a unified visual-language annotation pipeline using vision-language reasoning to infer object semantics, interaction constraints, and 3D-grounded cues, providing human-prior guidance for identifying meaningful interaction regions. Second, a fully automatic and massively parallel physics annotation pipeline grounds these priors in each asset's geometry and physical constraints through candidate generation...

论文介绍 原始3D资产缺乏机器人操作所需的语义和物理知识。本文提出通用自动标注框架AnnotateAnything,用于将被动3D资产转换为带结构化操作标签的「操作就绪」资产。该框架包含两个互补流程:一是基于视觉语言推理的流程,用于推断物体语义和交互区域;二是全自动、大规模的物理标注流程,将先验知识基于几何和物理约束落地,从而为仿真中的机器人操作提供可执行的标注。

DexLink Hand: A Compact, Affordable, 16-DOF Linkage-Driven Hand with Human-Like Dexterity

第一作者: Hao Wu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dexterous robotic hands face a longstanding trade-off among dexterity, compactness, and affordability. Particularly, high-degree-of-freedom designs typically demand complex actuation and transmission, hindering integration into human-scale forms. To address these challenges, this work presents a compact, low-cost linkage-driven anthropomorphic hand that achieves high dexterity, structural integration, and human-hand-like functionality. The hand integrates 20 joints driven by 16 independent actuators, with all actuation, sensing, and transmission components compactly embedded within a human-hand-sized structure. The resulting prototype weighs only 320g at a total cost below USD 400. To meet these objectives, a hybrid mechanical architecture combining planar and spatial linkage mechanisms is proposed, enabling decoupled multidirectional motion, biomimetic joint synergies, and...

论文介绍 为解决灵巧手在灵巧性、紧凑性与成本之间的权衡问题,本文提出DexLink手。这是一种紧凑、低成本的连杆驱动人形手,拥有20个关节和16个独立驱动器,所有组件均集成在类人手大小的结构中。其核心是提出了结合平面和空间连杆机构的混合机械架构,实现了多方向运动解耦和仿生关节协同,使得原型手重量仅320克,成本低于400美元。

Where Should Action Generation Begin? A Learnable Source Prior for Generative Robot Policies

第一作者: Meipo Dai · 方向: 机器人操作 · 来源: cs.RO

Abstract:Generative robot policies typically begin action generation from an observation-independent standard Gaussian distribution, leaving the choice of source distribution underexplored. This work asks a simple question: where should action generation begin? We propose LeaP, a Learnable source Prior that replaces the standard Gaussian with a proprioception-conditioned diagonal Gaussian over action chunks. Parameterized by a lightweight MLP, LeaP jointly predicts the mean and state-adaptive variance of the source distribution, while keeping the downstream generator architecture and inference solver unchanged. This design provides an observation-informed yet stochastic initialization, allowing the generator to focus on precise action refinement rather than transporting samples from an uninformed noise source. On 15 RoboTwin manipulation tasks, LeaP achieves an average success rate of...

论文介绍 生成式机器人策略通常从与观测无关的标准高斯分布开始动作生成。本文探讨「动作生成应从何处开始」这一问题,提出了可学习源先验LeaP。它用一个由本体感觉条件化的轻量级MLP来预测动作块的初始分布均值和方差,替代了标准高斯分布。这种设计为下游生成器提供了观测信息丰富的随机初始化,使其能专注于精确的动作修正,而非从无信息的噪声源传输样本。

Damage Adaptation in Seconds for Architected Materials

第一作者: James Avtges · 方向: 具身智能 · 来源: cs.RO

Abstract:Adaptation to damages and in-situ physical repairs is essential for long-term robot autonomy, yet challenging outside of narrowly defined and well-anticipated bounds. In this work we proprioceptively adapt to catastrophic damage in soft-actuated systems in under one minute. Architected materials are well equipped for adaptation: actuator failure occurs gradually rather than acutely, and damage can be described in a low-dimensional, discrete coordinate space. Surprisingly, latent damage representations plus a simple yet robust ensemble method is sufficient for adapting to unseen damage in real-time. Moreover, we identify conditions under which exponential sample complexity collapses to linear sample complexity for learned representations of architected materials, a concrete advantage over rigid components or continuum soft mechanisms. We demonstrate LEAP, our method for...

论文介绍 机器人适应未预料到的损伤是实现长期自主的关键。本文研究软驱动架构材料系统,利用其损伤渐进和低维离散坐标空间描述的特性,实现灾难性损伤后的快速自适应。核心方法LEAP基于潜在损伤表示和简单的集成方法,使得系统能在不到一分钟内适应未见过的损伤。该工作还识别了架构材料学习表示中样本复杂度从指数级降至线性级的条件。

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning

第一作者: Gaotian Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Internet videos constitute the largest reservoir of embodied human manipulation knowledge, yet converting arbitrary RGB footage into actionable robot training data remains a major bottleneck. Existing lab- or factory-collected datasets are narrow in scale and diversity, limiting open-world robot learning. Instead of proposing a static dataset, we introduce EgoInfinity, a universal 4D hand-object interaction data engine that enables web-scale data generation for robot retargeting and learning. EgoInfinity is a modular engine integrating perception, segmentation, reconstruction, interaction-aware refinement, and retargeting to automate this traditionally unscalable video-to-action problem without human-in-the-loop annotation. Its modular design lets the engine continuously benefit from advances in any incorporated component. With EgoInfinity, in-the-wild human manipulation...

论文介绍 该研究针对从网络视频提取机器人操作数据的瓶颈,提出了EgoInfinity这个模块化4D手物交互数据引擎。它集成了感知、分割、重建、交互感知优化和重定向等模块,能自动化地将野外人类视频转换为机器人可用的动作轨迹,旨在为机器人学习提供大规模、多样化的数据支持。

Contactless Respiratory Monitoring on Heterogeneous Mobile Robots: A Multimodal Edge-Computing Framework

第一作者: Milind Rampure · 方向: 多模态具身 · 来源: cs.RO

Abstract:Respiratory-rate (RR) monitoring is a critical component of remote triage and victim assessment in emergency response, disaster recovery, and infectious-disease scenarios, where minimizing physical contact can reduce responder risk and improve operational safety. However, field deployment of contactless RR monitoring remains challenging due to variable illumination, posture changes, platform heterogeneity, and the impracticality of wearable sensors in hazardous environments. In this paper, we present a modality-adaptive contactless RR monitoring framework for heterogeneous mobile robots with onboard edge computing. The proposed system combines brightness-adaptive sensor selection across RGB, thermal, near-infrared (NIR), and low-light cameras, keypoint-guided chest ROI extraction for posture-robust monitoring, and a signal-quality-index (SQI)-based filtering mechanism for...

论文介绍 该研究关注在灾难响应等场景中,通过移动机器人进行无接触呼吸监测的挑战。它提出了一个模态自适应的框架,能根据环境条件在RGB、热成像、近红外和低光相机间自动选择,并利用边缘计算实现鲁棒的呼吸率估计,以提升远程伤员评估的安全性与效率。

Transformer-Based Warm-Starting for Feasible and Optimal Terminal Approach to Tumbling Objects with Space Manipulators

第一作者: Yuji Takubo · 方向: 导航与运动 · 来源: cs.RO

Abstract:Real-time trajectory generation for on-orbit robotic servicing is challenging due to the nonlinear coupling between spacecraft bus motion, manipulator dynamics, visibility cone, and trajectory-level safety constraints. This paper studies learning-based warm-starting for sequential convex programming (SCP) in the terminal approach of a space manipulator toward a tumbling target. The proposed framework decomposes the problem into a system center-of-mass translational planning stage and a coupled attitude--manipulator torque-allocation stage, and applies a causal transformer warm-start to the latter, which constitutes the dominant computational bottleneck. Linear and flow matching action decoders are compared under different action-chunking and training dataset sizes, and the resulting warm-starts are evaluated under both cost-optimal and feasibility projection using SCP. Across...

论文介绍 该研究解决空间机械臂对翻滚目标的终端接近轨迹规划问题。它提出了一种基于因果Transformer的热启动方法,用于加速序列凸规划求解器。框架将问题分解为质心平移规划和姿态-机械臂力矩分配两个阶段,并通过学习提供初始猜测,旨在提升规划效率和可行性。

Abstention-Aware Personalized Object Rearrangement via Uncertainty-Guided LLM Assistance

第一作者: Sam Collin · 方向: 具身智能 · 来源: cs.RO

Abstract:Robotic assistance in household environments requires not only predicting where objects should be placed, but also reasoning about when objects should not be placed at all. Existing approaches to personalized object rearrangement primarily focus on placement decisions under the assumption of clean observations and complete actionability, limiting their applicability in realistic, cluttered, and partially erroneous settings. In this paper, we introduce APOLLO, a hybrid framework for abstention-aware personalized object rearrangement that combines a lightweight, personalized embedding model (PEM) with selective large language model (LLM) assistance. PEM is trained for each user-environment pair using a small number of demonstrations, operates entirely on CPU, and produces uncertainty estimates, which are used to selectively invoke LLM-based reasoning only for ambiguous...

论文介绍 该研究针对家庭环境中的个性化物体重新排列问题,引入了弃权感知机制。APOLLO框架结合了轻量级个性化嵌入模型和不确定性估计,仅当模型对重排决策不确定时,才选择性地调用大语言模型进行推理,以提高系统在杂乱、含错误观测的真实场景中的实用性与可靠性。

VISTA: Scale-Aware Visual Navigation via Action History Conditioning

第一作者: Maeva Guerrier · 方向: 导航与运动 · 来源: cs.RO

Abstract:Vision Navigation Foundation Models (VNMs) promise end-to-end learned navigation policies capable of zero-shot deployment across diverse embodiments and environments. To maintain generality, many vision-based navigation models predict normalized actions. However, this normalization introduces a critical deployment vulnerability: applying different scaling factors to the same normalized trajectory alters its physical geometry, which degrades navigation performance and increases collision risks. We address this vulnerability by conditioning the model on normalized action histories alongside image observations, providing explicit context on the relationship between the model's predictions and the robot's actual physical displacement. Furthermore, current VNMs often struggle in visually repetitive environments that lack distinct features. To resolve this issue, we integrate a...

论文介绍 该研究针对视觉导航基础模型中因动作归一化导致的尺度不一致问题,提出了VISTA方法。它通过条件化动作历史来提供实际物理位移的上下文,并集成记忆机制以增强在特征稀缺、视觉重复环境中的导航能力,旨在提升模型在不同机器人平台上的部署鲁棒性。

Contrastive Action-Image Pre-training for Visuomotor Control

第一作者: Yuvan Sharma · 方向: 具身智能 · 来源: cs.RO

Abstract:Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior work circumvents this data scarcity by turning to internet-scale image and language data or egocentric human video. While these models show promise, neither paradigm learns from paired vision and action data, which downstream visuomotor control policies require. However, robot trajectories, the most direct source of this paired signal, are not available at pre-training scale, motivating us to extract action signals from abundant human video instead. To this end, we introduce CAIP (Contrastive Action-Image Pre-training), a vision encoder that treats human hand poses from large-scale egocentric video as a proxy for end-effector actions. By extracting 3D hand keypoints, a representation that aligns naturally with downstream robot...

论文介绍 该研究解决机器人视觉运动控制预训练中配对视觉-动作数据稀缺的问题。CAIP方法利用大规模第一人称视频中的手部姿态作为末端执行器动作的代理,通过对比学习框架预训练视觉编码器,旨在使预训练特征更好地与下游机器人控制任务所需的表示对齐。

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

第一作者: Hao Li · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos provide complementary real-world supervision in pretraining. However, joint training on human and robot data remains challenging due to divergences in action spaces, embodiment structures, temporal dynamics, and supervision quality. We introduce ACE-EGO-0, a unified VLA pretraining framework jointly leveraging heterogeneous data sources. To extract large-scale pretraining supervision from egocentric human videos, we build a scalable egocentric video-to-action pipeline that converts raw human videos into robot-format pseudo-action trajectories. To make these labels comparable with robot demonstrations, ACE-EGO-0 uses a unified action representation based on...

论文介绍 该研究致力于利用大规模第一人称人类视频来辅助视觉-语言-动作模型的预训练。ACE-EGO-0框架通过可扩展的管道将人类视频转换为机器人格式的伪动作轨迹,并设计统一的动作表示,以实现人类和机器人异构数据的联合训练,旨在缓解机器人数据采集成本高的问题。

Extracting Semantics: LLM-Guided Automatic Population of Robot Ontology from URDF

第一作者: Bastien Dussard · 方向: 具身智能 · 来源: cs.RO

Abstract:While commonsense knowledge may suffice for virtual agents, embodied robots interacting with humans require grounded and semantically rich representations of both their environment and their own physical embodiment. In cognitive robotics, ontologies are effective for integrating such heterogeneous knowledge to enable explainable reasoning, even during continuous knowledge updates. Yet, their manual construction remains a bottleneck. We present a preliminary approach for the automatic generation of robot semantic abstractions by transforming Unified Robot Description Format (URDF) models into populated ontologies. Although URDF files provide structural and kinematic descriptions, their identifiers often require commonsense interpretation to recover meaningful semantics, a task at which Large Language Models (LLMs) excel. Our pipeline leverages LLMs to infer semantic...

论文介绍 该研究提出一种从机器人描述格式URDF文件自动生成语义本体论的初步方法。它利用大语言模型解读URDF中的标识符以推断其语义,并自动构建包含结构、功能等丰富语义信息的机器人本体论,旨在支持具身机器人的可解释推理和异构知识集成。

Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So

第一作者: Josef Liyanjun Chen · 方向: 机器人操作 · 来源: cs.RO

Abstract:A robot's flash endurance is a non-renewable stock: every persisted write spends one of a few thousand program/erase cycles and never refills, yet no fielded robot memory system prices which memories are worth an erase cycle. We treat embodied memory as depreciating capital and price that stock with a single endurance shadow price $\eta$, which makes cost-minimizing placement across a RAM / on-board NVM / cloud hierarchy a threshold in a wear-augmented per-byte index. The index is cost-optimal whatever the sign of the value-write association $\chi$; only when $\chi > 0$ does the optimum turn non-monotone, sending a robot's most valuable memories off its flash. The pivot is thus empirical, and we measure $\chi$ on real robot logs at a pre-specified gate: its sign is a property of the deployment regime -- positive on recurrent long-horizon manipulation ($\hat{\chi} \approx +1.0...

论文介绍 本研究将具身机器人的闪存耐久性视为一种会耗尽的非再生性资源。论文引入了一个「耐久性影子价格」η,将内存建模为折旧资本,并基于此在 RAM、本地非易失性存储和云存储的层次结构中,构建了一个考虑磨损的字节级索引,以实现成本最优的内存放置策略。该策略在内存价值与写入量呈正相关时,最优解可能表现为非单调,会将最有价值的记忆迁移到云端,这取决于具体部署场景。

Learn to Quantify Social Interaction with Constraints for Pedestrian Walking

第一作者: Xiaodan Shi · 方向: 导航与运动 · 来源: cs.RO

Abstract:Long-term human path forecasting in crowds is critical for autonomous moving platforms (like autonomous driving cars and social robots) to avoid collision and make high-quality planning. Although the current research take into account social interactions for prediction, they don't reveal the exact kinds of social interactions happened among people and how the social interactions affect the decision-making process of pedestrians, which further limits its robustness. Social interactions in pedestrian walking are intuitively massive and hard to label and quantify. In this paper, we explore creatively to quantify and interpret how pedestrians interact with others by proposing Learn to Cluster. Our clustering social interactions is probabilistic latent variable generative, learning directly from sequential trajectory observations, scalable to arbitrary number of pedestrians. Learn...

论文介绍 该研究聚焦于量化行人间的社交互动,以提升自主移动平台的长期路径预测。论文提出了一种名为「Learn to Cluster」的方法,它通过概率隐变量生成模型,直接从轨迹序列观测中学习聚类社交互动,从而解释不同类型的社交行为及其对行人决策的影响。该方法具有可扩展性,能够处理任意数量的行人,旨在为规划系统提供更鲁棒和可解释的社交信号。

GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning

第一作者: Haoyu Wang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience to plan reliable robot trajectories. GeneralVLA provides a hierarchical interface for converting language and RGB-D observations into 3D end-effector paths, but two bottlenecks remain. First, monocular SAM3D-style object reconstruction can hallucinate pose and unseen geometry, while manipulation benefits from stable object shape when calibrated multi-view observations are available. Second, the original KnowledgeBank mainly retrieves semantically similar snippets and appends new knowledge, which makes it difficult to control memory quality, conflicts, confidence, and geometric relevance. To address the first challenge, we introduce GeoFuse-MV3D, a geometry-prior-guided MV-SAM3D reconstruction branch that verifies external geometry cues with input-view masks, applies...

论文介绍 本文针对通用视觉语言动作模型在机器人规划中的局限进行了改进。针对单视角3D重建可能产生的几何误差问题,提出了 GeoFuse-MV3D 分支,利用几何先验和多视图信息来引导更准确的物体形状重建。同时,为解决现有知识库在记忆质量、冲突和相关性控制方面的不足,引入了受治理的记忆模块,旨在管理可重用的操作经验,以提升规划的可靠性和复用性。

WeaveLA: Event Driven Cross-Subtask Latent Memory Weaving for Repetitive Robot Manipulation

第一作者: Shoujing Zhu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) policies have achieved remarkable single-step manipulation, yet they remain brittle precisely where each stage depends on what was just completed. The core issue is structural: short-window VLAs lack an explicit channel for rouxting information across sub-task boundaries, and existing memory-augmented variants either write at every frame, retrieve from demonstration-time stages, or fire at sub-goal events without performing an explicit sub-task-to-sub-task hand-off into the action expert. We identify the sub-goal completion event as the natural temporal unit for cross-subtask memory hand-off, and present WeaveLA (Weave Latent memory for Vision-Language-Action policies), a cross-subtask memory interface that, on top of a frozen VLA backbone, compresses each completed segment into latent tokens via query-driven attention pooling and routes them...

论文介绍 视觉语言动作策略在处理多阶段重复操作任务时,因缺乏子任务间的信息传递机制而表现脆弱。本文提出 WeaveLA,一个事件驱动的跨子任务潜记忆接口。它以子目标完成事件为触发点,通过查询驱动的注意力池化将已完成的子任务压缩为潜表示,并将其路由至动作专家,从而在冻结的 VLA 主干网络上实现跨子任务的记忆交接,以增强策略的连续性和鲁棒性。

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models

第一作者: Xinglong Sun · 方向: 策略学习 · 来源: cs.RO

Abstract:Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-dependent. Commonly used rule-based driving metrics like EPDMS are interpretable but lack context-awareness, while recent VLMbased evaluations are context-aware but limited by ambiguous VLM outputs and weak physical grounding. To evaluate driving in a manner that is both interpretable and context-aware, we introduce DriveJudge. DriveJudge is a driving evaluation agent that combines rule-grounded evaluation with Vision-Language Model (VLM) reasoning and selectively invokes physically-grounded deterministic rule functions after interpreting the environmental context. To train and evaluate DriveJudge, we curate a large-scale dataset of 33,577 challenging driving samples with human annotations on whether...

论文介绍 本文提出 DriveJudge,旨在重新思考自动驾驶策略的评估方式。现有的规则评估指标缺乏情境感知,而基于视觉语言模型的评估则存在输出模糊和物理依据不足的问题。DriveJudge 作为一个评估智能体,结合了规则评估与视觉语言模型推理能力,能根据环境上下文有选择地调用物理基础的确定性规则函数,从而在保持可解释性的同时实现对驾驶质量的情境感知评估。

Kolmogorov Regression for Robust Diffusion Policies

第一作者: Lekan Molu · 方向: 策略学习 · 来源: cs.LG

Abstract:Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long-horizon performance (when deployed on physical systems). We introduce a backward Kolmogorov equation that lifts diffusion policies to a Cameron-Martin space -- a subset of the Hilbert space. Essentially, replacing stochastic score matching with a deterministic boundary-value PDE problem. Our core innovation thrives on Gaussian measure theory whereupon the diffusion noise covariance operator is realized from a colored noise distribution which prescribes a notion of regularity on samples from the model at inference time. We train the diffusion model with a derived precision-weighted Cameron- Martin loss and a Kolmogorov residual is introduced as a PDE diagnostic during inference. These substitutions yield (i) convergence guarantees where the bound's constants...

论文介绍 针对有限维扩散策略在长时间步执行中因离散化误差导致的性能漂移问题,本文提出一种基于柯尔莫哥洛夫回归的鲁棒方法。核心思想是将扩散策略提升至卡梅隆-马丁空间,通过求解一个确定性的偏微分方程边界值问题来替代随机的分数匹配。该方法利用高斯测度理论推导出的精度加权损失进行训练,并在推理时引入柯尔莫哥洛夫残差作为诊断,旨在为策略提供理论保证并增强其长期执行的鲁棒性。

市场总览

美股技术面整体呈现分化:主要ETF如SPY和QQQ当前价格低于SMA20但高于SMA50和SMA200,RSI分别在50.7和53.4的中性区域,趋势为bullish但近期回调,pct1Day跌幅在-1%左右。加密资产市场情绪极度恐慌,恐慌贪婪指数为15,总市值2.31万亿美元且24小时下跌1.47%,BTC主导率56.1%;技术指标显示BTC、ETH和SOL均呈空头排列,但出现MACD金叉信号,RSI在39-47区间,价格远低于长期均线。中概股普遍承压,BABA、PDD等RSI进入超卖或接近超卖,价格显著低于SMA,pct5Day跌幅较大,空头排列明显。商品外汇中,黄金期货RSI44.8、趋势中性并出现MACD金叉;原油期货RSI29.4超卖,价格大幅低于均线;美元指数DXY RSI64.9偏强,接近52周高,多头排列。

今日关注

BABA 阿里巴巴 (BABA)
偏下行

RSI为25,处于超卖状态,pct1Day -3.18%,pct5Day -6.88%,当前价格107.44远低于SMA20的122.28和SMA200的149.34,信号包括RSI超卖和空头排列,表明强劲下行技术压力。

CL=F WTI 原油期货
偏下行

RSI为29.4,处于超卖状态,pct1Day -1.13%,pct5Day -16.48%,当前价格75.19远低于SMA20的89.48和SMA50的94.64,信号显示RSI超卖,下行技术趋势持续。

DX-Y.NYB 美元指数 DXY
偏上行

RSI为64.9,接近52周高,pct1Day 0.74%,pct5Day 0.33%,当前价格100.28高于SMA20的99.52、SMA50的98.89和SMA200的98.67,信号包括MACD金叉、接近52周高和多头排列,显示上行动能。

BTC-USD Bitcoin
中性

RSI为39.3,pct1Day -1.39%,pct5Day 1.8%,当前价格64689.99低于SMA20的65578.79但高于近期低点,信号有MACD金叉和空头排列,技术面矛盾,呈现中性震荡格局。

全部资产

^VIX

VIX 恐慌指数

$18.44 +12.37%
5 日
-17.01%
距 52w 高
-47.8%
RSI(14)
51.5
趋势
中性
SMA 20 / 50 / 200
17.42 / 17.85 / 18.56
MACD / 信号
0.015 / -0.039
死叉(SMA50↓SMA200) (5 天前)

^TNX

10Y 美债收益率 (%)

$4.46 -0.53%
5 日
-1.96%
距 52w 高
-10.7%
RSI(14)
48.2
趋势
多头
SMA 20 / 50 / 200
4.52 / 4.42 / 4.21
MACD / 信号
0.016 / 0.026
多头排列

DX-Y.NYB

美元指数 DXY

$100.28 +0.74%
5 日
+0.33%
距 52w 高
-0.4%
RSI(14)
64.9
趋势
多头
SMA 20 / 50 / 200
99.52 / 98.89 / 98.67
MACD / 信号
0.290 / 0.269
MACD 金叉 (今天)接近 52 周高多头排列

SPY

S&P 500 ETF

$740.96 -1.25%
5 日
+2.14%
距 52w 高
-2.6%
RSI(14)
50.7
趋势
多头
SMA 20 / 50 / 200
746.80 / 728.24 / 687.83
MACD / 信号
4.034 / 5.922
接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$722.51 -1.01%
5 日
+4.15%
距 52w 高
-3.5%
RSI(14)
53.4
趋势
多头
SMA 20 / 50 / 200
725.50 / 690.49 / 627.76
MACD / 信号
8.804 / 11.508
多头排列

AAPL

Apple

$295.95 -1.10%
5 日
+1.50%
距 52w 高
-6.8%
RSI(14)
48.9
趋势
多头
SMA 20 / 50 / 200
303.61 / 287.96 / 267.85
MACD / 信号
1.330 / 3.724
多头排列

MSFT

Microsoft

$378.91 -3.79%
5 日
-4.64%
距 52w 高
-31.8%
RSI(14)
34.6
趋势
空头
SMA 20 / 50 / 200
415.23 / 412.87 / 451.98
MACD / 信号
-7.295 / -2.288
空头排列

NVDA

Nvidia

$204.65 -1.33%
5 日
+2.11%
距 52w 高
-13.5%
RSI(14)
45.3
趋势
中性
SMA 20 / 50 / 200
212.43 / 208.74 / 189.70
MACD / 信号
-1.301 / 0.126

GOOGL

Alphabet

$363.79 -2.53%
5 日
+2.08%
距 52w 高
-11.0%
RSI(14)
46.1
趋势
中性
SMA 20 / 50 / 200
372.67 / 366.36 / 310.32
MACD / 信号
-2.043 / -0.455

TSLA

Tesla

$396.38 -2.05%
5 日
+3.88%
距 52w 高
-20.5%
RSI(14)
45.3
趋势
空头
SMA 20 / 50 / 200
414.54 / 401.34 / 416.61
MACD / 信号
-2.618 / 0.082
空头排列

META

Meta

$567.58 -5.44%
5 日
-0.60%
距 52w 高
-28.7%
RSI(14)
39.3
趋势
空头
SMA 20 / 50 / 200
600.87 / 622.60 / 655.71
MACD / 信号
-11.365 / -9.463
空头排列
加密恐慌贪婪
15
极度恐慌
加密总市值
$2.31 T
-1.47% / 24h
BTC 主导率
56.1%
ETH 9.2%
24h 成交量
$86.6 B
活跃币 17,432

BTC-USD

Bitcoin

$64,689.99 -1.39%
5 日
+1.80%
距 52w 高
-48.7%
RSI(14)
39.3
趋势
空头
SMA 20 / 50 / 200
65,578.79 / 73,204.11 / 77,263.47
MACD / 信号
-2,427.802 / -3,049.862
MACD 金叉 (4 天前)空头排列

ETH-USD

Ethereum

$1,760.12 -1.69%
5 日
+5.70%
距 52w 高
-64.5%
RSI(14)
42.4
趋势
空头
SMA 20 / 50 / 200
1,767.49 / 2,036.31 / 2,389.43
MACD / 信号
-86.370 / -112.270
MACD 金叉 (3 天前)空头排列

SOL-USD

Solana

$72.46 -1.30%
5 日
+8.56%
距 52w 高
-71.4%
RSI(14)
47.1
趋势
空头
SMA 20 / 50 / 200
71.14 / 80.74 / 98.70
MACD / 信号
-2.874 / -4.114
MACD 金叉 (3 天前)空头排列

BABA

阿里巴巴 (BABA)

$107.44 -3.18%
5 日
-6.88%
距 52w 高
-44.2%
RSI(14)
25.0
趋势
空头
SMA 20 / 50 / 200
122.28 / 129.67 / 149.34
MACD / 信号
-5.987 / -4.472
RSI 超卖空头排列

PDD

拼多多 (PDD)

$79.86 -2.12%
5 日
-2.40%
距 52w 高
-42.7%
RSI(14)
32.5
趋势
空头
SMA 20 / 50 / 200
86.36 / 94.24 / 110.80
MACD / 信号
-4.032 / -3.905
接近 52 周低空头排列

JD

京东 (JD)

$27.91 -1.66%
5 日
-1.90%
距 52w 高
-24.3%
RSI(14)
37.6
趋势
空头
SMA 20 / 50 / 200
29.32 / 30.10 / 30.27
MACD / 信号
-0.602 / -0.476
空头排列

0700.HK

腾讯控股 (0700.HK)

HK$445.40 -0.45%
5 日
-4.34%
距 52w 高
-34.8%
RSI(14)
44.4
趋势
空头
SMA 20 / 50 / 200
449.79 / 470.32 / 566.82
MACD / 信号
-3.803 / -5.152
空头排列

GC=F

黄金期货

$4,344.00 +0.30%
5 日
+5.74%
距 52w 高
-22.2%
RSI(14)
44.8
趋势
中性
SMA 20 / 50 / 200
4,391.30 / 4,568.51 / 4,432.97
MACD / 信号
-90.344 / -92.799
MACD 金叉 (今天)

CL=F

WTI 原油期货

$75.19 -1.13%
5 日
-16.48%
距 52w 高
-37.1%
RSI(14)
29.4
趋势
中性
SMA 20 / 50 / 200
89.48 / 94.64 / 73.61
MACD / 信号
-4.695 / -2.960
RSI 超卖

USDCNY=X

美元 / 人民币

¥6.76 +0.00%
5 日
-0.23%
距 52w 高
-6.3%
RSI(14)
33.4
趋势
空头
SMA 20 / 50 / 200
6.77 / 6.80 / 6.96
MACD / 信号
-0.013 / -0.013
接近 52 周低空头排列
风险提示

本报告基于公开行情数据计算的技术指标进行客观解读,所有描述均反映当前指标读数。过去走势不代表未来表现,技术分析存在局限性,本内容仅供技术指标解读参考,不构成任何投资建议。

How the oil market shrugged off the Iran crisis

Fears of summer shortages and $200 oil have been replaced by a focus on looming gluts

中文摘要 石油市场的关注点已从对伊朗危机可能引发夏季供应短缺及油价升至200美元的担忧,转向对即将到来的供应过剩的预期。

Channel Buys Fidante to Grow Australia Fund Manager Services

Channel Capital will merge with Challenger Ltd.’s Fidante unit to create an Australian fund manager spanning a range of investment strategies with around A$150 billion ($105 billion) in assets.

中文摘要 Channel Capital将与挑战者有限公司旗下的Fidante部门合并,创建一家资产规模约1500亿澳元(1050亿美元)、涵盖多种投资策略的澳大利亚基金公司。

Short Seller Left Loses Mistrial Bid Over Court Error, for Now

Famed short seller Andrew Left, convicted on 13 counts of securities fraud earlier this month, is off to a rocky start in his effort to sway a judge to overturn his verdict due to an unusual courtroom error.

中文摘要 知名做空者安德鲁·莱夫特因一项罕见的法庭失误,其要求法官推翻判决的请求在现阶段受挫。他此前因13项证券欺诈罪名被定罪。

China’s Unusually Heavy Rains Fill Dams and Put Crops at Risk

It’s been a particularly wet start to southern China’s rainy season, and more unusually heavy downpours could be on the way over the summer.

中文摘要 中国南方雨季初期降雨异常充沛,导致水坝蓄水增加,并对农作物构成风险。气象预报显示夏季可能还会有更多强降雨。

Braskem Debt Talks Hit Snag, Boosting Odds of Emergency Order

Braskem SA and its new controlling shareholder, IG4 Capital, are struggling to win support from enough creditors to move forward with an out-of-court restructuring proposal amid disagreements over uneven treatment and collaterals, according to people familiar with the matter.

中文摘要 巴西石化公司Braskem及其新控股方IG4资本在推进庭外债务重组方案时遇到阻力,因债权人就差异化的处置方案和抵押品问题产生分歧,紧急司法令的可能性增加。

Traders Keep Faith in Malaysian Bonds in Face of Deficit Warning

Traders remain sanguine about Malaysia’s financial outlook even after officials warned the government may miss its deficit target this year, according to a closely watched financial-market metric.

中文摘要 尽管马来西亚官员警告政府可能无法实现本年度的赤字目标,但根据一项备受关注的金融市场指标,交易员对该国金融前景仍保持信心。

Asia Strategists Eye Yen Intervention, Tech Stocks Post-Warsh

Strategists say markets are on watch for yen intervention after the Federal Reserve took a hawkish stance at Kevin Warsh’s debut meeting as governor, triggering a drop in the currency to levels that have previously prompted Japan’s finance ministry to step in.

中文摘要 美联储在凯文·沃什首次主持的会议上采取鹰派立场后,日元跌至此前曾促使日本财务省干预的水平,市场策略师正密切关注日本当局可能进行的干预行动。

'We had to get out of the way': The backlash over delivery robots

As the delivery vehicles increasing take to US streets, bans and protest groups are springing up.

中文摘要 随着配送机器人在美国街道上日益增多,相关的禁令和抗议团体也相继出现,引发了公众反弹。

Interest rates expected to be held by Bank of England

The Bank last cut interest rates in December but upheaval in the Middle East has stalled any further reductions.

中文摘要 预计英国央行将维持当前利率水平不变。该行上次降息是在12月,但中东地区的动荡已使其进一步降息的进程陷入停滞。

OT GROUP CEO on Onitsuka Tiger Spinning off from Asics

Onitsuka Tiger is looking to bring its mark as an independent brand, separate from Asics. OT GROUP President and CEO Ryoji Shoda sat down for an interview with Bloomberg News' Shery Ahn in Tokyo to explain his vision for the brand going forward. (Source: Bloomberg)

中文摘要 鬼塚虎正致力于作为独立于亚瑟士的独立品牌发展其市场形象。OT集团总裁兼CEO庄田亮治阐述了他对该品牌未来发展的愿景。

Gold Advances as Peace Deal Optimism Counters Hawkish Fed

Gold rose, supported by the signing of an interim peace deal between the US and Iran, even as the Federal Reserve signaled a rate hike later in the year.

中文摘要 黄金价格上涨,美伊临时和平协议的签署提供了支撑,尽管美联储暗示今年晚些时候可能加息。

New Zealand’s Economy Accelerated Before War Sapped Momentum

New Zealand’s economy sparked into life in the first three months of the year, responding to low interest rates and a lift in spending that largely pre-dated the Iran war.

中文摘要 新西兰经济在第一季度增长0.8%,受低利率和消费提振而加速,但这一增长主要发生在伊朗战事削弱增长动力之前。

【CHY公益站】知道不稳定是咋回事了

昨天晚上让号池里的4.8翻了翻请求日志,没管它去睡觉了,今天一看,原来4.8直接请求上游是完全可用的状态,但是通过New-API中转就出各种bug 准备今天下午放学就自写聚合网关,佬友们等我好消息 18 个帖子 - 17 位参与者 阅读完整话题

「君の公益」 gpt 复活(今日签到618刀~)

仅限codex调用,能活多久未知 目前的约束措施主要有 限制调用客户端为codex cli/Claude code cli 对色情和政治敏感词汇进行关键词拦截 限制单ip请求速率为一分钟一百次,超过即封禁ip 限制新注册必须为LD用户,再也不敢开邮箱注册了 不定期数据库拉ip清单去查分发,封号 所谓有求而不得,人心欲壑,可填沧海~ 告诉大家一个好消息,告别单身了,嘿嘿 428 个帖子 - 395 位参与者 阅读完整话题

打野归来,小鸡毛的佬友们今晚可以安心蹬GPT5.5了

目前号池每个PLUS的成本大概是: 6月17日小鸡毛公益站大致调整: 公益推广申明 状态 我的项目是免费使用的,无收费(变相收费、赞助)部分 是 我的帖子已经打上 公益推广 标签 是 我的项目属于个人项目,与公司或商业机构无关 是 我的项目不存在 QQ、TG 等群组引流 是 我的项目不存在非运营必要的网站引流 是 我的项目不存在为他人推广、AFF 是 我的项目无关联的商业项目 是 我的站点存在登录,并已接入 LINUX DO Connect 是 我帖子内的项目介绍,AI 生成、润色内容部分已截图发出 是 以上选择我承诺是永久有效的,接受社区和佬友监督 是 24 个帖子 - 21 位参与者 阅读

【学生优惠】阿里云免费薅2核2G服务器教程

首先打开阿里云的学生优惠页面https://university.aliyun.com 找到“我是学生”,领取这里的300元抵扣券 然后进入ECS(自定义配置)购买页面https://ecs-buy.aliyun.com/ecs/#/custom/ 像我这样选择配置: 关键点: 实例规格:ecs.e-c1m1.large 流量计费方式:(CDT)按量计费 选择正确的话,最后显示的结账金额应为300元以下,结合优惠券可以直接抵扣 楼主这里开启了CDT计费方式,流量按量计费(前20G免费),挂点不怎么烧流量的小程序还是不错的 你们也可以像楼主一样设置一些报警规则,防止被打了(我这里设置的不对,感谢

DeepSeek Free Key - 20260617

[!info] https://2c2ch1u11-share-api-0.hf.space/v1 [!info] sk-cb3f3c12406241ee8aca238f01cd18d9 [!warning] 蹬完就关贴,没关就是还能蹬 [!tip] DeepSeek V4 API | LDC兑换额度 | 超便宜渠道 本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营

grok heavy公益站,免费体验grok-composer-2.5-fast

本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 之前本来打算拼车的 但是发现grok拼车存

求顶,让始皇和管理大哥们看到!!!!迫不得已用此方法,还望见谅

事先声明:我并非有意这样做,是因为个人害怕某一天或者换设备后无法登录到社区内,只是想让领导者们帮忙解决一下小佬弟的问题 问题:因为个人的问题哈,我的邮箱无法登录了,被风控了,想要更换邮箱,但是邮箱的话需要原邮箱收件后确认才可更换,但是已经无法登录,所以收不到,需要始皇修改一下我的邮箱即可,当然是因为我个人的问题,确实是很着急,怕哪天上不来了,发帖也是为了让百忙之中的领导者们能看到,绝无恶意。 通过站内私信和在TG上联系,TG上显示已读,我也不知道是不是机器人读的(因为没有给我回复) support.google.com 账号恢复指南 - Google 帐号帮助 获取与恢复 Google 账号或