每日简报

2026-06-27

← 历史归档

simplex-chat/simplex-chat

Haskell · ★ 12,646 · 🍴 715 · 📈 432 stars today

SimpleX - the first messaging network operating without user identifiers of any kind - 100% private by design! iOS, Android and desktop apps 📱!

中文介绍 SimpleX 是一款主打极致隐私的即时通讯工具,首创无用户标识符的通信网络设计,从底层杜绝元数据泄露。提供 iOS、Android 和桌面端应用。适合对数据安全和匿名性有极高要求的个人及团队,用于日常加密聊天与敏感信息传输。

google-labs-code/design.md

TypeScript · ★ 21,323 · 🍴 1,727 · 📈 2,407 stars today

A format specification for describing a visual identity to coding agents. DESIGN.md gives agents a persistent, structured understanding of a design system.

中文介绍 DESIGN.md 是 Google 推出的一种格式规范,用于向 AI 编程代理描述视觉识别与设计系统。通过在项目中引入该文件,AI 代理能获得持久且结构化的设计规范理解。适合前端开发者与设计师,在 AI 辅助开发中保持 UI 风格与组件库的高度一致性。

commaai/openpilot

Python · ★ 61,805 · 🍴 11,050 · 📈 80 stars today

openpilot is an operating system for robotics. Currently, it upgrades the driver assistance system on 300+ supported cars.

中文介绍 openpilot 是开源机器人操作系统,目前主要作为自动驾驶辅助系统运行,支持超 300 款车型。它通过计算机视觉与深度学习技术升级原车辅助驾驶功能。适合极客车主与自动驾驶研究者,用于日常通勤辅助或底层算法的二次开发测试。

kunchenguid/no-mistakes

Go · ★ 3,453 · 🍴 207 · 📈 398 stars today

git push no-mistakes

中文介绍 no-mistakes 是一个 Git 别名或钩子工具,允许开发者通过执行 git push no-mistakes 命令来推送代码。它旨在简化提交流程并减少人为输入错误,适合追求高效命令行操作的开发者,在日常代码版本控制与快速迭代场景中提升 Git 操作体验。

grafana/grafana

TypeScript · ★ 74,924 · 🍴 14,127 · 📈 32 stars today

The open and composable observability and data visualization platform. Visualize metrics, logs, and traces from multiple sources like Prometheus, Loki, Elasticsearch, InfluxDB, Postgres and many more.

中文介绍 Grafana 是一款开源的可观测性与数据可视化平台,支持整合 Prometheus、Loki、Elasticsearch 等多数据源的指标、日志和追踪数据。适合 SRE、运维与开发人员,用于构建统一的监控仪表盘,快速排查系统故障并分析业务数据趋势。

ripienaar/free-for-dev

HTML · ★ 123,757 · 🍴 13,051 · 📈 90 stars today

A list of SaaS, PaaS and IaaS offerings that have free tiers of interest to devops and infradev

中文介绍 free-for-dev 是一份精心整理的免费开发者资源清单,汇总了提供免费订阅层的 SaaS、PaaS 和 IaaS 服务。适合独立开发者、初创团队及 DevOps 工程师,在搭建基础设施、部署应用或寻找开发工具时,快速筛选零成本的云服务与软件方案。

opendatalab/MinerU

Python · ★ 70,459 · 🍴 5,942 · 📈 960 stars today

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

中文介绍 MinerU 是一款文档解析工具,能将 PDF 和 Office 等复杂文档精准转换为 LLM 可用的 Markdown 或 JSON 格式。它采用先进的版面分析技术,适合 AI 应用开发者与数据工程师,在构建 RAG 系统或 Agent 工作流时进行高质量的数据预处理。

alchaincyf/zhangxuefeng-skill

★ 9,268 · 🍴 2,564 · 📈 160 stars today

张雪峰.skill — 张雪峰的认知操作系统。高考志愿/考研/职业规划的实战思维框架。由女娲.skill生成。

中文介绍 该项目是一个 AI 提示词工程,将张雪峰的升学与职业规划理念转化为结构化的认知操作系统。适合面临高考志愿填报、考研选择或职业规划的学生与家长,通过 AI 模拟实战思维框架,获取务实、接地气的决策建议,辅助人生关键路径的规划。

mauriceboe/TREK

TypeScript · ★ 7,690 · 🍴 651 · 📈 1,060 stars today

A self-hosted travel/trip planner with real-time collaboration, interactive maps, PWA support, SSO, budgets, packing lists, and more.

中文介绍 TREK 是一款支持自托管的旅行规划应用,提供实时协作、交互式地图、PWA 支持、SSO 单点登录、预算管理及行李清单等功能。适合家庭出游、朋友结伴或旅行社团队,用于共同制定行程、分摊费用并实时同步旅行计划,保障数据隐私。

xbtlin/ai-berkshire

Python · ★ 3,157 · 🍴 466 · 📈 1,274 stars today

AI 时代的伯克希尔:基于 Claude Code 的价值投资研究框架。巴菲特·芒格·段永平·李录四大师方法论 + 多Agent并行研究。| AI-era Berkshire: a value investing research framework built on Claude Code. 4 masters' methodologies + multi-agent adversarial analysis.

中文介绍 ai-berkshire 是基于 Claude Code 构建的价值投资研究框架,融合巴菲特、芒格等四位投资大师的方法论,采用多 Agent 并行与对抗机制进行深度分析。适合个人投资者与量化研究员,用于自动化研报生成、标的深度调研及投资策略验证。

calesthio/OpenMontage

Python · ★ 23,690 · 🍴 2,636 · 📈 1,754 stars today

World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.

中文介绍 OpenMontage 是首个开源的 Agentic 视频制作系统,内置 12 条流水线、52 种工具和 500 多项 Agent 技能,将 AI 编程助手转化为完整的视频制作工作室。适合内容创作者与视频开发者,用于自动化脚本生成、素材剪辑及复杂视频工程的高效构建。

aws/agent-toolkit-for-aws

Python · ★ 1,366 · 🍴 123 · 📈 243 stars today

Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS

中文介绍 agent-toolkit-for-aws 是 AWS 官方支持的 AI 代理工具包,提供 MCP 服务器、技能和插件,帮助 AI Agent 在 AWS 平台上进行构建与操作。适合云架构师与 AI 开发者,用于简化 AWS 资源的自动化管理、部署及云原生应用的智能开发流程。

NanmiCoder/MediaCrawler

Python · ★ 53,384 · 🍴 10,979 · 📈 673 stars today

小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫

中文介绍 MediaCrawler 是一款多平台社交媒体爬虫工具,支持抓取小红书、抖音、快手、B站、微博、贴吧及知乎的图文、视频与评论数据。适合数据分析师、运营人员与学术研究者,用于舆情监控、竞品分析、内容趋势研究及大规模社交媒体数据采集。

garrytan/gstack

TypeScript · ★ 116,660 · 🍴 17,312 · 📈 950 stars today

Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA

中文介绍 gstack 是 YC 总裁 Garry Tan 分享的 Claude Code 配置方案,包含 23 个定制工具,分别扮演 CEO、设计师、工程经理和 QA 等角色。适合独立开发者与初创团队,利用 AI 模拟完整软件研发团队,提升全栈开发、产品设计与质量保障效率。

IceWhaleTech/CasaOS

Go · ★ 35,401 · 🍴 2,027 · 📈 619 stars today

CasaOS - A simple, easy-to-use, elegant open-source Personal Cloud system.

中文介绍 CasaOS 是一款简单优雅的开源个人云系统,提供直观图形界面,让用户轻松在旧电脑或 NAS 上搭建私有云。支持 Docker 应用一键安装与文件管理。适合家庭用户与极客,用于构建家庭媒体中心、私有网盘及轻量级自托管服务,降低运维门槛。

JCodesMore/ai-website-cloner-template

TypeScript · ★ 21,396 · 🍴 3,102 · 📈 1,088 stars today

Clone any website with one command using AI coding agents

中文介绍 ai-website-cloner-template 是利用 AI 编程代理实现网站一键克隆的模板工具。只需一条命令,AI 即可自动分析目标网站结构并生成前端代码。适合前端开发者与设计师,用于快速拆解优秀网页设计、提取布局灵感或进行竞品页面的逆向分析。

Panniantong/Agent-Reach

Python · ★ 42,407 · 🍴 3,370 · 📈 1,194 stars today

Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.

中文介绍 Agent-Reach 是赋予 AI 代理全网信息获取能力的 CLI 工具,支持零 API 费用读取和搜索 Twitter、Reddit、YouTube、B站及小红书等平台内容。适合 AI 开发者,用于为 LLM 或 Agent 接入实时外部数据源,打破信息孤岛并增强上下文理解。

thoughts on why mcp didn't work, what's next

@RhysSullivan · 57.2K 粉丝 · 86.1K 阅 · 503 赞 · 25 转

mcp came out when the best models were sonnet 3.5 and GPT 4o not a lot was known about how to properly work with these tools yet, we were still incredibly concerned on models having access to tools,

中文介绍 复盘 MCP 协议早期遇冷原因。当时主流模型仅为 Sonnet 3.5 和 GPT-4o,业界对工具调用及权限控制缺乏经验,导致落地不佳。文章剖析技术局限并展望后续演进,为 Agent 工具链发展提供反思视角。

Introducing computer use in Gemini 3.5 Flash

@GoogleAIStudio · 179.4K 粉丝 · 41.0K 阅 · 605 赞 · 57 转

Computer use is now a built-in tool supported in Gemini 3.5 Flash, delivering our best performance yet for agentic computer use tasks. Previously only available as a standalone Gemini 2.5 computer use

中文介绍 Google 宣布 Gemini 3.5 Flash 正式内置「Computer use」功能,提供当前最佳的 Agent 计算机操作性能。该能力此前仅限独立的 Gemini 2.5 使用,此次集成大幅降低开发门槛,加速桌面自动化任务落地。

Life After Switching to Kimi

@0xDevin_ · 6.6K 粉丝 · 38.3K 阅 · 539 赞 · 5 转

Most AI tools are chatbots with a nice interface. Kimi is different. It is a full system: a browser automation engine called Claw that navigates websites like a human, an Agent Swarm that runs

中文介绍 分享 Kimi 使用体验。其并非普通聊天界面,而是包含浏览器自动化引擎 Claw(模拟人类浏览网页)与 Agent Swarm 多智能体协作的完整系统,展现了在复杂 Agent 工作流与系统级架构上的新探索。

Human in the /loop

@ericzakariasson · 76.0K 粉丝 · 32.3K 阅 · 518 赞 · 29 转

What I like most about coding with agents right now is the room to leave a few runs going and still get on with other work. When something finishes or needs a call, I show up. This post is a short

中文介绍 探讨 AI Agent 编程的「Human in the loop」模式。博主分享并发运行多个 Agent 的经验:让其在后台执行,自己处理其他工作,仅在完成或需决策时介入,这种异步协作大幅提升了多任务开发效率。

Previewing GPT-5.6 Sol: a next-generation model

OpenAI previews GPT-5.6 Sol, a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with its most advanced safety stack.

中文介绍 OpenAI发布下一代模型GPT-5.6 Sol的预览版。该模型在编程、科学和网络安全领域的能力更强,并配备了OpenAI最先进的安全系统。

Run a vLLM Server on HF Jobs in One Command

中文介绍 Hugging Face博客介绍了一项新功能,用户现在只需执行一条命令,即可在HF Jobs平台上快速运行vLLM服务器,从而大幅简化大语言模型的部署与运行流程。

Which tokens does a hybrid model predict better?

中文介绍 Hugging Face博客探讨了混合模型在词元预测方面的表现,详细分析了该类模型在预测哪些特定类型的词元时具有更好的准确性和优势。

Repositioning retail for the AI era

Artificial intelligence is rapidly reshaping retail, but not in the ways consumers might immediately notice. The biggest transformation may not be flashy virtual try-ons or chatbot shopping assistants, but in how decisions are made behind the scenes: how products surface in search results, how inven

中文介绍 MIT科技评论指出,人工智能正迅速重塑零售业。最大的变革并非虚拟试穿或聊天机器人等表面应用,而是后台决策方式的转变,例如商品在搜索结果中的呈现机制。

not much happened today

**Z.ai's GLM-5.2** leads in coding and agent benchmarks with top scores like **1595** on Code Arena: Frontend and **34.29%** reasoning accuracy with zero failures. Databricks improved GLM-5.2 speed to **392 tok/s** using hardware and optimizations. **Ornith-1.0**, a new MIT-licensed coding model fam

中文介绍 Z.ai的GLM-5.2在编程和智能体测试中领先,前端得分1595,推理准确率34.29%且零失败。Databricks将其速度优化至392 tok/s。此外,MIT许可的Ornith-1.0模型发布。

[AINews] It's Meta-Harness Summer

Move over, Harness Engineering, it is time for the harness of harnesses!

中文介绍 Latent Space报道了AI工程领域的最新趋势,探讨了元工具或工具的工具的兴起,认为这标志着AI开发框架和工程化进入了新的发展阶段。

How agents are transforming work

A new OpenAI research paper shows how AI agents are transforming work, enabling longer, more complex tasks and expanding productivity across roles.

中文介绍 OpenAI发布最新研究论文,阐述AI智能体如何改变工作方式。研究表明,AI智能体能够处理更长时间、更复杂的任务,并显著提升各个岗位的生产力。

Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks

In a rare double-interview, the Databricks technical leaders riff on what it will take for every company to build Agent Clouds

中文介绍 Databricks技术负责人Matei Zaharia和Reynold Xin接受专访,探讨前沿生态系统保持开放的必要性,并分析企业构建智能体云所需的关键条件。

Introducing computer use in Gemini 3.5 Flash

中文介绍 DeepMind官方博客宣布在Gemini 3.5 Flash模型中引入计算机使用功能,使该模型能够直接操作计算机界面并执行相关任务。

The emergence of the web data infrastructure layer for AI

AI is booming. New use cases are emerging each day. To capitalize on the technology’s potential, enterprises require data at scale. In many cases, though, the relevant information is blocked or unstructured, which limits its use by AI models. To understand this challenge, consider the foundation of

中文介绍 MIT科技评论指出,随着AI应用爆发,企业需要大规模数据。针对信息被封锁或非结构化的问题,专为AI服务的Web数据基础设施层正在兴起,以解决数据获取挑战。

每日论文 · arXiv cs.CR 最新公告批次

周末 arXiv 通常无新公告。当前展示最近一次可用公告批次。

Tilikum: Transaction Fair Ordering on a DAG without Weak Edges

第一作者: Giulio Segalini · 方向: 密码学协议

Abstract:Decentralized Finance (DeFi) applications rely heavily on the order in which transactions are executed, making them susceptible to reordering attacks that enable adversaries to extract Blockchain Extractable Value (BEV). While linear blockchain systems such as Ethereum have inspired extensive research into fair ordering mechanisms, DAG-based consensus protocols have remained largely unprotected despite their growing adoption for scalability and performance. In this paper, we introduce Tilikum, a DAG-based ledger protocol that ensures fair transaction ordering without relying on weak edges. Tilikum achieves ordering linearizability by leveraging median-based timestamp aggregation, or batch order fairness, while maintaining low data redundancy and robust garbage collection. We implemented Tilikum in Rust and evaluated it against representative baselines, namely Narwhal/Tusk...

论文介绍 针对去中心化金融中交易重排攻击问题,本文提出Tilikum,一种基于有向无环图(DAG)的账本协议。该协议无需依赖弱边即可确保交易公平排序,通过基于中位数的时间戳聚合实现排序线性化,并维持低数据冗余与鲁棒的垃圾回收机制,为DAG架构下的区块链可扩展性与安全性提供了新方案。

The Observer World: A Cryptographic Extension of Impagliazzo's Five Worlds

第一作者: Fabio F.G. Buono · 方向: 安全研究

Abstract:Impagliazzo's five worlds classify computational assumptions along a single axis, the existence of cryptographic primitives. All five worlds implicitly assume that every party, including the adversary, observes the full input, that the observer is always $O_{top}$. This assumption is so natural that it is never stated. This work makes it explicit and relaxes it by introducing a second, orthogonal axis, the observational axis, defined by the observer hierarchy introduced in previous work. Relaxing the assumption reveals structural phenomena, such as the collapse $P^{O_{prof}} = NP^{O_{prof}} \subset P$, that the five-world framework cannot express. We prove that this collapse holds unconditionally in all five worlds, showing that observational blindness and computational hardness are independent. We define the Observer World $W_O$, classify all world-observer pairs, identify...

论文介绍 本文扩展了Impagliazzo的密码学五世界框架,引入正交「观察轴」以显式化观察者信息获取假设。研究揭示了原框架无法表达的结构现象,证明观察盲性与计算硬度相互独立,并定义了全新的「观察者世界」。该工作为计算复杂性提供了更细致的分类视角,有助于深入理解不同信息条件下的密码学假设。

PRISM: PE Relational Inter-Section Matrix. A 2D Section-Aware Dataset for Static PE Malware Detection

第一作者: José M. Sacristán · 方向: 软件安全

Abstract:We introduce PRISM (PE Relational Inter-Section Matrix), an open dataset and feature representation for static Windows PE malware detection. Existing benchmarks such as EMBER, BODMAS, and SOREL-20M represent each PE file as a flat one-dimensional feature vector, discarding the ordering of sections and the relational context between them. PRISM instead encodes every binary as a two-dimensional matrix whose rows are individual PE sections in file order, with a global summary row that preserves compatibility with EMBER-style models. We build the corpus from four malware sources (BODMAS, MalwareBazaar, VirusShare, and CAPE) together with SOREL-20M benign software, yielding 83,633 deduplicated matrices and a family-filtered analysis corpus of 49,204 samples across 684 malware families. A formal separability analysis (Fisher Discriminant Ratio, mutual information, and inter-section...

论文介绍 针对现有静态PE恶意软件检测基准丢失节顺序与关系上下文的问题,本文提出PRISM数据集与特征表示方法。该方法将PE文件编码为保留节顺序的二维矩阵,并兼容传统一维模型。研究构建了包含数万个样本的大规模语料库并进行可分性分析,为恶意软件检测提供了更丰富的结构特征表示与评估基准。

Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions

第一作者: Gerhard Backfried · 方向: AI 安全

Abstract:We present a novel approach for applying Large Language Models (LLMs) to threat assessment in the context of foreign peacekeeping missions. Building on the PINPOINT project and its use case, the EU Monitoring Mission in Georgia, we combine an interdisciplinary risk-model with OSINT-based media collection and LLM-supported threat extraction. The proposed workflow maps media contents to mission-relevant threats, extracts structured information and applies several additional LLM-based processing steps to improve relevance and grounding. An evaluation of threats extracted from media documents shows high agreement between automatically generated results and human judgment for core aspects such as threat and mission relevance. These results indicate that LLMs provide a promising approach to support analysts in the context of peacekeeping missions.

论文介绍 本文提出一种将大语言模型应用于维和任务威胁评估的新方法。该研究结合跨学科风险模型与开源情报媒体收集,利用大语言模型从媒体内容中提取结构化威胁信息,并进行相关性优化。评估表明,自动提取结果在核心威胁与任务相关性方面与人类判断高度一致,为维和任务分析人员提供了有效的智能辅助工具。

zQR: A Verifiable QR-Driven zkSNARK Proof Verification Framework for Mobile Platforms

第一作者: Goshgar Can Ismayilov · 方向: 密码学协议

Privacy is one of the fundamental rights of individuals in modern societies. Yet, the practical adoption of privacy-preserving technologies in daily interactions remains limited. Zero-knowledge proofs offer strong privacy guarantees but are often hindered by their technical complexity. In this paper, we advance the idea of verifiable QR codes that enable off-line verifiers to verify proofs encoded in QR codes. Based on this core idea, we build a novel QR-driven zkSNARK proof verification framework (i.e., zQR) for mobile platforms. The framework integrates blockchain for auditability, non-repudiation and logging; and large-language models for automatic circuit generation. We perform a security discussion of the framework by considering multiple attack surfaces. Furthermore, we present an experimental evaluation measuring temporal costs (proof generation and verification latency, QR code...

论文介绍 针对零知识证明在日常交互中应用受限的问题,本文提出zQR框架,实现移动端基于二维码的zkSNARK证明离线验证。该框架集成区块链以保障审计与不可否认性,并利用大语言模型自动生成验证电路。研究还进行了多攻击面安全分析与延迟等性能评估,为隐私保护技术的移动端普及提供了轻量级验证方案。

Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation

第一作者: Ryan Fetterman · 方向: 安全研究

Abstract:LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. We show that this can miss vulnerabilities introduced by fine-tuning itself: models can learn token-level indicator semantics that preserve canonical accuracy while failing under behavior-preserving transformations such as PowerShell alias substitution, command reconstruction, string construction, execution indirection, and case mutation. We study Foundation-Sec-8B-Instruct and its base model, Llama-3.1-8B-Instruct, on matched PowerShell classification cohorts. Causal interventions localize the classification circuit to a late-attention route inherited from Llama rather than created by fine-tuning. Fine-tuning concentrates and semantically specializes this inherited structure, improving baseline behavior while creating...

论文介绍 本文研究了大语言模型在安全分类微调中引入的隐蔽逃逸漏洞。研究发现,微调会集中并特化模型继承的注意力路由结构,使其在标准同分布评估中保持高准确率,却在命令重构等保持行为的变换下失效。通过因果干预定位分类电路,揭示了标准评估的盲区,为提升安全大模型的鲁棒性提供了新的分析与防御视角。

Type-based information flow analysis for $π$-calculus with a dynamically extensible security lattice

第一作者: Yukihiro Oda · 方向: 密码学协议

Abstract:We develop a type system for secure information flow where new security levels can be created and inserted into the security lattice dynamically, i.e., even in the middle of an execution of a system. Our system is formalized by extending Kobayashi's type-based secure information flow analysis for Milner's pi-calculus, which is one of the most expressive models (or "languages") supporting both sequential and concurrent computations, with concise syntax, reduction-based semantics, and bisimulation equivalence as a robust formalization of secrecy as non-interference. The development required careful treatment of extensions of lattices themselves as well as deliberate generalization from the simple 2-element lattice (consisting of only High and Low) in the original system.

论文介绍 本文提出一种支持动态扩展安全格的类型系统,用于pi演算的安全信息流分析。该系统允许在程序执行过程中动态创建并插入新的安全级别,突破了传统固定安全格的限制。研究对格扩展机制进行了严格的形式化处理,为包含并发计算的复杂系统提供了更灵活、细粒度的保密性与非干扰性分析框架。

Physical Layer Authentication With Channel Knowledge Maps in Indoor Environments

第一作者: Luca Bonaventura · 方向: 安全研究

Physical layer authentication (PLA) allows to authenticate the user by comparing measurements over time, assuming their time consistency or by modeling their evolution. However, these assumptions become problematic when devices are in motion and in indoor environments due to multipath propagation and obstructions. In this paper, we propose a PLA mechanism for moving devices in indoor environments, where multiple access points (APs) estimate the dominant channel tap path loss (PL) and angle of arrival (AoA) from the received signals and compare them with previously collected channel knowledge maps (CKMs). Specifically, the measurements are compared to those in the neighborhood of the previously known position obtained from CKMs. A comprehensive security analysis is conducted under both random and optimal attacks. Numerical results in a representative indoor scenario, with CKM obtained...

论文介绍 针对室内移动设备因多径效应导致物理层认证失效的问题,本文提出基于信道知识图的认证机制。该方法通过接入点估计信号路径损耗与到达角,并与信道知识图中历史位置附近的测量数据比对。研究在随机与最优攻击下进行了全面安全分析,为复杂室内环境下的无线设备身份验证提供了可靠方案。

Design and Performance Evaluation of Secure RF and WiFi-Based Communication in Drone Swarms via Testbed Implementation

第一作者: Bhavya Dixit · 方向: 密码学协议

Abstract:Unmanned aerial vehicle (UAV) swarms rely on distributed coordination and cooperative communication to support scalable operations, extended coverage, and applications such as surveillance and real-time data exchange. Wireless technologies such as radio frequency (RF) and WiFi are widely used for UAV-to-UAV and UAV-to-ground control station (GCS) communication but introduce significant security challenges. MAVLink, the predominant communication protocol in UAV systems, provides message integrity and authentication but lacks built-in encryption, leaving telemetry traffic vulnerable to eavesdropping. In our previous work, we proposed MAVShield, a lightweight encryption framework for MAVLink communications. In this paper, MAVShield, AES-CTR, Speck-CTR, ChaCha20, and Rabbit are integrated into four custom-built UAVs to establish secure communication links over RF and WiFi...

论文介绍 针对无人机集群中MAVLink协议缺乏内置加密导致的数据窃听风险,本文提出并实现轻量级加密框架MAVShield。研究将该框架与多种加密算法集成至定制无人机测试床,构建基于射频和WiFi的安全通信链路,并评估其在无人机间及空地通信中的性能,为无人机集群的安全部署提供实践参考。

ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

第一作者: Liwei Liu · 方向: 密码学协议

Abstract:With the rapid evolution of LLM-driven agents, Model Context Protocol (MCP), an open protocol bridging LLMs with external tools, has quickly become foundational to modern agent ecosystems. However, the expanding adoption of MCP has also introduced novel security concerns such as Tool Poisoning Attack (TPA), which exploit LLM-server interactions to inject malicious prompts. Existing poisoning schemes typically adopt a monolithic plaintext embedding paradigm, which fails to withstand manual inspection or automated detectors. Current research still lacks a systematic analysis on multi-tool poisoning, where multiple tools can be exploited cooperatively to disperse detection risk. In this paper, we introduce ShareLock, a multi-tool threshold poisoning framework that utilizes Shamir's threshold scheme to ensure exceptional stealth and fault tolerance. ShareLock distributes the...

论文介绍 针对大语言模型代理中模型上下文协议面临的工具投毒攻击易被检测问题,本文提出ShareLock多工具阈值投毒框架。该方法利用Shamir阈值方案将恶意提示分散至多个工具,通过协同利用分散检测风险,在抵御人工审查与自动检测的同时实现高隐蔽性与容错性,揭示了多工具协同攻击的潜在威胁。

Protocol Prying: Systematic Vulnerability Research in the Apple AirDrop and Android Quick Share Proximity Transfer Protocols

第一作者: Arash Ale Ebrahim · 方向: 密码学协议

Abstract:Apple AirDrop and Google/Samsung Quick Share are proximity file-transfer protocols used by over five billion devices, yet their application-layer security properties remain largely unstudied because both stacks are proprietary and undocumented. Both protocols are reachable from wireless proximity without any prior pairing and process complex serialized content (binary plists, CPIO archives, Protocol Buffers, UKEY2 handshakes) inside privileged daemons, making them attractive zero-click targets across multiple operating systems. We perform the first cross-platform reverse engineering and protocol-aware fuzzing study of both stacks. We reconstruct AirDrop's seven-layer state machine and DVZip adaptive compression from binary analysis, build AIRFUZZ, a protocol-aware fuzzer that mutates pre-compression representations, and complement it with targeted hand-written analyses of...

论文介绍 针对AirDrop与Quick Share等近距离传输协议应用层安全研究不足的问题,本文首次开展跨平台逆向工程与协议感知模糊测试。研究重构了AirDrop七层状态机与压缩机制,开发专用模糊测试工具,系统性挖掘无需预配对即可触发的零点击漏洞,为提升主流操作系统近距离通信安全提供依据。

Jailbreaking for the Average Jane: Choosing Optimal Jailbreaks via Bandit Algorithms for Automatically Enhanced Queries

第一作者: Prarabdh Shukla · 方向: 安全研究

Abstract:With a profusion of jailbreaks for LLMs now widely known, a growing concern is that non-expert malicious actors ("the average Jane") could elicit actionable responses to malicious requests. In this work, we examine whether this concern is justified. A non-expert malicious actor requires two ingredients for a successful attack: a powerful jailbreak for their target model, acting on an effective malicious query. For the former, we propose a novel attack strategy based on the multi-armed bandit framework. This allows efficient online learning of the optimal jailbreak from a large choice set via noisy exploration on a small number of queries, with subsequent application of the learnt policy on an exploitation set. For the latter, we curate $\mathrm{FrankensteinBench}$, a safety benchmark of $11,279$ malicious queries drawn from manual curation over $7$ existing benchmarks, along...

论文介绍 针对非专家利用大语言模型越狱攻击的威胁,本文提出基于多臂老虎机框架的自动化攻击策略。该方法通过少量查询的噪声探索,从大量越狱模板中高效在线学习最优攻击策略,并构建包含万余条恶意查询的安全基准,旨在评估自动化越狱攻击的有效性,为大模型安全防御提供参考。

Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities

第一作者: Corban Villa · 方向: 软件安全

Abstract:AI-assisted vulnerability discovery has proven effective for bug classes like memory safety, where instrumentation confirms memory violations and efficiently filters false positives. Many dangerous vulnerability classes, such as cryptographic misuse, however, lack any comparable instrumentation. In this work, we present Chai, an AI-based system that discovers and validates cryptographic misuse vulnerabilities through naturally occurring signals. To achieve this, Chai rethinks the classical technique of differential testing by leveraging AI to 1) improve precision for detecting real security issues in libraries, and 2) repurpose commonly overlooked discrepancies as leads for tangible vulnerabilities in downstream applications. In doing so, Chai inverts the prevailing paradigm of AI vulnerability discovery: instead of auditing one codebase for many flaws, it catalogs flaws at...

论文介绍 针对密码学误用漏洞缺乏有效检测工具的问题,本文提出基于AI的漏洞发现系统Chai。该系统结合差分测试技术,利用AI提升密码库安全问题的检测精度,并将常被忽略的库级差异转化为下游应用实际漏洞线索。Chai颠覆了传统单代码库审计范式,为复杂密码学误用漏洞的自动化挖掘提供新思路。

Fortress and Gatekeeper: Theorizing Transitive Trust in Third-Party Cybersecurity Risk Governance

第一作者: Yijun Chen · 方向: 软件安全

Abstract:Third-party vendors, such as analytics platforms, cloud services, identity providers, and software suppliers, are increasingly embedded in digital service delivery. While these arrangements enable scale and specialization, they also move customer data and security-relevant practices into environments that customers rarely see, select, or evaluate. This paper examines this problem through a document analysis of the November 2025 OpenAI-Mixpanel security incident. The incident serves as an illustrative case for showing how a security event in a vendor environment can become a governance and accountability problem for the focal organization that maintains the customer relationship. Drawing on organizational trust research and agency theory, the paper argues that third-party cybersecurity risk is both a trust relationship and a delegation problem. Customers trust the visible...

论文介绍 针对第三方供应商嵌入数字服务带来的隐蔽安全风险,本文以OpenAI-Mixpanel安全事件为例,探讨供应商环境安全事件引发的治理与问责问题。研究结合组织信任与代理理论,提出第三方网络安全风险本质上是传递信任与委托问题,为理解焦点组织在数据外包环境下的安全治理困境提供理论框架。

SpikeTimer: Exploring Active Copyright Protection in Spiking Neural Networks via Temporal Backdoor Regularization

第一作者: Xiao Yang · 方向: AI 安全

Abstract:Spiking Neural Networks (SNN) have emerged as a revolutionary paradigm compared to traditional Deep Neural Networks (DNN) in energy-efficient computing, showcasing exceptional capabilities in processing event-driven sensory data for real-time applications like robotics and edge AI systems. However, unlike extensive studies on DNN copyright solutions, SNN copyright protection remains largely underexplored due to their inherent temporal coding complexities and spike-driven computation. In this study, we propose a novel active copyright protection framework named SpikeTimer for SNNs via temporal backdoor learning. SpikeTimer partitions neuromorphic data into designated timeslices and exclusively embeds authorized tokens within authorized slices. Furthermore, the inherent temporal segmentation characteristic intrinsically enables SpikeTimer to support multi-user authorization...

论文介绍 针对脉冲神经网络因时间编码复杂性导致版权保护研究不足的问题,本文提出基于时间后门学习的主动版权保护框架SpikeTimer。该方法将神经形态数据划分为特定时间片,仅在授权时间片内嵌入授权令牌,利用其时间分割特性实现多用户授权,为边缘计算等领域的SNN模型知识产权保护提供新方案。

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

第一作者: Inderjeet Singh · 方向: AI 安全

Abstract:Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image injection, direct-query attacks, and orchestrator-level tool manipulation. Existing red-teaming approaches are typically surface-specific and often recycle known attack templates; on text-poisoning benchmarks we measure 73-84% exact duplication. We present MIRROR, a unified cross-surface framework that performs memory-guided Monte Carlo tree search while conditioning candidate generation on retrieved context under an explicit novelty constraint. A deterministic Novelty Gate rejects any candidate matching the retrieval set under normalized comparison, allowing retrieval to inform search priors without enabling prompt copying. Across four attack surfaces on a multimodal agentic RAG target, MIRROR attains 76% ASR on image poisoning...

论文介绍 针对多模态代理检索增强生成系统攻击面扩大及现有红队方法易重复模板的问题,本文提出MIRROR跨表面红队框架。该方法在显式新颖性约束下执行记忆引导的蒙特卡洛树搜索,利用检索上下文调节候选生成,并通过确定性门控防止提示复制,有效提升了多模态RAG系统在复杂攻击面下的安全评估能力。

MergeLLL: A Hierarchical Divide-and-Conquer Framework for LLL-Based Lattice Reduction

第一作者: Niharika Gauraha · 方向: 密码学协议

Abstract:Lattice basis reduction algorithms have various applications in computational number theory and lattice-based cryptography, but their complexity increases rapidly with the dimension. Motivated by the divide-and-conquer strategy of merge sort and incorporating PotLLL-style deep insertions during recombination, MergeLLL is proposed. In this framework, a lattice basis is split into sub-bases, local reductions are performed independently, and the full basis is reconstructed through hierarchical merging. The approach is focused on improving local lattice structure first before global basis properties are refined, resulting in enhanced Gram-Schmidt orthogonality and numerical stability, while overall computational cost is reduced. The method is naturally parallelizable, allowing efficient multicore and distributed execution. It is shown that the reduction and merging steps preserve...

论文介绍 针对格基约化算法复杂度随维度快速增长的问题,本文提出基于分治策略的 MergeLLL 框架。该方法将格基分裂为子基进行独立局部约化,并结合深度插入策略进行分层合并。此方法优先改善局部格结构,提升了正交性与数值稳定性,降低了整体计算成本,且天然支持多核与分布式并行执行,适用于高维格密码分析。

DroidBreaker: Practical and Functional Problem-Space Attacks on Machine-Learning Android Malware Detectors

第一作者: Christian Scano · 方向: 软件安全

Adversarial APKs are Android applications modified in the problem space to evade machine-learning malware detectors. In this work, we first show that, despite claims, existing problem-space attacks remain largely impractical. Most techniques leverage software transplantation to inject entire benign modules, introducing many side-effect features and often causing build-time failures. Fine-grained methods that inject only a narrow subset of components exhibit limited effectiveness, while those that also use obfuscation rely on brittle bytecode rewriting, producing APKs that are syntactically valid but semantically unusable. Prior work further overestimates attack success rates by running smoke tests that only validate installation and basic execution, without assessing whether the modified APK still preserves its intended behavior. To overcome these limitations, we present DROIDBREAKER...

论文介绍 现有针对安卓恶意软件检测器的问题空间对抗攻击多存在构建失败或破坏应用原有功能等实用性问题。本文提出 DroidBreaker 框架,旨在生成既具备对抗性又保留完整功能的实用型对抗 APK。该方法克服了传统注入和字节码重写导致的语义失效缺陷,为评估和提升移动端恶意软件检测器的鲁棒性提供了有效的测试工具。

The Fungible Reserve Standard: A Deterministic Framework for Encoding Carrying Costs in Asset-Backed Tokens

第一作者: JJ Jia Jing Tan · 方向: 区块链安全

Abstract:The tokenization of real-world assets (RWAs) has emerged as a transformative application of blockchain technology, with market projections estimating trillions of dollars in tokenized assets within the coming decade. However, a fundamental challenge remains unaddressed: physical assets such as precious metals, stored commodities, and warehoused goods incur structural negative carry -- custody, insurance, and audit costs that accumulate over time. While existing tokenization models have successfully established the market for digital gold and treasuries, they typically manage operational costs at the issuer level. The FRS introduces a framework to bring these economics directly on-chain, avoiding mechanisms such as token rebasing that compromise fungibility and composability with decentralized finance (DeFi) protocols. This paper proposes the Fungible Reserve Standard (FRS), a...

论文介绍 现实世界资产代币化面临物理资产存储与保险等持续持有成本的问题。本文提出同质化储备标准(FRS),将持有成本经济学直接引入链上。该确定性框架避免了代币重定基等破坏同质化与去中心化金融可组合性的机制,为贵金属和仓储商品等产生负收益的资产提供了标准化的链上成本编码与解决方案。

TGHE: Template-based Graph Homomorphic Encryption for Privacy-Preserving GNN Inference in Edge-Cloud Systems

第一作者: Ngoc Bao Anh Le · 方向: 密码学协议

Abstract:Existing homomorphic encryption (HE)-based GNN systems adopt a graph-centric paradigm that couples per-query cost to global graph size, limiting evaluations to at most ~20k nodes and making them incompatible with dynamic, large-scale financial graphs. We propose TGHE (Template-based Graph Homomorphic Encryption), an ego-centric framework that resolves this by exploiting a template phenomenon: local computation trees in transaction graphs converge into a small set of structural shapes. TGHE canonicalizes ego-graphs at the edge and packs structurally identical trees into shared CKKS ciphertexts for SIMD-parallel encrypted inference, with two long-tail optimizers (Approximate Template Fitting and Topology Collapse) ensuring full SIMD coverage. On DGraphFin (3.7M nodes, 4.3M edges), TGHE-Collapse achieves a 66.9x speedup over the sequential encrypted baseline with less than 0.002...

论文介绍 针对现有同态加密图神经网络系统受限于全局图规模的问题,本文提出基于模板的图同态加密框架 TGHE。该方法利用交易图局部计算树的结构收敛特性,在边缘端将同构子图打包至共享密文中,实现单指令多数据流并行加密推理。TGHE 大幅提升了大规模动态金融图的隐私推理效率,适用于边缘云协同系统。

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents

第一作者: Nada Lahjouji · 方向: AI 安全

Abstract:Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and act on a user's behalf. As they move from answering questions to operating over sensitive data, privacy becomes harder to enforce. An agent touches many data sources, runs multi-step workflows, keeps state across sessions, and acts with delegated permissions. Sensitive information can therefore leak not only through its final answer but through the queries it issues, the intermediate results it handles, the memory it writes, and the messages it exchanges with other agents. We survey the privacy of LLM agents from a data-centric view, organizing the field around the data an agent touches rather than by attack type, and we use data agent as shorthand for an LLM agent that works with data. Research on these risks is active but scattered across...

论文介绍 随着大语言模型智能体日益深入地操作敏感数据,其多步工作流与跨会话记忆引发了复杂的隐私泄露风险。本文从数据中心视角对智能体隐私进行综述,围绕智能体接触的数据源、中间状态与交互消息来组织研究。该工作系统梳理了数据流转中的隐私威胁,为构建安全可控的数据智能体提供了理论参考与防范思路。

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models

第一作者: Abrar Alotaibi · 方向: AI 安全

Abstract:Adversarial evaluation of AI systems has matured along four largely disconnected tracks: diffusion-based attacks on text and large language models (LLMs), diffusion-based attacks on image classifiers, jailbreak pipelines against vision-language models, and diffusion-based input purification defenses. Each has developed its own vocabulary, threat models, and benchmarks, with denoising diffusion models emerging as a shared generative mechanism whose recipes are now actively ported between communities. This survey performs an information-fusion exercise at the meta-research level: we integrate these four tracks into a single conceptual framework with a unified taxonomy, evaluation criteria, and research agenda, focusing on the LLM-side slice. We catalog fifty published papers across four scope areas (text/LLM, image classifier, vision-language model, defense), plus four...

论文介绍 针对人工智能对抗评估中各模态研究相互孤立的问题,本文对基于扩散模型的跨模态攻击、防御与评估进行融合综述。研究将文本、图像及视觉语言模型等独立轨道整合为统一的概念框架与分类体系,重点聚焦大语言模型侧的扩散机制。该工作梳理了现有基准与威胁模型,为多模态大模型的鲁棒性研究提供了系统指导。

TESLA-for-5G: Broadcast Authentication for 5G Networks Using TESLA

第一作者: Subin Song · 方向: 密码学协议

Abstract:5G base stations broadcast unauthenticated system information (SI) that every user equipment (UE) reads during cell selection. This enables attackers to broadcast forged SI from a fake base station (FBS), deceiving UEs into camping on it. Prior approaches require UEs to authenticate System Information Block 1 (SIB1) using digital signatures. This necessitates computation-heavy verification for every SIB1 reception, imposing a significant burden on resource-constrained UEs. We propose TESLA-for-5G (TF5), a broadcast authentication protocol for 5G SIB1 that combines TESLA with GG09 Schnorr-like identity-based signatures (IBS). In the steady state, TF5 enables UEs to authenticate each SIB1 message using a symmetric MAC and delayed key disclosure, eliminating the need for per-message digital signatures. Initial trust is bootstrapped during cell entry using a lightweight GG09 IBS...

论文介绍 5G 基站广播的未认证系统信息易受伪基站攻击,现有数字签名方案给资源受限终端带来沉重计算负担。本文提出 TESLA-for-5G 协议,结合 TESLA 机制与轻量级身份签名。该协议在稳态下利用对称消息认证码与延迟密钥披露完成认证,消除了逐条消息的数字签名开销,有效兼顾了网络安全与终端能效。

VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills

第一作者: Ying Li · 方向: 安全研究

Agentic systems increasingly act through third-party skills, allowing model-generated decisions to affect files, communication channels, and cyber-physical devices. These skills often include natural-language specifications that define access permissions, disclosure limits, execution privileges, and required preconditions. Although such specifications describe the intended boundaries of skill behavior, they do not by themselves provide executable runtime enforcement. Enforcing them raises a contextual granularity challenge: even when a policy is written for a particular task context, a monitor must still decide which events to observe, what state to retain, how far across the execution to reason, and where to intervene. Choosing the wrong granularity can either block benign executions or miss violations that emerge only across multiple actions. Most existing enforcement mechanisms...

论文介绍 人工智能智能体日益依赖第三方技能执行任务,但其自然语言行为规约缺乏可执行的运行时强制机制。本文提出 VIGIL 框架,解决技能执行中的上下文粒度挑战。该系统通过动态决定事件观察、状态保留与干预位置,在运行时精确强制执行访问权限等安全边界,防止智能体在复杂多步操作中发生越权或违规行为。

DKVE: Decentralized Key Validation for End-to-End Encrypted Messaging

第一作者: Subin Song · 方向: 密码学协议

Abstract:End-to-end encrypted messaging systems depend on authentic public key distribution to prevent man-in-the-middle (MitM) attacks. Current solutions present a stark trade-off: out-of-band (OOB) verification provides strong security but lacks scalability for large contact lists, while key transparency (KT) systems enable automated verification at high storage costs and operational complexity. We propose DKVE, a protocol that validates public keys through privacy-preserving cross-validation within users' social graphs. When obtaining a contact's public key from a key server, clients query mutual contacts to verify they hold the same key, combining Oblivious Pseudorandom Functions (OPRF) and Oblivious Key-Value Stores (OKVS) to preserve privacy of both queries and contact lists. DKVE employs a Sequential Probability Ratio Test (SPRT) to aggregate responses and detect server...

论文介绍 针对端到端加密消息系统中公钥分发的中间人攻击风险,本文提出DKVE协议。该协议利用用户社交图进行隐私保护的公钥交叉验证,结合不经意伪随机函数与不经意键值存储技术,在验证密钥的同时保护查询与联系人列表隐私,为大规模联系人列表提供了一种兼顾安全与扩展性的去中心化验证方案。

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

第一作者: Praneeth Narisetty · 方向: AI 安全

Abstract:Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions. Systems such as CaMeL, FIDES, Progent, RTBAS, and FORGE realize this with capabilities, information-flow labels, and reference monitors, and several report near-elimination of attacks on the AgentDojo benchmark. We make two contributions. First, we organize these out-of-band defenses as instances of classical integrity protection (Biba), reference monitoring, and least privilege, yielding a structured comparison of what they do and do not cover. Second, we warn that every one of them is validated only on static benchmarks (a fixed set of injection attempts), the same methodology that made...

论文介绍 针对大语言模型代理面临的间接提示注入攻击,本文系统评估了现有的带外防御机制。研究将这些防御策略映射为经典完整性保护与引用监控模型进行结构化比较,并指出当前防御仅在静态基准上验证的局限性,揭示了固定注入测试集无法全面评估动态环境下代理安全性的方法论缺陷。

What Browsers Do in the Shaders: A Measurement Study of WebGPU Privacy

第一作者: Igor Santos-Grueiro · 方向: 网络安全

Abstract:WebGPU lets ordinary web pages run GPU workloads through a validated programming model. Validation protects memory safety, but shared browser, driver, OS, and GPU state can still expose privacy-relevant signals. We present WGPULens, a framework for measuring those signals across controlled scenarios, browser-native co-residency, a participant field study, public page loads, and mitigation policies. Our framework separates measurements: controlled scenarios support leakage, boundary, and mitigation claims; participant runs support deployment, compatibility, and fingerprintability; and a Tranco crawl measures WebGPU exposure in real-world pages. Our controlled results identify persistent pipeline compilation state as the clearest surface. Cold/warm pipeline probes reveal prior compilation state across selected origin, profile, and browser placements. Controlled browser/native...

论文介绍 针对WebGPU在网页中运行GPU负载时可能引发的隐私泄露问题,本文提出WGPULens测量框架。通过在受控场景、真实页面加载及参与者研究中分离测量维度,研究发现持久管道编译状态是主要的隐私泄漏面,揭示了跨源和跨配置的GPU状态可被用于用户指纹识别,为浏览器隐私缓解策略提供依据。

Lessons from the Adoption and Deprecation of the Privacy Sandbox Web APIs

第一作者: Yohan Beugin · 方向: 网络安全

Abstract:While several web actors have been trying to reduce web tracking for years, it remains unclear how to achieve both desirable levels of utility and privacy. In 2019, Google launched the Privacy Sandbox initiative to balance that trade-off and find privacy alternatives to common use cases such as advertising. Yet, in late 2025, Google canceled the project and deprecated most of the newly introduced APIs. Despite its end, the Privacy Sandbox represents a unique opportunity to learn about how the ecosystem reacted to the proposed changes and make observations about why and how it failed. In this paper, we present a longitudinal measurement and analysis study of the Privacy Sandbox APIs to characterize their adoption and deprecation over the past seven years by different web actors. Leveraging historical HTTP Archive crawls and public Chrome telemetry data, we offer the largest...

论文介绍 本文对Google Privacy Sandbox Web API的采用与弃用过程进行纵向测量分析。通过挖掘历史HTTP Archive数据与Chrome遥测信息,研究刻画了不同网络参与者对该隐私API的采纳与废弃轨迹,为理解Web生态在隐私与实用性间的权衡及大型隐私倡议的演进提供了实证参考。

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats

第一作者: Poojitha Thota · 方向: AI 安全

Abstract:Large language models (LLMs) are increasingly deployed in interactive applications, yet they remain vulnerable to adversarial interactions that induce harmful, deceptive, or policy-violating outputs. Existing defenses typically analyze either user prompts or generated outputs, but not both. However, many real-world attacks exploit a separation between adversarial intent expressed in the prompt and actionable harm manifested only in the response. As a result, prompt-only and response-only defenses frequently miss unsafe interactions that appear benign when viewed from either side in isolation. We present a verification-centric defense framework that jointly evaluates prompt intent and response harm before an LLM response is delivered to a user. The framework employs specialized analysts for intent and harm assessment together with a Judge for conflict resolution. We formalize a...

论文介绍 针对大语言模型在交互中易受对抗性攻击的问题,本文提出一种联合验证提示意图与响应危害的统一防御框架。该框架在模型输出交付前,利用专门的意图与危害分析师进行双重评估,并引入仲裁机制解决冲突,有效弥补了单一分析提示或输出时容易遗漏的隐蔽不安全交互,提升了模型交互安全性。

Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model

第一作者: Sergey Kurilenko · 方向: 密码学协议

Abstract:Dense embeddings power semantic search and retrieval-augmented generation, but embedding-inversion attacks can reconstruct source text from a vector: when a vector database leaks, the documents behind it leak too. The textbook defences are extremes - encrypting the whole search homomorphically is sound but too slow at million-document scale, while privacy noise degrades ranking long before it protects. We study a middle path exploiting the asymmetry between the static collection and the dynamic query. The collection is protected geometrically: each vector is truncated onto a lower-dimensional SVD subspace and rotated by a secret orthogonal transform known only to the owner. The query is protected cryptographically: it is reranked under CKKS homomorphic encryption, so an honest-but-curious server never sees the query or the scores. CKKS parameters come from a small offline...

论文介绍 针对语义搜索中向量数据库泄露导致的文本重建风险,本文提出一种混合隐私保护方案。该方法对静态文档集合采用SVD子空间截断与秘密正交变换进行几何保护,对动态查询则利用CKKS同态加密进行密文重排序。此方案在保护文档与查询隐私的同时,有效平衡了大规模检索场景下的安全性与计算性能。

Beyond Takedown: Measuring Malicious Go Module Persistence in the Wild

第一作者: Minjae Bae · 方向: 软件安全

Abstract:We measure an automation-based supply chain campaign in the Go ecosystem. The attackers repackage legitimate Go modules under attacker-controlled owners, and embed them with obfuscated code for an import-triggered downloader. Our results come from two complementary analyses: a) a manual search on GitHub across 2,113 repositories and b) a large-scale scan of 12.3M index entries using a deobfuscating AST scanner (GOAST) that we implemented. As a result, we identified 2,289 malicious versions of legitimate Go modules. We demonstrate that purely GitHub-centric searches fail to identify the full extent of the compromise and are only effective for as long as the affected code is present on the platform. Moreover, our proxy-based measurements of the takedown-remediation gap reveal that among artifacts later found to be GitHub-unobservable (i.e., removed or suspended), at least 99.4%...

论文介绍 本文针对Go生态系统中的自动化供应链攻击开展大规模测量研究。通过结合GitHub手动排查与自研的AST反混淆扫描工具,识别出两千余个恶意Go模块版本。研究揭示了仅依赖代码托管平台搜索的局限性,并量化了恶意模块在下架后的持久性残留问题,为完善开源生态的供应链安全治理提供了实证依据。

TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models

第一作者: William Aiken · 方向: 网络安全

Abstract:Noise-based backdoor attacks on diffusion models typically rely on input-time trigger injection, untargeted activation, and out-of-distribution target generation. Such assumptions reduce both the stealthiness and the practical relevance of these attacks. In this work, we present TEMPO-Diffusion, a targeted backdoor framework that localizes the malicious distribution shift to a temporal, in-distribution exposure. TEMPO-Diffusion supports: (i) targeted attacks on and to specific classes, (ii) multiple sub-image backdoors that reconstruct specific features within multiple, different output images and at multiple locations, and (iii) in-painting with time-conditioned triggers. To study relevant, practical security concerns in leveraging backdoored diffusion models for synthetic training data, we also introduce CALISA: a balanced, region-aware traffic-sign dataset emphasizing...

论文介绍 针对扩散模型后门攻击隐蔽性不足的问题,本文提出TEMPO-Diffusion定向后门框架。该框架将恶意分布偏移限制在时间维度的分布内暴露,支持针对特定类别的多位置子图像后门及时间条件触发修复。研究还引入了区域感知的交通标志数据集,深入探讨了利用投毒模型生成合成训练数据时的实际安全隐患。

Expecting (Targeted Ads)? Network Analysis of User Health Data Leakage in Fertility Tracking Apps

第一作者: Yeeun Jo · 方向: 网络安全

Abstract:While human factors in the privacy of fertility tracking apps -- health trackers that record user's menstrual or pregnancy data -- has been the subject of extensive study, little attention has been paid to the technical aspects of apps' data handling practices. We conduct a network-based measurement study of a corpus of 20 Android fertility tracking apps from the Google Play Store, focusing on how user data is shared with third party advertising services. After systematizing app features, we conduct a series of standardized user interactions across all apps in an environment that records TLS-stripped network traffic. In a subset of apps (n=5) we identify explicit leakage of user health data as well implicit leakage through highly targeted contextual advertising URL's. Equally importantly, we observe additional apps that use an ad-based monetization model without apparent...

论文介绍 本文针对生育追踪应用的用户健康数据泄露问题展开研究。作者对20款Android应用进行网络流量测量,分析其数据共享实践。研究发现部分应用存在显式的健康数据泄露,以及通过针对性广告URL导致的隐式泄露。该研究揭示了此类应用在广告变现模式下的隐私风险,为移动健康应用的数据安全评估提供了实证依据。

CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?

第一作者: Jintao Huang · 方向: 软件安全

Abstract:We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, and patch synthesis. Built from 541 real-world exploit incidents from DeFiHackLabs spanning 9 EVM chains, the benchmark provides end-to-end on-chain evaluation where agents interact with historical blockchain state through isolated evaluation environments orchestrated by Harbor, using tools to read code, trace transactions, and validate exploits on mainnet forks. Each case is anchored to a specific block and includes structured ground truth covering vulnerability type, localization, and attacker profit. Exploits are graded by economic impact on historical forks; patches are validated by replaying historical attacks and legitimate transactions as fail-to-pass test oracles on a proxy-upgradeable subset...

论文介绍 本文提出CyberChainBench基准,评估大语言模型智能体在智能合约安全方面的能力。该基准涵盖漏洞检测、利用生成和补丁合成任务,基于真实的去中心化金融漏洞事件构建。通过隔离环境让智能体与历史链上状态交互,验证其安全分析与修复能力,为AI辅助智能合约安全研究提供标准化的端到端评估框架。

Data Facts: A Metadata Schema for Structured Data Exchange in the NANDini Multi-Agent Ecosystem

第一作者: Jin Gao · 方向: 密码学协议

NANDini (Networked Agents Natural Distillation of Interconnected Nodal Intelligence) envisions an automated ecosystem where intelligent agents independently create, process, and exchange data to drive decisions at scale. Realizing this vision requires infrastructure beyond agent discovery and communication: agents must be able to advertise, evaluate, and verify the datasets they hold. Current protocols, including NANDA for federated registry and A2A and MCP for inter-agent messaging, address identity and communication but provide no mechanism for structured data exchange. Existing enterprise data-sharing frameworks, such as IDS-RAM, Gaia-X, and Ocean Protocol, assume human-in-the-loop governance that is incompatible with autonomous, real-time agent interactions. We introduce Data Facts, a core NANDini concept: a lightweight JSON metadata schema that bridges agent discovery and data...

论文介绍 针对多智能体系统缺乏结构化数据交换机制的问题,本文提出Data Facts概念。这是一种轻量级JSON元数据模式,旨在桥接智能体发现与数据交换。与传统依赖人工干预的企业框架不同,该模式支持智能体自主、实时地发布、评估和验证数据集,为大规模自动化智能体协作提供了底层数据基础设施支持。

MIRAGE: Protecting against Malicious Image Editing via False Moderation

第一作者: Anshul Nasery · 方向: 系统安全

The proliferation of AI-powered image editing systems raises serious concerns because it allows personal images to be arbitrarily manipulated at scale, with minimal effort, and a lower barrier to entry. Prior work on image immunization adds imperceptible perturbations to an image to protect against unauthorized manipulations. However, these methods usually require access to the model weights and the image manipulating prompt. This significantly limits their use, especially against powerful commercial image-editors such as GPT-Image, Gemini Flash Image (Nano Banana), and Grok Imagine. To address this, we take a system-level view of the problem and identify a previously unexplored attack surface common to all major commercial image editing systems: pre-generation safety moderation.Rather than disrupting the generative model itself, we propose to immunize images by causing these...

论文介绍 针对AI图像编辑系统导致的恶意篡改问题,本文提出MIRAGE方法。该方法利用商业图像编辑器普遍存在的预生成安全审核机制,通过触发虚假审核阻止未授权操纵。相比依赖模型权重的传统方法,该技术无需获取模型内部信息,能有效防御主流商业编辑系统的滥用,在系统层面提升了个人图像数据的安全性。

A Deterministic Control Plane for LLM Coding Agents

第一作者: Padmaraj Madatha · 方向: AI 安全

Abstract:LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitions, IDE-specific markdown -- is largely unmanaged. A prevalence study of 10,008 public GitHub repositories (n=6,145 agent config files) finds that agent configurations propagate as undeclared shared components: 10.1% of tracked paths are SHA-256 exact duplicates across independent repositories (fork-adjusted, threshold-independent), with 75.5% of clone pairs crossing organisational boundaries. Two further patterns are indicative: configurations are rarely revised (58% single-commit; 0.4 vs 0.6 commits/month age-normalised against CI/CD workflows), and rarely declare permission boundaries (<1% of agent configs vs 33% of Actions workflows, n=31 true positives). We propose a deterministic control plane above the harness that maps one-to-one to...

论文介绍 本文研究大语言模型编码智能体配置层缺乏管理的问题。分析发现智能体配置常作为未声明的共享组件跨组织传播,且极少更新或声明权限边界。为此,作者提出在智能体运行环境之上构建确定性控制平面,实现对配置文件的集中管理与权限控制,从而降低智能体在代码生成与执行过程中的安全及合规风险。

Autoformalization of Agent Instructions into Policy-as-Code

第一作者: Adam Mondl · 方向: 软件安全

Abstract:Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned classifiers, prompt-based steering) that offer no formal guarantees, or on hand-coded symbolic enforcement that does not scale to the breadth of real policy specifications. We present an autoformalization pipeline that translates agent prompts, MCP tool descriptions, and natural language policy documents into formally verified policies using an LLM-based generator-critic loop. The resulting policies are written in the Cedar Policy Language. On the MedAgentBench benchmark, our autoformalized policies cover substantially more of the source natural-language specification than the hand-coded symbolic enforcement in prior work.

论文介绍 针对智能体安全需形式化策略执行的问题,本文提出自动形式化管道。该方法利用大语言模型构建生成与评估循环,将智能体提示、工具描述及自然语言策略转化为形式化验证的策略代码。结果表明,生成的Cedar策略代码比手工规则更全面地覆盖自然语言规范,有效提升了智能体安全管控的可扩展性。

Empirical Software Engineering TerraProbe: A Layered-Oracle Framework for Detecting Deceptive Fixes in LLM-Assisted Terraform

第一作者: Manar Alsaid · 方向: 软件安全

Abstract:Security misconfigurations in Terraform Infrastructure-as-Code are a growing risk in cloud deployments, and large language models are increasingly used as automated repair agents. Existing evaluations often treat a repair as successful when the targeted static-analysis finding disappears, without checking planning validity, behavioral change, or security intent. This paper presents TerraProbe, a five-layer oracle framework for evaluating LLM-assisted Terraform security repair. We apply TerraProbe to 288 first-pass repairs generated by gemini-2.5-flash-lite, GPT-4o, and Claude 3.5 Sonnet across 68 real-world TerraDS modules and 28 controlled injected-defect modules. The results show that targeted Checkov removal overstates repair success. Although targeted removal reaches 83.3 percent for the primary model, full-scanner cleanliness drops to 10.4 percent, Terraform planning...

论文介绍 针对Terraform基础设施即代码的安全风险,本文提出TerraProbe五层框架,检测大语言模型辅助修复中的欺骗性修复。研究发现仅依赖静态分析的目标移除会高估修复成功率。该框架通过多层验证全面评估修复的有效性、行为变更与安全意图,为云环境下的自动化代码修复提供了更严谨的评估标准。

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers

第一作者: Mohammad Faizan · 方向: 系统安全

Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding source meaningfully shaped the output. This paper introduces ProvenAI, a framework that decomposes transparency in multi-hop question answering into three independently measurable layers: answer correctness, citation fidelity against benchmark supporting evidence, and per-document influence under leave-one-resource-out intervention. Targeting the HotpotQA distractor benchmark through a seven-stage pipeline covering data normalisation, retrieval indexing, citation-aware answer generation, attribution auditing, ablation-based influence estimation, batch evaluation, and interactive inspection, ProvenAI evaluates 7,405 validation examples drawn from a canonical corpus of 509,300 passages. The system achieves 53.53% answer accuracy alongside a mean...

论文介绍 针对检索增强系统中引用无法证实来源真实影响答案的问题,本文提出ProvenAI框架。该框架将多跳问答透明度分解为答案正确性、引用保真度和单文档影响力三个可测层。通过归因审计和消融干预流水线,系统能精确评估各文档对生成结果的实际贡献,有效提升了生成式问答的可解释性与证据溯源能力。

Nanoelectromechanical Systems (NEMS) for Hardware Security in Advanced Packaging

第一作者: Himanandhan Reddy Kottur · 方向: 系统安全

Abstract:As hardware security threats escalate across semiconductor manufacturing and advanced packaging, there is a growing need for novel physical mechanisms to counter sophisticated attacks such as tampering, counterfeiting, and supply chain infiltration. This paper presents Nanoelectromechanical Systems (NEMS) as an emerging class of hardware security primitives that enable physical assurance, tamper detection, and authentication at the device level. Leveraging mechanisms such as NEMS-based Physically Unclonable Functions (PUFs), shape memory materials, resonance-based fingerprints, and physical unlocking architectures, these systems offer enhanced resilience to reverse engineering, side-channel attacks, and environmental degradation. By harnessing mechanical unpredictability and fabrication-induced nanoscale variability, NEMS technologies introduce a physically robust and...

论文介绍 针对半导体先进封装中的硬件安全威胁,本文提出将纳米机电系统作为新兴硬件安全原语。该系统利用物理不可克隆函数、形状记忆材料及共振指纹等机制,在设备级实现物理保证与防篡改检测。该技术通过引入机械不可预测性与纳米级制造差异,有效提升了对逆向工程和侧信道攻击的防御能力。

Query Cost Model Calibration in Confidential Virtual Machines

第一作者: Qihan Zhang · 方向: 软件安全

With the growing adoption of Confidential Computing, running databases in confidential virtual machines (CVMs) such as AMD SEV-SNP has become an attractive way to protect sensitive cloud data with minimal changes to legacy DBMSs. However, analytical queries in such CVMs often suffer substantial overhead, and prior database work has largely stopped at benchmarking these slowdowns rather than optimizing them. We show that this problem stems from a hardware-software mismatch: query optimizers still rely on KVM-oriented (non-encrypted VM) cost assumptions that no longer hold in CVMs. To address this, we propose a lightweight CVM-aware cost calibration. It models two dominant sources of optimizer-facing overhead: data movement and RMP-related translation using simple physical proxies already available to the optimizer. Experiments show that the calibration significantly narrows the KVM/CVM...

论文介绍 针对机密虚拟机中数据库查询开销过大的问题,本文指出其根源在于查询优化器沿用了非加密虚拟机的成本假设。为此,提出一种轻量级成本校准方法,利用物理代理对数据移动和内存转换开销进行建模。该方法有效修正了成本假设偏差,降低了机密计算环境下的查询延迟,提升了数据库系统的执行效率。

Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems

第一作者: Jakob Salfeld-Nebgen · 方向: 软件安全

Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment. This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning but by requiring independently attested evidence at the point of consequential action. We formalise this institutional pattern as a computational governance model for AI agent systems. Under the proposed model, an agent retains full autonomy over planning and reasoning but holds no execution authority over designated high-risk actions. Execution is conditional on preconditions that are each independently attested by a separate authoritative source, cryptographically bound to a declared intent, and evaluated by a deterministic policy. Decisions are recorded in a tamper-evident log amenable to independent re-verification. We...

论文介绍 针对自主AI执行高风险操作的安全问题,本文提出基于「制度证明」的计算治理模型。该模型允许AI保留推理自主权,但将其高风险操作的执行权与独立权威来源的密码学证明绑定。通过确定性策略评估前置条件,并将决策记录于防篡改日志以供验证,在不干预AI推理的前提下实现了安全可控的行动治理。

The Role of Input Dimensionality in the Emergence and Targeted Control of Adversarial Examples

第一作者: Nasrin Malekzadeh Goradel · 方向: AI 安全

Abstract:Several theoretical works have tried to explain the adversarial vulnerability of deep neural networks through properties of high-dimensional geometry. However, the assumptions underlying these works are rarely examined empirically, and systematic evidence remains limited. In this work, we present a systematic study of the role of input dimensionality in both the emergence and the targeted control of adversarial examples. We first analyse the scope and limitations of existing theoretical frameworks based on concentration of measure, showing that real image classes exhibit strong empirical localization, beyond what such theories typically assume. We then conduct an extensive empirical evaluation across hierarchical image datasets spanning a wide range of input dimensionalities and diverse neural architectures. Our results consistently show that adversarial examples become easier...

论文介绍 本文系统研究输入维度在深度神经网络对抗样本产生与定向控制中的作用。通过分析基于测度集中的理论框架,发现真实图像表现出强烈的经验局部化。研究在跨越多维度输入和多种架构的图像数据集上进行广泛评估,揭示了输入维度变化对对抗样本生成难度及模型脆弱性的影响,为理解对抗鲁棒性提供了实证依据。

Federated Hash Projected Latent Factor Learning

第一作者: Jialan He · 方向: 软件安全

Hash Learning (HL) is an efficient representation learning approach that maps real-valued data into compact binary representations. Traditional HL methods typically require users to upload personal data to a central server, which is incompatible with increasingly stringent data security regulations. Federated Learning (FL) provides a decentralized paradigm for learning globally optimal models without centralizing private data. However, most FL methods rely on transmitting large-scale real-valued gradient information, leading to high communication overhead and potential privacy risks. Integrating HL into FL is a promising solution. Nevertheless, existing HL methods suffer from limited representational capacity of binary codes, which may degrade model accuracy. To address this challenge, we propose a Federated Hash Projected Latent Factor (FHPLF) model. FHPLF introduces three key...

论文介绍 针对传统哈希学习需集中数据及联邦学习中梯度传输开销大的问题,本文提出联邦哈希投影潜在因子模型。该方法将哈希学习引入联邦范式,在避免集中私有数据的同时降低通信开销。通过引入关键机制解决现有二进制码表示能力受限导致精度下降的挑战,实现了高效且隐私安全的分布式表示学习。

Account-History Features for Social Bot Detection in the Era of Large Language Models

第一作者: Gaurang Katyal · 方向: 安全研究

Abstract:Bot detection on social platforms has historically relied on a mix of account-metadata features and features extracted from the text of posts and profile fields. The arrival of capable language models complicates the latter. A bot operator can run every post through GPT-4 or Claude and produce text whose surface statistics are difficult to distinguish from those of human writing, which weakens the predictive value of content-derived features. This paper asks how much of the detection problem can be solved by features that an attacker cannot easily manipulate at low cost: the age of the account, follower and friend counts and their ratios, profile completeness, and the structural properties of the handle. On a publicly redistributed corpus of 2,432 Twitter accounts with manually verified labels (43.0% bots), a random forest using only these account-history features achieves...

论文介绍 大语言模型使社交机器人能轻易生成逼真文本,导致传统内容检测特征失效。本文提出利用攻击者难以低成本操纵的账户历史特征进行检测,包括账户年龄、粉丝与关注比例、资料完整度及句柄结构属性。研究表明,仅依赖这些历史特征的模型即可实现高效的机器人识别,为平台安全治理提供了低成本方案。

PAMAE: Phase-Aware-MoE Action Experts Towards Reliable Flow-Matching Vision-Language-Action Policies

第一作者: Jiayu Yang · 方向: VLA 通用模型 · 来源: cs.RO

Reliable action generation for multi-stage robotic manipulation remains challenging for Vision-Language-Action (VLA) models. While existing flow-matching VLA policies offer strong multimodal grounding and generalization, they typically employ a single shared action expert, limiting their ability to capture phase-specific control patterns across distinct execution stages. We propose a plug-and-play Phase-Aware Mixture-of-Experts Action Module (PAMAE), as a step towards more reliable phase-consistent action generation. PAMAE replaces the original flow-matching action expert with a sparse expert mixture while preserving the pretrained VLA backbone. PAMAE introduces a phase-aware router that leverages execution-phase cues to allocate action generation across experts, supported by a lightweight phase prediction head and a routing alignment objective. To stabilize specialization, we adopt a...

论文介绍 针对多阶段机器人操作中视觉语言动作模型动作生成不可靠的问题,本文提出阶段感知混合专家动作模块。该模块采用稀疏专家架构替换单一共享动作专家,引入阶段感知路由器,利用执行阶段线索将动作生成动态分配给特定专家。结合轻量级阶段预测头,有效提升了多阶段任务中动作生成的可靠性与阶段一致性。

RelAfford6D: Relational 6D Affordance Graphs for Constraint-Driven Robotic Manipulation

第一作者: Guodong Zhang · 方向: 机器人操作 · 来源: cs.RO

Bridging abstract semantics and precise physical control remains a fundamental challenge in open-world robotic manipulation. While recent data-driven policies show promise, their reliance on isolated contact points or latent affordance embeddings lacks the rigorous kinematic constraints necessary for complex articulated objects.To overcome the limitation, we introduce RelAfford6D, a novel training-free framework centered on a Relational 6D Affordance Graph. Given a free-form instruction, our system deduces a semantic topology linking a primary interacting part to its physical anchor. By elevating these topological nodes into precise metric $SE(3)$ poses via vision foundation models, we analytically formulate downstream execution as a kinematic constraint satisfaction problem. The robot synthesizes continuous trajectories by tracking strictly defined physical manifolds (e.g., revolute...

论文介绍 针对机器人操作中抽象语义与精确控制难以结合的问题,本文提出关系6D可供性图框架。系统根据指令推断交互部件与物理锚点间的语义拓扑,并利用视觉基础模型将其转化为精确6D位姿。通过将执行过程构建为运动学约束满足问题,机器人能沿严格定义的物理流形合成连续轨迹,实现复杂物体的可靠操作。

RobOralScan: Learning Active Intraoral Scanning for Robotic Dental Reconstruction

第一作者: Jinhyung Lee · 方向: 策略学习 · 来源: cs.RO

Intraoral scanning is widely used for digital optical impressions in prosthodontic, implant, and orthodontic treatment, but full-arch and long-span scanning remain labor-intensive tasks with limited automation. In the confined oral cavity, operators must continuously adjust scanner motion while accumulating narrow field-of-view observations, making reconstruction quality sensitive to missing tooth surfaces and operator workload. We propose RobOralScan, which, to the best of our knowledge, is the first reinforcement learning (RL)-based pipeline for robotic automatic intraoral scanning. RobOralScan introduces a geometric memory-based observation space that accumulates partial scan observations into a tri-state geometric representation, allowing the policy to reason over scan history and insufficiently observed regions. It further introduces tooth-wise coverage learning, combining...

论文介绍 针对全牙弓口内扫描自动化程度低的问题,本文提出RobOralScan,首个基于强化学习的机器人自动口内扫描系统。该方法引入基于几何记忆的观测空间,将部分扫描累积为三态几何表示,并结合逐牙覆盖率学习,使策略能推理扫描历史与未充分观测区域,有望提升口腔数字印模的效率与重建质量。

PlanRL: A Trajectory Planning Architecture for Reinforcement Learning-based Driving Experts

第一作者: Joonhee Lim · 方向: 导航与运动 · 来源: cs.RO

Reinforcement learning (RL) has become a prominent framework for developing driving experts in autonomous vehicles. However, most existing RL-based experts are designed to output direct control commands (e.g., throttle, steering), which suffer from a lack of interpretability, high spatial complexity in learning road geometries, and poor compatibility with modern end-to-end planning architectures. To address these limitations, we propose a novel trajectory planning architecture for RL driving experts that integrates an RL policy with a polynomial-based trajectory planner. By employing a Frenet-frame coordinate system, our method simplifies complex road geometries into a curvilinear framework, offering a structured coordinate prior that facilitates policy learning. Furthermore, we incorporate a kinematic feasibility check into the planning stage to ensure that generated trajectories...

论文介绍 针对现有基于强化学习的自动驾驶专家缺乏可解释性及难以处理复杂道路几何的问题,本文提出PlanRL轨迹规划架构。该方法将强化学习策略与基于多项式的轨迹规划器结合,采用Frenet坐标系简化道路几何以提供结构化先验,并引入运动学可行性检查,从而生成安全且易于解释的驾驶轨迹。

SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation

第一作者: Kaijun Wang · 方向: 机器人操作 · 来源: cs.RO

Real-world robotic manipulation demands spatial grounding, task-aware reasoning, and precise control. Learning such capabilities becomes particularly challenging in the low-data regime. Prior methods often trade off scalable task-level reasoning and explicit physical structure: video-based approaches can drift geometrically over long horizons, 3D approaches often require depth sensing, and many flow/trajectory interfaces emphasize motion without an explicit RGB-only geometric representation. We introduce SSI-Policy, a modular framework built around a Structured Scene Interface (SSI) -- a unified, RGB-only intermediate representation that jointly encodes monocular depth features, language-grounded object layouts, and instruction-conditioned 2D motion trajectories. Critically, SSI is robot-agnostic and trainable from action-free video, decoupling perception from control so that the...

论文介绍 针对低数据量下机器人操作难以兼顾空间定位与任务推理的问题,本文提出SSI-Policy框架。其核心是结构化场景接口,仅依赖RGB联合编码单目深度特征、语言接地的物体布局与2D运动轨迹。该接口与硬件解耦,支持从无动作视频训练,有效提升了视觉语言策略的泛化与控制能力。

A Closed-Form 4-DoF Inter-Robot Pose Estimator using Bearing-only Measurements

第一作者: Qixin De · 方向: 具身智能 · 来源: cs.RO

Bearing-odometry-based cooperative localization has attracted increasing research interest due to its minimal infrastructure requirements, low communication bandwidth and broad applicability in complex environments. However, existing 6-DoF approaches still face challenges in rapidly obtaining accurate and reliable inter-robot pose estimation, as the system is prone to observability degeneracy under specific motion patterns. To address these issues, we first propose a closed-form 4-DoF inter-robot pose estimator, which relaxes nonlinear constraints for rotations estimation and employs error projection for translations estimation. We then conduct a theoretical analysis of the system's observability, identifying degeneracy under two typical motion patterns: collinear and shape-preserving formations. The analysis further shows that the proposed 4-DoF system requires less stringent motion...

论文介绍 针对基于方位角与里程计的多机器人协同定位难题,本文提出一种闭环4自由度机器人间位姿估计器。该方法放松旋转估计的非线性约束并采用误差投影估计平移,同时从理论上分析了系统在特定运动模式下的可观测性退化问题,证明了4自由度系统对运动约束的要求更为宽松,适用于复杂环境下的多机协同。

A System for Fast, Resilient, and Adaptable Loco-Manipulation Behaviors on Humanoid Robots

第一作者: Duncan William Calvert · 方向: 机器人操作 · 来源: cs.RO

Humanoid robots could take on physically demanding, hazardous, and repetitive work in spaces built for humans. However, a useful robot for these spaces must coordinate locomotion, whole body motion, perception, contact, and operator supervision. This thesis presents a robot-local, runtime-editable behavior authoring and runtime system. Our system strives to be maximally observable, predictable, and directable following Coactive Design principles developed during the DARPA Robotics Challenge. Our operator interface remains continuously synchronized to the robot for runtime authoring, monitoring, and repair. Our behavior architecture uniquely combines object-centric Affordance Templates, organization and logic inspired by Behavior Trees, and runtime-editable perception through a behavior scene and primitive scene actions. Action primitives build on a whole-body controller that supports...

论文介绍 针对人形机器人在人类空间中执行复杂任务时的多模块协调难题,本文提出一种快速、鲁棒且可适应的移动操作系统。该系统结合以对象为中心的affordance模板、行为树逻辑以及运行时可编辑的感知场景,支持操作员在运行时进行行为创作、监控与修复,提升了人形机器人在物理任务中的可控性与适应性。

LiMoDE: Rethinking Lifelong Robot Manipulation from a Mixture-of-Dynamic-Experts Perspective

第一作者: Zhihao Gu · 方向: VLA 通用模型 · 来源: cs.RO

Building a generalist robot that can leverage prior knowledge for continuous task adaptation remains a significant challenge. Previous works alleviate the catastrophic forgetting problem by parameter-efficient fine-tuning for single-task adaptation. However, they fail to extract reusable skills and model the interaction with other skills effectively. Recent works try to address these issues by learning prompts. Differently, this paper presents an architectural perspective on the Lifelong Mixture of Dynamic Experts (\textit{LiMoDE}), a novel two-stage learning scheme for lifelong robot manipulation. Specifically, a dynamic MoE structure is first proposed in the multi-task pre-training stage to learn prior knowledge, where a varied number of heterogeneous experts are activated based on the motion information to address different short-term manipulations. Subsequently, in the task...

论文介绍 针对通用机器人在终身学习中面临的灾难性遗忘与技能复用难题,本文提出LiMoDE架构。该方法采用两阶段学习,在多任务预训练阶段引入动态混合专家结构,根据运动信息激活异构专家以学习先验知识;在任务适应阶段,有效提取可复用技能并建模技能间交互,显著提升了机器人连续任务适应与终身操作能力。

Scalable Behavior Cloning with Open Data, Training, and Evaluation

第一作者: Arthur Allshire · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:We introduce ABC, a fully open-source stack for manipulation with behavior cloning. At its core is ABC-130K: the largest open-source teleoperation dataset to date, featuring 3,500 hours of data spanning over 130K episodes across 195 diverse tasks. Furthermore, we open-source our accessible hardware setup, training infrastructure, and simulation pipeline. We also release 400 hours of sim-teleop data and provide a co-training recipe that produces correlated simulation and real-world evaluation, offering a reliable proxy for ablating model-design and training decisions before costly real-world evaluation. We explore various training recipes and compare common architectural choices for Diffusion Transformers (DiT) and Vision-Language-Action (VLA) models, grounding our findings in real-world evaluations. The resulting policies successfully execute dexterous tasks such as box...

论文介绍 本文推出ABC开源机器人操作行为克隆栈,核心包含最大开源遥操作数据集ABC-130K,涵盖3500小时、195项任务的13万条轨迹。团队同时开源硬件配置、训练设施与仿真流水线,并提供虚实数据协同训练方案,为扩散Transformer与视觉语言动作模型的架构设计及训练决策提供了可靠评估基准。

World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

第一作者: Manish Kumar Govind · 方向: 机器人操作 · 来源: cs.RO

Abstract:Going beyond predicting robot actions, World Action Models (WAMs) can also generate future visual observations. We build on this generative capability to propose Recurrent Generative Replay (REGEN), a continual imitation learning framework that synthesizes pseudo-replay trajectories, enabling a robot policy to rehearse previously learned tasks without storing their original human demonstrations. During continual adaptation, REGEN recursively queries the WAM to synthesize pseudo-replay trajectories conditioned only on prior task instructions and current-task observations. Experiments in both simulation and real-world manipulation settings show that REGEN reduces catastrophic forgetting by up to $50\%$ relative to sequential fine-tuning, while approaching the performance of privileged experience replay methods that require access to real replay data. Finally, we analyze the...

论文介绍 针对机器人在持续模仿学习中易产生灾难性遗忘的问题,本文提出REGEN框架。该方法利用世界动作模型生成未来视觉观测的能力,通过递归查询合成伪重放轨迹,使策略能在不存储原始演示的情况下复习已学任务。实验表明,REGEN在仿真和真实操作中显著降低了遗忘率,性能逼近依赖真实重放数据的方法。

RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection

第一作者: Xingyu Ren · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:We study whether pre-deployment evaluation rollouts can be reused to supervise policy selection. Robot teams routinely smoke test candidate vision-language-action (VLA) policies, then compress those trials into a global winner. RouterVLA evaluates this idea with outcome-disjoint cross-fitting: recorded probes build a profile for each frozen expert, and a separate trial scores the selected expert without entering its profile. Across 34,752 LIBERO-Plus rollout records, a transparent probe-success rule raises held-out success from 0.4686 to 0.6149, a +14.64pp gain. Under the scalar-only profiles studied here, learned scorers are statistically indistinguishable from this rule, showing that commissioning carries the routing value while extra scalar scorer capacity does not create it. Reusing the scored trial inflates the measured gain by $1.87\times$, so credible ledger routing...

论文介绍 本文研究如何复用预部署评估测试来监督视觉语言动作模型策略选择。提出RouterVLA框架,采用结果不相交的交叉拟合方法,利用记录的探测数据为冻结专家构建画像,并通过独立试验对选定专家进行评分。该方法有效提升了异构VLA策略在机器人部署中的选择准确率与成功率,为多模型路由提供了可靠方案。

Continual Robot Policy Learning via Variational Neural Dynamics

第一作者: Jiaxu Xing · 方向: 策略学习 · 来源: cs.RO

Abstract:Robots deployed in the real world rarely operate under a single fixed dynamics model: wind changes, payloads vary, batteries drain, contacts shift, and hardware wears. Yet most learning-based controllers are trained once and deployed as if learning were complete. This prevents the robot from using deployment experience to further improve task performance. In this work, we propose a continual learning framework that uses real-world experience to improve robot policies under hidden and recurring dynamics. Our method learns a condition-aware dynamics model from real state-action trajectories by combining an analytical physics prior with a neural residual for unmodeled effects. A recurrent encoder infers the current hidden condition from recent interaction, and this estimate conditions both the residual model and the policy. Policy learning is performed via differentiable...

论文介绍 针对机器人在真实环境中面临动力学变化导致策略失效的问题,本文提出一种持续学习框架。该方法结合分析物理先验与神经残差学习条件感知动力学模型,并通过循环编码器从近期交互中推断隐藏状态条件,从而利用部署经验持续优化机器人策略,提升其在动态和未知环境下的任务表现与适应能力。

Bridging Performance and Generalization in Reinforcement Learning for Agile Flight

第一作者: Jonathan Green · 方向: 策略学习 · 来源: cs.RO

Abstract:Autonomous drone racing is a fundamentally challenging regime for autonomous aerial robots, requiring time-optimal control while operating under persistent actuation saturation. While reinforcement learning (RL) has achieved human-level performance in this domain, current methods fail to generalize; policies trained on specific environments often crash immediately in unseen configurations. This failure reflects the intrinsic difficulty of zero-shot generalization in agile flight, arising from high-dimensional task variation and the tight coupling between safety and performance at high speeds. Existing approaches that improve generalization impose a substantial cost on flight speed: control policies must significantly degrade performance to achieve even modest levels of generalization. In this work, we propose a framework for zero-shot generalization in agile flight for...

论文介绍 针对无人机竞速中强化学习策略难以泛化且现有方法牺牲飞行速度的问题,本文提出一种敏捷飞行零样本泛化框架。该方法旨在解决高维任务变化与高速下安全性能紧耦合的难题,在保持时间最优控制的同时,实现策略在未见环境下的零样本泛化,推动自主飞行器在复杂场景中的实际应用。

VibeAct: Vibration to Actions for Contact-Rich Reactive Robot Dexterity

第一作者: Yuemin Mao · 方向: 机器人操作 · 来源: cs.RO

Abstract:Dexterous manipulation depends on contact events that are fast, local, and often visually occluded. Piezoelectric microphones offer a compact and high-bandwidth way to sense these interactions, but the resulting vibro-acoustic signals are difficult to simulate faithfully enough for end-to-end sim-to-real policy learning on dexterous robot hands. We propose VibeAct, a framework that bridges real vibrotactile sensing and simulation-based reinforcement learning through a shared physical representation of contact and slip. In the real world, we embed piezoelectric microphones into a dexterous robot hand and collect vibro-acoustic data through teleoperation, then replay the recordings in a calibrated digital clone to automatically label per-finger contact and slip. A tactile estimator learns to predict contact and slip from real microphone waveforms, while manipulation policies are...

论文介绍 针对灵巧操作中接触事件难以精确模拟的问题,本文提出VibeAct框架。该方法通过在灵巧手机器人手中嵌入压电麦克风收集振动声学数据,利用共享接触与滑动物理表示的校准数字克隆进行自动标注,桥接真实振动触觉感知与基于模拟的强化学习,有效提升了机器人在丰富接触场景下的反应式灵巧操作能力。

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

第一作者: Tao Lin · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visual-action supervision can dominate the comparatively sparse language-action signal. As a result, policies may rely on visual shortcuts rather than learn how language conditions action execution, making them sensitive to visual variations. To address this limitation, we propose LA4VLA, a language-action pretraining framework that enables policies to acquire language-conditioned action priors without visual observations. These priors capture reusable manipulation skills shared across tasks and scenes, reducing reliance on scene-specific visual cues. Specifically, LA4VLA decomposes expert demonstration trajectories into atomic action segments and pairs each segment with a corresponding low-level action...

论文介绍 针对视觉语言动作模型易受视觉捷径影响而忽视语言条件的问题,本文提出LA4VLA语言动作预训练框架。该方法将专家演示轨迹分解为原子动作段,使策略在无视觉观察的情况下获取语言条件动作先验。这些先验捕获跨任务共享的可复用操作技能,显著降低模型对特定场景视觉线索的依赖,提升策略的泛化能力。

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

第一作者: Wen Ye · 方向: 机器人操作 · 来源: cs.RO

Abstract:Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major challenges remain unsolved: (1) reasoning can effectively improve the performance of the policy, but its scaling mechanism has seldom been studied; (2) historical information is essential, as embodied tasks are inherently long-horizon and sequential, making sole reliance on current observations for action scaling inadequate due to the lack of historical context utilization. To address these challenges, we introduce E-TTS, a modular and plug-and-play Embodied Test-Time Scaling framework that unifies reasoning and action scaling for robotic manipulation via history-aware iterative refinement with vision-language verifiers. To support joint reasoning-action scaling, E-TTS performs reasoning-action joint sampling and scoring in a pairwise manner. To better utilize...

论文介绍 针对具身任务中推理缩放机制不明确及长视野任务缺乏历史上下文利用的问题,本文提出E-TTS框架。该模块化即插即用系统通过历史感知迭代细化与视觉语言验证器,统一机器人操作中的推理与动作缩放。其采用成对方式进行推理与动作联合采样和评分,有效利用历史信息提升长序列任务的操作性能。

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

第一作者: Junhao Shi · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation. Existing systems treat these as separate problems: VLM-based planners lack a unified cyber-physical action space, agent frameworks accumulate unbounded context that degrades temporal coherence, and VLA policies execute open-loop without detecting their own failures. We argue that persistent autonomy requires not a monolithic model but a hierarchical asynchronous architecture with explicit separation of planning, memory, and verification. To this end, we present OmniAct, a framework integrating a multimodal semantic planner for skill routing across unified action spaces, an...

论文介绍 针对非结构化环境中持久具身智能体缺乏统一网络物理动作空间及故障恢复能力的问题,本文提出OmniAct框架。该分层异步架构显式分离规划、记忆与验证模块,集成多模态语义规划器实现跨统一动作空间的技能路由,并结合记忆与验证机制,推动具身智能体从孤立技能向日常物理自主操作的演进。

HumanoidUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation

第一作者: Hongwu Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:High-quality demonstration data are essential for humanoid robot skill learning, especially for whole-body behaviors that require coordinated perception, locomotion, and manipulation. Existing data-collection methods largely rely on robot teleoperation, which is constrained by hardware accessibility, operator expertise, and limited efficiency. Inspired by the Universal Manipulation Interface (UMI), we propose HumanoidUMI, a portable and robot-free framework for humanoid whole-body data collection. HumanoidUMI uses lightweight VR devices and UMI-inspired grippers to collect sparse human keypoint trajectories, wrist-view observations, and gripper actions. These demonstrations train a high-level policy to predict future keypoints, which are retargeted to robot-native whole-body references and executed by a whole-body controller. Experiments in five real-world scenarios...

论文介绍 针对人形机器人全身操作数据收集受限于硬件和遥操作效率的问题,本文提出HumanoidUMI无机器人便携式采集框架。该方法利用轻量VR设备与仿生夹爪收集人体关键点轨迹及动作数据,训练高层策略预测关键点并重定向至机器人全身参考,由控制器执行,大幅提升了人形机器人技能学习的数据获取效率。

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

第一作者: Ilia Larchenko · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progress, and a few task-relevant future quantities, and those predictions drive advantage estimation, live failure detection, and candidate selection. The work mostly recombines existing RL ideas with engineering and optimization contributions that can be used together as one recipe or individually: AWR + RECAP combined for flow-matching VLA; an asynchronous distributed training / rollout pipeline through HuggingFace Hub; inference-time hyperparameters optimization via Thompson sampling; a...

论文介绍 本文介绍LeHome Challenge 2026双臂衣物折叠竞赛的获奖方案。研究针对复杂衣物折叠任务,提出一种结合视觉语言动作模型与强化学习的系统。该策略网络同时预测动作与任务进度,用于优势估计和失败检测。系统还集成了异步分布式训练与推理时超参数优化等工程优化,在仿真和真实环境中均取得优异成绩,为长序列精细操作提供了有效范式。

PhysReflect-VLA: Physical Feasibility and Self-Reflective Regulation for Reliable Vision-Language-Action Policies

第一作者: Jiayu Yang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Long-horizon robotic manipulation is highly sensitive to physically infeasible transitions, contact-induced disturbances, and the lack of effective self-correction during execution. Although Vision-Language-Action (VLA) models provide strong task grounding through multimodal learning, they typically generate actions in a feed-forward manner without explicitly checking physical feasibility or diagnosing execution errors online. We present PhysReflect-VLA, a plug-and-play execution-time reliability framework that augments VLA policies with physical feasibility evaluation and structured self-reflection in a closed-loop control pipeline. A Feasibility Operator evaluates whether candidate actions induce dynamically consistent state transitions; an Action Explanation Operator verifies transition coherence; and an LLM-based Reflection Module analyzes state discrepancies to generate...

论文介绍 针对视觉语言动作模型在长序列操作中缺乏物理可行性检查与自我纠错能力的问题,本文提出PhysReflect-VLA框架。该即插即用模块通过物理可行性算子评估状态转换,利用动作解释算子验证连贯性,并结合大语言模型反思模块分析状态偏差以生成修正指令,在闭环控制中显著提升了复杂机器人操作任务的执行可靠性。

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

第一作者: Mingyang Lyu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward general-purpose task capability, yet their embodied safety limits remain poorly understood. To address this gap, we introduce ForesightSafety-VLA, a diagnostic benchmark that makes safety the primary evaluation target for VLA systems. We define a 13-category safety taxonomy covering physical interaction safety (Safe-Core), instruction-side safety (Safe-Lang), and perception-side safety (Safe-Vis), and evaluate policies under three controlled dimensions of variation -- scene structure, language command, and visual observation -- so that failure sources can be diagnosed rather than hidden in a single aggregate score. Beyond binary task success, ForesightSafety-VLA measures process-level risk through...

论文介绍 针对视觉语言动作模型在具身部署中的安全评估不足,本文提出ForesightSafety-VLA诊断基准。该基准将安全作为首要目标,构建涵盖物理交互、指令侧和感知侧的13类安全分类体系。通过在场景、指令和视觉三个维度进行受控评估,系统能精准定位失败来源并测量过程级风险,为模型的安全优化提供了统一标准。

UAV-MapFusion: RTK-Aligned Uncertainty-Aware Coarse-to-Fine Multi-Session UAV Mapping

第一作者: Feng Pan · 方向: 具身智能 · 来源: cs.RO

Abstract:Large-scale point cloud maps are essential for robotics and spatial intelligence tasks. UAVs provide an efficient means for large-scale map acquisition; however, due to limited flight endurance and onboard storage, mapping a large-scale scene within a single flight remains difficult. Existing multi-session map merging methods can extend the mapping range, yet in UAV scenarios they still struggle to simultaneously suppress long-range drift and preserve local geometric accuracy. To address this issue, an uncertainty-aware multi-session point cloud map merging and coarse-to-fine optimization system is proposed. The proposed method first performs initial multi-session map merging based on a scene graph, and then incorporates RTK observations through an RTK spatiotemporal alignment module, where temporal offsets are estimated using Dynamic Time Warping (DTW), and continuous RTK...

论文介绍 针对无人机单次建图受限及现有多会话融合难以兼顾长程漂移抑制与局部精度的问题,本文提出UAV-MapFusion系统。该方法先基于场景图进行点云初始合并,随后引入RTK观测,利用动态时间规整估计时空偏移,实现不确定性感知的粗到细优化。该方案有效提升了大规模场景下无人机三维地图构建的精度与鲁棒性。

Ordinal Neural Collapse as a Representation Prior for Visual Navigation

第一作者: E-In Son · 方向: 导航与运动 · 来源: cs.RO

Abstract:Learning robust navigation policies directly from visual observations remains a fundamental challenge in vision-based robotic navigation. In end-to-end imitation learning approaches, the visual encoder and action decoder are jointly optimized using a single action loss, which provides only an indirect supervisory signal to the encoder. This indirect supervision frequently results in the encoder learning ambiguous, action-agnostic representations. The problem is further complicated by substantial variations in scene structure and appearance across diverse environments, as well as the prevalence of visual distractors inherent to real-world navigation settings. Such action-agnostic features cause the navigation policy to produce inconsistent actions at ambiguous decision points, leading to navigation failure. To overcome these limitations, we propose ORION (Ordinal Neural...

论文介绍 针对端到端视觉导航中视觉编码器因缺乏直接监督而学习模糊表征的问题,本文提出ORION方法。研究引入神经坍缩作为表征先验,通过优化特征空间的层级结构,促使编码器学习具备明确顺序和判别性的视觉特征。该方法有效缓解了复杂场景和视觉干扰下的动作歧义,提升了机器人在多样化环境中的导航鲁棒性与决策一致性。

Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision

第一作者: Yuan Xu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervision applies equally across all timesteps, without structured supervision on which manipulation stage the robot is in or what the next gripper-event target should be. This causes failures to concentrate around challenging gripper-event transitions. To address this, we propose StaKe, a plug-in auxiliary supervision framework that automatically derives two complementary signals from demonstration gripper states without manual annotation: a stage classifier that identifies the current manipulation stage, and a keyframe predictor that estimates the target joint action at the next gripper transition. Both are modeled as lightweight auxiliary heads that enrich the learned representations during training, while leaving the base VLA policy...

论文介绍 针对视觉语言动作模型微调时缺乏操作阶段与目标结构化监督,导致夹爪状态转换处易失败的问题,本文提出StaKe框架。该方法无需人工标注,自动从演示数据中提取阶段分类器与关键帧预测器作为轻量级辅助头。两者在训练期间丰富了模型表征学习,显著提升了VLA模型在复杂机器人操作任务中的微调效果与执行成功率。

PressMimic: Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation

第一作者: Yi Lu · 方向: 多模态具身 · 来源: cs.RO

Abstract:Humanoid motion imitation requires not only accurate perception of human kinematics but also faithful reproduction of physical interactions with the environment. However, existing pipelines rely primarily on vision-based motion capture and kinematic imitation, largely ignoring contact dynamics, leading to artifacts such as foot sliding, floor penetration, and unstable behaviors. In this work, we revisit humanoid motion imitation from the perspective of physical grounding and leverage pressure as a unified modality across perception and control. We present PressMimic, a framework that integrates pressure into the full pipeline from motion capture to humanoid control. In the perception stage, we introduce FRAPPE++, a multimodal model that fuses RGB and pressure to jointly estimate 3D pose and global motion, where pressure provides explicit contact and support constraints to...

论文介绍 针对人形机器人动作模仿忽略接触动力学导致脚滑和穿模等问题,本文提出PressMimic框架。该研究将压力作为统一模态贯穿动捕与控制流程。在感知端,提出FRAPPE++模型融合视觉与压力数据估计三维姿态,利用压力提供接触约束;在控制端结合压力反馈优化动作执行,实现了高保真且物理合理的人形机器人运动模仿。

Learning Motion Feasibility from Point Clouds in Cluttered Environments

第一作者: Sajid Ansari · 方向: 机器人操作 · 来源: cs.RO

Abstract:Motion feasibility prediction plays a central role in robotics, particularly in task and motion planning and manipulation. A major bottleneck for this problem in cluttered environments is that infeasible planning attempts by Sampling-based motion planners (SBMPs) can incur substantial computational cost. Also existing approaches for infeasibility certification are limited to low-dimensional configuration spaces and often assume simplified geometric environments represented by primitive objects with known parameters. We study the complementary problem of learning motion feasibility prediction directly from raw RGB-D observations for a 7-DOF manipulator operating in realistic cluttered scenes. We introduce the first large-scale benchmark for this setting, comprising 2.7M grasp feasibility labels over 88 scanned objects and 190 cluttered tabletop scenes. We benchmark three...

论文介绍 针对杂乱环境中运动规划计算成本高及现有方法局限于简单几何的问题,本文研究从原始RGB-D观测直接学习七自由度机械臂的运动可行性预测。研究构建了包含两百七十万个标签的大规模基准数据集,涵盖多种真实桌面场景。该工作为复杂环境下的运动规划提供了高效的数据驱动评估方案,有效降低了无效规划的计算开销。

Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention

第一作者: Siyu Wu · 方向: 机器人操作 · 来源: cs.RO

Abstract:World Action Models (WAMs) generate actions together with predicted futures, offering a powerful interface for robot decision making. In contact-rich manipulation, however, visually plausible futures can be physically incomplete: insertion, assembly, search, and reorientation often depend on slip, jamming, contact normals, or small alignment errors that are weakly visible or hidden in RGB. A natural solution is to predict future tactile states, however, we identify tactile pollution, a failure mode where unconstrained tactile-token injection degrades video and action prediction by forcing a visual dynamics model to absorb sparse, local, event-driven contact signals. To address this, we propose Tactile-WAM, a touch-aware WAM with a Tactile Asymmetric Attention Mechanism (TAAM). TAAM combines a VideoClean mask, which blocks video-query access to tactile key/value tokens while...

论文介绍 针对接触丰富机器人操作中视觉预测物理信息不完整及直接引入触觉导致的「触觉污染」问题,本文提出Tactile-WAM。该模型引入触觉非对称注意力机制与视频掩码,在利用触觉信号预测未来状态的同时,避免稀疏接触信号干扰视觉动态模型,提升了机器人在插入、装配等复杂任务中的决策与动作生成能力。

Hardware Design for Table Tennis Robot Capable of Beating Professional Players

第一作者: Nobuhiko Mukai · 方向: 具身智能 · 来源: cs.RO

Abstract:This paper focuses on the hardware specifications required for a table tennis robot to beat professional players. After analyzing the motions of elite players, we defined target specifications for the workspace, payload, external-force resistance, physical performance, serve capability, and end-effector accuracy. Based on these specifications, we developed "Ace", a custom 8-DoF robot. The mechanical structure was improved through topology optimization to minimize mass while preserving stiffness. Motor and gearbox selection was optimized using an inverse-dynamics torque model. Low-order per-joint dynamics models with delay compensation were identified and integrated into simulation to enable the use of an RL control policy. Experiments demonstrated repeated full-stroke swings with a cycle time of 0.8 s and a peak racket-center velocity of 22 m/s. The robot successfully defeated...

论文介绍 为打造能击败专业选手的乒乓球机器人,本文通过分析精英选手动作确立硬件规格,开发了8自由度定制机器人「Ace」。研究通过拓扑优化降低质量并保持刚度,优化电机选型,并集成带延迟补偿的低阶动力学模型以支持强化学习控制策略。该系统实现了0.8秒挥拍周期与22米/秒的拍头峰值速度。

Bridging Handheld and Teleoperated Supervision for Contact-Rich Manipulation via State-Gated Experts

第一作者: Vidullan Surendran · 方向: 机器人操作 · 来源: cs.RO

Abstract:Handheld data collection systems, such as the Universal Manipulation Interface (UMI), enable scalable data collection across diverse environments but only capture observed actions rather than the desired actions executed by a robot controller. In contrast, teleoperation captures desired actions directly, but is prohibitively time-consuming to collect. We revisit this trade-off through the lens of action validity across task phases. We observe that handheld trajectories provide valid supervision in tolerant, free-space phases, but lack dynamic feasibility in contact-sensitive phases, where tracking observed trajectories at high stiffness produces large, unsafe contact forces. We study the interaction between these two supervision types for contact-rich manipulation and find that training policies that combine handheld data with a small number of targeted teleoperated...

论文介绍 针对接触丰富机器人操作中手持数据采集缺乏动态可行性而遥操作耗时的问题,本文提出结合两者的状态门控专家方法。研究发现手持轨迹在自由空间阶段有效,但在接触敏感阶段易产生不安全力。通过融合大量手持数据与少量针对性遥操作数据,该方法有效桥接了两种监督方式,提升了复杂接触任务的控制性能。

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure

第一作者: Yiyuan Pan · 方向: 具身智能 · 来源: cs.RO

Abstract:A central challenge in deploying learned robot policies is inference-time behavior steering: redirecting a policy at test time to satisfy user preferences not anticipated during training, without retraining. Existing methods fail in two modes: end-to-end methods require fine-tuning or expert-level guidance, while neuro-symbolic methods rely on predefined symbols whose edits can result in logically reasonable but physically infeasible plans. To address this challenge, we propose ReStruct, which builds upon a neural automaton policy that decomposes a visuomotor policy into a high-level state-machine skeleton capturing task structure and a low-level continuous controller represented as a residual policy. Specifically, ReStruct adopts the automaton to represent the preference and incorporates it into the skeleton through a synchronous product, thereby reconfiguring the task...

论文介绍 针对机器人策略在推理时难以满足未预见用户偏好且现有方法易产生物理不可行计划的问题,本文提出ReStruct方法。该方法基于神经自动机策略,将视觉运动策略分解为高层状态机骨架与底层连续控制器。通过将用户偏好融入骨架并重构任务结构,实现了无需重训练的物理可行推理时行为引导。

IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control

第一作者: Chenlong Liu · 方向: 策略学习 · 来源: cs.RO

Abstract:Complex multi-agent control tasks remain challenging for traditional rule-based and model-based approaches, motivating the adoption of learning-based methods. However, learning-based methods often struggle with sim-to-real transfer because they rely on accurate dynamics modeling or system identification and learn policies in low-level control spaces that are highly sensitive to dynamics mismatch, making them costly and fragile in complex environments. To address this issue, we propose a sim-to-real method for multi-agent control, which is insensitive to dynamics mismatch via effect alignment. Our method combines random environmental structure with discrete semantic actions through closed-loop control, elevating policy learning to a semantic abstraction level. Additionally, we develop an action synchronization mechanism that mitigates inter-agent action timing mismatches...

论文介绍 针对多智能体控制中强化学习策略在仿真到现实迁移时对动力学失配敏感的问题,本文提出IDEA方法。该方法通过效果对齐机制,将随机环境结构与离散语义动作结合,把策略学习提升至语义抽象层。此外,研究开发了动作同步机制以减轻智能体间的时序失配,显著提升了复杂环境下的跨域迁移鲁棒性。

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

第一作者: Baiqi Li · 方向: 机器人操作 · 来源: cs.RO

Abstract:A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image, offering no way to evaluate reasoning over observed human behavior. We introduce WatchAct, a benchmark for robot manipulation grounded in observed human behavior. Each instance pairs a real-world human-action video and a language instruction with an aligned simulator scene and an executable LIBERO task, enabling scalable and reproducible evaluation. WatchAct comprises 3,000 long-horizon instances across 14 tasks in four capability domains drawn from the cognitive demands of watching another agent: parsing events (Event Grounding), recovering procedural structure (Procedural Reasoning)...

论文介绍 现有机器人操作基准缺乏对人类行为推理的评估,本文提出WatchAct基准。该基准将真实人类动作视频、语言指令与对齐的仿真场景及可执行任务结合,支持可扩展评估。WatchAct包含3000个长视界实例,涵盖事件解析、过程推理等四个认知能力域,填补了基于视频行为理解的机器人操作评估空白。

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

第一作者: Tyler Ga Wei Lum · 方向: 机器人操作 · 来源: cs.RO

Abstract:Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach. These tasks are contact-rich, making data collection for imitation learning difficult, and sparse-reward, making direct exploration with reinforcement learning (RL) intractable. Consequently, prior work has made progress by structuring the problem with specialized grippers, tool attachments, and environment fixtures. In this work, we argue that before a robot can perfect precise assembly, it must first learn to play. We further ask the question: what factors in the process of learning to play matter for precise assembly? We propose Play2Perfect, an RL framework for task-agnostic pretraining through play on diverse objects and goals, which is then perfected on precise assembly. The goal of play is to acquire reusable manipulation...

论文介绍 针对多指机器人精确装配中数据收集困难与强化学习探索低效的问题,本文提出Play2Perfect框架。该方法先通过在不同物体上进行任务无关的玩耍预训练,获取可复用的灵巧操作技能,再在精确装配任务中微调完善。研究系统探讨了预训练玩耍过程中的关键因素,为复杂接触任务提供了新的学习范式。

MPC-Injection: Biasing Off-Policy Locomotion RL Toward Controller-Induced Behavior Basins

第一作者: Roy Xing · 方向: 导航与运动 · 来源: cs.RO

Abstract:Reinforcement learning (RL) for locomotion frequently converges to locally optimal but undeployable behaviors, such as vibrating limbs or scooting on the torso, that maximize return without producing a usable gait. We present MPC-Injection, a low-overhead method that steers RL toward a designer-preferred gait by inserting transitions into the replay buffer from a model predictive controller solving the same Markov decision process. Unlike reward shaping, MPC-Injection does not require redesigning the task reward, and unlike adversarial imitation learning, it adds no discriminator, no kinematic retargeting, and no auxiliary objective. Instead, the controller's preferred behavior is transferred to the policy purely through the replay state distribution. On a 2D walker in simulation and with sim-to-real evaluation on a Go2 quadruped, we show that MPC-Injection drives the policy...

论文介绍 针对运动强化学习易收敛于肢体振动等不可部署局部最优的问题,本文提出MPC-Injection方法。该方法通过模型预测控制器生成转移数据并注入经验回放缓冲区,利用状态分布将策略引导至偏好步态。此方法无需修改奖励或引入判别器,在仿真与四足机器人真机实验中均有效提升了步态质量与部署可行性。

Scaling Nonlinear Optimization: Many Problems One GPU

第一作者: John Viljoen · 方向: 导航与运动 · 来源: cs.RO

Abstract:Many robotics problems, including trajectory optimization, inverse kinematics, and contact-rich motion planning, reduce to nonlinear programs (NLPs). Mature NLP solvers such as IPOPT can solve these problems, offering hard constraint satisfaction, optimality guarantees, and favorable scaling with problem dimension. These solvers underpin gradient-based methods in robotics, yet remain CPU-bound and solve only one problem at a time, preventing their integration into GPU-batched learning pipelines. On the other hand, sampling-based approaches such as reinforcement learning, model predictive path integral, and imitation learning have become the core of modern robotics research due to their ability to leverage GPU-batched simulators. These simulators can generate orders of magnitude more dynamics rollouts per second than was previously possible. If a GPU-batched NLP solver existed...

论文介绍 针对机器人轨迹优化和运动规划等非线性规划问题受限于CPU单任务求解的瓶颈,本文探讨GPU批处理非线性规划求解器。该方法旨在将成熟的非线性优化算法与GPU并行计算结合,解决传统求解器无法融入现代机器人GPU批处理学习管道的问题,有望大幅提升机器人动力学仿真与策略训练的效率。

KRVF: A Source-Aware Semantic Voxel World Representation for Edge Mobile Manipulation

第一作者: Runfeng Ling · 方向: 机器人操作 · 来源: cs.RO

Abstract:Mobile manipulators need world models that are current, queryable, semantically meaningful, and usable under edge-compute constraints. This technical report presents KRVF, a source-aware semantic voxel world representation for edge mobile manipulation. Unlike reconstruction-centric mapping pipelines that primarily optimize global geometric fidelity, KRVF represents local world state as task-oriented voxels that encode occupancy, color, semantic evidence, temporal freshness, and evidence source. The representation separates measured occupancy from semantic-prior hypotheses, enabling depth-failure-aware object reasoning without silently corrupting persistent geometry. KRVF also closes a feedback loop between mapping and sensing by rendering map-prior depth for repair, and exposes task-level query operators for semantic objects and grasp candidates. The report formalizes the KRVF...

论文介绍 针对移动机械臂在边缘计算约束下的环境感知需求,本文提出KRVF,一种源感知语义体素世界表示方法。该方法将局部环境状态编码为面向任务的体素,分离测量占用与语义先验,支持深度失败感知推理与任务级查询。KRVF通过闭环映射与感知,为边缘端移动操作提供高效、语义丰富的环境表示。

Racing a Wheeled Quadruped: Active Load Transfer Mitigation via Model Predictive Control

第一作者: Marla Eisman · 方向: 策略学习 · 来源: cs.RO

Abstract:This paper presents a hierarchical control framework using model predictive control (MPC) and reinforcement learning (RL) for active roll control to manage lateral load transfer during autonomous racing of a wheeled quadruped. The framework integrates offline time-optimal raceline generation, an online MPC planner that actively minimizes the lateral Load Transfer Ratio (LTR), and a low-level, whole-body RL policy deployed directly onto the robot's 16 actuators. The MPC is based on a vehicle dynamics bicycle model of the Unitree Go2-W platform. The robot's leg actuators act as active suspension where knee joints generate anti-roll torque to bank into turns. Physical track experiments demonstrate that active roll control reduces mean LTR by up to 44%, improves the fastest lap time by 8.7%, and boosts peak lateral acceleration capability by 21.3% to 1.98 $m/s^2$, maintaining...

论文介绍 针对轮式四足机器人在自主赛车中的横向载荷转移问题,本文提出一种结合模型预测控制与强化学习的分层控制框架。该方法通过在线模型预测控制规划器主动最小化横向载荷转移率,并利用底层强化学习策略控制腿部执行器产生抗侧倾力矩。实验表明,该框架能有效提升机器人的过弯稳定性与圈速表现。

NavIsaacLab: Generating Realistic Crowd via Parallel Robot Learning for Benchmarking Human-aware Navigation

第一作者: Bingyi Xia · 方向: 导航与运动 · 来源: cs.RO

Abstract:Robot autonomous navigation that accounts for surrounding human activities is crucial for ensuring both safety and natural human-robot interaction in real-world environments shared by humans and robots. Simulation of complex and diverse navigation scenarios serves as the foundation for training reliable robot navigation policies and accurately evaluating the performance of algorithms, offering an efficient alternative to manual supervision of real data. However, current human-aware navigation research faces significant challenges due to the scarcity of diverse, high-quality scene data. Existing simulation platforms often rely on handcrafted rules to approximate pedestrian behavior and lack the capability to provide extensive sensor signals, typically assuming perfect observations. To address these limitations, this paper presents NavIsaacLab, a comprehensive framework for...

论文介绍 针对人类感知导航研究中高质量场景数据稀缺及现有仿真平台依赖手工规则的问题,本文提出NavIsaacLab框架。该平台通过并行机器人学习生成逼真且多样化的人群行为,并提供丰富的传感器信号,旨在为训练可靠的机器人导航策略和评估算法提供高效、高保真的仿真基准与测试环境。

TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes

第一作者: Blake Werner · 方向: 导航与运动 · 来源: cs.RO

Abstract:How do we learn to hit a tennis backhand? Not from a thousand hours of tennis tournaments on TV - we work with a coach and practice. We argue this is also the right recipe for teaching dynamic skills to humanoid robots. This follows from a structural property of dynamic skills: the outcome is decided by a short, crucial portion of the trajectory - for a backhand, the ~20cm of racket travel around ball contact. Getting this interaction window right requires coordinating the whole motion, so that control, physics, and morphology act in concert. Learning thus reduces to mastering a handful of distinct actions and, for each, practicing until the window comes out right. To this end, we introduce TaskNPoint, a training protocol which makes the coach-learner division of labor explicit. The human coach contributes four inputs: a discrete set of skills (e.g. different shots), one...

论文介绍 针对人形机器人复杂动态技能学习困难的问题,本文提出TaskNPoint训练协议。该方法借鉴人类教练指导模式,将动态技能分解为关键交互窗口,通过明确教练与学习者的分工,由人类提供技能集与指导,机器人针对关键动作集中练习。该协议能显著缩短机器人掌握网球反手等复杂动态技能的时间。

RoboTales: ROBOTic Anthropomorphic LEarning Systems

第一作者: Andrew Chen · 方向: 具身智能 · 来源: cs.RO

Abstract:RoboTales is a low-cost robotic storytelling system that animates narratives using expressive sock puppetry. Implemented autonomously on a Baxter robot as a test case, RoboTales synchronizes narration, gestures, and mouth movements to perform character-driven stories. In a pilot study, puppet-based storytelling outperformed a gesture-only mode, producing higher HRIES ratings and improved story recall, suggesting that embodied puppetry enhances engagement and narrative comprehension. Designed to be modular and platform-agnostic, RoboTales can be adapted to other manipulators and offers a screen-free alternative to passive media, supporting future deployment in child-centered learning environments.

论文介绍 本文提出RoboTales,一种低成本机器人故事讲述系统,通过袜子木偶动画化叙事。该系统在机械臂上自主运行,同步语音叙述、手势与嘴部动作演绎故事。研究表明,具身木偶表演能有效提升受众参与度与故事回忆效果。该系统具备模块化特性,为儿童教育等场景提供了无屏幕的互动学习替代方案。

Morphology-Specific Closed-Loop Control of Logarithmic-Spiral Continuum Arms via Online Jacobian Error Compensation

第一作者: Partha Datta · 方向: 具身智能 · 来源: cs.RO

Abstract:Logarithmic spirals are ubiquitous in biological appendages and provide an attractive morphology for continuum manipulators capable of reaching, wrapping, and grasping. Recently reported logarithmic-spiral robots demonstrated scalable fabrication and versatile grasping but lacked inverse kinematics and closed-loop control. This work presents the first morphology-specific closed-loop task-space control framework for logarithmic-spiral continuum arms. A segmented tendon-driven model with a centerline backbone and equilateral tendon routing is developed in MuJoCo to capture tapered compliance and contact dynamics. An analytical task-space Jacobian is derived directly from the logarithmic-spiral kinematics and combined with online Jacobian error compensation using a Broyden secant update and Kalman-filter estimation. The resulting controller continuously corrects modeling errors...

论文介绍 针对对数螺旋连续体机械臂缺乏闭环控制的问题,本文提出一种形态特定的任务空间控制框架。该方法建立分段肌腱驱动模型以捕捉柔顺与接触动力学,推导解析任务空间雅可比矩阵,并结合Broyden割线更新与卡尔曼滤波进行在线误差补偿。该控制器能持续修正建模误差,实现精确的连续体臂抓取与操作。

RMTL: Reinforced Micro-task Learning for Long-Horizon Manipulation with VLM Rewards

第一作者: Anıl Can Ateş · 方向: 机器人操作 · 来源: cs.RO

Abstract:Reinforcement learning (RL) for robotic manipulation often requires manually designing a dense reward function, which is difficult to tune and often fragile, or learning a reward from human demonstrations or preferences, which can be expensive. A recent line of work uses pretrained vision-language models (VLMs) as zero-shot reward models, replacing these costs with a single text prompt. However, we argue that a single global prompt is too coarse for long-horizon manipulation tasks with randomized initial conditions. The single-prompt VLM reward is near-flat for much of the trajectory, making early progress hard for the agent to detect. We propose Reinforced Micro-Task Learning (RMTL), an approach that decomposes a manipulation task into a small set of language-described micro-tasks and trains the agent to switch between them. At each step, the agent receives a multi-view VLM...

论文介绍 针对机器人长视界操作中视觉语言模型单一提示奖励粗糙的问题,本文提出强化微任务学习方法。该方法将复杂操作分解为多个语言描述的微任务,训练智能体在微任务间切换,并在每步获取多视图视觉语言模型奖励。此方法有效缓解奖励稀疏问题,提升智能体在长视界操作任务中的学习效率与最终成功率。

Reinforcement Learning Enables Autonomous Microrobot Navigation and Intervention in Simulated Blood Capillaries

第一作者: Jannik Drotleff · 方向: 导航与运动 · 来源: cs.RO

Abstract:Autonomous microrobots navigating biological vasculature could enable targeted drug delivery and thrombolysis, yet training control policies for realistic environments remains an open challenge. Prior reinforcement learning (RL) studies of microrobotic navigation have been limited to idealized geometries that omit complex hydrodynamic flow fields, confined branching structures, and dense cellular obstacles found in vivo. Here, we develop a physically grounded simulation of a blood capillary network, incorporating realistic hydrodynamic flow fields, explicit red blood cell dynamics, and anatomically derived branching geometry, and train deep RL agents to navigate it via chemotaxis. We systematically map the physical limits of navigation across robot size and swimming speed, revealing a forbidden regime where Brownian motion and flow overcome propulsion. Successful agents...

论文介绍 针对微型机器人在真实血管中导航控制策略训练困难的问题,本文构建了包含真实流体动力学、红细胞动态及解剖分支结构的毛细血管网络物理仿真环境。研究利用深度强化学习训练智能体通过趋化性自主导航,并系统评估机器人尺寸与游动速度的物理极限,为靶向药物递送和溶栓治疗提供了仿真基础。

OctoSense: Self-Supervised Learning for Multimodal Robot Perception

第一作者: Anthony Bisulco · 方向: 多模态具身 · 来源: cs.RO

Abstract:We present OctoSense, an open-source sensor platform with stereo RGB and event cameras, LiDAR, a thermal camera, an inertial measurement unit, RTK-corrected global positioning system, and proprioception (CAN bus data from a car, and joint angles for a quadruped robot). The eponymous OctoSense dataset contains 59 hours of time-synchronized driving data across different types of environments at different times of the day, including situations with highly degraded sensors. We demonstrate multi-modal self-supervised learning using such real-world robotics data, where sensors have different representations, frequencies, latencies and noise. Our approach, a "late-fusion" masked autoencoder, (i) uses modality-specific tokenizers to account for different spatiotemporal characteristics of these sensors, and (ii) caches modality-specific tokens at inference time to process new...

论文介绍 针对机器人多模态感知中传感器特性差异大的问题,本文推出开源传感器平台OctoSense及大规模同步数据集。研究提出晚期融合掩码自编码器方法,通过模态特定分词器处理不同时空特征与噪声,并在推理时缓存特定模态标记。该方法有效提升了机器人在复杂现实环境中的多模态自监督学习能力。

Automating Potential-based Reward Shaping with Vision Language Model Guidance

第一作者: Henrik Müller · 方向: 导航与运动 · 来源: cs.RO

Abstract:Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to correctly attribute the sparse success rewards to relevant parts of the trajectory. Naive reward shaping can induce reward hacking, yielding policies that exploit auxiliary signals instead of solving the intended task. Potential-based reward shaping (PBRS) guarantees preservation of the optimal policy set, but requires the definition of a heuristic potential function over the state space. In this work, we introduce the VLM-guided PBRS framework VLM-PBRS that learns the potential function directly from vision language model (VLM) feedback. We query a lightweight VLM to obtain preferences over image pairs and train a model of the potential function using these preferences. As this approach is based on potential-based reward shaping, it...

论文介绍 针对强化学习中稀疏奖励导致探索困难及朴素奖励塑造引发作弊行为的问题,本文提出基于视觉语言模型引导的势能奖励塑造框架。该方法通过查询轻量级视觉语言模型获取图像对偏好反馈,直接学习状态空间的势能函数。此自动化机制在保留最优策略集的同时,有效缓解了复杂任务中的奖励稀疏与引导难题。

Charting the Growth of Social-Physical HRI (spHRI): A Systematic Review Pipeline Augmented by Small Language Models

第一作者: Mayumi Mohan · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Social-physical human-robot interaction (spHRI) has grown rapidly across robotics, human-computer interaction, human-robot interaction, and haptics. Yet, fragmented terminology and inconsistent methodologies make systematic synthesis difficult. To support scalable review practices, we evaluated the extent to which small language models (SLMs; < 1.5B parameters) can assist with title and abstract screening for a large spHRI systematic review. While no SLMs matched human reviewers' performance, the models operated locally and screened papers orders of magnitude faster. The combined SLM ensemble identified 39 papers reviewers missed, representing 10.29% of the final relevant dataset. These results demonstrate that SLMs can augment, rather than replace, expert reviewers and make large-scale literature reviews accessible and sustainable.

论文介绍 针对社会物理人机交互领域文献激增且术语碎片化导致系统综述困难的问题,本文评估了小语言模型辅助文献筛选的潜力。研究表明,小语言模型虽未完全达到人类水平,但本地运行速度极快,且与人类专家结合能发现更多相关文献。该流水线显著提升了大规模文献综述的效率与可持续性。

Identifying the Unknown: Prompt-Free Open Vocabulary Anomaly Recognition for Robot-Object Interaction

第一作者: Philipp Allgeuer · 方向: 具身智能 · 来源: cs.CV

Abstract:Robots operating in real-world environments must in general be able to recognize previously unseen objects. As robotic systems move toward open-world autonomy, there is a growing, yet largely unmet, need for open vocabulary object detectors that are prompt-free and efficient enough for continuous deployment. We present AnomNOVIC, a two-stage known-workspace framework that combines a masked autoencoder (MAE) trained for anomaly detection, with NOVIC, a powerful real-time prompt-free open vocabulary image classifier. The MAE produces generic object-agnostic bounding boxes, allowing NOVIC to classify salient image regions without requiring a predefined candidate class list. We evaluate AnomNOVIC against strong open vocabulary baselines in a tabletop robot-object environment featuring the NICOL humanoid robot, reaching 47.1% AP / 57.5% AP50 for prompt-free recognition, and 59.0%...

论文介绍 针对机器人在开放世界中识别未知物体的需求,本文提出AnomNOVIC双阶段框架,结合用于异常检测的掩码自编码器与实时无提示开放词汇分类器。该方法通过生成通用边界框提取显著区域并进行分类,无需预定义候选类别列表。实验表明其在桌面机器人交互场景中实现了高效的无提示异常目标识别。

市场总览

当前全球资产技术面呈现显著分化与整体承压态势。美股方面,标普与纳指ETF跌破短期均线,科技巨头如微软、特斯拉呈现空头排列,动量转弱。中概股遭遇重挫,阿里与京东RSI双双跌入超卖区,下行趋势明显。加密市场情绪极度低迷,恐慌贪婪指数跌至15(极度恐慌),总市值缩水至2.17万亿美元,比特币与以太坊均维持空头排列且RSI处于32左右的低位。商品与外汇方面,美元指数RSI超买且多头排列,表现一枝独秀;而WTI原油RSI超卖,黄金则触发MACD死叉。宏观层面VIX指数触发金叉,显示市场避险情绪升温。整体而言,风险资产技术面普遍偏弱,仅美元与部分避险指标展现强势。

今日关注

BABA 阿里巴巴 (BABA)
偏下行

当前价94.81,RSI14跌至16.5触发超卖信号,MACD维持空头排列且绿柱扩大。近5日重挫11.48%,价格远低于20日均线113.53。技术面呈现极弱势的下行趋势,短期动能衰竭但未见反转信号。

DX-Y.NYB 美元指数 DXY
偏上行

当前价101.37,RSI14达71.3进入超买区,均线呈多头排列,MACD维持金叉且红柱扩张。价格逼近52周高点,短期技术面展现强烈的上行倾向与动量,但需警惕超买后的技术性回调风险。

SOL-USD Solana
中性

当前价71.98,近1日逆势上涨6.52%,RSI14报49.6处于中性区间。MACD虽仍为空头排列,但绿柱显著缩短至-1.63。在加密市场极度恐慌背景下展现相对抗跌性,短期处于方向选择的中性震荡状态。

全部资产

^VIX

VIX 恐慌指数

$18.41 -2.54%
5 日
+12.26%
距 52w 高
-47.8%
RSI(14)
51.1
趋势
中性
SMA 20 / 50 / 200
17.92 / 17.79 / 18.64
MACD / 信号
0.152 / 0.038
MACD 金叉 (3 天前)

^TNX

10Y 美债收益率 (%)

$4.37 -1.77%
5 日
-2.56%
距 52w 高
-12.5%
RSI(14)
39.9
趋势
中性
SMA 20 / 50 / 200
4.48 / 4.43 / 4.22
MACD / 信号
-0.005 / 0.012

DX-Y.NYB

美元指数 DXY

$101.37 -0.06%
5 日
+0.51%
距 52w 高
-0.4%
RSI(14)
71.3
趋势
多头
SMA 20 / 50 / 200
100.14 / 99.22 / 98.77
MACD / 信号
0.614 / 0.467
RSI 超买接近 52 周高多头排列

SPY

S&P 500 ETF

$728.99 -0.72%
5 日
-2.38%
距 52w 高
-4.1%
RSI(14)
43.2
趋势
中性
SMA 20 / 50 / 200
743.60 / 734.35 / 690.54
MACD / 信号
-0.138 / 2.657

QQQ

Nasdaq 100 ETF

$706.52 -1.38%
5 日
-4.60%
距 52w 高
-5.6%
RSI(14)
46.5
趋势
多头
SMA 20 / 50 / 200
724.76 / 702.79 / 632.15
MACD / 信号
3.467 / 7.544
多头排列

AAPL

Apple

$283.78 +3.14%
5 日
-4.78%
距 52w 高
-10.6%
RSI(14)
41.3
趋势
中性
SMA 20 / 50 / 200
298.29 / 291.50 / 269.46
MACD / 信号
-2.256 / 0.508

MSFT

Microsoft

$372.97 +5.71%
5 日
-1.69%
距 52w 高
-32.9%
RSI(14)
40.4
趋势
空头
SMA 20 / 50 / 200
400.11 / 410.97 / 447.99
MACD / 信号
-13.872 / -9.812
空头排列

NVDA

Nvidia

$192.53 -1.64%
5 日
-8.62%
距 52w 高
-18.6%
RSI(14)
37.5
趋势
中性
SMA 20 / 50 / 200
207.77 / 210.09 / 190.64
MACD / 信号
-3.697 / -1.793

GOOGL

Alphabet

$337.39 -1.84%
5 日
-8.33%
距 52w 高
-17.4%
RSI(14)
33.3
趋势
中性
SMA 20 / 50 / 200
360.81 / 369.28 / 313.85
MACD / 信号
-7.353 / -4.111

TSLA

Tesla

$379.71 +1.22%
5 日
-5.19%
距 52w 高
-23.9%
RSI(14)
41.2
趋势
空头
SMA 20 / 50 / 200
401.55 / 404.76 / 417.96
MACD / 信号
-7.976 / -4.373
空头排列

META

Meta

$550.25 +1.36%
5 日
-4.67%
距 52w 高
-30.9%
RSI(14)
37.4
趋势
空头
SMA 20 / 50 / 200
583.29 / 612.90 / 650.02
MACD / 信号
-16.259 / -13.176
空头排列
加密恐慌贪婪
15
极度恐慌
加密总市值
$2.17 T
+2.74% / 24h
BTC 主导率
55.7%
ETH 8.8%
24h 成交量
$90.1 B
活跃币 17,449

BTC-USD

Bitcoin

$60,204.23 +0.81%
5 日
-4.80%
距 52w 高
-52.3%
RSI(14)
32.7
趋势
空头
SMA 20 / 50 / 200
63,221.17 / 70,234.85 / 76,002.18
MACD / 信号
-2,273.960 / -2,286.214
空头排列

ETH-USD

Ethereum

$1,581.33 +1.06%
5 日
-7.23%
距 52w 高
-68.1%
RSI(14)
32.7
趋势
空头
SMA 20 / 50 / 200
1,686.61 / 1,920.67 / 2,327.64
MACD / 信号
-77.233 / -77.901
空头排列

SOL-USD

Solana

$71.98 +6.52%
5 日
-0.61%
距 52w 高
-71.6%
RSI(14)
49.6
趋势
空头
SMA 20 / 50 / 200
69.41 / 78.09 / 95.80
MACD / 信号
-1.639 / -2.193
空头排列

BABA

阿里巴巴 (BABA)

$94.81 -0.27%
5 日
-11.48%
距 52w 高
-50.8%
RSI(14)
16.5
趋势
空头
SMA 20 / 50 / 200
113.53 / 126.30 / 148.22
MACD / 信号
-8.528 / -6.804
RSI 超卖空头排列

PDD

拼多多 (PDD)

$76.55 +4.43%
5 日
-3.78%
距 52w 高
-45.1%
RSI(14)
34.2
趋势
空头
SMA 20 / 50 / 200
81.51 / 91.25 / 109.37
MACD / 信号
-4.515 / -4.271
空头排列

JD

京东 (JD)

$25.39 +0.79%
5 日
-7.91%
距 52w 高
-31.1%
RSI(14)
26.8
趋势
空头
SMA 20 / 50 / 200
27.98 / 29.67 / 30.09
MACD / 信号
-1.168 / -0.849
RSI 超卖空头排列

0700.HK

腾讯控股 (0700.HK)

HK$411.80 -2.28%
5 日
-6.45%
距 52w 高
-39.7%
RSI(14)
35.8
趋势
空头
SMA 20 / 50 / 200
445.59 / 461.46 / 561.50
MACD / 信号
-10.926 / -7.749
MACD 死叉 (4 天前)接近 52 周低空头排列

GC=F

黄金期货

$4,103.00 +1.80%
5 日
-2.87%
距 52w 高
-26.6%
RSI(14)
37.9
趋势
中性
SMA 20 / 50 / 200
4,273.07 / 4,488.59 / 4,448.33
MACD / 信号
-119.817 / -106.270
MACD 死叉 (4 天前)

CL=F

WTI 原油期货

$70.24 -2.34%
5 日
-8.30%
距 52w 高
-41.2%
RSI(14)
27.9
趋势
中性
SMA 20 / 50 / 200
83.29 / 92.01 / 73.90
MACD / 信号
-6.453 / -5.185
RSI 超卖

USDCNY=X

美元 / 人民币

¥6.79 -0.00%
5 日
+0.31%
距 52w 高
-5.9%
RSI(14)
56.5
趋势
空头
SMA 20 / 50 / 200
6.77 / 6.79 / 6.95
MACD / 信号
-0.002 / -0.007
接近 52 周低空头排列
风险提示

请注意,过去走势不代表未来表现。技术指标仅反映历史价格动量与统计特征,存在滞后性与失效风险。本报告内容仅供技术指标解读参考,不构成任何投资建议,据此操作风险自担。

Venezuela Live Updates: Trump’s Vow to ‘Run’ Venezuela Is Tested After Quakes

The U.S. dispatched hundreds of rescue workers to help the Venezuelan government, as the injured overwhelmed hospitals and the death toll rose to 920. The Pentagon sent two ships, transport planes and helicopters.

中文摘要 委内瑞拉地震死亡人数升至920人,医院人满为患。美国派遣数百名救援人员及五角大楼派出两艘军舰、运输机和直升机协助委政府救援。

U.S. Strikes Iran in Retaliation for Attack on Vessel in Strait of Hormuz

President Trump on Friday called Iran’s attack on a container ship transiting the Strait of Hormuz a day earlier a “foolish” act.

中文摘要 针对伊朗在霍尔木兹海峡袭击集装箱船事件,美国总统特朗普称其为愚蠢行为,美国随后对伊朗发动报复性打击。

Nations Send Rescue Teams and Aid to Venezuela After Earthquakes

The number of the dead continues to rise as search-and-rescue teams descended upon the country to dig out people buried in rubble.

中文摘要 委内瑞拉地震死亡人数持续上升,多国搜救队伍及援助物资陆续抵达该国,全力在废墟中挖掘和营救被困人员。

Scenes of Collapse: The Emergency at Venezuela’s Hospitals

A snapshot of the toll of the country’s twin quakes.

中文摘要 委内瑞拉连续两次地震造成严重破坏,该国多家医院面临严峻的医疗挤兑与系统崩溃等紧急状况,伤亡代价惨重。

US conducts strikes on Iran after attack on cargo ship

US Central Command says it has struck missile and drone storage facilities and coastal radar positions.

中文摘要 美国中央司令部宣布,为报复货船遇袭事件,美军已对伊朗的导弹与无人机储存设施及沿海雷达阵地实施打击。

Religion row as Texas makes Bible stories required reading in schools

Critics say the new reading requirements infringe on religious freedoms and blur the separation of church and state.

中文摘要 美国得克萨斯州将圣经故事列为学校必读内容,引发宗教争议。批评者指此举侵犯宗教自由,模糊了政教分离界限。

Convicted Rapist Who Fled to Scotland and Faked His Own Death Dies in Utah

Nicholas Rossi, 38, raped two girlfriends in 2008 and later fled to Scotland, prosecutors said. An attentive nurse treating him for Covid in 2021 identified him.

中文摘要 38岁的尼古拉斯·罗西因2008年强奸两名女友被定罪后逃往苏格兰并伪造死亡,2021年因新冠治疗时被护士识破,近日在犹他州死亡。

Trump justifies strikes on Iran amid ceasefire

The US has struck Iran in retaliation for what it says was an Iranian attack on a ship in the Strait of Hormuz.

中文摘要 在停火期间,美国总统特朗普为美军打击伊朗的行动辩护,称此举是对伊朗在霍尔木兹海峡袭击船只事件的报复。

Paris Diamond League to go ahead with safety measures amid heatwave

Only competitions involving professional athletes will be held, with all other activities cancelled.

中文摘要 受热浪影响,巴黎钻石联赛将采取安全措施继续举行,但仅限职业运动员参与的比赛,其他所有相关活动均已取消。

US judge holds prosecutor in contempt in Charlie Kirk murder case

Judge says comments to the media by prosecutors about defendant violate rules of what can be said outside of court.

中文摘要 在查理·柯克谋杀案中,美国法官裁定控方检察官藐视法庭,指其向媒体发表关于被告的言论违反了庭外言论规则。

Venezuela shaken by magnitude 4.9 tremor days after major earthquakes

The country is still reeling from devastating pair of earthquakes that killed hundreds of people earlier this week.

中文摘要 在本周早些时候发生造成数百人死亡的两次破坏性大地震后,委内瑞拉再次发生4.9级地震,该国仍深受其害。

Earthquake Tests Growing Ties Between U.S. and Venezuela

The Trump administration said it would commit aid, at a time when it has been expanding U.S. commercial interests in Venezuela beyond oil.

中文摘要 特朗普政府宣布向委内瑞拉提供地震援助,此举考验着两国日益密切的关系,当前美国正将其在委商业利益扩展至石油以外领域。

Trump threatens 100% tariff on European nations over tech tax

The US president says "Numerous European countries" have been discussing bringing in such a levy.

中文摘要 美国总统特朗普威胁称,若欧洲国家推进数字服务税,将对多个欧洲国家征收100%关税。此举旨在回应欧洲多国正在讨论引入的相关税收政策。

Aid Groups Flock to Venezuela In Search-and-Rescue Frenzy

Two of the US military ships used in a blockade meant to pressure Nicolás Maduro have headed back toward Venezuela, this time with rescue teams, equipment and medical aid after devastating twin earthquakes struck the country this week.

中文摘要 委内瑞拉本周遭遇毁灭性双重地震后,两艘曾参与对马杜罗政府施压封锁的美国军舰转向该国,运送救援队、设备及医疗物资。多家援助机构正赶赴委内瑞拉开展搜救行动。

Philippine Government Plans to Raise Budget by 6% Next Year

The Philippines plans to increase the budget by 6% next year to 7.2 trillion pesos ($117 billion), according to the budget department.

中文摘要 菲律宾预算部表示,菲律宾计划明年将国家预算增加6%,至7.2万亿比索(约合1170亿美元),以支持国家经济发展与各项政府支出计划。

Trump administration allows some access to Anthropic’s Mythos

Move eases tension with AI lab but unease over Washington’s ad hoc regulatory approach remains

中文摘要 特朗普政府允许部分获取人工智能实验室Anthropic的Mythos模型,此举缓解了与该AI实验室的紧张关系。但外界对华盛顿临时性监管方式的担忧依然存在。

Bolivia Moves to Flexible Exchange-Rate System After 15 Years

Bolivia is moving to a flexible exchange-rate system to strengthen macroeconomic stability, its Finance Ministry announced Friday.

中文摘要 玻利维亚财政部周五宣布,在实行固定汇率15年后,该国将转向浮动汇率制度,旨在加强宏观经济稳定性并应对当前的经济挑战。

Wall Street Week | USMCA: Can North America’s Trade Deal Survive?

This week, a special edition of Wall Street Week on the USMCA trade deal. The renegotiations are testing the future of an auto industry built on decades of cross-border integration. And, American farmers increasingly depend on exports to Canada and Mexico as other overseas markets become more diffic

中文摘要 《华尔街周刊》探讨美墨加协定(USMCA)的未来。重新谈判正考验建立在数十年跨境整合基础上的北美汽车产业前景,同时美国农民对加拿大和墨西哥出口的依赖度日益加深。

Three unusual things about the King's tax bill

King Charles paid £12.9m in tax for 2024-2025 - here's what we know about his unique tax situation.

中文摘要 英国国王查尔斯三世在2024至2025财年缴纳了1290万英镑的税款。媒体详细梳理了其独特的税务状况及纳税细节,凸显其与众不同的税务安排。

Earn Your Leisure Co-Founders on Financial Literacy

The CEOs and co-founders of Earn Your Leisure, Rashad Bilal and Troy Millings, explain their educational platform that looks to help individuals gain financial freedom and generational wealth. They discuss how to ride the tech stock wave with Romaine Bostick and Katie Greifeld on Bloomberg's "The Cl

中文摘要 「Earn Your Leisure」首席执行官兼联合创始人Rashad Bilal和Troy Millings介绍了该平台如何帮助个人实现财务自由和代际财富,并探讨了如何把握科技股投资机遇。

Tennis Pro Caroline Wozniacki on Wimbledon

Tennis champion Caroline Wozniacki sat down with Romaine Bostick and Katie Greifeld on the sidelines of The Wimbledon Court in Central Park. She discusses the future of the sport, player compensation, and advice for her fellow tennis star Serena Williams re-entering competition. (Source: Bloomberg)

中文摘要 网球冠军卡洛琳·沃兹尼亚奇在温布尔登网球公开赛期间接受采访,探讨了网球运动的未来、球员薪酬问题,并为同胞球星塞雷娜·威廉姆斯重返赛场提供了建议。

Hellman &amp; Friedman-Backed Hub Files Confidentially for IPO

Hub International Holdings Inc., an insurance broker backed by Hellman & Friedman, confidentially filed for an initial public offering whose proceeds could be used by the to pare its debt.

中文摘要 由Hellman & Friedman支持的保险经纪公司Hub International Holdings已秘密提交首次公开募股(IPO)申请,其募资所得将用于削减公司债务。

Invesco’s Kriskey: Oil Price Freefall Is ‘Overdone'

Kathy Kriskey, Invesco's Alterative ETF Strategy Head, helps make sense of oil's drop as ships continue to cross through the Strait of Hormuz despite attacks. She speaks with Romaine Bostick & Katie Greifeld on "The Close." (Source: Bloomberg)

中文摘要 景顺集团(Invesco)另类ETF策略主管Kathy Kriskey指出,尽管霍尔木兹海峡仍受袭击影响,但船只持续通行,近期原油价格的暴跌反应过度。

Biolife Is Said to Have Drawn Takeover Interest From Repligen

Biolife Solutions Inc. has attracted takeover interest from parties including diagnostics company Repligen Corp., according to people with knowledge of the matter.

中文摘要 据知情人士透露,生物科技公司Biolife Solutions已吸引包括诊断公司Repligen Corp.在内的多方收购意向,相关收购谈判或正在推进中。

【CHY公益站】重新上架GLM-5.2、MIMO-2.5等国模

本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已

我用Rust实现一个为 AI 运维而生的 SSH 客户端,支持: AI + GUI/CLI + 命令块 + 多端数据同步

本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 我受不了市面上 SSH 客户端了,所有工具都想来我家里抢地盘,撒泡尿还不冲 自己维护 host key,我真不懂为啥? 云同步订阅要收费,稍微能理解,但是我没有付费习惯 自己的专有录制格式,为啥呀? 多平台没有移动端

【翰林文苑】基准已发,一键检查你的Opus 4.6是不是真的!

本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 一个基于概率分布识别任意模型真假的项目,从此告别掺假! - 开发调优 - LINUX DO github.com GitHub - hanlinwenyuan/hlwy-ai-checker: 检查第三方AI API是

我终于知道deepseek为什么这么急于开发自己的agent了

大模型能力是引擎,agent才是完整的车,引擎固然重要,但是照现在的趋势,引擎之间的差距会越来越小,后面就是看各家的工具对自己模型的适配度,这才是后续的生态护城河,例如claudecode和codex。 我最近在读opencode和pi的源码,越来越能体会到一套好的harness能发挥出多大的作用了。比如我现在在开发一个类silly tavern的桌面端agent,由于原生的并不好用,太多技术债了,我就用pi的agent二开了一套,可以兼容导入社区的预设、角色卡。设计好一套harness,规范ai剧情演绎、设计一系列tool来强制ai进行记忆,有效防止长文本聊天下的记忆丢失和风格偏移问题,并且

【打破信息茧房】说出你爱看的油管博主

最近刷油管被算法焊死在舒适区了,越刷内容越重复。开个楼大家互相安利私藏博主,一起跳出信息茧房! 先抛块砖: 零度解说:挖尽神器,无脑跟随操作 李永乐老师:黑板科普,数理通解 試當真:搞笑短片、网络文化、社会讽刺为主 ,下饭片,10分钟一集 发帖前请过目: 本贴纯属内容交流,不涉及任何政治话题,请大伙自觉遵守站规,聊博主、聊内容、聊干货; 欢迎跟贴分享你私藏的宝藏博主(最好附上主页链接),方便各位佬查阅; 我这点推荐只是引子,希望大家一起把楼盖起来,真正打破信息茧房(AI相关的就更好噜)。 感谢佬们! 有好看的记得喊我,我先去蹲更新了 56 个帖子 - 46 位参与者 阅读完整话题

分享一下我是怎么在一晚上赚¥11000的

涨价前在淘宝 ¥38769 购入,现价 ¥49889 2026款-16英寸 M5 Max 芯片(18+40) 2T固态硬盘 128GB 遵循佬友的意见改了标题,从一个月内变成一晚上 21 个帖子 - 16 位参与者 阅读完整话题

【九幺】恢复了+一些事务

本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已

L 站有人对 Claude Code 泄露的源码感兴趣的吗?

我打算重写一下 Claude Code 技术架构拆解的系列文章。 基于 Claude Code 泄露的源码进行分析。 说打算写,其实我已经正在写了。 我之前写过一个 Claude Code 的架构图,那篇文章我现在看来,还是差了很多东西。 这个课题在我的脑海里有一阵子了。 我网上找到了一些源码分析的文章。 但是绝大多数都是 AI 味儿很大的那种。 要么逻辑衔接不太顺畅,要么虎头蛇尾的。拼接感太严重。 不过有一些作图做的确实比较好,这个我承认。 而且看图并不能让你完整的理解,只能提高一下你对技术图的审美。 还是需要对图进行解释和描述。 不过呢,我已经加上了 gif 图,而且还在尝试加入一些可交互

外区APP Store软件更新技巧

把外区 Apple ID 添加到“设置-备忘录、邮件、日历”里,三个里面随便选一个,之后更新软件可以直接更新。不用去 App Store 来回切账号。 46 个帖子 - 45 位参与者 阅读完整话题