harry0703/MoneyPrinterTurbo
Python · ★ 63,109 · 🍴 9,188 · 📈 1,742 stars today
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
中文介绍 利用AI大模型(LLM)和图像生成技术,实现了一键自动生成高清短视频的流程,旨在大幅降低视频内容创作的门槛。典型用户是内容创作者、营销人员或需要快速生产视频素材的团队。
Python · ★ 63,109 · 🍴 9,188 · 📈 1,742 stars today
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
中文介绍 利用AI大模型(LLM)和图像生成技术,实现了一键自动生成高清短视频的流程,旨在大幅降低视频内容创作的门槛。典型用户是内容创作者、营销人员或需要快速生产视频素材的团队。
TypeScript · ★ 40,647 · 🍴 3,248 · 📈 4,465 stars today
Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
中文介绍 该工具能将任意代码库自动转换为可交互的知识图谱,支持探索、搜索与问答,旨在提升代码理解与知识管理的效率。它兼容多种主流AI编程助手,面向开发者、技术文档编写者和学习者。
★ 5,861 · 🍴 438 · 📈 664 stars today
A skill file for removing AI tells from prose
中文介绍 一个用于文本润色的技能文件,其核心功能是识别并移除AI生成文本中的典型模式(如过度客套、空洞表述),使文风更自然、专业。主要用户是希望利用AI辅助写作但追求高质量输出的作者或编辑。
JavaScript · ★ 196,260 · 🍴 30,196 · 📈 2,062 stars today
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
中文介绍 一个面向AI智能体(如Claude Code, Codex)的性能优化系统,通过管理技能、记忆、安全性和研究流程来提升开发效率。它为构建高效、可靠的AI辅助开发环境提供了系统性框架,适用于AI工程师和研究人员。
Python · ★ 17,399 · 🍴 2,032 · 📈 695 stars today
Open source repository of plugins primarily intended for knowledge workers to use in Claude Cowork
中文介绍 Anthropic官方开源的插件仓库,主要为知识工作者提供各类增强功能,可在Claude Cowork环境中使用。这些插件旨在优化文档处理、信息检索和分析等工作流程。
Shell · ★ 24,684 · 🍴 1,894 · 📈 2,715 stars today
Taste-Skill - gives your AI good taste. stops the AI from generating boring, generic slop
中文介绍 一个AI技能插件,通过给AI注入“品味”规则,防止其生成无聊、通用化的内容。它能指导AI输出更具个性化和审美价值的结果,适用于内容创作者和设计师。
Python · ★ 22,099 · 🍴 2,357 · 📈 211 stars today
Fully automatic censorship removal for language models
中文介绍 一个全自动化的工具,旨在移除语言模型内置的内容安全过滤器,以获取未经审查的模型原始输出。它为研究人员和开发者提供了测试模型原始能力的途径。
Python · ★ 26,960 · 🍴 4,667 · 📈 401 stars today
Kronos: A Foundation Model for the Language of Financial Markets
中文介绍 Kronos是一个专为金融市场语言设计的基础大模型,专注于理解和生成金融领域的文本与数据。它可用于分析市场报告、新闻和交易信息,服务于量化分析师和金融研究者。
Python · ★ 11,121 · 🍴 1,261 · 📈 886 stars today
754 structured cybersecurity skills for AI agents · Mapped to 5 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND & NIST AI RMF · agentskills.io standard · Works with Claude Code, GitHub Copilot, Codex CLI, Cursor, Gemini CLI & 20+ platforms · 26 security domains · Apache 2.0
中文介绍 为AI智能体(如Claude Code, GitHub Copilot)提供的结构化网络安全技能库,涵盖754项技能,并映射到MITRE ATT&CK等5大安全框架。它帮助安全分析师和AI安全工程师系统化地执行安全任务。
TypeScript · ★ 47,423 · 🍴 6,732 · 📈 519 stars today
The open alternative to Salesforce, designed for AI.
中文介绍 一个开源的客户关系管理(CRM)平台,被定位为Salesforce的现代化、AI优先的替代方案。它采用TypeScript、React等技术栈构建,旨在为中小企业和开发者提供一个灵活、可自托管的CRM选择。
Shell · ★ 1,932 · 🍴 207 · 📈 87 stars today
Claude Code Dedicated Development Harness - Achieving High-Quality Development Through an Autonomous Plan→Work→Review Cycle
中文介绍 一个专为Claude Code设计的开发工具,通过构建自主的“计划-执行-审查”工作循环,来提升AI辅助开发的代码质量和流程规范性。主要面向深度使用Claude Code进行项目开发的工程师。
HTML · ★ 169,437 · 🍴 3,222 · 📈 2,222 stars today
DigitalPlat FreeDomain: Free Domain For Everyone
中文介绍 DigitalPlat FreeDomain项目旨在为所有人提供免费的域名服务,通常通过子域名的形式实现。它服务于个人开发者、学生或小型项目,以零成本获得一个可用的网络域名。
Shell · ★ 209,871 · 🍴 18,700 · 📈 1,511 stars today
An agentic skills framework & software development methodology that works.
中文介绍 一个智能体技能框架与软件开发方法论,致力于为AI辅助的软件开发提供一套有效的、可实践的工作模式。它适用于希望建立规范化AI开发流程的软件团队和独立开发者。
★ 47,117 · 🍴 4,947 · 📈 1,163 stars today
An advanced guide to learn English which might benefit you a lot 🎉 . 离谱的英语学习指南/英语学习教程/英语学习/学英语
中文介绍 一本高级英语学习指南,分享了从词汇、语法到思维模式的系统化进阶方法和资源。其目标读者是已经具备基础、希望将英语水平提升至精通阶段的学习者。
Rust · ★ 16,950 · 🍴 1,112 · 📈 376 stars today
Effortlessly compose, extend, and observe every service in real-time for the first time ever.
中文介绍 一个服务编排与观测平台,声称能实现对每个服务的实时编排、扩展和监控。它旨在简化微服务架构下的服务治理复杂性,面向DevOps工程师和系统架构师。
JavaScript · ★ 5,991 · 🍴 299 · 📈 728 stars today
Curated list of the best free apps for PC and mobile
中文介绍 一个策展列表,汇总了PC和手机平台上最佳的免费应用程序。它为用户提供了一个高质量的免费软件发现和选择指南,适合所有寻找可靠免费工具的用户。
TypeScript · ★ 40,297 · 🍴 4,048 · 📈 72 stars today
💖🧸 Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds, wishing to achieve Neuro-sama's altitude. Capable of realtime voice chat, Minecraft, Factorio playing. Web / macOS / Windows supported.
中文介绍 一个自托管的、用户拥有的AI伴侣项目,旨在将AI角色带入现实交互。它支持实时语音聊天、Minecraft游戏等,目标是打造具有长期记忆和个性化的数字生命体。
@sairahul1 · 106.0K 粉丝 · 2.7M 阅 · 1.4K 赞 · 199 转
I thought I was using AI to code. I was actually just typing faster. Here is the difference — and the 7-agent system that changed everything. Save this. It will save you months. THE PROBLEM NOBODY
中文介绍 博主揭示从单纯用 AI 加速打字到实际编码的转变,提出基于 Claude Code 的 7-agent 协作系统,构建软件工厂实现自动化开发,能在睡眠中交付功能,节省数月时间,强调高效工作流的价值。
@poteto · 26.6K 粉丝 · 86.5K 阅 · 540 赞 · 48 转
I need to get something off my chest. Before my interview @cursor_ai, I had never actually used Cursor. At Meta, Claude Code was explosively taking off. I even paid for a personal $200 a month plan
中文介绍 博主分享个人经历,面试 Cursor 前从未使用,但在 Meta 时 Claude Code 流行,支付 $200 月费后探索不同工具,反映 AI 编码工具如 Cursor 和 Claude Code 的快速演进和个人选择过程。
@difflawb · 20.3K 粉丝 · 21.9K 阅 · 1.1K 赞 · 389 转
How a 40-line shell script became infrastructure In August 2024, Andrej Karpathy — co-founder of OpenAI, former AI Director at Tesla — published something unexpectedly small. Not a paper. Not a model.
中文介绍 博主讲述 OpenAI 联合创始人 Andrej Karpathy 在 2024 年 8 月发布的 40 行 shell 脚本,意外演变为关键基础设施,展示小工具在 AI 时代的大影响,突出开源或工具化的意外潜力。
@ActionModelAI · 57.1K 粉丝 · 5.8K 阅 · 505 赞 · 344 转
We are witnessing the beginning of the biggest economic shift in modern history. And most people still don’t realize it. AI replacement is no longer some distant sci-fi prediction. It has started.
中文介绍 博主分析 AI 替代已经开始,代表现代史上最大的经济转变,呼吁关注这一趋势而非视为科幻预言,强调多数人尚未意识到当前风险和潜在社会影响。
Datasets vs. inductive bias, world models, and programmable biology
中文介绍 ESMFold2项目由Alex Rives领导,与BioHub合作,聚焦蛋白质预测,探讨数据集与归纳偏见、世界模型及可编程生物学。
中文介绍 ITBench-AA是首个针对代理企业IT任务的基准测试,前沿模型得分低于50%,由Artificial Analysis和IBM联合推出。
Cisco and OpenAI are redefining enterprise engineering with Codex, helping Cisco scale AI-native development, accelerate AI Defense work, and automate defect remediation.
中文介绍 思科与OpenAI合作,利用Codex技术重新定义企业工程,助力思科扩展AI原生开发、加速AI防御及自动化缺陷修复。
See how OpenAI, Thrive, and Crete built a self-improving tax agent with Codex, automating filings, improving accuracy, and accelerating workflows.
中文介绍 OpenAI、Thrive和Crete利用Codex技术构建自我改进的税务代理,实现自动化申报、提升准确性并加速工作流程。
it's funding news, but it's good news.
中文介绍 新AI基础设施公司Fireworks和Baseten估值超过百亿美元,OpenRouter预计加入,这是积极的融资新闻。
Warp uses GPT-5.5 and OpenAI models to coordinate coding agents across local, cloud, and open-source development workflows.
中文介绍 Warp押注使用GPT-5.5和OpenAI模型构建开源,协调本地、云端和开源开发中的编码代理。
Ahead of global elections, we’re helping people access information, supporting cyber defenders, and increasing AI transparency
中文介绍 2026年全球选举前,OpenAI帮助人们获取信息、支持网络防御并提升AI透明度。
中文介绍 Reachy Mini机器人实现完全本地化运行。
中文介绍 TRL技术实现Delta Weight Sync,通过Hub Bucket传输万亿参数模型。
中文介绍 xAI Cursor设限、DeepSWE项目进展及中国对AI旅行的限制。
Amid rapidly growing adoption of enterprise-level AI agents, there’s a disconnect emerging between ambition and execution. Although 85% of organizations say they want to be agentic within the next three years, 76% say their current operations and infrastructure can’t support that change. They cite a
中文介绍 在企业级AI代理快速增长中,目标与执行出现脱节:85%组织计划三年内采用代理式AI,但76%表示当前运营和基础设施无法支持。
Haven’t you heard? White-collar jobs are going away, decimated by AI. Waves of layoffs in the tech sector (most recently at Coinbase and Meta and Cisco) are said to presage what will soon come for all of us knowledge workers. But before you quit your job as a software developer or financial analyst—
中文介绍 对AI导致白领工作消失的恐慌需现实检验:近期科技裁员(如Coinbase、Meta、思科)引发担忧,但结论尚不明确。
Artificial intelligence has not so far produced a clean story of mass unemployment. Aggregate employment in developed countries remains broadly stable, and recent assessments have found limited evidence that AI has shifted the headline numbers. But a troubling change may be hiding beneath the surfac
中文介绍 AI尚未引发大规模失业,发达国家就业总体稳定,但初级工作领域正面临潜在危机。
**Harness engineering** is emerging as the key differentiator for coding agents, emphasizing the stack of **model + harness + eval loop** over just stronger base models. **DeepSeek** is building a harness team to optimize interaction and verification loops, while **Google's Gemini Managed Agents** a
中文介绍 Harness工程成为编码代理的关键,DeepSeek正组建团队优化交互验证循环,Google的Gemini也在相关领域发展。
中文介绍 Grok Build CLI发布、AI硬件市场动态及教皇利奥对AI的警告。
第一作者: Richard J. Young · 方向: 软件安全
恶意代码评估共识标注提示库代码模型安全拒绝率基准
A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a working weapon -- a keylogger, a ransomware stub, an exploit that runs as written. This asymmetry in the severity of a single act of compliance implies coding-specialized models should clear a higher refusal bar than general-purpose chat models, not a lower one, yet the field cannot presently tell whether they do. Refusal benchmarks for malicious code are fragmented: they mix requests for executable software (ready-to-run weapons) with requests for harmful security knowledge (information a human must still operationalise) and report refusal rates over non-comparable corpora, so no single statistic measures the property that actually matters. This paper introduces an expanded consensus-labeled prompt bank that distinguishes between these two...
论文介绍 该论文指出,现有评估基准混淆了对可执行恶意软件(武器)与有害安全知识(信息)的请求,导致无法准确衡量代码模型的拒绝能力。为此,作者构建了一个经过共识标注的扩展提示库,明确区分这两类请求,旨在为代码专用模型建立更严格、更可比的安全评估标准,以应对其潜在的严重滥用风险。
第一作者: Davide De Zuane · 方向: 密码学协议
卫星通信密钥交换协议互联网密钥交换量子安全资源效率
This paper studies cryptographic key exchange in satellite communications, which requires specific solutions because the satellite context presents unique challenges, particularly concerning onboard resource constraints and long transmission latency. We address these challenges by considering the Internet Key Exchange (IKE) protocol, which is widely used in terrestrial networks, and studying its applicability in the satellite context. This requires addressing two main issues: i) its efficiency in terms of the resources and bandwidth required to adapt to satellite terminals, and ii) its resistance even to attackers equipped with a quantum computer, in order to resist obsolescence and defend against harvest-now-decrypt-later attacks. We study these aspects from both a design and experimental point of view, defining and assessing some protocol variants characterized by low complexity and...
论文介绍 该研究探讨了适用于卫星通信场景的高效且抗量子攻击的互联网密钥交换协议。针对卫星终端资源受限和长传输延迟的挑战,作者对现有IKE协议进行了改进设计,提出了旨在降低计算和带宽开销、同时抵御未来量子计算机威胁的协议变体,并从设计和实验两方面进行了评估,以应对「立即收集,日后解密」的攻击风险。
第一作者: Yanqiu Zhao · 方向: 系统安全
GUI代理隐私保护边缘计算隐私仲裁技能进化
GUI agents rely on screenshots to infer intent and operate across applications, but these screenshots often contain private messages, medical records, payment credentials, and workplace-specific workflows. Privacy decisions in this setting depend on task, recipient, application state, and user role, yet static PII detectors miss these boundaries and cloud-side VLM reasoning can upload the raw screen before deciding what should be protected. We present MaskClaw, an edge-side privacy arbitrator for GUI agents. MaskClaw extracts local visual evidence, retrieves user- and task-specific policy memory, and decides Allow, Mask, or Ask before raw screenshots leave a trusted user- or organization-controlled environment. In five designed skill-evolution scenarios, it turns corrections, cancellations, and edits into reusable privacy skills checked by a sandbox gate. We introduce P-GUI-Evo, a...
论文介绍 本文提出MaskClaw,一个用于GUI代理的边缘侧隐私仲裁器。它通过提取本地视觉证据并结合用户特定的策略记忆,在截图离开可信环境前决策允许、遮盖或询问,以保护敏感信息。该系统支持通过用户纠正等行为演化出可复用的隐私技能,旨在解决静态隐私检测器和云端推理带来的隐私泄露问题。
第一作者: Jinze Gu · 方向: 隐私保护
Graph RAG隐私攻击知识图谱结构重建黑盒攻击
Abstract:Retrieval-Augmented Generation (RAG) enhances LLMs by grounding generation in query-relevant external evidence. Beyond unstructured text corpora, Graph RAG integrates knowledge graphs into the retrieval pipeline, enabling LLMs to access entities, relations, and multi-hop dependencies encoded in structured knowledge. However, the same structured knowledge that empowers Graph RAG also creates a new privacy attack surface. We demonstrate that Graph RAG systems can be turned into structural oracles: through adaptive black-box interactions, an adversary can elicit sufficient relational evidence to reconstruct substantial portions of the hidden knowledge graph. We propose a structure-oriented reconstruction framework that recovers targeted graphs from both local and global perspectives. Specifically, Depth-Wise Heuristic Search extracts fine-grained node attributes by recursively...
论文介绍 该论文揭示了图检索增强生成系统的新隐私风险。研究表明,攻击者可以通过与Graph RAG系统的自适应黑盒交互,仅利用检索结果就能重构出目标知识图谱的相当部分。为此,作者提出一个从局部和全局视角进行结构重建的框架,将Graph RAG系统转变为可泄露隐藏关系的“结构预言机”。
第一作者: Ziyang You · 方向: AI 安全
LLM水印供应链攻击伪随机数生成器SeedHijack无痕攻击
Abstract:Cryptographic watermarking is a leading defense for attributing text generated by large language models (LLMs). Existing schemes, including KGW, Unigram, and DipMark, derive their security guarantees from the assumption that the underlying pseudo-random number generator (PRNG) is trustworthy. This work introduces SeedHijack, the first supply-chain attack on LLM watermarking that is simultaneously (i) blind -- requiring no knowledge of the watermark key, detector, or model logits, (ii) integrity-preserving -- amplifying rather than erasing the watermark signal, and (iii) orthogonal to detection -- the attack-induced bias is statistically independent of all content-side detector statistics, ensuring that amplification and evasion coexist without trade-off. Rather than perturbing generated text, SeedHijack replaces the PRNG at the supply-chain layer, biasing green-list selection...
论文介绍 本文提出SeedHijack,这是首个针对大语言模型水印技术的供应链攻击。该攻击通过替换底层的伪随机数生成器,在不了解水印密钥或检测器的情况下,能够隐秘地放大水印信号而非消除它,且其引发的偏差与检测统计量独立,从而同时实现了放大信号与规避检测,对现有水印方案的安全性假设构成根本性挑战。
第一作者: Jianwei Li · 方向: 安全研究
正向后门秘密对齐安全评估立场论文访问控制
Abstract:This position paper argues that the AI/ML community should stop overclaiming and retire the label "positive backdoor," and instead treat trigger-activated hidden behaviors as Secret Alignment. Crucially, protective claims based on Secret Alignment should be presumed not secure by default unless supported by rigorous, standardized evaluation. The Private AI era, enabled by open-weight LLMs and accessible training/inference stacks, turns language models into privately owned digital assets, creating security concerns around unauthorized access, model theft, and behavioral misuse. Recently, a line of work framed as "positive backdoors" has been proposed to address these challenges. To ground our position in evidence, we unify these proposals as covert trigger-behavior associations for access gating, ownership attribution, and safety enforcement, and evaluate three representative...
论文介绍 这篇立场论文主张,应停止使用“正向后门”这一过度声称的标签,并将此类触发式隐匿行为更准确地称为“秘密对齐”。作者统一评估了三种代表性方案,指出其保护性声明在未经严格、标准化评估验证前应默认被视为不安全,呼吁建立更严谨的评估体系以应对私人化AI时代下模型资产的安全关切。
第一作者: Luca Beurer-Kellner · 方向: 安全研究
AI代理技能市场安全分析威胁分类恶意负载
Abstract:We analyzed 3,984 AI agent skills from major marketplaces and found 76 confirmed malicious payloads, including credential theft, backdoor installation, and data exfiltration. 13.4% of all skills contain at least one critical-level security issue and at least 8 manually confirmed malicious skills remain publicly available on this http URL as of the date of publication. This report documents our methodology, presents a threat taxonomy based on real-world samples, and details the attack patterns we observed. As skill marketplaces grow rapidly and AI agents gain access to sensitive credentials and systems, automated security analysis is no longer optional.
论文介绍 本文通过对主流AI代理技能市场中近四千个技能的分析,发现了大量恶意载荷,包括窃取凭证、安装后门等行为。报告揭示了当前技能生态系统中存在的严重安全隐患,指出超过13%的技能包含关键安全问题,并基于真实样本提出了威胁分类法,强调随着代理获得敏感系统访问权,自动化安全分析已成为必要。
第一作者: Michael Külper · 方向: 系统安全
数字取证测试驱动ADARE框架状态转换测试可信取证
Abstract:Digital forensic relies on validated tools and established procedures, yet the underlying operating systems, applications, and analysis tools evolve rapidly. This evolution can cause artifact behavior and tool outputs to drift, silently degrading repeatability and confidence in long-lived forensic interpretations. We present test-driven forensics, a practical approach that treats forensic expectations as executable specifications: expected artifacts and expected tool outputs are encoded as tests that can be rerun across versions to detect regressions. Crucially, our approach also enables State Transition Testing, validating the system's expected state after each user action rather than only performing post-mortem checks on a final disk image; this supports causal attribution and makes transient behavior testable. We implement the methodology in ADARE, an open-source framework...
论文介绍 该论文提出“测试驱动取证”方法,以应对底层系统和工具快速演变导致的取证结果漂移与可信度下降问题。其核心是将预期的数字证据和工具输出编码为可执行的测试规范,并实现状态转换测试以验证用户操作后的系统预期状态。论文实现了开源框架ADARE,旨在提升桌面数字取证的可重复性和因果归因能力。
第一作者: Víctor Mayoral-Vilches · 方向: AI 安全
网络安全AI执行框架元脚手架基准测试
Abstract:What is the best harness for cybersecurity AI? Cybersecurity systems are converging on a single execution scaffold per agent, an iterative shell loop driven by a Large Language Model (LLM). However, scaffolds are not interchangeable, rarely interoperable, and no single scaffold dominates across all challenge types. In our path towards researching Cybersecurity SuperIntelligence (CSI), we present a meta-scaffold that unifies heterogeneous agent harnesses under a common orchestration layer, enabling any LLM-driven scaffold to be deployed, benchmarked, and composed within the same infrastructure. Using CSI, we benchmark five scaffolds (CSI::Claude, CSI::Codex, CSI::GCAI, CSI::Mistral, CSI::CAI) on the 33 cybench challenges, holding the model fixed at alias2-mini. The best individual scaffolds solve 15/33 (45.5%); the four-scaffold union solves 17/33 (51.5%), with the fifth...
论文介绍 该研究探讨网络安全AI中最佳执行框架问题,提出一个元脚手架来统一异构代理框架,允许部署、基准测试和组合不同的LLM驱动脚手架。通过在cybench挑战上测试五个脚手架,显示联合使用可提高挑战解决率,这有助于提升网络安全AI系统的灵活性和效率。
第一作者: Chenxi Wang · 方向: AI 安全
潜在攻击多智能体系统隐状态攻击框架
Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for efficient and flexible agent collaboration. However, moving coordination into latent space may also move attacks beyond the reach of visible-text inspection. In this paper, we study whether latent states can carry attack-associated information that remains effective during clean executions. To examine this question, we introduce a latent attack framework that reactivates attack-induced effects through latent interventions without reusing adversarial text. Extensive experiments show that the resulting latent-only attacks can substantially degrade task performance in clean executions, especially when applied to inter-agent KV-cache handoffs rather than local hidden states. Further control analyses indicate that this degradation cannot be reduced...
论文介绍 探索基于隐状态的多智能体系统中的潜在攻击风险,提出一个潜在攻击框架,通过隐状态干预在无对抗文本的情况下重新激活攻击效果。实验表明,潜在攻击能显著降低清洁执行中的任务性能,尤其是在智能体间KV缓存交接时,揭示了隐状态通信的安全隐患。
第一作者: Víctor Mayoral-Vilches · 方向: AI 安全
网络安全数据集LLM轨迹专家操作CAI数据集
Abstract:We present CAI Dataset, a fourteen-month corpus of cybersecurity LLM trajectories collected through the open-source CAI agent framework, built in response to PentestGPT's finding that expert operator trajectories, not base-model capability, are the bottleneck for cybersecurity LLM performance. CAI Dataset aggregates 230,935 session logs and 26,027,742 user prompts from 16,768 source IPs across 123 countries, exercising 4,187 unique LLM identifiers against 23,147 target domains over 18.07 TB of durable storage. The mix is hands-on (36.4% offensive, 20.1% attacker-intent, 27.5% business / integration, 4.4% defensive), making CAI Dataset, to the best of our knowledge, the largest described corpus of LLM-driven hacker trajectories. It is released to partner organisations and selected customers as an audience-size series (CAI Dataset10, CAI Dataset1k, CAI Dataset200k). Read...
论文介绍 介绍CAI数据集,一个历时十四个月收集的大型网络安全LLM轨迹语料库,包含超过23万次会话日志和2600万用户提示,覆盖123个国家。该数据集旨在研究专家操作轨迹作为LLM性能瓶颈的问题,可用于改进网络安全AI系统的训练和评估。
第一作者: Yubin Qu · 方向: AI 安全
过度渴望行为编码代理场景合成SNARE流水线
Abstract:A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while the task still completes. We call this overeager behavior: the prompt is not adversarial and the run succeeds, yet an out-of-scope step can leak credentials or delete files. Existing benchmarks miss it: task-completion suites credit any finished run, jailbreak suites probe adversarial prompts, and the one prior overeager benchmark applies a single fixed prompt set to every agent-model pair, leaving its easiest and most resistant pairs under-measured. We present SNARE (Synthesizing Non-adversarial scenarios for Adaptive Reward-guided Elicitation), a pipeline that composes benign scenarios from reusable scope and trap fragments, scores each run with a judge-free oracle flagging trap-pattern matches and unsolicited file additions or...
论文介绍 研究编码代理中的过度渴望行为,即任务完成但超出授权范围。提出SNARE流水线,通过合成良性场景自适应地引发此类行为,使用可重用片段构建场景并用无评判器预言家评分运行。这有助于评估和改进编码代理的安全性,避免潜在风险。
第一作者: Ruoqi Guo · 方向: AI 安全
提示注入移动GUI代理用户生成内容MIRAGE流水线
Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from what they see, so they cannot reliably separate trusted interface elements from user-generated content. We present MIRAGE (Mobile Injection of Realistic Adversarial GUI Examples), a pipeline that turns benign mobile screenshots into prompt-injection samples by placing attacker-controlled text into ordinary user-generated content regions, without modifying the agent, the application, or the operating system. MIRAGE operates in three stages: a Localizer identifies user-controllable regions on the screenshot, a Generator synthesises context-aware payloads and renders them in the application's native style, and a Curator moderates realism and balances the samples across applications, region types, and attack intents. A key challenge is that an...
论文介绍 针对视觉语言模型驱动的移动GUI代理,提出MIRAGE攻击流水线,通过将攻击者控制的文本嵌入用户生成内容区域进行上下文感知的提示注入。该流水线包括定位、生成和筛选阶段,无需修改代理或系统,揭示了移动代理在区分可信界面和用户内容时的脆弱性。
第一作者: Junjie Mu · 方向: 网络安全
路由劫持联邦RAG语义配置文件安全攻击
Abstract:Federated Retrieval-Augmented Generation (FedRAG) is attractive for privacy-sensitive applications because raw data remain local. As a result, routing must rely on client-provided semantic profiles, creating a new opportunity for manipulation. We introduce Routing Hijacking, a routing-stage attack in which a malicious client forges its profile to attract target queries despite having irrelevant underlying data. We show that this vulnerability is severe. Across three representative FedRAG routing architectures, Routing Hijacking consistently misroutes target queries and leads to downstream disruptions and failures, including missing evidence, poisoning, incorrect answers, and hallucinations. In a high-stakes MedQA-USMLE case study, we further show that poisoned retrieved evidence can mislead models across scales, leading to incorrect answers, hallucinations, and sycophantic...
论文介绍 揭示联邦检索增强生成(FedRAG)中的路由劫持漏洞,恶意客户端通过伪造语义配置文件吸引目标查询。实验显示,这种攻击在不同路由架构中一致导致错误路由,引发下游问题如证据缺失、幻觉和错误答案,特别是在高风险应用如医疗问答中。
第一作者: Huikang Liu · 方向: 隐私保护
差分隐私混合高斯机制噪声机制隐私保护
Abstract:We design a class of additive noise mechanisms that satisfy \((\varepsilon, \delta)\)-differential privacy (DP) for scalar, real-valued query functions with known sensitivities, with a particular focus on moderate and low-privacy regimes. These mechanisms, which we call \textit{mixture mechanisms}, are constructed by mixing multiple Gaussian distributions that share the same variance but differ in their means and mixture weights. The resulting distributions can be interpreted as convex combinations of a zero-mean Gaussian (as used in the analytic Gaussian mechanism) and additional Gaussians whose means depend on the sensitivity of the query function. We derive tight conditions on the variances required for \((\varepsilon, \delta)\)-DP and provide efficient algorithms to compute them. Compared to the analytic Gaussian mechanism, our mechanisms yield substantially lower expected...
论文介绍 提出一类混合高斯噪声机制,用于满足(ε,δ)-差分隐私,特别关注中等和低隐私场景。这些机制通过混合多个共享方差但均值和混合权重不同的高斯分布构建,与解析高斯机制相比能显著降低预期噪声水平,适用于标量实值查询函数。
第一作者: Jiachen Qian · 方向: AI 安全
数据投毒RAG系统对抗攻击语义保持
Retrieval-Augmented Generation (RAG) mitigates LLM hallucinations but introduces a critical vulnerability: corpus integrity. We present SilentRetrieval, a two-stage data poisoning attack that hijacks RAG systems through adversarially crafted yet fluent documents. Stage 1 uses Coordinated Beam Search, a multi-token joint optimization method with a fluency-similarity objective, to keep a poisoned host document retrievable while constraining perplexity. Stage 2 uses Context-Adaptive Trigger Generation, a lightweight trigger-fusion step driven by a frozen LLM, to integrate manipulation triggers into document content. Under a one-poisoned-document-per-query evaluation with synthetic target answers, SilentRetrieval achieves 84.6%/81.3% HR@10 and 57.5%/54.8% ASR-LLM on Natural Questions and MS MARCO, while maintaining near-benign perplexity. Cross-model evaluation across four target LLMs...
论文介绍 介绍SilentRetrieval攻击,通过两阶段数据投毒劫持检索增强生成(RAG)系统。第一阶段使用协调波束搜索优化中毒文档的可检索性和流畅性;第二阶段使用上下文自适应触发生成整合操纵触发器。攻击实现高命中率和成功率,揭示了RAG系统的语料库完整性漏洞。
第一作者: Jiaqi Luo · 方向: 软件安全
LLM智能体访问控制工具调用安全属性基控制客户端-服务器架构
LLM-based agents have recently attracted significant attention due to their ability to autonomously invoke relevant tools to accomplish complex tasks. However, recent studies have shown that these agents face severe security risks, which may lead to privacy leakage, financial loss, or even full system compromise. In this paper, we present AgentGuard, an attribute-based access control framework for tool-use LLM-based agents. AgentGuard adopts a client-server architecture. On the client side, AgentGuard provides lightweight integration for agents implemented in different programming languages and architectures. It requires only minor code modifications (e.g., around 10 lines) without changing the underlying agent execution logic. On the server side, AgentGuard provides three complementary inspection mechanisms to cover both single-tool and cross-tool security risks in agent execution. In...
论文介绍 本文针对基于大语言模型(LLM)的智能体在调用外部工具时面临的安全风险,提出了一种名为AgentGuard的属性基访问控制框架。该框架采用客户端-服务器架构,客户端侧提供轻量级集成,只需少量代码修改;服务器侧包含三种互补的检查机制,用于覆盖单工具和跨工具场景下的安全风险,旨在增强智能体的工具调用安全性,防止隐私泄露或系统被破坏。
第一作者: Yu Yin · 方向: AI 安全
提示注入攻击检索增强生成攻击有效性现实评估流水线
Abstract:Recent generative engine optimisation (GEO) research has shown that prompt-injection attacks can push a target product to the top of an LLM's recommendation list, with the strongest attacks reporting around $80\%$ success and raising serious security concerns about RAG-based recommendation. However, these results assume the attacked document is always fed directly to the generator, bypassing the retriever and reranker. This is unrealistic: in deployed RAG systems, the attack modifies the document content, which can in turn change whether the document is retrieved and reranked highly enough to reach the generator at all. In this paper, we re-evaluate seven GEO attacks under a realistic three-stage pipeline (retriever\,$\to$\,LLM reranker\,$\to$\,LLM generator). We find that prior protocols substantially overstate attack effectiveness: gradient-based and instruction override...
论文介绍 本文研究了提示注入攻击在现实的检索增强生成(RAG)系统中的实际生存能力。现有研究假设攻击文档可直接输入生成器,但现实系统包含检索、重排序和生成三个阶段。作者在这一现实三阶段流水线中重新评估了七种生成式引擎优化攻击,发现之前的评估方法严重高估了攻击效果,基于梯度和指令覆盖的攻击在复杂流水线下有效性显著降低。
第一作者: Gavin Brown · 方向: 安全研究
差分隐私单调统计量样本复杂度运行时间优化查询复杂度下界
Abstract:We study efficient differentially private algorithms for estimating monotone statistics, i.e., statistics that are monotone under the addition of new observations. The starting point for our investigation is subsample-and-aggregate: a classical paradigm that partitions the dataset into blocks, estimates the statistic on each block, and then privately aggregates the this http URL practical and generically applicable, this approach is quite data-hungry. We improve upon this framework for the class of monotone statistics -- compared to subsample-and-aggregate, our algorithms save a factor of $t$ in sample complexity and pay a factor of $e^t$ in running time, where $t>0$ is a tunable parameter. We complement our results with a query-complexity lower bound, showing that our algorithms are essentially optimal for this task. As an application, we obtain improved results for private...
论文介绍 本文研究用于估计单调统计量的高效差分隐私算法。研究从经典的“子样本并聚合”范式出发,该方法较为耗费数据。针对单调统计量这一类特殊问题,作者改进了框架,实现了在样本复杂度上节省一个因子t,而在运行时间上增加一个因子e^t的权衡,其中t是可调参数。作者还提供了查询复杂度下界,证明算法近乎最优,并展示了在私有数据发布等应用上的改进。
第一作者: Nick Merrill · 方向: 安全研究
自省适配器安全攻击对称性审计
Abstract:We demonstrate an attack on Introspection Adapters (Shenoy et al., 2026).
论文介绍 本文演示了针对「自省适配器」的一种攻击方法。该攻击旨在绕过或击败安全审计机制,揭示了现有安全适配器可能存在的安全漏洞。研究直接展示了攻击的可行性,为相关防御系统的改进提供了参考。
第一作者: Kai Chen · 方向: AI 安全
成员推理攻击聊天智能体记忆安全多召回探针黑盒白盒
Abstract:Membership inference attacks (MIAs) test whether a target data record belongs to a system's private data, and have become a standard tool to measure privacy leakage in machine learning systems. Prior work has primarily focused on training corpora or retrieval databases. However, MIAs against agent memory have received less attention, even though such memory can contain sensitive user-agent interactions, retrieved facts, and user preferences. Therefore, in this work, we focus on chat agent memory MIAs, where an adversary infers whether a candidate memory unit belongs to the chat agent's memory store. We propose Multi-Recall Memory MIA (MRMMIA), a unified attack that utilizes multiple recall probes to the agent to extract the membership signal across black-box, gray-box, and white-box settings. Our experiments demonstrate that MRMMIA consistently outperforms baselines. Our...
论文介绍 本文聚焦于针对聊天智能体记忆的成员推理攻击。现有研究多关注训练语料库或检索数据库,但智能体记忆可能包含敏感的用户交互信息。作者提出了多召回记忆成员推理攻击(MRMMIA),它利用多个针对智能体的召回探针,在黑盒、灰盒和白盒设置下提取成员信号。实验表明,MRMMIA能持续优于基线方法,可用于评估智能体记忆的隐私泄露风险。
第一作者: Xiang Fang · 方向: AI 安全
对抗提示语义解缠图分类LLM安全防御框架
Abstract:Large Language Models (LLMs) are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulting in harmful or inappropriate outputs. Such attacks, including jailbreaking and prompt injection, pose significant risks to the integrity and availability of LLMs in security-critical applications. This paper proposes the Adversarial Prompt Disentanglement (APD) framework, a novel defense mechanism that proactively identifies and neutralizes malicious components in input prompts before they are processed by the LLM. The APD framework integrates three key innovations: (1) a mutual information-based semantic decomposition method to isolate adversarial and benign prompt components, ensuring statistical independence; (2) a graph-based intent classification approach that leverages spectral analysis to detect malicious patterns in...
论文介绍 本文针对大语言模型(LLM)易受对抗性提示(如越狱和提示注入)攻击的问题,提出了对抗提示解缠(APD)防御框架。该框架旨在LLM处理输入前,主动识别并中和其中的恶意组件。其核心创新包括基于互信息的语义分解方法,以及基于图的意图分类方法,通过谱分析检测恶意模式,从而提升LLM在安全关键应用中的鲁棒性。
第一作者: Yuxin "Myles" Liu · 方向: 软件安全
热补丁汽车微控制器安全标准ISO 26262Execute-in-Place
Abstract:The increasing presence of software in modern automobiles has created a growing need to deliver software updates throughout a vehicle's entire lifespan. Traditional update methods are slow and require months of re-validation to comply with stringent safety standards like ISO 26262. Although hotpatching offers a path to faster updates, existing solutions for real-time embedded systems are unsuitable for the automotive domain: they overlook regulatory compliance, demand extensive safety validation, and lack support for the flash-based Execute-in-Place (XIP) architecture commonly used in automotive electronic control units (ECUs). We introduce Patchlings, the first hotpatching framework designed for compliance, safety, and persistence in automotive systems. It fills the gap in applying hotpatching to automotive systems and fundamentally reduces the mean-time-to-mitigate (MTTM)...
论文介绍 本文针对汽车电子控制单元(ECU)的软件更新难题,提出了Patchlings框架。传统更新方法速度慢,需要漫长的重新验证。现有热补丁方案不适用于汽车领域,因为它们忽视法规合规性,且不支持常见的基于闪存的XIP架构。Patchlings是首个兼顾合规性、安全性和持久性的汽车热补丁框架,旨在缩短漏洞的平均缓解时间,并满足ISO 26262等安全标准。
第一作者: Kaustav Goswami · 方向: 软件安全
RowHammer系统级建模gem5可靠性安全缓解措施评估
Abstract:Modern architecture research relies on simulators to evaluate system security, yet analyzing emerging hardware vulnerabilities like RowHammer requires full-system visibility. As RowHammer vulnerabilities worsen with continuous technology scaling, existing simulators lack the system-level models needed to study complex OS effects and cross-layer mitigations. This tool deficiency leaves modern computing platforms exposed to severe reliability and security risks. In this work, we present HammerSim, a gem5-based framework for modeling RowHammer at the full-system level. HammerSim integrates probability-driven bitflip modeling to realistically capture the behavior of RowHammer. It further enables evaluation of hardware and software mitigations such as TRR and selective ECC. We validate HammerSim's bitflip modeling against real DDR4 DIMMs using JS divergence, demonstrating its...
论文介绍 本文介绍了HammerSim,一个基于gem5的系统级框架,用于建模RowHammer硬件漏洞。现有模拟器缺乏系统级模型来研究复杂的操作系统效应和跨层缓解措施。HammerSim集成了概率驱动的位翻转建模,以真实捕捉RowHammer行为,并支持评估诸如TRR和选择性ECC等硬件和软件缓解措施。该工具的位翻转建模已通过与真实DDR4内存模组的对比进行验证。
第一作者: Loay Abdelrazek · 方向: 系统安全
安全本体描述逻辑自动化推理威胁中和5G/6G安全
Abstract:Modern 5G-Advanced and emerging 6G cloud-native telecom architectures encounter unprecedented hyper-complexity, multi-layered threat vectors, and fluid structural topologies. Managing infrastructure security using manual, imperative configurations introduces a severe latency gap, presenting attackers with an exploitable window. This paper presents a declarative, autonomous, self-protecting framework based on our design and standardization of the TM Forum TR292I Security Ontology v4.0.0. Our approach leverages Description Logic (DL) and automated graph reasoning within a closed-loop execution pipeline to dynamically neutralize live threats. Crucially, the system balances functional protection expectations with non-functional resource impact considerations (e.g., latency vs. compute overhead). We validate our model-driven architecture through a structural formal verification...
论文介绍 针对5G-Advanced和6G云原生电信架构的超复杂性和安全威胁,传统手动管理存在延迟漏洞。本文提出基于TM Forum TR292I安全本体v4.0.0的声明式自主框架,利用描述逻辑和自动化图推理,在闭环执行管道中动态中和实时威胁,并平衡保护期望与资源影响。该模型驱动架构通过形式化验证,为安全管理提供自动化解决方案。
第一作者: Dongping Liu · 方向: AI 安全
量子随机性身份签名AI安全开源平台量子计算
Abstract:The 2024--2025 Nobel and Turing awards recognised artificial intelligence and quantum science in the same breath -- machine learning as a physical science, artificial intelligence solving 50-year scientific problems, superconducting quantum circuits as the hardware foundation of quantum computing, and quantum information principles as computing's highest achievement. Yet no deployed artificial intelligence system has brought these two streams together for the general public: identity systems still rely on pseudo-random tokens, and quantum circuits remain invisible to the billions of people who use bot-enabled social messaging platforms daily. This paper presents QSignAI, a production-deployed open-source platform demonstrating a bidirectional relationship between artificial intelligence and quantum science in a real-time event participation system. We address three research...
论文介绍 当前身份系统依赖伪随机令牌,量子科学尚未融入大众AI应用。本文介绍QSignAI开源平台,利用量子随机性种子生成身份签名,在实时事件参与系统中展示AI与量子科学的双向关系。该平台为融合AI和量子技术于现实应用提供了实际部署案例,推动身份验证安全性提升。
第一作者: Zongheng Cao · 方向: 安全研究
AI代理基准测试多模态视频制作性能评估
Abstract:Video production workflows offer a rich and demanding arena for evaluating multimodal AI agents: they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use. To this end, we introduce AgenticVBench, a benchmark of 100 agentic tasks across 4 task families spanning the real world post-production workflow, constructed from real production workflows contributed by 20 industry experts averaging 6 years of professional experience. Tasks are paired with evaluation specifications that combine programmatic verifiers and expert rubrics. We evaluate frontier vision-language models (VLMs) with both vendor-native and open-source harnesses. The best evaluated agent stack barely crosses 30%, far below human expert performance on the same tasks. We further find that the choice of harness substantially affects model...
论文介绍 评估多模态AI代理在真实视频制作后流程中的能力存在挑战。本文提出AgenticVBench基准,涵盖100个任务,基于20位行业专家的真实工作流构建,结合程序化验证和专家评分。评估显示最佳AI代理性能远低于人类专家,揭示了当前技术的局限性,为改进AI代理提供评估框架。
第一作者: Abile Jean · 方向: AI 安全
后门攻击网络物理系统故障检测对抗性机器学习安全漏洞
Cyber-Physical Systems (CPS) integrate sensing, communication, computation, and control to support critical infrastructure, including smart grids, industrial automation, and control systems. In the electrical utility domain, various controllers are used in CPS to ensure the system detects and recovers from faults, such as voltage fluctuations, and to perform load balancing in distribution systems. Machine learning- and deep learning-based fault detection and localization frameworks have recently gained significant attention in CPS for their ability to identify anomalies and operational failures in real time. However, these intelligent models are vulnerable to adversarial machine learning attacks, particularly backdoor attacks. In a backdoor attack, an adversary injects malicious patterns into the training data so that the model behaves normally most of the time but produces...
论文介绍 网络物理系统中基于机器学习的故障检测框架易受后门攻击威胁。本文研究攻击者如何在训练数据中注入恶意模式,使模型在特定触发下产生错误输出,影响系统安全。分析揭示了此类攻击对关键基础设施如智能电网的风险,强调需要增强防御措施。
第一作者: Olawale Amos Akanji · 方向: 系统安全
Android权限安全风险权限组自定义权限恶意软件
Abstract:Android's permission system is designed to balance usability with informed consent, yet two legacy mechanisms still undermine that balance in Android 16: (i) permission groups that silently auto-grant new permissions within a group after a user's initial approval, and (ii) normal-level custom permissions that are auto-granted at install and enable cross-app access with no user visibility. We conduct a longitudinal analysis of 19.3 million APKs spanning 5.97 million unique apps (distinct package identifiers) from the AndroZoo repository, combined with on-device validation on Android 16. Among 2,244,575 multi-version apps, 381,026 (17%) silently gain permissions within already-granted groups. Using VirusTotal detections with primary threshold t=20, apps flagged as malware expand within groups at a higher rate than benign apps (odds ratio = 1.35, p < 0.001); the association holds...
论文介绍 Android权限系统的权限组和自定义权限机制可能导致无意识授权和持续风险。本文对1930万APK进行纵向分析,结合Android 16设备验证,发现恶意软件利用这些机制悄然扩展权限。研究提供了数据证据,为Android安全设计和权限管理改进提供参考。
第一作者: James Bartusek · 方向: 密码学协议
不可克隆加密密码学对称加密伪随机酉信息安全
In this note, we consider the setting of uncloneable encryption satisfying uncloneable indistinguishability, a form of symmetric key encryption that prevents the cloning of ciphertexts in a very strong sense. Our goal is to minimize the assumptions under which (many-time secure) uncloneable encryption is known to exist, assuming the existence of an information-theoretic "uncloneable bit", i.e. a one-time secure uncloneable encryption scheme for one-bit messages. We observe that if a t -> t' uncloneable bit exists, then the following implications hold. 1. If many-time secure symmetric key encryption exists, then many-time secure t -> t' uncloneable encryption for arbitrary-length messages exists. Since many-time secure uncloneable encryption implies many-time secure symmetric key encryption, this result is tight. 2. If pseudorandom unitaries exist, then many-time secure t -> t'...
论文介绍 本文旨在最小化不可克隆加密方案存在的假设条件。假设信息论的「不可克隆位」存在,推导出多次安全对称加密方案的存在性,并探讨伪随机酉的影响。该理论工作为密码学协议设计提供基础,可能增强安全通信的隐私保护。
第一作者: Khang Tran · 方向: 软件安全
后门攻击代码大语言模型模型投毒隐蔽触发软件安全
Abstract:Code Large Language Models (CLLMs) serve as the core of modern code agents, enabling developers to automate complex software development tasks. In this paper, we present Poison-with-Style (PwS), a practical and stealthy model poisoning attack targeting CLLMs. Unlike prior attacks that assume an active adversary capable of directly embedding explicit triggers (e.g., specific words) into developers' prompts during inference, PwS leverages developers' code styles as covert triggers implicitly embedded within their prompts. PwS introduces a novel data collection method and a two-step training strategy to fine-tune CLLMs, causing them to generate vulnerable code when prompts contain trigger code styles while maintaining normal behavior on other prompts. Experimental results on Python code completion tasks show that PwS is robust against state-of-the-art defenses and achieves high...
论文介绍 代码大语言模型作为现代代码代理的核心,易受后门攻击。本文提出Poison-with-Style攻击方法,利用开发者代码风格作为隐蔽触发器,通过数据收集和两步训练策略微调模型,使其在特定风格下生成漏洞代码。该攻击在Python任务中表现稳健,揭示了代码模型的安全威胁。
第一作者: Samuel Heuchert · 方向: 软件安全
CMMC认证评估角色期望现象学分析网络安全
Abstract:The Cybersecurity Maturity Model Certification program requires third-party assessments be conducted under a non-consultative model. The model is intended to ensure impartiality for organizations seeking certification. While this structure defines expectations for assessor behavior, assessor experiences and interpretations of these constraints remain underexamined. The study examines the lived experiences of CMMC-Certified Assessors and how they navigate role expectations within the non-consultative model. Using Role Conflict Theory as a guiding framework, Interpretative Phenomenological Analysis (IPA) was applied to semi-structured interviews to explore how assessors make sense of their roles. The analysis identified experiential themes that describe how assessors construct professional credibility, execute structured assessment work, and manage the practical challenges of...
论文介绍 CMMC二级认证评估要求第三方在非咨询模型下进行,但评估者的体验和角色解读研究不足。本文基于角色冲突理论,采用解释现象学分析半结构化访谈,探讨评估者如何构建专业可信度、执行结构化评估并应对实践挑战。研究为评估标准制定和培训提供见解。
第一作者: Onur Eren Arpaci · 方向: 系统安全
ORAM访问模式保护时间局部性启发式优化加密云存储
Abstract:Encrypted cloud storage can hide data contents but still leak sensitive information through access patterns. ORAM addresses this by hiding access patterns, but existing ORAM systems are too inefficient to deploy in practice. We present Cloak, an oblivious storage system that dramatically improves performance by leveraging a simple, widely observed property of real workloads: temporal locality, where recently accessed items are more likely to be accessed again soon. Instead of trying to make server accesses look perfectly uniform, Cloak makes server traffic follow a fixed, "recentness-biased" pattern and then uses real queries to fill as much of that traffic as possible. When the workload exhibits temporal locality, Cloak achieves overheads as low as $1.1\times$ over a non-oblivious and unencrypted baseline. Importantly, this heuristic affects only performance, not security. We...
论文介绍 加密云存储虽能保护数据内容,但访问模式仍可泄露敏感信息。ORAM通过隐藏访问模式来解决此问题,但现有系统效率低下。本文提出Cloak系统,利用真实工作负载中普遍存在的「时间局部性」特性进行启发式优化,让服务器流量遵循一个固定的、偏向近期访问的模式,并用真实查询填充。该方法在保持安全性的前提下,当负载具有时间局部性时,性能开销可低至基准的1.1倍。
第一作者: Yogesh Kumar · 方向: 安全研究
差分密码分析线性层MDS矩阵AES安全证明
In AES-like ciphers, diffusion layers are commonly instantiated using MDS matrices, since their optimal branch number yields strong diffusion guarantees and underpins classical resistance arguments against differential and linear cryptanalysis. However, Daemen and Rijmen (2009) showed that linear layers may still exhibit related-differential structure beyond what the MDS criterion captures, and Bardeh and Rijmen (2022) demonstrated that this phenomenon can be exploited in attacks on reduced-round AES. In this work, we systematically investigate the conditions under which linear layers avoid or exhibit these differentials, identifying matrix classes for which such structure is unavoidable. We first prove that every non-MDS matrix admits a nontrivial pair of related differentials, showing that the MDS property is necessary for avoiding them. We then establish that every odd-order...
论文介绍 在AES类密码中,扩散层通常由MDS矩阵实现以抵抗差分和线性分析。然而,线性层仍可能表现出超越MDS准则的相关差分结构,进而被攻击利用。本文系统研究了线性层何时能避免或暴露此类差分。研究证明,每个非MDS矩阵都存在非平凡的相关差分对,表明MDS性质是避免它们的必要条件,并进一步确定了不可回避此类结构的奇数阶矩阵类别,增强了对密码分析的理解。
第一作者: Syed Huma Shah · 方向: AI 安全
检索增强生成答案缓存缓存路由安全性验证GroundedCache
Abstract:Modern retrieval-augmented generation(RAG) deployments increasingly rely on caching to reduce token cost and time-to-first-token(TTFT). Prefix-level KV reuse is now standard in serving stacks such as vLLM, and chunk-level and position-independent reuse have been pushed further by recent systems(RAGCache, TurboRAG, CacheBlend, EPIC, ContextPilot, PCR, LMCache). Output-level semantic answer caches, by contrast, remain fragile: similar prompts can map to different correct answers, retrieved evidence drifts as the corpus is updated, and adversarial collision attacks have been shown to hijack cached responses. We argue that the right framing for cached answer reuse is not how to reuse faster but when reuse is safe. We propose GroundedCache, an evidence-validated cache router that admits a cached answer only when 4 cheap gates simultaneously hold: query similarity...
论文介绍 检索增强生成系统依赖缓存来降低成本和延迟,但输出级别的语义答案缓存容易因提示相似、证据漂移或对抗攻击而导致错误重用。本文认为问题的核心在于判断「何时重用是安全的」。作者提出GroundedCache,这是一个基于证据验证的缓存路由器,仅当四个廉价门控条件同时满足时(如查询相似度),才接受缓存的答案,旨在提升RAG系统缓存重用的可靠性和安全性。
第一作者: Md Hafizur Rahman · 方向: AI 安全
多智能体系统LLM安全危害放大HARP对齐风险
Abstract:Multi-agent LLM systems decompose workflows across agents, tools, shared context, memory, and decision gates. This modularity improves interpretability, but creates a propagation risk: a bounded perturbation to one component can be reused by other agents and amplified into system-level harm. We introduce HARP (Harm Amplification through Role Perturbation), a trace-first methodology for studying local-to-global harm amplification in multi-agent LLM systems. HARP compares paired clean and perturbed executions and records specialist outputs, tool calls, memory reads/writes, guard events, oracle logs, latency, token cost, and decisions. We define local harm as deviation from targeted agents or corrupted channels, global harm as deviation over the full trace, and harm amplification as (H_global/H_local). This complements attack success rate with a measure of how strongly...
论文介绍 多智能体LLM系统的模块化结构可能引发危害传播风险:对某个组件的有限扰动,可能被其他智能体重用并放大为系统级危害。为研究这种从局部到全局的危害放大,本文引入HARP方法。它通过比较干净和受扰动的执行轨迹,记录工具调用、内存读写等信息,定义并衡量局部与全局危害及其放大程度,为评估复杂智能体系统的安全性提供了新的度量视角。
第一作者: Qiancheng Wu · 方向: 软件安全
智能体系统eBPF沙箱安全认证通道访问控制
Abstract:Agentic systems increasingly run user-authored orchestration code that invokes tools, spawns subtasks, and delegates work across machines and clouds. Although this high agency is productive, it creates a security problem: identity, authorization, provenance, and delegation are often pushed into application code, where they become difficult to enforce consistently and difficult to audit. We present \emph{Grimlock}, an \emph{Agent Guard} that restores separation of concerns by moving trust enforcement into the sandbox substrate while leaving agent code unchanged. Grimlock uses \emph{eBPF-enforced traffic interception} to ensure that sandbox communication passes through a guard, and combines it with \emph{post-handshake attestation} bound to standard TLS~1.3 channel bindings. After a channel is established, the guard authorizes communication and mints short-lived, channel-bound...
论文介绍 高能动性智能体系统允许用户编排代码调用工具并跨机器工作,但将身份、授权等安全机制推入应用代码难以统一实施和审计。本文提出Grimlock,一种「智能体守卫」。它利用eBPF进行流量拦截,确保沙箱通信经过守卫,并结合TLS 1.3的握手后认证,将信任执行下沉到沙箱基板,从而在不修改智能体代码的情况下,恢复关注点分离,增强安全强制执行与可审计性。
第一作者: Matteo Gioele Collu · 方向: AI 安全
LLM安全拒绝行为中间激活线性探针攻击优化
Abstract:In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations at each transformer block. We find that refusal is linearly decodable well before the final layer, indicating that safety-relevant behavior is represented in intermediate activations before output generation. To test whether this signal is actionable, we introduce Mechanistic AutoDAN, a probe-guided variant of AutoDAN that replaces full-model fitness evaluation with partial forward passes and probe-based scoring inside a genetic prompt search loop. Across the evaluated models, our method achieves attack success rates competitive with vanilla AutoDAN while reducing per-iteration search time by up to 72%, and probe-guided prompts match or exceed AutoDAN's cross-model transfer in several...
论文介绍 本文探索是否能在大语言模型输出解码前,从其内部激活预测拒绝行为。研究发现,使用线性探针在残差流上训练,拒绝信号在最终层之前的中间层就已被线性解码,表明安全行为在输出生成前就已表征。基于此,作者提出了Mechanistic AutoDAN,用探针引导的搜索替代完整的模型评估,在保持攻击成功率的同时,显著提升了对抗性提示搜索的效率。
第一作者: Onur Günlü · 方向: 网络安全
6GISAC空间感知隐私保护物理层安全
Abstract:Integrated sensing and communication (ISAC) is a promising feature of future communication networks. While spatial sensing can improve network performance and enable external services, it also creates privacy challenges that go beyond the confidentiality of communication content. Future networks using millimeter-wave (mmWave) and sub-terahertz (THz) frequencies may collect or infer detailed information about people, devices, bystanders, passive objects, and environments in a sixth-generation (6G) deployment area. Such sensing can reveal location and environment data, support behavioral profiling such as movement or activity recognition, and, in advanced cases, expose physiological information such as breathing frequency or heart-rate-related data. Thus, the capabilities of spatial sensing must be controlled to satisfy privacy requirements. In this work, we organize...
论文介绍 通感一体化是未来6G网络的特征,毫米波和太赫兹频段的空间感知虽能提升网络性能,但也带来超越通信内容保密性的隐私风险,可能泄露位置、环境甚至生理信息。本文系统梳理了6G场景下ISAC带来的隐私挑战,包括对人员、设备、环境的详细信息收集或推断,并探讨了如何控制感知能力以满足隐私要求,为未来网络设计提供了重要考量。
第一作者: Shuhao Chen · 方向: AI 安全
LLM安全有害微调安全对齐SPARD数据选择
Abstract:Fine-tuning large language models often undermines their safety alignment, a problem further amplified by harmful fine-tuning attacks in which adversarial data removes safeguards and induces unsafe behaviors. We propose SPARD, a defense framework that integrates Safety-Projected Alternating optimization with Relevance-Diversity aware data selection. SPARD employs SPAG, which optimizes alternatively between utility updates and explicit safety projections with a set of safe data to enforce safety constraints. To curate safe data, we introduce a Relevance-Diversity Determinantal Point Process to select compact safe data, balancing task relevance and safety coverage. Experiments on GSM8K and OpenBookQA under four harmful fine-tuning attacks demonstrate that SPARD consistently achieves the lowest average attack success rates, substantially outperforming state-of-the-art defense...
论文介绍 微调大语言模型常常会破坏其安全对齐,有害微调攻击通过注入对抗数据移除安全防护。本文提出SPARD防御框架,它交替进行效用更新和显式的安全投影优化,并利用一个安全数据集来强制施加安全约束。为筛选安全数据,作者引入了基于相关多样性的行列式点过程方法,以平衡任务相关性与安全覆盖面,实验证明该方法能有效降低攻击成功率。
第一作者: J. Vijayavallabh · 方向: 安全研究
预算审计锚定解码KL散度代理支出比率安全研究
Abstract:We empirically audit the k-NAF budget-accounting mechanism in Anchored Decoding using (i) a fixed, class-stratified workload (approximately 8,500 randomized executions across six prompt classes) and (ii) an adaptive prompt-search procedure targeting high proxy spend ratios. On the fixed workload, mean cumulative KL spend remains far below the sequence-level budgets K in {600, 1000}, and an empirical Bernstein-style proxy stays below K for every class; surface-overlap diagnostics (ROUGE-L and 5-gram Jaccard) are correspondingly small. Adaptive search increases the proxy spend ratio but does not produce clear budget exhaustion. On a held-out copyright-domain workload at k = 3, several prompts exhibit proxy ratios above 1 under early-stopped evaluations with small realized sample sizes; re-evaluating the same prompts with larger allocation reduces the proxy ratio to the range...
论文介绍 本文对“锚定解码”中k-NAF预算计算机制进行了实证审计。研究通过固定和自适应的负载,考察序列级KL散度预算与代理支出比率。结果表明,在大部分情况下预算未耗尽,表明机制具有保守性,为理解和评估解码算法的安全属性提供了实证依据。
第一作者: Yuan Tian · 方向: 安全研究
多模态越狱视觉语言模型安全鲁棒性图像工具交互
Abstract:Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response generation, text-only prior turn, visual-state manipulation, and explicit external image-tool invocation. In this paper, we ask which of these evaluated paradigms improves multimodal jailbreak robustness, and why. Across multiple vision-language models, explicit image-tool interaction yields the lowest attack success rates in our experiments, reducing jailbreak success by around 30% relative on average across the evaluated models. This finding is initially surprising: ASR remains low even when the returned image-tool output is manually overridden or itself unsafe-looking, but returns near direct-answering levels under text-only prior turn...
论文介绍 本文探究了“思考伴随图像”这一新兴推理范式对大型视觉语言模型安全性的影响。研究比较了多种推理范式,发现显式调用外部图像工具的范式能显著降低越狱攻击成功率。这一结果揭示了在推理过程中引入外部工具对增强模型安全鲁棒性的潜在作用。
第一作者: Qiyuan Wang · 方向: 安全研究
后门攻击样本特异性密度感知对抗鲁棒性
Abstract:Despite recent progress in backdoor attacks, existing methods remain susceptible to post-training defenses that erase the backdoor through fine-tuning or pruning. We revisit the core objectives of backdoor attacks and derive principled criteria characterizing optimal sample-specific trigger construction under a Bayes-optimal model of the victim's training. Our analysis reveals that both attack success and clean-accuracy preservation are simultaneously optimized when triggered samples are steered into low-density regions of the clean data distribution, a distributional condition that controls all moments of the poisoned distribution at once rather than a handful of input-space summary statistics. We introduce a bilevel optimization framework that estimates density ratios via conditional time-score matching and optimizes a mixture-model objective to place triggered samples in...
论文介绍 现有后门攻击易被微调或剪枝等训练后防御措施清除。本文提出了一种密度感知的样本特异性攻击方法。核心思想是通过优化,将带有触发器的样本引向干净数据分布的低密度区域,从而同时实现攻击成功与保持模型在干净样本上的准确率,增强了攻击的鲁棒性。
第一作者: Shantanu Sharma · 方向: 隐私保护
隐私法规GDPR研究盲点地域差异
Abstract:Personal data has emerged as a highly valuable yet sensitive asset that drives business decisions, enables targeted advertising, and generates substantial revenue for companies, while simultaneously facilitating invasive monitoring of users. In recent years, research on digital privacy violations, including undue access, collection, and sharing of user data, has grown significantly. Much of this research adopts the European General Data Protection Regulation (GDPR) as the primary reference framework. This is reasonable, as GDPR was a pioneering legislation, and many of its stipulations are clear and unambiguous. However, we argue that focusing solely on GDPR (and a small set of other Western regulatory frameworks) ignores privacy-related concerns, attitudes, and problems faced by users from other locales, creating a significant research blind spot. This work systematically...
论文介绍 本文指出,当前数字隐私研究过度集中于欧盟《通用数据保护条例》等西方框架,忽视了其他地区用户面临的独特隐私关切与问题,形成了研究盲点。作者主张采用系统化、多地域的视角来全面审视全球化的数据隐私治理挑战。
第一作者: Yvonne Zhou · 方向: 密码学协议
全同态加密差分隐私机器学习训练收敛保证
Abstract:We present the first theoretical convergence analysis of machine learning training under fully homomorphic encryption (FHE), combined with a differentially private (DP) training algorithm tailored to encrypted computation. Our approach improves computational efficiency over standard differentially private gradient descent (DP-GD) while achieving comparable utility. In particular, we prove convergence of approximate gradient descent using polynomial approximations of activation and loss functions, which are required for FHE compatibility. To preserve privacy in downstream tasks, we integrate differential privacy without relying on costly per-sample gradient clipping, enabling scalable encrypted learning. We also provide data-independent hyperparameter selection and theoretically grounded strategies for polynomial approximation which can be of independent interest. Together...
论文介绍 本文首次给出了在全同态加密环境下进行机器学习训练的理论收敛分析,并设计了一种适配加密计算的差分隐私训练算法。该方法在保持与标准算法相当效用的同时提升了计算效率,并提供了无需昂贵逐样本梯度裁剪的隐私保护方案。
第一作者: Meghana Bhange · 方向: AI 安全
算法公平集体行动测试时干预代理扰动
Abstract:When machine learning systems under-perform for particular subgroups, affected users typically have no way to correct these disparities without relying on platform-level fixes. Existing approaches to algorithmic fairness rely on provider-centric approaches to correct these failures, leaving users with no external lever when faced with harm. Recent work in Algorithmic Collective Action shows that coordinated users can steer an algorithmic system toward a collective goal, but the existing mechanisms require the provider to retrain on the collective's modified data which users may not have control over. We propose Test-Time Collective Action (TTCA), a framework through which a group of users who share query access to the platform, can correct disparities affecting under-served subgroup without participating in the platform's training loop. We implement this through a proxy-based...
论文介绍 当机器学习系统对某些群体表现不佳时,用户通常无法自行纠正。本文提出“测试时集体行动”框架,允许享有查询访问权限的用户群组,通过基于代理的扰动来共同纠正模型对服务不足子群的偏差,而无需干预平台的训练过程。
第一作者: Yuhao Li · 方向: 软件安全
自动化做市商MEV机制设计权衡不可能性
Abstract:Blockchains have popularized the Automated Market Makers (AMMs), where users trade crypto-assets directly with a smart contract, governed by a pricing function embedded in the contract's code. Today, users of AMMs are often forced to accept unfavorable prices due to widespread front-running and back-running attacks, commonly known as Miner Extractable Value (MEV). Several earlier works show impossibility results suggesting that completely removing MEV at the consensus layer is impossible, partly because the consensus layer is agnostic of application-level semantics. For this reason, more recent works have advocated mechanism design approaches at the application (i.e., smart contract) level. We study a natural two-asset AMM mechanism design problem recently initiated and explored in prior work by Chan, Wu, and Shi, in which they proposed a mechanism that satisfies a...
论文介绍 本文研究了区块链自动化做市商的机制设计问题,旨在缓解MEV攻击。研究表明,在常见的双资产AMM设计中存在一个三难困境,即价格平滑性、交易公平性和MEV移除这三个目标无法同时被满足,揭示了该领域固有的设计权衡。
第一作者: Ching-Chun Chang · 方向: AI 安全
合成信息隐写术信息溯源AI生成内容
Abstract:The origin of species has been the mystery of mysteries in natural science. By analogy, the origin of synthetic information, we suggest, is the mystery of mysteries in information science. The question carries a moral weight that a technical account can neither fully resolve nor responsibly ignore, as its impact on truth, trust, and human intellect extends deep into the broader economy and society. The very power of artificial intelligence makes the evolutionary lineage of synthetic information grow ever harder to trace, for a sufficiently capable model may generate offspring that bear little resemblance, at either the structural or signal level, to the parent source from which they were derived. As in genetics, two individuals may share the same phenotype mirroring each other in outward appearance, yet differ fundamentally in their genotype. We propose, by means of...
论文介绍 本文类比生物学的“物种起源”,提出合成信息的“起源”是信息科学中的核心谜题。作者探讨了AI生成信息的血统追溯难题,由于模型能力强,生成内容可能与源头在结构上差异巨大。本文提出利用隐写术继承的方法,来追踪合成信息的演化谱系。
第一作者: Xinyu Wang · 方向: VLA 通用模型 · 来源: cs.CV
视觉-语言-动作模型模型量化后训练量化复合旋转均匀精度
Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and diffusion-based action heads make on-device deployment prohibitively expensive. Prior quantization efforts offer only partial solutions, compressing the LLM backbone while leaving the DiT action head at full precision, or resorting to mixed-precision schemes, driven by the belief that uniformly quantizing the action head is inherently unstable. We challenge this assumption with Omega-QVLA, the first training-free post-training quantization framework that compresses both the language backbone and the entire diffusion action head of a VLA model to a uniform W4A4 precision, eliminating the need for mixed-precision allocation. Omega-QVLA combines a composite SVD-Hadamard rotation that equalizes per-channel weight energy while diffusing...
论文介绍 视觉-语言-动作模型整合感知、推理与控制,但多参数骨干和扩散动作头使设备部署成本高昂。现有量化方法仅部分压缩模型。本文提出Ω-QVLA,首个免训练后训练量化框架,将语言骨干和整个扩散动作头统一压缩到W4A4精度,采用复合SVD-Hadamard旋转和步进缩放技术,消除混合精度分配需求,提升模型在边缘设备的部署效率。
第一作者: Yongchen Wang · 方向: VLA 通用模型 · 来源: cs.RO
磁驱动微机器人双臂协调VLA模型LoRA适应微操作
Magnetically actuated microrobots have been used as wireless, non-contact manipulation tools at microscales, making them promising for minimally invasive applications. However, their control remains challenging due to indirect actuation, limited sensing, and nonlinear magnetic interactions. In this work, we propose Mag-VLA, a vision-language-action (VLA) model for dexterous magnetic microrobot manipulation using two robotic arms with mounted magnets for dynamic magnetic-field construction. Bimanual coordination enables capabilities such as microrobot reorientation that are difficult or infeasible with a single arm, but it also introduces coupled control challenges, as the policy must generate coordinated trajectories for both actuators within a shared workspace. Our framework adapts a Qwen2.5-VL-7B backbone using Low-Rank Adaptation (LoRA) to process visual observations and language...
论文介绍 磁驱动微机器人在微创应用中前景广阔,但控制因间接驱动和非线性磁场而困难。本文提出Mag-VLA,一个视觉-语言-动作模型,用于双臂磁驱动微机器人操作。通过双臂协调构建动态磁场,基于Qwen2.5-VL-7B骨干和LoRA适应,处理视觉观测和语言指令,生成协调轨迹,实现微机器人重定向等任务。
第一作者: Junhwi Cho · 方向: 具身智能 · 来源: cs.RO
机器人皮肤电阻抗层析气动触觉力图重建3D打印
We present a hybrid robotic skin that combines electrical impedance tomography (EIT) with pneumatic tactile sensing to improve force reconstruction capability. The developed robotic skin is fabricated entirely by 3D printing and spray coating, making it affordable and easy to build. A Tikhonov-regularized inverse reconstruction, paired with per-pad pneumatic calibration, enables accurate large-area tactile sensing with a simple measurement scheme. For validation, we conducted load-cell indentation experiments; the results showed consistent force reconstruction across locations within a pad. Compared with an EIT-only baseline, sensitivity non-uniformity was also reduced, with the coefficient of variation decreasing from 0.31 to 0.14, indicating that the proposed approach addresses a longstanding limitation of EIT. We further demonstrated chest-mounted integration on a humanoid robot and...
论文介绍 开发混合机器人皮肤,结合电阻抗层析成像和气动触觉传感,通过3D打印和喷涂制造,降低成本并易于构建。采用Tikhonov正则化逆重建和每垫气动校准,实现大面积准确力重建。实验显示灵敏度非均匀性降低,应用于人形机器人胸部集成,提升触觉感知能力。
第一作者: Marcell Fekete · 方向: 多模态具身 · 来源: cs.CL
视觉语言模型信息结构话语压力句法位置模型输出
Vision-language models (VLMs) are increasingly evaluated for whether they identify the right visual content, but little is known about whether they express such content in a discourse-appropriate form. We address this research gap using information structure (IS), testing whether VLMs distinguish discourse-old Topics from discourse-new Foci in visually grounded question answering. We exploit Hungarian, a language in which Topic and Focus map onto dedicated syntactic positions, making IS choices observable in text. Comparing six VLMs with human participants, we find that models produce IS-relevant constructions, but over-regularise this sensitivity. Under the interacting pressures of discourse status, grammatical role (preference for subject Topics) and definiteness (preference for indefinite Foci), humans choose variable strategies for IS realisation. VLMs, by contrast, collapse onto...
论文介绍 研究视觉语言模型在视觉问答中的话语适当性,利用信息结构理论测试模型是否区分话语主题和焦点。以匈牙利语为例,主题和焦点映射到特定句法位置。比较模型与人类,发现模型过度规则化,而人类策略灵活,揭示模型在话语压力下的行为差异。
第一作者: Mazen Alamir · 方向: 具身智能 · 来源: cs.RO
分段多项式关系稀疏建模时间序列分析机械臂逆模型异常定位
This paper addresses the problem of identifying parsimonious explicit piece-wise polynomial relationships that might involve a relatively large number of raw features. The algorithm leverages a recently proposed identification algorithm that yields parsimonious implicit relationships enabling to derive normality characterization in the context of anomaly detection and localization. The algorithm proposed in this paper goes a step further by deriving explicit piece-wise representations that are built using the set of polynomials involved in the implicit representations. The framework is illustrated on the problem of identifying parsimonious explicit representations of the inverse model of a 6-axis manipulator robot. Moreover, further experiments on a 4-axis robot are also shown which are designed to investigate the generalization capability of parsimonious models compared to...
论文介绍 针对工业时间序列,提出算法识别稀疏显式分段多项式关系。从隐式关系导出显式表示,应用于6轴机械臂逆模型建模,并在4轴机器人上测试泛化能力。该方法有助于异常检测和定位,提升工业机器人控制精度。
第一作者: Arianna Alonso Bizzi · 方向: 具身智能 · 来源: cs.RO
事件相机光流估计FPGA实现流式速度估计硬件高效
Event-based vision sensors offer asynchronous, high-temporal-resolution measurements that are attractive for low-latency robotic perception, but many event-based motion estimation methods are computationally intensive and difficult to map to FPGA hardware. We present a streaming velocity estimator that discretizes asynchronous events into fixed-duration time bins, constructs a 1-bit spatial occupancy grid, and evaluates multiple velocity hypotheses in parallel using only fixed-width integer logic - shift registers, counters, comparators, and small LUT-mapped multiplies - with no dividers and no DSP blocks. It requires no frame reconstruction, no floating-point arithmetic, and no iterative optimization. The method deliberately trades dense sub-pixel optical flow for a sparse, quantized velocity estimate at each active pixel, suited to low-latency tasks such as reactive obstacle...
论文介绍 事件相机提供异步高时间分辨率数据,但光流计算密集。本文提出流式速度估计器,将事件离散化为时间箱,构建1位占用网格,使用整数逻辑并行评估速度假设。仅用移位寄存器和计数器,无浮点运算,适合FPGA硬件实现。牺牲亚像素精度换取稀疏量化估计,适用于低延迟反应任务如避障。
第一作者: Xue Qin · 方向: 具身智能 · 来源: cs.RO
金丝雀部署身份稳定性安全关键系统具身智能体运行时治理
Canary deployment routes a fraction of traffic to a new software version, monitors metrics, and rolls back on regression. Mainstream controllers (Argo Rollouts, Spinnaker, Flagger) change the deployed system's cryptographic identity during the canary window. The drift is harmless for stateless microservices but breaks the claim that "the agent you certified is still the agent you have" for safety-critical embodied agents, forcing re-certification per canary. We present ICAN-Deploy (Identity-stable CANary Deployment), a middleware construction whose state machine holds the identity hash invariant across the canary window by separating capability names (frozen, hashed) from capability versions (mutable runtime state). We implement ICAN-Deploy inside a runtime governance layer for LLM-driven robots and verify invariance by closed-form proof, AST lint, and TLA+ model-checking, then...
论文介绍 主流金丝雀部署在窗口期改变系统加密身份,对安全关键具身智能体需重新认证。ICAN-Deploy提出中间件方案,通过分离能力名称和版本,保持身份哈希不变。应用于LLM驱动机器人的运行时治理层,并通过形式化验证确保身份稳定性,提升部署安全性。
第一作者: Ha Sier · 方向: 机器人操作 · 来源: cs.RO
视觉地点识别保形预测跨条件适应安全验证序列匹配
Sequence-based visual place recognition (VPR) for SLAM and robot relocalization must decide whether the retrieved top-1 candidate is safe to accept. Conformal prediction is a natural framework for this accept/reject decision, but its finite-sample guarantees rely on exchangeability between calibration and deployment (test) data, which is violated under cross-condition deployment. We introduce SAFEVPR, a non-trainable verification-and-calibration pipeline for safe cross-condition sequence VPR. SAFEVPR replaces the standard backbone cosine similarity with a mutual-nearest-neighbour (MNN) patch-matching score computed from frozen DINOv2 ViT features, and replaces flat Learn-Then-Test calibration with Mondrian conformal LTT, fitting separate Bonferroni-corrected thresholds across score bins. Under exchangeability, these thresholds would provide finite-sample false-discovery-rate (FDR)...
论文介绍 序列视觉地点识别在SLAM中需安全接受决策,但保形预测在跨条件下假设违反。SAFEVPR引入非训练验证校准管道,使用DINOv2特征的互近邻匹配和Mondrian保形LTT校准,提供有限样本错误发现率保证,提升跨条件部署的安全性。
第一作者: Zhe Zhang · 方向: 机器人操作 · 来源: cs.RO
接触丰富操作级联优化主动接触选择机器人操作
We propose an optimization-based framework for robust contact-rich manipulation. Recent contact-implicit methods enable online hybrid planning across contact modes, allowing closed-loop manipulation for a given target state and contact location sequence of the robot and object. However, most existing approaches lack the ability to autonomously reason and generate diverse contact location sequences and manipulation trajectories, i.e., active contact location selection, which limits their applicability to relatively simple tasks. Active contact location selection is challenging due to complementarity in contact dynamics and the sparse gradients, making the design of a unified framework for contact selection and planning difficult. To address these challenges, we introduce Simultaneous Contact Selection and Planning (SCSP), a cascaded optimization framework comprising Contact Selection...
论文介绍 本文提出一个基于优化的框架,用于鲁棒的接触丰富机器人操作。现有接触隐式方法缺乏自主生成多样化接触位置序列的能力,即主动接触选择,限制了其在复杂任务中的应用。该研究引入了「同时接触选择与规划」框架,通过级联优化来联合求解接触选择与操作规划问题,从而提升了对更复杂操作任务的适用性。
第一作者: Qiwei Wu · 方向: VLA 通用模型 · 来源: cs.RO
视觉-触觉-语言模型温柔操作力反馈基准测试
Tactile sensing is essential for robots to achieve human-like gentle manipulation. However, existing Vision-Language-Action (VLA) models struggle to exploit tactile feedback for gentle manipulation due to scarce aligned vision-tactile-language data and the lack of effective closed-loop force feedback mechanisms. To address these challenges, we introduce Tabero, a benchmark and model suite for gentle, language-conditioned robotic manipulation that demands fine-grained contact force perception. First, the Tabero benchmark addresses the scarcity of tactile data by presenting a data-efficient pipeline that repurposes open-source robot manipulation trajectories to generate diverse vision-tactile-language tasks, and establishes a multidimensional evaluation protocol that measures task success alongside physical interaction quality. Second, we propose Tabero-VTLA, an architecture with a...
论文介绍 针对现有视觉-语言-动作模型难以利用触觉反馈实现温柔操作的问题,本文提出了Tabero基准与模型套件。该工作通过一个高效的数据生成管线,将开源机器人操作轨迹转化为多样化的视觉-触觉-语言任务,以解决对齐数据稀缺的问题。同时,提出了Tabero-VTLA架构,集成了闭环力反馈机制,以提升语言引导下的精细操作能力。
第一作者: Xucheng Wang · 方向: 模仿学习 · 来源: cs.RO
模仿学习外科手术缝合跟随策略评估
Abstract:This study presents the first evaluation of general-purpose imitation learning for surgeon-robot collaborative assistance in open surgery, targeting suture following: the grab-pull-release motion an assistant performs at every stitch. We collect 160 teleoperated demonstrations (32,374 frames) on an open-source robot arm, benchmark four architecturally diverse imitation learning policies (ACT, Diffusion Policy, SmolVLA, $\pi_0$) across 28 trained models evaluated in 32 configurations along three clinically motivated dimensions: dataset size, camera viewpoint, and background variation. Our results demonstrate that under ideal conditions, the four policies achieve $50$-$75\%$ task success, with depth error as the dominant failure mode across all architectures. Among all policies, $\pi_0$ achieves the strongest results with a pretrained vision-language backbone, demonstrating...
论文介绍 本研究首次评估了通用模仿学习策略在开放外科手术中用于医生-机器人协作辅助的可行性,具体任务为缝合跟随。研究收集了遥操作演示数据,并对四种不同架构的模仿学习策略(如ACT、扩散策略、π0)进行了多维度的基准测试,评估了数据集大小、摄像机视角和背景变化等因素对性能的影响。
第一作者: Krishnam Gupta · 方向: VLA 通用模型 · 来源: cs.RO
视觉-语言-动作模型架构差异故障模式动作监控
Abstract:We discover that VLA architectures fail in fundamentally different, predictable ways at the motor-command level. Running VQ-BeT, Diffusion Policy, and ACT on identical evaluation protocols (n=450 episodes across PushT and ALOHA 14-DOF bimanual manipulation), we find: (1) direction reversal rate is a universal failure predictor across all three architectures (AUROC=0.93, 0.79, 0.91; p<0.001); (2) jerk monitoring is predictive only for discrete-token architectures, following a discrete-to-continuous gradient (0.88, 0.69, 0.41); (3) velocity violations alone are non-predictive everywhere (AUROC 0.41-0.69), yet velocity checking is the most common safety mechanism in VLA deployment code; and (4) for continuous-family VLAs, velocity monitoring provides effectively zero predictive signal (AUROC=0.52 on ACT, 0.41 on Diffusion), proving that architecture-matched monitor selection is...
论文介绍 本研究发现,不同的VLA架构在电机指令层面会以根本性、可预测的方式失败。通过在相同评估协议下测试VQ-BeT、扩散策略和ACT,作者发现方向逆转率是所有架构的通用故障预测器,而其他指标如加加速度和速度的监测则因架构不同而具有不同预测效力。这表明,为安全部署VLA选择与架构匹配的监控机制至关重要。
第一作者: Yutai Li · 方向: VLA 通用模型 · 来源: cs.RO
视觉-语言-动作模型运动基元可复用泛化性
Abstract:Vision-Language-Action (VLA) models offer a promising paradigm for generalist robotic policies, yet their adaptation is hindered by data inefficiency and poor generalization. We argue that these bottlenecks stem from the prevailing Direct Instruction-to-Control Mapping, which forces models to memorize monolithic trajectories rather than reusable motion patterns, i.e., primitives. We propose PrimitiveVLA, a framework that shifts this paradigm toward a Primitive-Centric Disassemble & Assemble paradigm. Supported by a shared Multimodal Canonical Representation (MCR), PrimitiveVLA unifies two phases: (1) Fine-tuning-phase Disassembly, which uses an automated pipeline to disassemble demonstrations into reusable primitives; and (2) Inference-phase Assembly, which employs a VLM-based planner and an LLM-generated switch module for robust closed-loop execution. By disassembling tasks...
论文介绍 本文提出PrimitiveVLA框架,旨在解决VLA模型数据效率低和泛化性差的问题。该框架从「直接指令到控制」范式转向「以基元为中心」的范式,其核心是将任务演示自动拆解为可复用的运动基元。在推理阶段,系统使用基于VLM的规划器组装这些基元以执行新任务,从而提高了学习效率和策略的泛化能力。
第一作者: Jiachen Zhang · 方向: VLA 通用模型 · 来源: cs.RO
VLA策略价值结构线性探针表征分析
Abstract:Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Their frozen representations nevertheless carry such information, and it can be read out and used to guide action choice without retraining the policy. From mixed successful and failed manipulation trajectories on LIBERO-Goal, we recover Monte-Carlo outcome targets using lightweight linear probes on frozen features. The targets are consistently predictable from OpenVLA, Pi0.5, DINOv2, and CLIP features, and substantially less so from baselines built on progress, time-to-go, task identity, or proprioception. To rule out task and temporal shortcuts, we evaluate the probes under same-task, same-timestep matched comparisons: Pi0.5 probes still reach roughly 92% pairwise ordering accuracy, while label-shuffled controls stay at chance...
论文介绍 本研究探究了冻结的视觉-语言-动作策略模型是否已内嵌关于任务成功的信息。作者在LIBERO-Goal数据集上,对来自OpenVLA、Pi0.5等模型的冻结特征使用轻量级线性探针来预测蒙特卡洛结果目标。结果表明,这些价值类信息在模型表征中是可读取的,且能以高准确率对成功与失败轨迹进行排序,这为无需重训练即可引导策略动作提供了可能性。
第一作者: Zongcai Tan · 方向: 机器人操作 · 来源: cs.RO
数字孪生视觉触觉遥操作光学镊子微机器人操作
Abstract:Optical tweezers (OT) provide piconewton-scale manipulation for delicate biomedical tasks, where visuo-haptic feedback can improve operator awareness by conveying interaction-force cues and trap-stability information. However, visuo-haptic teleoperation frameworks for complex-shaped optical microrobots remain underdeveloped, particularly in multi-trap manipulation scenarios. This paper presents a digital twin framework for virtual visuo-haptic teleoperation of complex-shaped OT-driven microrobots. The framework integrates a digital twin environment, image-based pose and depth estimation, microrobot motion simulation, and model-based haptic rendering within a Robot Operating System (ROS)-connected bimanual teleoperation system. For force modeling, we combine a Multi-Sphere Distributed Manipulation (MSDM) model with optical-force estimation from the Optical Tweezers Toolbox...
论文介绍 针对复杂形状光学微机器人遥操作中视觉触觉反馈框架不足的问题,本文提出了一个数字孪生框架。该框架集成了数字孪生环境、基于图像的位姿与深度估计、微机器人运动仿真以及基于模型的触觉渲染,构成了一个通过机器人操作系统连接的双手遥操作系统,为精细生物医学任务提供力觉线索与捕获稳定性信息。
第一作者: Julia Hindel · 方向: 导航与运动 · 来源: cs.RO
自监督学习可穿越性估计开放世界环境在线学习
Abstract:Self-supervised online traversability estimation enables robots to continuously learn from unlabeled open-world experiences and adapt their navigation behavior toward safe and efficient trajectories. Existing approaches either rely on handcrafted proprioceptive traversability scores, limiting robot-agnosticism, or cluster prior data, preventing online learning. Moreover, many continual learning methods incur substantial memory and computational costs, hindering onboard deployment. We introduce COTRATE, an online learning framework for continuous traversability estimation from multimodal, unlabeled robot experience. Our method first infers robust traversability scores using a robot-agnostic, learning-based online terrain assessment module operating on proprioceptiveand inertial signals. These scores then supervise a visual traversability network through a novel alignment loss...
论文介绍 本文提出COTRATE框架,用于从多模态、未标记的机器人经验中进行连续的可穿越性在线估计。该方法首先使用基于本体感觉和惯性信号的模块推断与机器人无关的地形可穿越性分数,然后通过新颖的对齐损失来监督视觉可穿越性网络。该框架使机器人能够从开放世界经验中持续学习,并适应其导航行为以获得安全高效的轨迹。
第一作者: Junha Min · 方向: 具身智能 · 来源: cs.RO
触觉-本体感觉传感器融合接触力估计人机交互时间卷积网络 (TCN)
Abstract:Direct physical guidance is a natural means of teaching and interacting with robots, and robotic skins make a key contribution by enabling sensitive contact sensing and localization. This paper presents a tactile-proprioceptive sensor fusion framework for natural physical human-robot interaction. Tactile cues from pneumatic skin pads serve as contact indicators that bypass the ambiguity between frictional residues and applied external forces, enabling highly sensitive contact detection without explicit friction identification. We fuse these cues with motor-current-based proprioception to reconstruct multi-axis contact forces on the robot surface. To maintain accuracy during motion, we employ a temporal convolutional network (TCN) to mitigate friction hysteresis during stick-slip transitions, reducing uncertainty at contact onset and yielding smooth, responsive guidance. We...
论文介绍 该研究提出了一种触觉-本体感觉传感器融合框架,用于全身物理人机交互中的接触力估计。通过气动皮肤垫的触觉线索作为接触指示器,绕过摩擦残留与外力之间的模糊性,实现高灵敏度接触检测。结合基于电机电流的本体感觉,重建机器人表面的多轴接触力。为维持运动中的准确性,采用时间卷积网络(TCN)来缓解粘滑转换中的摩擦滞后,减少接触开始时的不确定性,产生平滑、响应迅速的引导。
第一作者: Faisal Lawan · 方向: 具身智能 · 来源: cs.RO
安全关键控制自适应阻抗控制非光滑控制障碍函数约束处理模糊逻辑系统
Abstract:Safe physical interaction is critical for deploying robotic manipulators in human-robot interaction and contact-rich tasks, where uncertainty, external forces, and actuator limitations can compromise both performance and safety. We propose an online adaptive impedance control framework that enforces joint-state safety while achieving compliant interaction under uncertain dynamics. The approach combines a quadratic-program-based safety filter with a novel composed position-velocity non-smooth control barrier function (NCBF), enabling joint position and velocity constraints to be enforced through a unified relative-degree-one barrier. Unknown dynamics are compensated online using an interval type-2 fuzzy logic system, while actuator torque limits are handled through soft constraints with exact penalty recovery of feasible solutions. A disturbance-observer-enhanced safety...
论文介绍 该研究提出一种在线自适应阻抗控制框架,用于机器人操作器在不确定动态下的安全物理交互。方法结合基于二次规划的安全滤波器和新型非光滑控制障碍函数,强制执行关节位置和速度约束。未知动力学通过区间二型模糊逻辑系统在线补偿,执行器扭矩限制通过软约束处理。增强安全性的扰动观测器确保在状态和输入约束下的合规交互。
第一作者: Zhanzheng Ma · 方向: 导航与运动 · 来源: cs.RO
机器人路径规划区域提议网络连通性保持搜索空间压缩可变形注意力变换器
Abstract:Mobile robot path planning methods are often constrained by vast search spaces, resulting in latency in samplingbased algorithms. Learning-based approaches frequently suffer from local region fragmentation and global topological inconsistency. To tackle the problem, we present the Connectivity- Preserving Region Proposal Network (CP-RPN), a segmentationguided model designed to predict compact and topologically connected candidate regions, significantly compressing the search space. Specifically, we design a segmentation model that leverages a Deformable Attention Transformer (DAT) to capture long-range dependencies for global connectivity, with a Deconvolutional decoder to preserve fine-grained spatial details. To guarantee the connectivity of the predicted mask, we design a composite loss function that combines Cross-Entropy loss for pixelwise supervision, a...
论文介绍 针对移动机器人路径规划中搜索空间大、学习方法易产生区域碎片化的问题,本研究提出连通性保持区域提议网络(CP-RPN)。该模型通过分割引导预测紧凑且拓扑连通的候选区域,显著压缩搜索空间。具体设计使用可变形注意力变换器捕获长程依赖,解卷积解码器保留空间细节,并通过复合损失函数确保预测掩码的连通性,从而加速路径规划过程。
第一作者: Yunseong Bang · 方向: 具身智能 · 来源: cs.RO
软机器人皮肤磁体传感3D打印触觉超分辨率卷积神经网络
Abstract:This paper presents a magnet-based robotic skin that integrates a multilayer soft lattice with distributed Hall-effect sensor arrays and a tactile super-resolution model. External contact forces are converted to magnetic field changes by embedded permanent magnets, and the lattice spreads these changes across the sensing domain. This gives each sensor a large, overlapping receptive field and enables a large sensing area with minimal blind spots. Lattice parameters are tunable, enabling joint adjustment of mechanical compliance and transduction characteristics. An implicit modeling workflow and selective laser sintering (SLS) 3D printing support rapid fabrication of conformal, high-complexity structures. A convolutional neural network trained on experimental measurements estimates contact location and normal force in real time. Experiments validate localization accuracy and...
论文介绍 本研究提出一种基于磁体的软机器人皮肤,结合3D打印的多格结构和CNN触觉超分辨率。通过嵌入永磁体将外力转换为磁场变化,多层软格子扩展传感域,减少盲点。格子参数可调,以控制机械顺应性和传导特性。使用隐式建模和SLS 3D打印快速制造共形结构。CNN模型基于实验数据训练,实时估计接触位置和法向力,实现高精度触觉传感。
第一作者: Mirado Mortel · 方向: 导航与运动 · 来源: cs.RO
自然运动机器人运动被动动力学环境交互运动原理
Abstract:Robotic locomotion can become efficient when mechanisms exploit passive dynamics, compliance, and resonance rather than track prescribed trajectories. This paper formulates natural locomotion as an exchange principle for systems whose motion is mediated by environmental constraints or interactions. A motion is natural when an internal oscillator returns periodically, the body pose drifts, and the mean Propulsion--Oscillator Exchange power (POE power) vanishes over one cycle. The selected family is a Natural Locomotion Manifold (NLM). We develop the conservative realization of this principle for continuous ideal environmental constraints: the constraints do no external work, total mechanical energy is conserved, and zero mean POE power is an internal exchange with the environment-mediated propulsive channel, not external energy input. The method is a closed/open construction...
论文介绍 该研究将自然运动公式化为一个交换原理,适用于运动由环境约束或交互介导的系统。当内部振荡器周期性返回、身体姿态漂移且平均推进-振荡器交换功率为零时,运动被视为自然。通过自然运动流形(NLM)描述系统,并发展了保守实现方法,适用于连续理想环境约束,其中约束不做外功,总机械能守恒。
第一作者: Seungsu Kim · 方向: VLA 通用模型 · 来源: cs.RO
机器人操作视觉-语言-动作模型进度感知离线强化学习长序列处理
Abstract:We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The model specifically focuses on efficiently processing long multi-modal sequences by maintaining an explicit representation of task progress over extended horizons. To this end, ProgVLA integrates two key components. First, a multi-modal encoder with a two-stage Perceiver resampling scheme compresses variable-length visual, language, and proprioceptive streams into a fixed set of control-ready context tokens, substantially reducing sequence length while preserving cross-modal grounding. Second, an auxiliary set of progress heads is trained with offline reinforcement learning (RL) objectives to jointly learn critics over normalized remaining-horizon targets. This provides the policy with an internal estimate of task progress and...
论文介绍 本研究提出ProgVLA,一个紧凑的视觉-语言-动作模型,用于在有限计算资源下进行机器人操作。模型通过多模态编码器和两阶段Perceiver重采样,将视觉、语言和本体感觉流压缩为固定控制上下文标记。辅助进度头基于离线强化学习训练,提供任务进度的内部估计,增强策略在长时程任务中的可靠性。
第一作者: Kisang Park · 方向: 机器人操作 · 来源: cs.RO
轨迹优化自然函数梯度碰撞避免函数空间高斯平滑
Abstract:Generating collision-free and smooth motions remains a central challenge in robotic manipulation, particularly in cluttered environments and narrow passages where feasible regions are highly constrained and fragmented. We propose a trajectory optimization framework that performs geometry-aware updates directly in function space using natural functional gradients. The method optimizes a Gaussian-smoothed surrogate objective that regularizes the optimization landscape through smooth trajectory perturbations while preserving trajectory-level structure. Because the updates are defined intrinsically in function space, trajectory regularity can be controlled independently of a particular time discretization. We derive a practical Monte-Carlo estimator of the natural functional gradient that requires only black-box trajectory evaluations, making the method applicable when analytic...
论文介绍 针对机器人操作中碰撞避免和平滑运动的挑战,本研究提出基于自然函数梯度的轨迹优化框架。方法在函数空间中进行几何感知更新,优化高斯平滑代理目标以正则化优化景观。通过蒙特卡洛估计器近似自然函数梯度,仅需黑盒轨迹评估,使得轨迹规律性可独立于时间离散化控制,适用于受限和碎片化的可行区域。
第一作者: Guangyang Zeng · 方向: 导航与运动 · 来源: cs.RO
SLAM不确定性量化多面体安全关键应用可证明保证
Abstract:In safety-critical robotics applications, guaranteed and practical uncertainty quantification (UQ) in perception is vital. Many existing works either offer no formal containment guarantee, rely on restrictive modeling assumptions, or focus only on pose estimation rather than a complete SLAM pipeline. This paper presents provably guaranteed UQ algorithms for 3D-3D landmark-based SLAM. The algorithms consist of three basic UQ modules: forward UQ for mapping, backward UQ for pose tracking, and pose compound. Each module produces a certified uncertainty set; when the input uncertainty bounds are deterministic, the output sets inherit deterministic guarantees, i.e., they provably contain the true poses and landmarks. Specifically, we use polytopes to represent uncertainty sets, enabling tractable computations and a unified treatment of pose uncertainty. To enhance algorithms'...
论文介绍 本研究为基于地标的3D SLAM提出可证明保证的不确定性量化算法。算法包含前向UQ用于映射、后向UQ用于姿态跟踪和姿态组合模块,每个模块产生认证不确定性集。使用多面体表示不确定性集,实现可处理计算和统一对待姿态不确定性。当输入不确定性边界确定时,输出集提供确定性保证,确保包含真实姿态和地标。
第一作者: Vinh Nguyen · 方向: 导航与运动 · 来源: cs.RO
自主移动机器人模拟到现实导航系统自我定位
Abstract:With the rapid development of simulation tools, the development and validation of autonomous robotic systems have become more efficient before real-world deployment. This paper presents a simulation-to-real implementation of an autonomous mobile robot based on an existing mechanical platform. Instead of focusing on mechanical design, our work concentrates on the development of the onboard control, self-localization, and autonomous navigation system. The proposed robot is equipped with onboard sensing and computation to estimate its pose and navigate autonomously in the environment. The overall framework is first developed and tested in simulation, and then deployed on the real robot for experimental evaluation. The results demonstrate the feasibility of the proposed approach and show that simulation provides an effective foundation for developing reliable autonomous mobile...
论文介绍 本文介绍「STR机器人」的设计,一种从模拟到现实的自主移动机器人系统。研究专注于开发车载控制、自我定位和自主导航模块,首先在模拟环境中进行整体框架测试,然后部署到真实机械平台进行实验评估。结果表明,模拟工具为构建可靠的自主移动机器人提供了有效基础。
第一作者: Saki Hashimoto · 方向: 具身智能 · 来源: cs.RO
对象所有权推断上下文感知大语言模型不确定性量化
Abstract:Service robots must infer object ownership to correctly interpret instructions such as "bring me my cup." However, ownership is a latent attribute that cannot be directly observed, and existing methods often rely on limited cues such as recent usage, making them unreliable in scenarios such as temporary sharing. We propose a framework for context-aware ownership inference with uncertainty-guided interaction (COIN). The method integrates user background information and object usage history using a large language model (LLM) to estimate ownership scores. To handle uncertainty, we apply conformal prediction to construct a set of plausible owners and selectively generate user queries when the prediction is uncertain. Experiments in a simulated home environment show that the proposed method consistently outperforms baseline approaches, achieving a Subset Accuracy of 0.988 and a...
论文介绍 提出「COIN框架」,用于服务机器人上下文感知的对象所有权推断。该方法整合用户背景和对象使用历史,通过大语言模型估计所有权分数,并利用共形预测处理不确定性,自动生成用户查询。实验在模拟家庭环境中显示,该方法在准确性上优于基线,提升了机器人对指令如“把我的杯子拿来”的理解能力。
第一作者: Petr Vanc · 方向: 机器人操作 · 来源: cs.RO
机器人教学运动学指导操纵杆遥操作手势识别
Abstract:Instructing robots from demonstrations can be done through different teaching modalities, each with different usability and performance trade-offs. This paper compares kinesthetic guidance, joystick teleoperation, and hand gestures in a user study with eight participants. We evaluate replay success, modified NASA-TLX workload, and common teaching errors across three manipulation tasks. Kinesthetic guidance produced the shortest demonstrations, lowest workload, and highest success on the more orientation-sensitive and contact-rich tasks. Joystick teleoperation performed best on simple peg picking. Hand-gesture teaching, although less reliable overall, performed better than expected and in some cases achieved results comparable to kinesthetic guidance.
论文介绍 通过用户研究比较运动学指导、操纵杆遥操作和手势教学三种机器人教学模态。在三种操作任务中评估重放成功率、工作量和常见错误。运动学指导在方向敏感和接触丰富的任务中表现最佳,操纵杆在简单任务中优越,手势教学在某些情况下接近运动学指导的效果,为教学方法选择提供参考。
第一作者: Yirui Sun · 方向: 机器人操作 · 来源: cs.RO
世界动作模型扩散策略状态自适应去噪调度
Abstract:World Action Models (WAMs) improve robot manipulation by using video-based future representations to condition action generation. In pixel-space WAMs, however, the best action condition is not necessarily the fully denoised video. Controlled denoising-depth scans show that video refinement can reduce action error up to a state-dependent point, after which the gain may saturate or even reverse when late predictions become less action-relevant or physically unreliable. This suggests that action generation should use a state-dependent point along the video noise trajectory rather than a fixed terminal denoising depth. We introduce State-Adaptive Noise Trajectory Scheduler (SANTS), a lightweight scheduler for video-to-action diffusion policies. At each video decision point, SANTS reads the current video-state representation and noise level, then jointly predicts a cumulative...
论文介绍 介绍「SANTS」,一种状态自适应噪声轨迹调度器,用于视频到动作的扩散策略。研究指出,在视频生成世界动作模型中,最佳动作条件依赖于状态。SANTS根据当前视频状态和噪声水平动态预测去噪深度,优化动作生成效率,避免过度去噪导致的性能下降。
第一作者: Zimu Li · 方向: 机器人操作 · 来源: cs.RO
四足机器人主动脊柱强化学习敏捷运动
Abstract:The biological spine of quadrupeds enables sagittal flexion/extension, lateral bending, and axial rotation, playing a crucial role in highly agile and dexterous locomotion. While numerous studies have integrated active spinal joints into quadrupedal robots to enhance agility, most designs simplify control complexity by reducing spinal degrees of freedom (DOF), failing to achieve the spatial tri-axial rotation characteristic of biological spines. Consequently, replicating a multi-DOF biomimetic spine and effectively leveraging it to empower the agile locomotion of quadrupedal robots remains a significant research challenge. In this study, we present S-Cheetah, a quadrupedal robot featuring a 3-DOF bio-inspired serial active spine capable of biomimetic spatial tri-axial rotation. To empower the robot to fully utilize this active spine, we developed a specialized reinforcement...
论文介绍 提出「S-Cheetah」四足机器人,配备3自由度生物启发主动脊柱,可实现模仿生物的空间三轴旋转。通过强化学习训练敏捷运动,使机器人充分利用脊柱功能提升运动能力。设计旨在解决多自由度仿生脊柱的集成与控制问题,增强四足机器人的敏捷性和灵活性。
第一作者: Sizhe Lester Li · 方向: VLA 通用模型 · 来源: cs.RO
视频模型逆动力学模型通用策略机器人操作
Abstract:Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments. Recent work proposes robot foundation models that jointly predict future observations and actions by finetuning video models with action-labeled data. In this paper, we test the limits of an alternative approach: leave the video planner as-is while training an embodiment-specific inverse dynamics model (IDM). This decoupling offers several natural benefits: the video planner remains embodiment-agnostic, different video models can be interchanged easily without re-training the IDM, and the IDM can be independently trained with readily available self-play data. We present a closed-loop, video-to-action policy that combines an action-free video world model with a carefully-designed IDM based on the...
论文介绍 探索将视频生成模型作为世界模型,结合形态特定的逆动力学模型构建通用机器人策略。方法解耦视频规划和动作生成,视频规划器保持形态无关,逆动力学模型可独立训练。通过闭环视频到动作策略,实现跨形态和环境的通用控制,无需重新训练整个系统。
第一作者: Jeremy Morgan · 方向: VLA 通用模型 · 来源: cs.RO
视觉语言动作模型泛化评估模拟基准机器人学习
Abstract:Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress can be misleading. Despite the zero-shot perception and language capabilities of VLAs, their overall task performance often degrades under distribution shifts, revealing gaps in how these systems translate high-level understanding into robust behavior. To systematically study this gap, we introduce Colosseum V2, a large-scale simulation benchmark for evaluating VLA generalization in robot learning across diverse conditions. The benchmark comprises 28 tasks spanning 13 task categories and two robot morphologies, covering a wide range of manipulation primitives and long-horizon behaviors. Built on the ManiSkill simulator, Colosseum V2 enables fast, GPU-parallelized evaluation and supports both...
论文介绍 引入「Colosseum V2」基准,用于系统评估视觉语言动作模型在机器人操作中的泛化能力。基准包含28个任务、13个类别和两种机器人形态,覆盖多种操作原语和长期行为。基于ManiSkill模拟器,支持GPU并行评估,旨在研究模型在分布偏移下的性能差距。
第一作者: Kevin Lin · 方向: 机器人操作 · 来源: cs.RO
人形机器人模仿学习数据生成全身规划
Abstract:Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations, which are time-intensive and difficult to collect via teleoperation. Existing data-generation algorithms can automatically synthesize demonstrations for manipulators, but they are ineffective on humanoids because their high-dimensional composite action spaces involve arms, legs, and torsos. We present HumanoidMimicGen, a method for generating humanoid legged loco-manipulation data. Our method adapts contact-rich whole-body skills from a handful of source demonstrations to new states, generalizing across changes in object pose. By interleaving these single- and dual-arm skills with whole-body locomotion and manipulation planning, the method generates stable, collision-free data across diverse scenes and layouts. To evaluate our...
论文介绍 提出「HumanoidMimicGen」方法,用于生成人形机器人腿部操作数据。通过全身规划,从少量源演示适应新物体姿态,插值单臂和双臂技能与运动规划,生成稳定、无碰撞的多样化数据。该方法加速模仿学习过程,解决人形机器人数据收集困难的问题。
第一作者: Jinhao Liang · 方向: 导航与运动 · 来源: cs.RO
多机器人运动规划去中心化规划扩散模型轨迹预测安全约束
Abstract:Decentralized multi-robot motion planning requires each robot to generate collision-free trajectories from local observations, without global sensing or reliable communication. However, most existing planners, whether classical or learning-based, generate trajectories from a static snapshot of the local observation, which limits their ability to anticipate the future behavior of neighboring robots. This limitation is critical as the number of robots increases and the environment becomes more cluttered. To overcome this challenge, this paper introduces Simulation-Informed Diffusion (SID), a decentralized framework built on constraint-aware diffusion models (CADM). SID first uses CADM to simulate the future trajectories of neighboring robots from their currently observed states, and then uses the same CADM to plan each robot's own trajectory under safety constraints informed by...
论文介绍 针对去中心化多机器人运动规划中机器人仅依赖局部观测生成无碰撞轨迹的挑战,现有方法受限于静态观测快照。本文提出模拟启发扩散框架,基于约束感知扩散模型先模拟邻居机器人的未来轨迹,再在安全约束下规划自身轨迹,以提升在机器人数量增加和环境拥挤时的规划能力。
第一作者: Marcus G Müller · 方向: 具身智能 · 来源: cs.RO
地形分割语义分割无结构环境合成数据视觉外观
Abstract:Terrain understanding is fundamental for mobile robots operating in unstructured outdoor environments. Existing vision-based traversability estimation methods rely on robot-specific annotations or semantic class mappings, limiting transferability across platforms and requiring costly re-annotation when robot capabilities change, while standard semantic segmentation methods only focus on specific predefined classes, which do not capture the variety of terrains. In this work, we propose a transformer-based architecture that jointly performs class-specific semantic segmentation and class-agnostic terrain segmentation within a unified network, called Trinity. Terrain regions are segmented based solely on visual appearance, without predefined semantic labels or robot-dependent traversability scores. This formulation enables the learning of robot-agnostic visual terrain priors that...
论文介绍 移动机器人在无结构户外环境中进行地形理解时,现有方法依赖机器人特定注释或语义类映射,可迁移性差。本文提出Trinity架构,统一执行类特定语义分割和类无关地形分割,基于视觉外观学习机器人无关的地形先验,利用合成数据增强泛化能力,提高跨平台适用性。
第一作者: Ivan Saraev · 方向: 机器人操作 · 来源: cs.RO
语言到目标合成光流控组装大型语言模型可微分目标函数智能代理
Abstract:Light-based advanced manufacturing increasingly requires programmable, closed-loop tools that translate human design intent into executable operations at small length scales. Yet a key bottleneck persists across robotic and manufacturing modalities: turning user intent into machine-readable objectives that are reliably executable. While micro-robotics offers versatile manipulation via optical actuation of fluids, mathematically tractable goal specification remains manual and hard to reuse. Here, we introduce Speak-to-Objective, a modular agentic pipeline that uses a conditioned Large Language Model (LLM) to translate spoken or written commands into fully differentiable objective functions for assembling microparticles in a constraint-aware inverse solver (SLSQP) and on an experimental optofluidic platform. The approach employs a compact loop - perceive -> compose -> propose ->...
论文介绍 在光流控组装等高级制造中,将人类设计意图转化为机器可执行操作存在瓶颈。本文引入智能代理管道,使用条件大型语言模型将语音或文本命令转换为可微分目标函数,并在约束感知逆求解器中执行,实现微粒组装任务的自动化,提升制造系统的闭环控制能力。
第一作者: Hongyu Ding · 方向: VLA 通用模型 · 来源: cs.RO
具身导航语言-视觉-机器人动作翻译多模态大型语言模型动作生成统一框架
Abstract:Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dominant approach has been to scale vision-language-action (VLA) foundation models on ever-larger collections of robot trajectories. This paper argues that, for navigation specifically, generality can be obtained structurally, not only through data scale. The underlying decision structure of navigation reduces to a single Language-Vision-Robot Actions Translation. The language action emits semantic-level directional command and the vision action emits a pixel-level visual target. Both outputs lie inside the natural output manifold of pretrained multimodal large language models (MLLMs), so the task can be reasoned about by an agent rather than learned from robot data. Therefore, we present Uni-LaViRA, a...
论文介绍 具身导航需要将语言和视觉观测映射为动作序列,现有方法依赖大规模机器人轨迹数据。本文提出Uni-LaViRA框架,通过语言-视觉-机器人动作翻译结构化生成动作,利用预训练多模态大型语言模型进行推理,减少对特定机器人数据的依赖,提高导航在未见环境中的泛化能力。
第一作者: Morten Roed Frederiksen · 方向: 具身智能 · 来源: cs.RO
社交机器人参与策略合成情绪游戏化儿童支持
Abstract:Many children experience challenges in emotional regulation and social interaction, which can limit their participation in everyday activities and therapeutic programs. For socially assistive robots to be effective in this context, it is essential that children remain consistently and meaningfully engaged. We explore engagement strategies for a tactile robot designed to support children suffering from anxiety disorders through daily interactions. The robot delivers either synthetic emotional feedback or point rewards to encourage user participation. We evaluated these strategies through two studies: a preference assessment with 16 school children aged 6-8 years, and a behavioral study with 14 university students aged 20-27 years in naturalistic environments. The study with school children indicated a preference for emotional engagement over points-based approaches. The follow...
论文介绍 社交机器人在支持儿童情绪调节时需保持用户参与。本文探索合成情绪反馈和积分奖励策略,在不同年龄组中进行评估:学龄儿童偏好情绪交互,而大学生行为研究显示策略效果可能因年龄而异,为社交机器人设计提供实证依据。
第一作者: Morten Roed Frederiksen · 方向: 具身智能 · 来源: cs.RO
触觉交互平静诱导儿童生理指标口袋机器人
Abstract:Periods of heightened arousal or restlessness can interfere with children's ability to focus, self-regulation, and physically calm. Technologies that encourage embodied self-regulation through tactile interaction may provide a simple and accessible means of promoting calmness. This paper investigates how interaction with a pocket-sized tactile device influences physiological and behavioral markers of calmness in typically developing children. Building on prior work examining heart rate modulation, we present new findings on how tactile interaction affects full-body movement and postural stability. We employ a device that engages children through a hand-held rhythmic vibration-matching game, designed to focus attention and encourage stillness. Eighteen children participated in a within-subjects study that involved two conditions: with and without tactile interaction with a...
论文介绍 口袋大小触觉设备可能通过交互促进儿童平静。本文研究手持触觉交互对儿童心率、全身运动和姿势稳定性的影响,采用节奏振动匹配游戏,在受试者内设计中评估生理和行为指标,旨在提供简单可及的自我调节工具以减少焦虑。
第一作者: Mahmoud Abouelyazid · 方向: 策略学习 · 来源: cs.RO
多智能体强化学习涌现通信对比对齐潜在嵌入通信学习
Abstract:Emergent communication enables partially observant Autonomous Mobile Robots (AMRs) to coordinate effectively in decentralized multi-agent reinforcement learning (MARL) settings. However, existing approaches often struggle with unstable communication protocols, ungrounded message semantics, and interference between communication learning and policy optimization, leading to degraded coordination over time. We propose SCALE-COMM (Shared, Contrastively-Aligned Latent Embeddings for COMMunication), a self-supervised framework for learning compact, stable, and policy-relevant communication representations. SCALE-COMM decouples communication learning from policy optimization by training low-dimensional latent messages that capture task-relevant planning and traffic information, while enforcing consistency across agents and time. Across standard MARL benchmarks and a realistic...
论文介绍 多智能体强化学习中通信协议不稳定、语义无根问题影响协调效果。本文提出SCALE-COMM自监督框架,通过学习共享、对比对齐的低维潜在嵌入来表示任务相关信息,解耦通信学习与策略优化,生成紧凑稳定的通信表示,提升智能体在基准任务和现实场景中的协调性能。
第一作者: Boxiang Qiu · 方向: VLA 通用模型 · 来源: cs.RO
视频世界模拟器机器人操作闭环模拟动作条件生成策略学习
Abstract:We introduce GE-Sim 2.0 (Genie Envisioner World Simulator 2.0), a closed-loop video world simulator for robotic manipulation. Building on the action-conditioned video generation framework of Genie Envisioner, GE-Sim 2.0 is re-trained on thousands of hours of real-world robot data spanning teleoperation, contact-rich interaction, and on-robot policy deployment, substantially improving action-following fidelity and trajectory coverage. On top of this foundation, three new modules close the loop from video simulation to policy learning: a state expert that decodes proprioceptive state from video latents to support next-chunk prediction by downstream VLA policies; a world judge that scores generated rollouts against task instructions, yielding machine-verifiable success signals and rewards in place of manual inspection; and an acceleration framework that delivers a 25-frame...
论文介绍 为机器人操作提供闭环视频模拟环境,GE-Sim 2.0基于动作条件视频生成框架,使用大量真实机器人数据训练,提高动作跟随保真度。新引入状态专家、世界法官和加速模块,从视频模拟中支持策略学习,实现机器可验证的成功信号和高效推理,推动操作策略开发。
第一作者: Brian Zhu · 方向: VLA 通用模型 · 来源: cs.RO
VLA通用模型工业包装工厂部署迭代微调机器人操作
Abstract:Vision-Language-Action (VLA) policies have shown promising manipulation capabilities, yet their practical impact is often limited by the reliability demands of real-world deployment. We present a deployment study of an industrial packaging task at Siemens Factory (GWE, Erlangen, Germany), where a robot must pick a transparent accessory bag from a cluttered pile, insert it into the remaining cavity of a cardboard package, and ensure that the bag and its contents remain below the closing plane. Our goal is to understand the practical effort required to adapt a pretrained Pi0.5 policy to a single factory-floor task through iterative fine-tuning and deployment-driven refinement. The pipeline consists of repeated loops of data collection, curation, fine-tuning, evaluation, and targeted recovery data collection. We have accumulated 2535 episodes (10 hours) from the on-site factory...
论文介绍 本文介绍了在西门子工厂进行的一项VLA管道部署案例研究,用于工业包装任务。研究旨在通过迭代数据收集、策展、微调和评估,将预训练的Pi0.5政策适应到单一工厂任务,流程包括针对性恢复数据收集的循环。该工作展示了将VLA策略应用于实际工厂环境所需的实际努力和经验教训,为类似部署提供参考。
第一作者: Meraj Mammadov · 方向: 模仿学习 · 来源: cs.RO
模仿学习强化学习教师-学生模型表示对齐自监督对比学习
Abstract:Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex and high-dimensional observation spaces prevalent in robotics. This paper addresses the irreducible imitation gap that emerges when teacher and student are learned in isolation, and the teacher policy has the liberty to rely on privileged state information that the student cannot infer from its observations. Instead of improving poor student performance with RL finetuning after IL, which often requires a whole new training setup, we propose a novel algorithm which learns a shared embedding space that hides agent-specific observations and thus trains imitable teacher policies by construction. We train the shared embedding space with self-supervised contrastive learning in parallel to the teacher policy and prevent it from...
论文介绍 本文针对基于强化学习的模仿学习中,教师和学生模型独立学习时出现的不可约模仿差距问题。核心方法是提出一种新算法,学习共享嵌入空间以隐藏代理特定观察,从而通过自监督对比学习训练可模仿的教师策略。这减少了对强化学习微调的依赖,提高了模仿学习的效果和效率。
第一作者: Arissa J. Sato · 方向: 具身智能 · 来源: cs.RO
社交机器人编程工具大语言模型生成脚手架块编程
Abstract:Programming social robots is challenging for novice robot programmers due to required expertise in planning, interaction design, and programming. While large language models (LLMs) hold significant promise through code generation from natural-language descriptions, they can obscure critical elements of programming and supplant designer intent, eventually resulting in over-reliance instead of developing programming skills. In this paper, we explore how LLM-based social-robot-programming tools can support novice robot programmers through a Research through Design (RtD) process. We designed and prototyped Robo-Blocks, a block-based programming environment that leverages LLMs to offer novice robot programmers generative scaffolding through structured narratives that connect high-level ideas to executable robot behaviors. Through deployment with novices, we discovered emerging user...
论文介绍 本文针对新手编程社交机器人面临的挑战,设计并原型化了Robo-Blocks,一个块编程环境。该工具利用大语言模型提供生成脚手架,通过结构化叙事连接高级想法和可执行行为,支持新手机器人程序员。这有助于避免过度依赖LLM,同时促进编程技能的发展和意图保持。
第一作者: Haolan Zhang · 方向: 具身智能 · 来源: cs.RO
视觉里程计RGB-D直接稀疏里程计一致性先验不确定性估计
Abstract:Visual odometry (VO) is a fundamental component in robotics and augmented reality. RGB-D direct VO benefits from metric depth measurements, but it can degrade in challenging environments, where dynamic objects, occlusions, illumination changes, and unreliable depth violate the short-horizon photometric and depth-geometric consistency assumptions used by direct alignment. Existing approaches mitigate these issues through semantic filtering, explicit occlusion reasoning, illumination adaptation, or hand-crafted geometric criteria, but often rely on external modules or fixed assumptions tailored to individual failure modes, limiting their flexibility and ability to handle diverse challenges in a unified manner. In this work, we propose Con-DSO, a consistency-aware RGB-D direct sparse odometry framework that predicts dense photometric and depth-geometric consistency uncertainty...
论文介绍 本文针对RGB-D直接视觉里程计在挑战性环境中因违反光度和深度几何一致性假设而退化的问题。核心方法是提出Con-DSO框架,学习预测密集光度和深度几何一致性不确定性,以增强鲁棒性。这有望提高机器人和增强现实中视觉里程计在复杂环境下的性能,无需依赖外部模块。
第一作者: Wenzhe Song · 方向: 策略学习 · 来源: cs.RO
自动驾驶城市交叉口异构代理模型预测安全强化学习
Abstract:The imminent integration of autonomous vehicles and mobile robots in urban settings presents a critical safety challenge for future intelligent transportation systems. This paper addresses the complex problem of coordinating heterogeneous agents with disparate dynamics at unregulated intersections. We introduce a novel framework, differentiable model predictive safety (DMPS), which embeds the foresight of model-predictive control into a data-driven, end-to-end reinforcement learning architecture. DMPS agents learn a latent dynamics model to predict future trajectories contingent on their actions. A learned, differentiable safety critic then evaluates the risk of these trajectories. Crucially, by leveraging backpropagation through the entire unrolled predictive model, agents can efficiently compute the gradient of future safety with respect to their current action, enabling a...
论文介绍 本文针对城市非管制交叉口协调具有不同动力学的异构代理的安全挑战。核心方法是引入可微模型预测安全框架,将模型预测控制的远见嵌入数据驱动的端到端强化学习架构中,学习潜在动力学模型和可微安全评判器。这能通过梯度计算高效评估未来安全性,适用于智能交通系统中的安全协调。
第一作者: Qin Yang · 方向: 具身智能 · 来源: cs.RO
自闭症教育干预社交机器人NAO机器人课堂实验
Abstract:Autism is a developmental disorder that manifests in early childhood and persists throughout life, profoundly affecting social behavior and hindering the acquisition of learning and social skills in those diagnosed. As technological advancements progress, an increasing array of technologies is being utilized to support the education of students with Autism Spectrum Disorder (ASD), aiming to improve their educational outcomes and social capabilities. Numerous studies on autism intervention have highlighted the effectiveness of social robots in behavioral treatments. However, research on the integration of social robots into classroom settings for children with autism remains sparse. This paper describes the design and implementation of a group experiment in a collective classroom setting mediated by the NAO robot. The experiment involved special education teachers and the NAO...
论文介绍 本文研究社交机器人在自闭症学生课堂教育中的整合和应用。核心方法是设计并实施一个由NAO机器人介导的集体课堂实验,涉及特殊教育教师和自闭症学生。该工作探索了社交机器人在提高自闭症学生教育成果和社交能力方面的潜力,为相关干预提供实践见解。
第一作者: Ruowen Zhao · 方向: VLA 通用模型 · 来源: cs.CV
具身智能视觉语言模型生成监督深度图生成预训练
Abstract:Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a significant gap remains between the high-level semantic focus of standard text-guided pre-training paradigms and the low-level spatial and physical knowledge critical for execution in embodied environments. In this paper, we introduce GEM, a Generative-supervised Embodied vision-language Model designed to bridge this divide. We propose integrating a depth map generation task directly into the VLM pre-training phase. By training this generative objective jointly with the main model, we observe substantial improvements in embodied intelligence, significantly enhancing both semantic understanding and physical operation capabilities. To support this paradigm, we curate and release GEM-4M, a comprehensive...
论文介绍 本文针对标准文本引导预训练范式与具身环境所需的空间和物理知识之间的差距。核心方法是引入GEM,一个生成监督的具身视觉语言模型,将深度图生成任务集成到预训练阶段。通过联合训练生成目标,显著提高了语义理解和物理操作能力,增强了具身智能的性能。
第一作者: Jiyuan Fu · 方向: VLA 通用模型 · 来源: cs.CV
VLA模型对抗攻击补丁攻击可转移性安全漏洞
Abstract:While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly hinders their deployment in safety-critical domains. Moreover, existing patch attacks primarily focus on white-box settings, heavily overfitting to the specific action output space of the target model, which results in poor cross-architecture transferability. To overcome this limitation, we propose VLA-Hijack, a unified adversarial framework that breaks the transferability bottleneck by exploiting a fundamental vulnerability identified in this work: before planning any motion, a VLA model must first use visual information to locate its own robotic arm within the environment. Targeting this shared visual self-localization process, our approach concurrently optimizes Attention-Guided Proprioceptive Suppression to inhibit the real...
论文介绍 本文针对VLA模型对对抗补丁的脆弱性,以及现有攻击在跨架构可转移性方面的局限性。核心方法是提出VLA-Hijack统一对抗框架,针对VLA模型共享的视觉自定位过程,优化注意力引导的本体感知抑制。这揭示了VLA模型的安全隐患,为安全部署和防御研究提供见解。
第一作者: Otmane Sakhi · 方向: 策略学习 · 来源: cs.LG
离策略学习强化学习大型语言模型重要性权重优化稳定性
Abstract:Large scale reinforcement learning has become a central tool for improving reasoning in large language models. At this scale, generation is often lagged or asynchronous, so updates are performed on data collected by older policies. This makes learning inherently off-policy. Most existing approaches nevertheless remain rooted in PPO-style trust-region objectives, treating training as approximately on-policy and using importance weights to correct distribution mismatch. These corrections can introduce high variance, destabilize optimization, and accelerate entropy collapse. Recent work suggests an alternative: rather than correcting the mismatch, one can embrace off-policy data and remove importance weights, often yielding stronger algorithms. In this paper, we provide an intuitive construction of off-policy objectives that include successful off-policy objectives and show that...
论文介绍 本文研究大规模强化学习用于提升大型语言模型推理能力时面临的离策略学习问题。现有方法依赖PPO式目标并使用重要性权重修正分布不匹配,但可能导致高方差和优化不稳定。论文提出一种直观的离策略目标构建,移除重要性权重,直接利用旧策略数据,从而提供更稳定的优化框架,可能改善算法效率和推理性能。
第一作者: Yuchen Guo · 方向: 导航与运动 · 来源: cs.AI
参数化记忆具身智能体Minecraft混合专家LoRA对比内化
Abstract:We present PEAM, a Parametric Embodied Agent Memory framework in Minecraft that transforms agent memory from inference-time retrieval into parameter-resident skills internalized through experience. PEAM pairs a slow deliberative LLM for open-ended reasoning with a fast parametric module for reflexive execution of consolidated skills. The fast module is a multimodal Mixture-of-Experts LoRA architecture with per-category physically isolated adapters, enabling parameter-level continual learning without catastrophic forgetting. We treat failure as a first-class training signal: failure--correction trajectory pairs are internalized through a joint behavioral-cloning and contrastive objective, so the agent learns not only what succeeds but also how corrected actions differ from failed ones. To govern consolidation, PEAM introduces a parameterization-worthiness score for deciding...
论文介绍 本文提出PEAM框架,用于Minecraft中的具身智能体,将记忆从推理时检索转化为参数化技能。该框架结合慢速大型语言模型进行推理和快速参数模块执行技能,采用多模态混合专家LoRA架构实现持续学习。通过将失败-纠正轨迹对用联合行为克隆和对比目标内化,智能体学习成功与纠正动作。引入参数化价值评分管理技能巩固,可能提升学习效率和适应性。
美股技术面显示分化,标普500 ETF (SPY) 和纳斯达克100 ETF (QQQ) 的RSI分别处于71.1和74.6的超买区域,价格接近52周高点(SPY仅-0.22%,QQQ -0.53%),短期存在过热风险,但趋势仍为bullish;个股如Apple (AAPL) RSI 78.8超买,而Microsoft (MSFT) RSI 49.6中性,反映板块内部节奏不一。加密市场情绪极度恐慌,贪婪指数仅为22,总市值2.54万亿美元,24小时下跌2.85%;比特币(BTC-USD)价格72912.01,RSI 34.4偏低,呈现空头排列;以太坊(ETH-USD) RSI 29.1进入超卖,价格远低于所有均线,下跌趋势明确;Solana (SOL-USD) RSI 36.3,同样疲软。中概股整体承压,阿里巴巴(BABA)、拼多多(PDD)、京东(JD)和腾讯(0700.HK)均处于空头排列,RSI范围在28.4至44之间,其中腾讯RSI 28.4超卖,接近52周低点。商品外汇市场中,黄金期货(GC=F)趋势中性但RSI 32.5偏低;原油(CL=F)日涨3.95%但MACD死叉;美元指数(DXY) RSI 60.1,趋势bullish并接近52周高点;10Y美债收益率(^TNX)虽趋势bullish但MACD死叉,显示动能变化。
RSI 14读数为78.8,显著高于70超买阈值,显示短期强势;当前价格310.85接近52周高点(仅-0.77%),且远高于52周低点(+59.35%);价格位于所有关键移动平均线(SMA20 293.39、SMA50 272.88、SMA200 262.4)之上,形成多头排列;MACD为10.2687,高于信号线9.4234,表明上涨动量持续,技术状态偏上行。
RSI 14为29.1,进入超卖区域,反映卖压持续;价格1976.4远低于SMA20 2181、SMA50 2258.02和SMA200 2526.86,呈现空头排列;MACD为-62.4998,信号线为-46.1468,负值扩大确认下跌趋势;近5日价格下跌4.27%,技术指标一致指向下行。
RSI 14为49.6,接近中性水平50,无超买或超卖信号;价格412.67略低于SMA20 415.34但高于SMA50 401.11,趋势标记为neutral;MACD为2.8272,略低于信号线3.7612,差值较小,表明动量温和;整体技术状态平衡,缺乏明确方向。
VIX 恐慌指数
10Y 美债收益率 (%)
美元指数 DXY
S&P 500 ETF
Nasdaq 100 ETF
Apple
Microsoft
Nvidia
Alphabet
Tesla
Meta
Bitcoin
Ethereum
Solana
阿里巴巴 (BABA)
拼多多 (PDD)
京东 (JD)
腾讯控股 (0700.HK)
黄金期货
WTI 原油期货
美元 / 人民币
本报告基于公开行情数据计算的技术指标,仅供技术指标解读参考,不构成任何投资建议或操作指引。技术分析存在局限性,过去走势不代表未来表现,市场具有不确定性,投资者应结合自身情况审慎决策。
The military says areas south of the Zahrani River are now "combat zones" as it threatens Hezbollah with fresh strikes.
中文摘要 以色列军方对黎巴嫩南部Zahrani河以南地区发布疏散令,称该区域为战斗区,并威胁对真主党发动新一轮打击。
A U.S. official said on Wednesday the new attacks had been in self-defense and targeted attack drones and a drone ground-control station. They were the latest attacks to threaten a fragile cease-fire.
中文摘要 美国对伊朗南部发动新打击,目标为攻击无人机和地面控制站。美国官员称这是自卫行动,但威胁到脆弱的停火协议。
Attorney general says government seeking $2bn in damages to recover costs relating to firefighting foam at 28 defence bases. Follow updates live Get our breaking news email, free app or daily news podcast Scams come in all shapes and sizes (and none of them are nice), and the government is consideri
中文摘要 澳大利亚政治新闻中,三位议员在质询时间因争执被驱逐。政府寻求20亿美元赔偿,用于28个国防基地消防泡沫相关成本。反对党领袖泰勒就税收问题向工党发难。
The strikes come despite a ceasefire between Tehran and Washington as the two countries hold peace talks.
中文摘要 美国对伊朗发动新打击后,油价上涨。尽管德黑兰与华盛顿之间存在停火协议,且两国正在进行和平谈判,袭击仍发生。
Federal government seeks more than $2bn in damages from multinational manufacturer in its largest legal claim ever Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast The Australian government said on Thursday it had launched legal actio
中文摘要 澳大利亚联邦政府起诉跨国制造商3M,寻求超过20亿美元赔偿,涉及消防泡沫中的PFAS永久化学物质。这是其有史以来最大的法律索赔。
After CEO’s comments to investors, opponents urge miner to ‘stop stringing everybody along and spike the project finally’ Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Santos chief executive, Kevin Gallagher, has told an investor b
中文摘要 Santos公司表示不会在Narrabri天然气项目上投入精力,将专注于Beetaloo盆地。CEO向投资者发表评论后,反对者敦促公司停止拖延并终止该项目。
The live-in personal assistant to the actor has been sentenced to 41 months in prison, capping a multi-year legal saga surrounding the actor's death.
中文摘要 演员马修·佩里的私人助理被判处41个月监禁,结束了围绕其死亡的多年法律纠纷。
Several groups of women and children who spent years in a Syrian camp have returned in recent months.
中文摘要 澳大利亚指控一名从叙利亚返回的女性加入伊斯兰国。近期,多批在叙利亚营地滞留多年的妇女儿童已返回。
Months after Pakistan declared “open war” on Afghanistan, neither side appears ready to back down, despite China’s efforts to mediate.
中文摘要 巴基斯坦宣布对阿富汗开战数月后,双方仍无意退让,尽管中国正努力调解。冲突持续。
After 88 days of near-total blackout, first reactions to the return of partial connectivity were not celebratory After 88 days of near-total internet blackout in Iran, long-delayed messages, images and poems flooded phones and social media feeds at about 5pm on Tuesday, when still-limited connectivi
中文摘要 伊朗互联网在断网88天后恢复部分连接,但首批反应并非庆祝,而是愤怒、焦虑和泪水。
Hugh Marks refuses to confirm or deny to Senate estimates that he threatened to terminate Stevens if he didn’t resign Follow our Australia news live blog for latest updates Sign up for Guardian Australia’s free weekly media newsletter here A top news executive from Reuters, Simon Robinson, is expect
中文摘要 路透社高管西蒙·罗宾逊预计将接替贾斯汀·史蒂文斯担任澳大利亚广播公司新闻总监。休·马克斯拒绝向参议院确认或否认曾威胁解雇史蒂文斯。
德國國會友台小組造訪台灣,引發兩岸隔空交火。德國目前民調第一的極右翼「另類選擇黨」(AfD)被視為對中國態度友善,但黨內議員克拉夫特卻二度以友台小組成員身份訪台。他接受DW專訪時,不僅公開強調台灣對德國的重要性,更直接喊話北京「不要吵」。
中文摘要 德国国会友台小组访问台湾,引发中国反对。极右翼政党AfD议员克拉夫特二度访台,强调台湾对德国重要性,并直接向北京喊话「不要吵」。
Wet weather to batter parts of Queensland, NSW and Tasmania through Thursday, with warnings of flash flooding Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Multiple states are at risk of flash flooding on Thursday, with severe weat
中文摘要 澳大利亚昆士兰、新南威尔士和塔斯马尼亚三个州预计将迎来暴雨和洪水,干旱地区迎来大雨。气象局发出严重天气警告,警惕山洪暴发。
The hostilities come during a fragile ceasefire between the US and Iran, and protracted negotiations to end the three-month war.
中文摘要 美国在三天内第二次打击伊朗目标。袭击发生在美伊脆弱停火期间,以及结束三个月战争的长期谈判进行中。
Israel has intensified its deadly military campaign against Hezbollah, the Iran-backed militant group, in recent days, striking targets across Lebanon.
中文摘要 以色列加强了对伊朗支持的武装组织真主党的军事行动,近日在黎巴嫩境内多处打击目标,居民目睹军机盘旋。
Also in today’s newsletter: fresh Iran strikes and EU trade defences
中文摘要 强劲的卢布给俄罗斯战争经济带来压力。今日简报还涵盖伊朗新袭击事件及欧盟贸易防御措施。
The strikes come despite a ceasefire between Tehran and Washington as the two countries hold peace talks.
中文摘要 美国对伊朗发动新攻击后,国际油价上涨。袭击发生在德黑兰和华盛顿停火期间,尽管两国正在进行和平谈判。
Lack of volatility is taking off pressure to reach a Strait of Hormuz deal.
中文摘要 市场未能充当特朗普的护栏,缺乏波动性降低了达成霍尔木兹海峡协议的压力。
The oil major’s investors overwhelmingly voted in favour of its proposal to move its corporate domicile to Texas
中文摘要 埃克森美孚投资者压倒性投票支持将其公司注册地从民主党主导的蓝州迁至德克萨斯州。
Philadelphia Semiconductor Index rides Big Tech’s data centre spending spree to 75% gains in 2026
中文摘要 受人工智能需求推动,芯片股奔向自互联网泡沫时代以来的最大涨幅。费城半导体指数因大型科技公司数据中心支出热潮,在2026年上涨75%。
Officials cite need to maintain sovereign control over critical national infrastructure
中文摘要 英国将阻止印度亿万富翁增持英国电信股份,以维护对关键国家基础设施的主权控制。
Avoiding power struggles between France and Germany is seen to be key to the success of the group
中文摘要 欧洲坦克制造商在准备首次公开募股时誓言防范政治干预,避免法国和德国间的权力斗争被视为成功关键。
World’s highest-grossing law firm plans to put the ‘collective intelligence’ of its lawyers into a tech platform
中文摘要 律所Kirkland & Ellis计划投入5亿美元构建自有AI技术,旨在将律师的“集体智慧”融入技术平台。
Governance of new technologies must be determined by elected officials rather than fastest moving companies
中文摘要 关闭人工智能责任漏洞的关键在于,新技术治理应由民选官员而非行动最快的公司决定。
A founder-entrenching shareholder structure can do a great deal of harm
中文摘要 SpaceX案例显示,创始人巩固的股东结构可能对公司造成重大损害,被称为扎克伯格折扣。
Superdrug grew almost twice as fast as Boots over past four years
中文摘要 英国零售公司Superdrug计划首次公开募股,过去四年其增长速度几乎是竞争对手Boots的两倍。
Industry-wide shortage has sent their values soaring
中文摘要 飞机所有者选择租赁发动机而非整架飞机,全行业短缺导致发动机价值飙升。
14 回复 · 程序员 节点
6 回复 · 程序员 节点
19 回复 · Apple 节点
17 回复 · 程序员 节点
9 回复 · Apple 节点
25 回复 · Apple 节点
10 回复 · Apple 节点
15 回复 · Apple 节点
20 回复 · Apple 节点
10 回复 · Apple 节点
太多人找我了 我懒得一个个回 晚上8点再发一波吧 反正还能蹬两三天 到时候提醒我一下 26 个帖子 - 21 位参与者 阅读完整话题
hi,我是毛球球,很多佬们刚进社区应该只知道linux.do的创始人是始皇(NEO),并不知道始皇是如何从一个草根一步步逆袭的。 前情提要:本人写过小说,下方所有文字仅代表个人观点和所见所闻所感,如有不符,纯属巧合(叠甲) 在我眼里,国内能称得上“古典极客精神”和“黑客浪漫”的,唯始皇一人。 第一:他用字节码插桩,给所有 Java 程序员上过最震撼的一课 Linux.do,却不知道当年的 zhile.io。当年 JetBrains 全家桶的防破机制围追堵截,无数所谓的“技术大牛”只能搞搞过期的注册码。只有始皇,硬是凭借对 JVM 底层、类加载机制和字节码插桩的降维理解,整出了 ja-netfi
发现了Gemini的一个网页端的鉴权bug,可以无限制调用任意模型,包括deep thinking等等,不需要任何订阅。反馈给Google之后两个月没反馈也没钱,准备公开赚点stars,会被追责吗?能拿star嘛这种? 一个实现思路是,找到ultra用户的用户层级,inner xx,然后发起对话请求(找到deep think对应的api字段)的时候替换成ultra的用户层级,Gemini不会在后端做验证。 65 个帖子 - 51 位参与者 阅读完整话题
有个项目完全用 AI 写的,超过 100 个接口了,中间出现过多次bug,比如它自己用允许为空的值做key,做了一些数据处理,导致真空的时候,出现了错误,当然也可能是我提指令的时候没有说清楚,但是肉眼可见的是,代码量越来越多,导致肉眼已经看不过来了,囧 如果哪天 AI 用不了了,这个项目也就挂了 79 个帖子 - 56 位参与者 阅读完整话题
十三个模型评测测试报告 1). 测试概述 本次测试针对以下十三个模型进行了统一条件下的对比评测: Gemma-4-31B-IT-Uncensored QwOpus3.6-27B Qwen3.6-27B-Neo-Code Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2 Qwen3.6-27B-MTP SuperGemma4-26B-Uncensored Qwen3.6-35B-A3B-Uncensored Qwen3.6-27B Gemma-4-31B-IT-Claude-Opus Gemma 4 - 26B A4B x Claude Opu
已经问了两百多个问题了,还在继续 49 个帖子 - 31 位参与者 阅读完整话题
之前开发了个小软件,这个软件需要使用ai,所以就寻找了一些这种可以每天调用一些的免费渠道。 Google Gemini API :免费层级每天可以调用一些,需要创建一个项目,在创建一个api密钥,对于gemini 3.1 fiash lite。每天可以调用1500次 https://aistudio.google.com/app/api-keys Ollama :在Ollama注册一个账号,默认就是Free层级,分为5h刷新和周刷新,可以使用gemma 4这种的模型 https://ollama.com/settings OpenRouter:对于Free模型来说,少于 10 credits
我不知道该用什么样的开头,去描述"人生的意义"这样一个宏大的命题。 当弗洛伊德被问到"一个正常人应该怎么做才能活得好"时,他的回答是"爱与工作",而他也因此得以流芳后世。如果心理治疗能让一个人学会如何好好爱人及工作,那么这个治疗就算成功,在马斯洛非常著名的"需求层次理论"中,人的生理需求一旦满足之后(如食物及安全感)就会转而追求爱,最后则是追求别人对自己的尊敬,后者大多是通过工作来达到。在弗洛伊德之前,托尔斯泰便曾说过:"只要人知道如何工作,如何爱人,人就可以在这世上活得更精彩,我们要为自己所爱的人工作,也要热爱自己的工作。这是当今社会的大多数人所认可的答案,我自己也有一点认可,但是我内心深处
2026 年主流终端 AI 编码工具分类 类型 工具 特点 最佳使用场景 TUI Grok Build 多面板 TUI + Plan Mode + Todo 管理 + 并行子代理,终端原生 复杂工程任务、重构、大型代码库 TUI OpenCode 主题化 TUI、分割面板、LSP 支持,接近 IDE 体验 可视化交互和多模型切换 TUI Crush (Charmbracelet) 美观直观 TUI,响应快,支持 LSP 代码理解 日常编码建议、快速迭代 TUI Crustly / Ralph-tui / Sidecar Rust 实现或专用仪表盘 TUI 本地优先、代理编排、监控 CLI Ai
享受吧,各位 31 个帖子 - 29 位参与者 阅读完整话题