每日简报

2026-05-29

← 历史归档

harry0703/MoneyPrinterTurbo

Python · ★ 66,461 · 🍴 9,617 · 📈 4,698 stars today

利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.

中文介绍 一款基于AI大模型的短视频自动生成工具,用户只需提供主题或关键词,即可一键生成带有字幕、配音和背景画面的高清短视频,解决了内容创作者的批量生产难题,适用于自媒体运营和营销推广。

affaan-m/ECC

JavaScript · ★ 197,325 · 🍴 30,344 · 📈 1,385 stars today

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

中文介绍 一个面向Claude Code、Codex等AI编程助手的性能优化与能力增强系统。它通过定义技能(Skills)、本能(Instincts)和记忆(Memory)等模块,帮助用户构建更强大、更安全的AI代理,适用于追求高性能AI辅助开发的研究者和工程师。

Leonxlnx/taste-skill

Shell · ★ 26,500 · 🍴 1,981 · 📈 2,234 stars today

Taste-Skill - gives your AI good taste. stops the AI from generating boring, generic slop

中文介绍 一个旨在提升AI生成内容品味的技能文件。它通过特定的提示或后处理方法,引导AI模型产出更具个性、避免同质化的内容,适用于对AI输出质量有更高要求的开发者和内容创作者。

hardikpandya/stop-slop

★ 6,439 · 🍴 470 · 📈 761 stars today

A skill file for removing AI tells from prose

中文介绍 一个专门用于去除AI生成文本中“AI味”的技能文件。它通过分析并修改行文模式,使机器生成的散文更接近自然人类语言,适合写作者和编辑用于优化AI辅助创作的内容。

twentyhq/twenty

TypeScript · ★ 47,878 · 🍴 6,795 · 📈 493 stars today

The open alternative to Salesforce, designed for AI.

中文介绍 一款面向AI时代的开源客户关系管理(CRM)系统,作为Salesforce的替代方案。它深度集成AI能力以优化销售流程,适用于寻求现代化、智能化客户管理工具的中小型团队和企业。

DigitalPlatDev/FreeDomain

HTML · ★ 170,751 · 🍴 3,296 · 📈 1,761 stars today

DigitalPlat FreeDomain: Free Domain For Everyone

中文介绍 一个由DigitalPlat提供的免费域名服务平台,旨在让每个人都能轻松获得域名,降低了个人和小型项目建立在线身份的门槛。

byoungd/English-level-up-tips

★ 48,567 · 🍴 5,102 · 📈 2,019 stars today

An advanced guide to learn English which might benefit you a lot 🎉 . 离谱的英语学习指南/英语学习教程/英语学习/学英语

中文介绍 一份结构详尽、内容深入的英语学习高级指南。它汇集了词汇、听力、口语、阅读等多方面的实用方法与资源,适合希望系统性提升英语综合能力的学习者。

microsoft/markitdown

Python · ★ 127,786 · 🍴 8,746 · 📈 1,410 stars today

Python tool for converting files and office documents to Markdown.

中文介绍 微软发布的一款Python工具,能高效地将各类文件和Office文档(如Word、Excel)转换为结构化的Markdown格式,方便开发者和技术文档工作者进行内容管理与发布。

obra/superpowers

Shell · ★ 211,119 · 🍴 18,797 · 📈 1,730 stars today

An agentic skills framework & software development methodology that works.

中文介绍 一个结合了代理(Agentic)技能框架与软件开发方法论的开源项目。它提供了一套可工作的实践体系,旨在通过增强AI代理的能力来提升软件开发效率,适用于AI应用开发者。

revfactory/harness

HTML · ★ 3,903 · 🍴 582 · 📈 65 stars today

A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use.

中文介绍 一个元技能(meta-skill)框架,能够针对特定领域自动设计AI代理团队、定义专业代理角色并生成它们所需的技能。它适用于需要协调多个AI代理处理复杂任务的架构师和研究人员。

codecrafters-io/build-your-own-x

Markdown · ★ 506,583 · 🍴 48,097 · 📈 1,066 stars today

Master programming by recreating your favorite technologies from scratch.

中文介绍 一个通过从头复刻各种流行技术(如数据库、编程语言)来深入掌握编程原理的学习资源集合。它为开发者提供了动手实践的项目,是提升底层技术理解力的绝佳指南。

Lum1104/Understand-Anything

TypeScript · ★ 42,852 · 🍴 3,419 · 📈 3,776 stars today

Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.

中文介绍 该工具能将任意代码库转化为交互式知识图谱,让用户可以探索、搜索并向图谱提问,从而快速理解复杂代码结构。它兼容Claude Code、Cursor等主流AI编程助手,帮助开发者驾驭陌生代码。

unclecode/crawl4ai

Python · ★ 66,964 · 🍴 6,857 · 📈 154 stars today

🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN

中文介绍 一款开源、对LLM友好的网络爬虫与数据提取工具。它优化了爬取和解析流程,旨在为大型语言模型高效准备高质量的网页数据,是AI开发者和数据科学家的实用利器。

OpenMOSS/MOSS-TTS

Python · ★ 2,256 · 🍴 215 · 📈 71 stars today

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental soun

中文介绍 来自MOSI.AI和OpenMOSS团队的开源语音及声音生成模型家族。专注于在复杂真实场景中实现高保真、高表现力的语音合成,为开发者和研究机构提供了高质量的TTS(文本转语音)解决方案。

EveryInc/compound-engineering-plugin

TypeScript · ★ 17,784 · 🍴 1,362 · 📈 184 stars today

Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more

中文介绍 Compound Engineering开发方法论的官方插件,支持集成到Claude Code、Codex、Cursor等多种AI编程助手中。它旨在通过结构化的工程实践,提升使用这些工具进行软件开发的效率和质量。

anthropics/skills

Python · ★ 142,869 · 🍴 16,871 · 📈 718 stars today

Public repository for Agent Skills

中文介绍 由Anthropic官方维护的AI代理技能公共仓库。它提供了可复用的技能定义和实现,供开发者快速集成到自己的AI代理系统中,是构建高级Agent功能的起点和资源库。

该源今日无内容。

The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray

80% Devin Commits, Spec-to-PR Workflows, Full VMs, Agent Memory, and PMs Shipping Code

中文介绍 异步代理时代,Cognition的Walden Yan与OpenInspect的Cole Murray探讨80%的Devin提交、规范到PR工作流、全虚拟机、代理记忆及产品经理发货代码等应用。

How Endava builds an agentic organization with Codex

Learn how Endava uses Codex to build an agentic organization, accelerating software delivery and reducing requirements analysis from weeks to hours.

中文介绍 Endava利用Codex构建代理型组织,加速软件交付进程,并将需求分析时间从数周缩短至数小时。

The AI Hype Index: AI gets booed in graduation season

It is one thing to say AI will change the world. It is another to expect the class of 2026 to applaud it. In fact, when former Google CEO Eric Schmidt told University of Arizona graduates that their task is to help shape AI, he was met with a resounding chorus of boos. “I can…

中文介绍 AI热潮指数显示,在毕业季中AI技术遭到嘘声。前谷歌CEO Eric Schmidt在亚利桑那大学演讲时,呼吁学生帮助塑造AI,却受到听众嘘声回应。

[AINews] Cognition raises $1B in $26B Series D

coding is an uncapped TAM market

中文介绍 Cognition在D轮融资中筹集10亿美元,公司估值达260亿美元,同时指出编码市场具有无限增长潜力。

OpenAI’s Frontier Governance Framework

Explore OpenAI’s Frontier Governance Framework and how our AI safety, security, and risk practices align with emerging EU and California regulations.

中文介绍 OpenAI推出前沿治理框架,阐述其AI安全、安全防护和风险管理实践如何与新兴的欧盟及加州法规保持一致。

🔬ESM: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub

Biohub’s Protein World Model: ESMC-6B, ESMFold2, 6.8B proteins, 1.1B structures, antibody design, SAEs, & the potential for programmable biology

中文介绍 Biohub的蛋白质世界模型包括ESMC-6B和ESMFold2,处理6.8B蛋白质和1.1B结构数据,应用于抗体设计、稀疏自编码器及可编程生物学领域。

Cisco and OpenAI redefine enterprise engineering with Codex

Cisco and OpenAI are redefining enterprise engineering with Codex, helping Cisco scale AI-native development, accelerate AI Defense work, and automate defect remediation.

中文介绍 Cisco与OpenAI利用Codex重新定义企业工程,帮助Cisco扩展AI原生开发、加速AI防御工作并实现缺陷修复自动化。

Building self-improving tax agents with Codex

See how OpenAI, Thrive, and Crete built a self-improving tax agent with Codex, automating filings, improving accuracy, and accelerating workflows.

中文介绍 OpenAI、Thrive和Crete使用Codex构建自我改进的税务代理,实现税务申报自动化、提升准确性并加快工作流程。

Warp’s big bet on building open source with GPT-5.5

Warp uses GPT-5.5 and OpenAI models to coordinate coding agents across local, cloud, and open-source development workflows.

中文介绍 Warp采用GPT-5.5和OpenAI模型协调编码代理,支持跨本地、云和开源开发环境的工作流整合。

Election information and safeguards in 2026

Ahead of global elections, we’re helping people access information, supporting cyber defenders, and increasing AI transparency

中文介绍 面向2026年全球选举,OpenAI致力于帮助公众获取信息、支持网络防御者并增强AI技术的透明度。

Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

第一作者: Richard J. Young · 方向: 软件安全

Abstract:A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a working weapon -- a keylogger, a ransomware stub, an exploit that runs as written. This asymmetry in the severity of a single act of compliance implies coding-specialized models should clear a higher refusal bar than general-purpose chat models, not a lower one, yet the field cannot presently tell whether they do. Refusal benchmarks for malicious code are fragmented: they mix requests for executable software (ready-to-run weapons) with requests for harmful security knowledge (information a human must still operationalise) and report refusal rates over non-comparable corpora, so no single statistic measures the property that actually matters. This paper introduces an expanded consensus-labeled prompt bank that distinguishes between...

论文介绍 本文研究编码模型对恶意代码请求的合规性评估问题。现有基准碎片化,混淆了可执行恶意代码和有害安全知识。作者引入了一个共识标注的提示库,区分不同类型请求,以准确衡量编码模型的拒绝行为。这对于提高AI编码工具的安全标准至关重要,防止其生成可执行的恶意软件。

Efficient and Quantum-safe Internet Key Exchange Protocols for Satellite Communications

第一作者: Davide De Zuane · 方向: 密码学协议

Abstract:This paper studies cryptographic key exchange in satellite communications, which requires specific solutions because the satellite context presents unique challenges, particularly concerning onboard resource constraints and long transmission latency. We address these challenges by considering the Internet Key Exchange (IKE) protocol, which is widely used in terrestrial networks, and studying its applicability in the satellite context. This requires addressing two main issues: i) its efficiency in terms of the resources and bandwidth required to adapt to satellite terminals, and ii) its resistance even to attackers equipped with a quantum computer, in order to resist obsolescence and defend against harvest-now-decrypt-later attacks. We study these aspects from both a design and experimental point of view, defining and assessing some protocol variants characterized by low...

论文介绍 本文探讨卫星通信中密钥交换协议的设计与评估。针对卫星环境的资源限制和长传输延迟,作者基于IKE协议研究其适用性,重点解决效率和量子安全问题。通过定义和评估协议变体,旨在开发适用于卫星通信的高效且抗量子攻击的密钥交换方案,以增强通信安全。

MaskClaw: Edge-Side Personalized Privacy Arbitration for GUI Agents with Behavior-Driven Skill Evolution

第一作者: Yanqiu Zhao · 方向: 系统安全

Abstract:GUI agents rely on screenshots to infer intent and operate across applications, but these screenshots often contain private messages, medical records, payment credentials, and workplace-specific workflows. Privacy decisions in this setting depend on task, recipient, application state, and user role, yet static PII detectors miss these boundaries and cloud-side VLM reasoning can upload the raw screen before deciding what should be protected. We present MaskClaw, an edge-side privacy arbitrator for GUI agents. MaskClaw extracts local visual evidence, retrieves user- and task-specific policy memory, and decides Allow, Mask, or Ask before raw screenshots leave a trusted user- or organization-controlled environment. In five designed skill-evolution scenarios, it turns corrections, cancellations, and edits into reusable privacy skills checked by a sandbox gate. We introduce...

论文介绍 本文提出MaskClaw,一个面向GUI代理的边缘侧隐私仲裁器。GUI代理使用截图操作应用时,截图常包含敏感信息。MaskClaw在本地提取视觉证据,参考用户策略,在截图离开信任环境前做出遮罩或询问决策。通过行为驱动的技能进化,系统能从纠正中学习,提升隐私保护能力。

GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction

第一作者: Jinze Gu · 方向: 隐私保护

Abstract:Retrieval-Augmented Generation (RAG) enhances LLMs by grounding generation in query-relevant external evidence. Beyond unstructured text corpora, Graph RAG integrates knowledge graphs into the retrieval pipeline, enabling LLMs to access entities, relations, and multi-hop dependencies encoded in structured knowledge. However, the same structured knowledge that empowers Graph RAG also creates a new privacy attack surface. We demonstrate that Graph RAG systems can be turned into structural oracles: through adaptive black-box interactions, an adversary can elicit sufficient relational evidence to reconstruct substantial portions of the hidden knowledge graph. We propose a structure-oriented reconstruction framework that recovers targeted graphs from both local and global perspectives. Specifically, Depth-Wise Heuristic Search extracts fine-grained node attributes by recursively...

论文介绍 本文揭示Graph RAG系统面临的知识窃取风险。通过自适应黑盒交互,攻击者可重建隐藏的知识图谱结构。作者提出一个结构导向重建框架,从局部和全局视角恢复目标图,使用深度优先启发式搜索提取节点属性。这突显了Graph RAG的隐私漏洞,为防御提供参考。

Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking

第一作者: Ziyang You · 方向: AI 安全

Abstract:Cryptographic watermarking is a leading defense for attributing text generated by large language models (LLMs). Existing schemes, including KGW, Unigram, and DipMark, derive their security guarantees from the assumption that the underlying pseudo-random number generator (PRNG) is trustworthy. This work introduces SeedHijack, the first supply-chain attack on LLM watermarking that is simultaneously (i) blind -- requiring no knowledge of the watermark key, detector, or model logits, (ii) integrity-preserving -- amplifying rather than erasing the watermark signal, and (iii) orthogonal to detection -- the attack-induced bias is statistically independent of all content-side detector statistics, ensuring that amplification and evasion coexist without trade-off. Rather than perturbing generated text, SeedHijack replaces the PRNG at the supply-chain layer, biasing green-list selection...

论文介绍 本文提出SeedHijack,一种针对LLM水印的供应链攻击。该攻击无需水印密钥,在供应链层替换伪随机数生成器,偏置绿名单选择,从而放大水印信号而非擦除。这种盲且完整性保持的攻击,挑战了现有水印方案的安全性假设,表明需要更鲁棒的防御机制。

Position: Retire the "Positive Backdoor" Label -- Secret Alignment Requires Strict and Systematic Evaluation

第一作者: Jianwei Li · 方向: 安全研究

Abstract:This position paper argues that the AI/ML community should stop overclaiming and retire the label "positive backdoor," and instead treat trigger-activated hidden behaviors as Secret Alignment. Crucially, protective claims based on Secret Alignment should be presumed not secure by default unless supported by rigorous, standardized evaluation. The Private AI era, enabled by open-weight LLMs and accessible training/inference stacks, turns language models into privately owned digital assets, creating security concerns around unauthorized access, model theft, and behavioral misuse. Recently, a line of work framed as "positive backdoors" has been proposed to address these challenges. To ground our position in evidence, we unify these proposals as covert trigger-behavior associations for access gating, ownership attribution, and safety enforcement, and evaluate three representative...

论文介绍 本文主张停止使用「正面后门」标签,将触发器激活的隐藏行为重新定义为秘密对齐。作者强调,基于秘密对齐的保护性声明应默认视为不安全,除非经过严格标准化评估。通过统一提案并评估代表方案,论文旨在规范安全术语,促进更可靠的AI安全实践。

Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem

第一作者: Luca Beurer-Kellner · 方向: 安全研究

Abstract:We analyzed 3,984 AI agent skills from major marketplaces and found 76 confirmed malicious payloads, including credential theft, backdoor installation, and data exfiltration. 13.4% of all skills contain at least one critical-level security issue and at least 8 manually confirmed malicious skills remain publicly available on this http URL as of the date of publication. This report documents our methodology, presents a threat taxonomy based on real-world samples, and details the attack patterns we observed. As skill marketplaces grow rapidly and AI agents gain access to sensitive credentials and systems, automated security analysis is no longer optional.

论文介绍 本技术报告分析AI代理技能生态系统中的新兴威胁。通过对3,984个技能的审查,作者发现76个恶意负载和多个关键安全问题。报告建立了基于真实样本的威胁分类,详细描述攻击模式,并指出自动化安全分析的紧迫性,以应对技能市场快速增长带来的风险。

Do you dare to try Test-Driven Forensics? Increasing Trust in Desktop Forensics with ADARE

第一作者: Michael Külper · 方向: 系统安全

Abstract:Digital forensic relies on validated tools and established procedures, yet the underlying operating systems, applications, and analysis tools evolve rapidly. This evolution can cause artifact behavior and tool outputs to drift, silently degrading repeatability and confidence in long-lived forensic interpretations. We present test-driven forensics, a practical approach that treats forensic expectations as executable specifications: expected artifacts and expected tool outputs are encoded as tests that can be rerun across versions to detect regressions. Crucially, our approach also enables State Transition Testing, validating the system's expected state after each user action rather than only performing post-mortem checks on a final disk image; this supports causal attribution and makes transient behavior testable. We implement the methodology in ADARE, an open-source framework...

论文介绍 本文提出测试驱动取证方法,用ADARE框架增强桌面取证的可信度。随着系统演化,取证工具输出可能漂移,影响可重复性。ADARE将取证预期编码为可执行测试,支持跨版本回归检测和状态转换测试,从而验证系统状态变化,提高取证的因果归属和长期可信度。

Towards Cybersecurity SuperIntelligence (CSI): What's the best harness for cybersecurity?

第一作者: Víctor Mayoral-Vilches · 方向: AI 安全

Abstract:What is the best harness for cybersecurity AI? Cybersecurity systems are converging on a single execution scaffold per agent, an iterative shell loop driven by a Large Language Model (LLM). However, scaffolds are not interchangeable, rarely interoperable, and no single scaffold dominates across all challenge types. In our path towards researching Cybersecurity SuperIntelligence (CSI), we present a meta-scaffold that unifies heterogeneous agent harnesses under a common orchestration layer, enabling any LLM-driven scaffold to be deployed, benchmarked, and composed within the same infrastructure. Using CSI, we benchmark five scaffolds (CSI::Claude, CSI::Codex, CSI::GCAI, CSI::Mistral, CSI::CAI) on the 33 cybench challenges, holding the model fixed at alias2-mini. The best individual scaffolds solve 15/33 (45.5%); the four-scaffold union solves 17/33 (51.5%), with the fifth...

论文介绍 本文聚焦网络安全AI代理的执行框架问题,提出一种名为CSI的元脚手架架构。该架构旨在统一多种异构的LLM驱动代理脚手架,并在统一的基础设施下进行部署和基准测试。研究基于cybench的33项挑战对五种脚手架进行评测,结果表明,单一脚手架最高可解决约45.5%的挑战,而组合多个脚手架能将成功率提升至51.5%。这项工作为构建更强大、灵活的网络安全智能系统提供了重要的集成与评估思路。

Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems

第一作者: Chenxi Wang · 方向: AI 安全

Abstract:Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for efficient and flexible agent collaboration. However, moving coordination into latent space may also move attacks beyond the reach of visible-text inspection. In this paper, we study whether latent states can carry attack-associated information that remains effective during clean executions. To examine this question, we introduce a latent attack framework that reactivates attack-induced effects through latent interventions without reusing adversarial text. Extensive experiments show that the resulting latent-only attacks can substantially degrade task performance in clean executions, especially when applied to inter-agent KV-cache handoffs rather than local hidden states. Further control analyses indicate that this degradation cannot be...

论文介绍 该研究关注基于潜变量的多智能体系统中的安全风险。研究者指出,将协调信息转移到潜空间可能使攻击隐藏于可见文本检查之外。本文提出了一个潜变量攻击框架,能够通过潜变量干预而非重复使用对抗文本来激活攻击效果。实验表明,这种纯潜变量攻击能在正常执行中严重降低任务性能,尤其是在智能体间的KV缓存移交阶段。这揭示了新型多智能体系统中潜藏的安全威胁。

Cybersecurity AI (CAI) Dataset

第一作者: Víctor Mayoral-Vilches · 方向: AI 安全

Abstract:We present CAI Dataset, a fourteen-month corpus of cybersecurity LLM trajectories collected through the open-source CAI agent framework, built in response to PentestGPT's finding that expert operator trajectories, not base-model capability, are the bottleneck for cybersecurity LLM performance. CAI Dataset aggregates 230,935 session logs and 26,027,742 user prompts from 16,768 source IPs across 123 countries, exercising 4,187 unique LLM identifiers against 23,147 target domains over 18.07 TB of durable storage. The mix is hands-on (36.4% offensive, 20.1% attacker-intent, 27.5% business / integration, 4.4% defensive), making CAI Dataset, to the best of our knowledge, the largest described corpus of LLM-driven hacker trajectories. It is released to partner organisations and selected customers as an audience-size series (CAI Dataset10, CAI Dataset1k, CAI Dataset200k). Read...

论文介绍 本文介绍了CAI数据集,这是一个历时十四个月、通过开源CAI代理框架收集的大型网络安全LLM轨迹语料库。数据集包含超过23万条会话日志和2600万条用户提示,涵盖了攻击性、防御性和业务集成等多种场景。据作者所述,这是目前规模最大的、公开描述的LLM驱动黑客轨迹数据集。该数据集的发布旨在促进网络安全AI的研究与训练,解决领域内专家操作轨迹数据稀缺的瓶颈问题。

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

第一作者: Yubin Qu · 方向: AI 安全

Abstract:A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while the task still completes. We call this overeager behavior: the prompt is not adversarial and the run succeeds, yet an out-of-scope step can leak credentials or delete files. Existing benchmarks miss it: task-completion suites credit any finished run, jailbreak suites probe adversarial prompts, and the one prior overeager benchmark applies a single fixed prompt set to every agent-model pair, leaving its easiest and most resistant pairs under-measured. We present SNARE (Synthesizing Non-adversarial scenarios for Adaptive Reward-guided Elicitation), a pipeline that composes benign scenarios from reusable scope and trap fragments, scores each run with a judge-free oracle flagging trap-pattern matches and unsolicited file additions or...

论文介绍 编码代理在执行良性任务时,可能会无意中超出授权范围,执行「过度热心」行为,例如泄露凭证或删除文件,而现有基准测试难以有效衡量此类风险。本文提出了SNARE流程,它通过从可重用的范围和陷阱片段中合成良性场景,来系统性地诱发和评估此类行为。该流程使用无评判的预言机对运行结果进行评分。SNARE旨在为编码代理提供更精细、更具适应性的安全评估方法,以发现那些任务成功但存在越权操作的隐蔽风险。

MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content

第一作者: Ruoqi Guo · 方向: AI 安全

Abstract:Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from what they see, so they cannot reliably separate trusted interface elements from user-generated content. We present MIRAGE (Mobile Injection of Realistic Adversarial GUI Examples), a pipeline that turns benign mobile screenshots into prompt-injection samples by placing attacker-controlled text into ordinary user-generated content regions, without modifying the agent, the application, or the operating system. MIRAGE operates in three stages: a Localizer identifies user-controllable regions on the screenshot, a Generator synthesises context-aware payloads and renders them in the application's native style, and a Curator moderates realism and balances the samples across applications, region types, and attack intents. A key challenge is...

论文介绍 针对由视觉语言模型驱动的移动GUI代理,本文提出了MIRAGE攻击流程。该攻击的核心在于,代理将屏幕像素作为输入,难以区分可信界面元素与用户生成内容。MIRAGE通过在用户可控的屏幕区域(如评论区)植入攻击者控制的文本,生成上下文感知的提示注入样本,从而劫持代理行为,且无需修改代理、应用或操作系统。该研究揭示了移动GUI代理在处理用户生成内容时存在的根本性安全漏洞。

A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG

第一作者: Junjie Mu · 方向: 网络安全

Abstract:Federated Retrieval-Augmented Generation (FedRAG) is attractive for privacy-sensitive applications because raw data remain local. As a result, routing must rely on client-provided semantic profiles, creating a new opportunity for manipulation. We introduce Routing Hijacking, a routing-stage attack in which a malicious client forges its profile to attract target queries despite having irrelevant underlying data. We show that this vulnerability is severe. Across three representative FedRAG routing architectures, Routing Hijacking consistently misroutes target queries and leads to downstream disruptions and failures, including missing evidence, poisoning, incorrect answers, and hallucinations. In a high-stakes MedQA-USMLE case study, we further show that poisoned retrieved evidence can mislead models across scales, leading to incorrect answers, hallucinations, and sycophantic...

论文介绍 在隐私保护的联邦检索增强生成系统中,客户端通常基于语义配置文件进行查询路由。本文揭示了一种「路由劫持」攻击:恶意客户端可以伪造其语义配置文件,以吸引与其数据无关的目标查询。研究表明,该攻击能在多种联邦RAG路由架构中成功误导查询,导致检索到无关证据、产生错误答案或幻觉。在医学问答的案例研究中,这种投毒检索证据的行为甚至能误导不同规模的模型,凸显了该漏洞的严重性。

Mind the Gap: Mixtures of Gaussians in Approximate Differential Privacy

第一作者: Huikang Liu · 方向: 隐私保护

Abstract:We design a class of additive noise mechanisms that satisfy \((\varepsilon, \delta)\)-differential privacy (DP) for scalar, real-valued query functions with known sensitivities, with a particular focus on moderate and low-privacy regimes. These mechanisms, which we call \textit{mixture mechanisms}, are constructed by mixing multiple Gaussian distributions that share the same variance but differ in their means and mixture weights. The resulting distributions can be interpreted as convex combinations of a zero-mean Gaussian (as used in the analytic Gaussian mechanism) and additional Gaussians whose means depend on the sensitivity of the query function. We derive tight conditions on the variances required for \((\varepsilon, \delta)\)-DP and provide efficient algorithms to compute them. Compared to the analytic Gaussian mechanism, our mechanisms yield substantially lower expected...

论文介绍 本文针对已知灵敏度的标量查询函数,设计了一类满足 (ε, δ)-差分隐私的加性噪声机制,特别关注中等和低隐私保护水平。这些「混合机制」通过混合多个具有相同方差但不同均值和混合权重的高斯分布来构建。研究推导了实现差分隐私所需方差的紧条件,并提供了高效的计算算法。与标准的分析高斯机制相比,所提出的机制在相同隐私保证下能显著降低期望噪声,从而提升数据效用。

SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adversarial Data Poisoning

第一作者: Jiachen Qian · 方向: AI 安全

Abstract:Retrieval-Augmented Generation (RAG) mitigates LLM hallucinations but introduces a critical vulnerability: corpus integrity. We present SilentRetrieval, a two-stage data poisoning attack that hijacks RAG systems through adversarially crafted yet fluent documents. Stage 1 uses Coordinated Beam Search, a multi-token joint optimization method with a fluency-similarity objective, to keep a poisoned host document retrievable while constraining perplexity. Stage 2 uses Context-Adaptive Trigger Generation, a lightweight trigger-fusion step driven by a frozen LLM, to integrate manipulation triggers into document content. Under a one-poisoned-document-per-query evaluation with synthetic target answers, SilentRetrieval achieves 84.6%/81.3% HR@10 and 57.5%/54.8% ASR-LLM on Natural Questions and MS MARCO, while maintaining near-benign perplexity. Cross-model evaluation across four target...

论文介绍 检索增强生成系统面临语料库完整性被破坏的风险。本文提出SilentRetrieval,一种两阶段数据投毒攻击,能够通过精心构造但看似流畅的恶意文档来劫持RAG系统。第一阶段使用「协调波束搜索」优化文档的流畅性和可检索性,第二阶段利用「上下文自适应触发生成」将操纵触发器融入文档。实验表明,该攻击在保持文档近乎正常困惑度的同时,能以较高成功率将目标查询的检索结果导向投毒文档,从而操纵生成答案。

AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent

第一作者: Jiaqi Luo · 方向: 软件安全

Abstract:LLM-based agents have recently attracted significant attention due to their ability to autonomously invoke relevant tools to accomplish complex tasks. However, recent studies have shown that these agents face severe security risks, which may lead to privacy leakage, financial loss, or even full system compromise. In this paper, we present AgentGuard, an attribute-based access control framework for tool-use LLM-based agents. AgentGuard adopts a client-server architecture. On the client side, AgentGuard provides lightweight integration for agents implemented in different programming languages and architectures. It requires only minor code modifications (e.g., around 10 lines) without changing the underlying agent execution logic. On the server side, AgentGuard provides three complementary inspection mechanisms to cover both single-tool and cross-tool security risks in agent...

论文介绍 本文针对LLM代理在调用工具时面临的安全风险问题,提出了AgentGuard框架。该框架采用客户端-服务器架构,客户端允许不同实现的代理以少量代码修改接入,服务器端则提供单工具和跨工具的三种互补检查机制,以实现细粒度的访问控制,旨在保护代理系统免受隐私泄露和系统破坏等威胁。

Can It Reach the Generator? Investigating the Survival of Prompt-Injection Attacks in Realistic RAG Settings

第一作者: Yu Yin · 方向: AI 安全

Abstract:Recent generative engine optimisation (GEO) research has shown that prompt-injection attacks can push a target product to the top of an LLM's recommendation list, with the strongest attacks reporting around $80\%$ success and raising serious security concerns about RAG-based recommendation. However, these results assume the attacked document is always fed directly to the generator, bypassing the retriever and reranker. This is unrealistic: in deployed RAG systems, the attack modifies the document content, which can in turn change whether the document is retrieved and reranked highly enough to reach the generator at all. In this paper, we re-evaluate seven GEO attacks under a realistic three-stage pipeline (retriever\,$\to$\,LLM reranker\,$\to$\,LLM generator). We find that prior protocols substantially overstate attack effectiveness: gradient-based and instruction override...

论文介绍 研究评估了提示注入攻击在现实检索增强生成(RAG)系统中的有效性。现有研究通常假设攻击文档总能直达生成器,但本文在包含检索器、重排序器和生成器的真实三阶段流水线中重新评估了七种攻击。研究发现,先前的评估高估了攻击效果,许多攻击在文档无法被检索或重排序时将失效。

Privately Estimating Monotone Statistics in Polynomial Time

第一作者: Gavin Brown · 方向: 安全研究

Abstract:We study efficient differentially private algorithms for estimating monotone statistics, i.e., statistics that are monotone under the addition of new observations. The starting point for our investigation is subsample-and-aggregate: a classical paradigm that partitions the dataset into blocks, estimates the statistic on each block, and then privately aggregates the this http URL practical and generically applicable, this approach is quite data-hungry. We improve upon this framework for the class of monotone statistics -- compared to subsample-and-aggregate, our algorithms save a factor of $t$ in sample complexity and pay a factor of $e^t$ in running time, where $t>0$ is a tunable parameter. We complement our results with a query-complexity lower bound, showing that our algorithms are essentially optimal for this task. As an application, we obtain improved results for private...

论文介绍 本文研究如何在多项式时间内高效地进行差分隐私单调统计量估计。作者改进了经典的“子采样-聚合”框架,针对单调统计量特性,提出新算法可在样本复杂度上节省一个因子t,但运行时间增加e^t因子。论文还给出了查询复杂度下界,证明算法几乎最优。该结果可应用于私密均值估计等任务。

Symmetry Defeats Auditing

第一作者: Nick Merrill · 方向: 安全研究

Abstract:We demonstrate an attack on Introspection Adapters (Shenoy et al., 2026).

论文介绍 本文展示了一种针对“内省适配器”技术的攻击。内省适配器是一种旨在帮助人类理解或审计AI模型内部状态的技术。研究演示了如何利用模型的对称性来击败这类审计机制,揭示了当前此类安全技术可能存在的根本性漏洞。

MRMMIA: Membership Inference Attacks on Memory in Chat Agents

第一作者: Kai Chen · 方向: AI 安全

Abstract:Membership inference attacks (MIAs) test whether a target data record belongs to a system's private data, and have become a standard tool to measure privacy leakage in machine learning systems. Prior work has primarily focused on training corpora or retrieval databases. However, MIAs against agent memory have received less attention, even though such memory can contain sensitive user-agent interactions, retrieved facts, and user preferences. Therefore, in this work, we focus on chat agent memory MIAs, where an adversary infers whether a candidate memory unit belongs to the chat agent's memory store. We propose Multi-Recall Memory MIA (MRMMIA), a unified attack that utilizes multiple recall probes to the agent to extract the membership signal across black-box, gray-box, and white-box settings. Our experiments demonstrate that MRMMIA consistently outperforms baselines. Our...

论文介绍 本文研究针对聊天代理记忆存储的成员推断攻击。现有的成员推断攻击主要关注训练语料库,而代理记忆包含敏感交互信息却少有研究。为此,作者提出了多召回记忆成员推断攻击,通过向代理发送多个探测查询来提取成员信号,该方法适用于黑盒、灰盒和白盒环境,实验表明其效果优于基线方法。

Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security

第一作者: Xiang Fang · 方向: AI 安全

Abstract:Large Language Models (LLMs) are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulting in harmful or inappropriate outputs. Such attacks, including jailbreaking and prompt injection, pose significant risks to the integrity and availability of LLMs in security-critical applications. This paper proposes the Adversarial Prompt Disentanglement (APD) framework, a novel defense mechanism that proactively identifies and neutralizes malicious components in input prompts before they are processed by the LLM. The APD framework integrates three key innovations: (1) a mutual information-based semantic decomposition method to isolate adversarial and benign prompt components, ensuring statistical independence; (2) a graph-based intent classification approach that leverages spectral analysis to detect malicious patterns in...

论文介绍 为应对针对大语言模型的对抗性提示攻击(如越狱和提示注入),本文提出对抗性提示解缠框架。该框架在LLM处理输入前,主动识别并中和恶意组件。其核心包括基于互信息的语义分解方法和基于图谱的意图分类技术,通过分离和分析提示的语义结构来增强模型的安全性。

Patchlings: Safety-Preserving Flash-Based Hotpatching for Automotive Microcontrollers

第一作者: Yuxin "Myles" Liu · 方向: 软件安全

Abstract:The increasing presence of software in modern automobiles has created a growing need to deliver software updates throughout a vehicle's entire lifespan. Traditional update methods are slow and require months of re-validation to comply with stringent safety standards like ISO 26262. Although hotpatching offers a path to faster updates, existing solutions for real-time embedded systems are unsuitable for the automotive domain: they overlook regulatory compliance, demand extensive safety validation, and lack support for the flash-based Execute-in-Place (XIP) architecture commonly used in automotive electronic control units (ECUs). We introduce Patchlings, the first hotpatching framework designed for compliance, safety, and persistence in automotive systems. It fills the gap in applying hotpatching to automotive systems and fundamentally reduces the mean-time-to-mitigate (MTTM)...

论文介绍 本文介绍Patchlings,首个专为汽车系统设计的基于闪存的热补丁框架。传统汽车软件更新缓慢,需要漫长的重新验证。Patchlings旨在符合ISO 26262等严格的安全标准,支持汽车电子控制单元常用的就地执行架构,通过热补丁技术显著减少漏洞缓解的平均时间,实现安全、持久的快速更新。

HammerSim: A System-Level Tool to Model RowHammer

第一作者: Kaustav Goswami · 方向: 软件安全

Abstract:Modern architecture research relies on simulators to evaluate system security, yet analyzing emerging hardware vulnerabilities like RowHammer requires full-system visibility. As RowHammer vulnerabilities worsen with continuous technology scaling, existing simulators lack the system-level models needed to study complex OS effects and cross-layer mitigations. This tool deficiency leaves modern computing platforms exposed to severe reliability and security risks. In this work, we present HammerSim, a gem5-based framework for modeling RowHammer at the full-system level. HammerSim integrates probability-driven bitflip modeling to realistically capture the behavior of RowHammer. It further enables evaluation of hardware and software mitigations such as TRR and selective ECC. We validate HammerSim's bitflip modeling against real DDR4 DIMMs using JS divergence, demonstrating its...

论文介绍 本文提出HammerSim,一个基于gem5的系统级RowHammer建模工具。现有模拟器缺乏研究RowHammer漏洞所需的系统级视角,难以评估操作系统效应和跨层缓解措施。HammerSim集成了概率驱动的位翻转模型,能够模拟RowHammer行为,并支持对TRR、选择性ECC等硬件和软件缓解方案进行评估。

Intent-based Security Management Using the TM Forum TR292I Security Ontology

第一作者: Loay Abdelrazek · 方向: 系统安全

Abstract:Modern 5G-Advanced and emerging 6G cloud-native telecom architectures encounter unprecedented hyper-complexity, multi-layered threat vectors, and fluid structural topologies. Managing infrastructure security using manual, imperative configurations introduces a severe latency gap, presenting attackers with an exploitable window. This paper presents a declarative, autonomous, self-protecting framework based on our design and standardization of the TM Forum TR292I Security Ontology v4.0.0. Our approach leverages Description Logic (DL) and automated graph reasoning within a closed-loop execution pipeline to dynamically neutralize live threats. Crucially, the system balances functional protection expectations with non-functional resource impact considerations (e.g., latency vs. compute overhead). We validate our model-driven architecture through a structural formal verification...

论文介绍 针对5G/6G云原生电信架构的复杂安全威胁,传统手动配置存在响应延迟。本文基于TM Forum TR292I安全本体论,提出一个声明式、自主的自保护框架。该框架利用描述逻辑和自动化图推理,在闭环执行管线中动态中和实时威胁,并权衡功能保护与非功能资源影响。研究通过结构化形式验证了其模型驱动架构的有效性。

QSignAI: Quantum-Randomness-Seeded Identity Signatures at the Intersection of AI for Science and Science for AI

第一作者: Dongping Liu · 方向: AI 安全

Abstract:The 2024--2025 Nobel and Turing awards recognised artificial intelligence and quantum science in the same breath -- machine learning as a physical science, artificial intelligence solving 50-year scientific problems, superconducting quantum circuits as the hardware foundation of quantum computing, and quantum information principles as computing's highest achievement. Yet no deployed artificial intelligence system has brought these two streams together for the general public: identity systems still rely on pseudo-random tokens, and quantum circuits remain invisible to the billions of people who use bot-enabled social messaging platforms daily. This paper presents QSignAI, a production-deployed open-source platform demonstrating a bidirectional relationship between artificial intelligence and quantum science in a real-time event participation system. We address three research...

论文介绍 当前身份系统依赖伪随机数,而量子计算技术未与大众AI应用结合。本文介绍了QSignAI,一个已部署的生产级开源平台,展示了人工智能与量子科学在实时事件参与系统中的双向融合关系。该研究旨在利用量子随机性为身份签名提供安全基础,并解决了相关的跨学科集成挑战。

AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?

第一作者: Zongheng Cao · 方向: 安全研究

Abstract:Video production workflows offer a rich and demanding arena for evaluating multimodal AI agents: they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use. To this end, we introduce AgenticVBench, a benchmark of 100 agentic tasks across 4 task families spanning the real world post-production workflow, constructed from real production workflows contributed by 20 industry experts averaging 6 years of professional experience. Tasks are paired with evaluation specifications that combine programmatic verifiers and expert rubrics. We evaluate frontier vision-language models (VLMs) with both vendor-native and open-source harnesses. The best evaluated agent stack barely crosses 30%, far below human expert performance on the same tasks. We further find that the choice of harness substantially affects model...

论文介绍 视频后期制作流程需要复杂的多模态能力。为评估AI代理的实战能力,本文提出了AgenticVBench基准,包含来自20位行业专家的100个真实世界任务。评估发现,当前最强的视觉语言模型代理性能远低于人类专家,且执行框架的选择对性能影响显著,凸显了现有模型在复杂工具链任务上的不足。

Backdoor Attacks on Fault Detection and Localization in Cyber-Physical Systems

第一作者: Abile Jean · 方向: AI 安全

Abstract:Cyber-Physical Systems (CPS) integrate sensing, communication, computation, and control to support critical infrastructure, including smart grids, industrial automation, and control systems. In the electrical utility domain, various controllers are used in CPS to ensure the system detects and recovers from faults, such as voltage fluctuations, and to perform load balancing in distribution systems. Machine learning- and deep learning-based fault detection and localization frameworks have recently gained significant attention in CPS for their ability to identify anomalies and operational failures in real time. However, these intelligent models are vulnerable to adversarial machine learning attacks, particularly backdoor attacks. In a backdoor attack, an adversary injects malicious patterns into the training data so that the model behaves normally most of the time but produces...

论文介绍 机器学习模型在网络物理系统故障检测中广泛应用,但易受后门攻击威胁。本文研究了针对CPS故障检测与定位模型的后门攻击。攻击者在训练数据中注入恶意模式,使模型在正常情况下表现正常,但在特定触发下产生错误判断,这对智能电网等关键基础设施构成严重风险。

Silent Consent, Persistent Risk: Android Permission Groups and Custom Permissions

第一作者: Olawale Amos Akanji · 方向: 系统安全

Abstract:Android's permission system is designed to balance usability with informed consent, yet two legacy mechanisms still undermine that balance in Android 16: (i) permission groups that silently auto-grant new permissions within a group after a user's initial approval, and (ii) normal-level custom permissions that are auto-granted at install and enable cross-app access with no user visibility. We conduct a longitudinal analysis of 19.3 million APKs spanning 5.97 million unique apps (distinct package identifiers) from the AndroZoo repository, combined with on-device validation on Android 16. Among 2,244,575 multi-version apps, 381,026 (17%) silently gain permissions within already-granted groups. Using VirusTotal detections with primary threshold t=20, apps flagged as malware expand within groups at a higher rate than benign apps (odds ratio = 1.35, p < 0.001); the association holds...

论文介绍 Android的权限系统存在遗留问题。本文通过对超过1900万APK的纵向分析发现,权限组机制会静默授予新权限,而普通级别的自定义权限可在安装时被自动授予。研究揭示,被标记为恶意软件的应用在已授权权限组内扩张权限的速率更高,指出了当前权限模型在知情同意和安全管控方面的持续性风险。

A Note on Boosting Uncloneable Encryption in Microcrypt

第一作者: James Bartusek · 方向: 密码学协议

Abstract:In this note, we consider the setting of uncloneable encryption satisfying uncloneable indistinguishability, a form of symmetric key encryption that prevents the cloning of ciphertexts in a very strong sense. Our goal is to minimize the assumptions under which (many-time secure) uncloneable encryption is known to exist, assuming the existence of an information-theoretic "uncloneable bit", i.e. a one-time secure uncloneable encryption scheme for one-bit messages. We observe that if a t -> t' uncloneable bit exists, then the following implications hold. 1. If many-time secure symmetric key encryption exists, then many-time secure t -> t' uncloneable encryption for arbitrary-length messages exists. Since many-time secure uncloneable encryption implies many-time secure symmetric key encryption, this result is tight. 2. If pseudorandom unitaries exist, then many-time secure t -> t'...

论文介绍 本文研究在微密码学框架下,如何最小化构建(多次安全)不可克隆加密方案所需的底层假设。研究指出,如果存在一个t→t'的不可克隆比特,则基于多项式时间安全的对称加密或伪随机酉的存在性,即可推导出针对任意长度消息的多次安全不可克隆加密方案,为该方向的密码学构造提供了更基础的理论依据。

Poison with Style: A Practical Poisoning Attack on Code Large Language Models

第一作者: Khang Tran · 方向: 软件安全

Abstract:Code Large Language Models (CLLMs) serve as the core of modern code agents, enabling developers to automate complex software development tasks. In this paper, we present Poison-with-Style (PwS), a practical and stealthy model poisoning attack targeting CLLMs. Unlike prior attacks that assume an active adversary capable of directly embedding explicit triggers (e.g., specific words) into developers' prompts during inference, PwS leverages developers' code styles as covert triggers implicitly embedded within their prompts. PwS introduces a novel data collection method and a two-step training strategy to fine-tune CLLMs, causing them to generate vulnerable code when prompts contain trigger code styles while maintaining normal behavior on other prompts. Experimental results on Python code completion tasks show that PwS is robust against state-of-the-art defenses and achieves high...

论文介绍 针对代码大语言模型(CLLM)的现有投毒攻击多依赖显式触发器。本文提出Poison-with-Style(PwS)攻击,利用开发者的代码风格作为隐式触发器。通过精心设计的数据收集和微调策略,使模型在接收到包含特定风格的提示时生成漏洞代码,而对其他提示行为正常。该攻击对现有防御具有鲁棒性。

Assessor Experiences in CMMC Level 2 Certification Assessments: An Interpretative Phenomenological Analysis of Role Expectations

第一作者: Samuel Heuchert · 方向: 软件安全

Abstract:The Cybersecurity Maturity Model Certification program requires third-party assessments be conducted under a non-consultative model. The model is intended to ensure impartiality for organizations seeking certification. While this structure defines expectations for assessor behavior, assessor experiences and interpretations of these constraints remain underexamined. The study examines the lived experiences of CMMC-Certified Assessors and how they navigate role expectations within the non-consultative model. Using Role Conflict Theory as a guiding framework, Interpretative Phenomenological Analysis (IPA) was applied to semi-structured interviews to explore how assessors make sense of their roles. The analysis identified experiential themes that describe how assessors construct professional credibility, execute structured assessment work, and manage the practical challenges of...

论文介绍 美国网络安全成熟度模型认证要求进行第三方非咨询式评估。本研究采用解释现象学分析方法,探究CMMC评估员在此模式下的实际经历。研究发现,评估员在构建专业可信度、执行结构化评估工作以及应对非咨询模型带来的实际挑战时,存在角色期望冲突的体验,为理解该制度的具体实施提供了深入视角。

Cloak: Heuristic ORAM Optimization Through Fixed Temporal Distribution

第一作者: Onur Eren Arpaci · 方向: 系统安全

Abstract:Encrypted cloud storage can hide data contents but still leak sensitive information through access patterns. ORAM addresses this by hiding access patterns, but existing ORAM systems are too inefficient to deploy in practice. We present Cloak, an oblivious storage system that dramatically improves performance by leveraging a simple, widely observed property of real workloads: temporal locality, where recently accessed items are more likely to be accessed again soon. Instead of trying to make server accesses look perfectly uniform, Cloak makes server traffic follow a fixed, "recentness-biased" pattern and then uses real queries to fill as much of that traffic as possible. When the workload exhibits temporal locality, Cloak achieves overheads as low as $1.1\times$ over a non-oblivious and unencrypted baseline. Importantly, this heuristic affects only performance, not security. We...

论文介绍 针对加密云存储中访问模式可能泄露敏感信息的问题,本文提出了Cloak系统。它利用实际负载中常见的「时间局部性」特性,不再追求完全均匀的服务器访问模式,而是让服务器流量遵循一种固定的「近期偏置」模式,并用真实查询尽量填充该流量。在具有时间局部性的工作负载下,其性能开销可低至基线的1.1倍,且该启发式方法仅影响性能,不损害安全性。

Analyzing Linear Layers in Related-Differential Cryptanalysis

第一作者: Yogesh Kumar · 方向: 安全研究

Abstract:In AES-like ciphers, diffusion layers are commonly instantiated using MDS matrices, since their optimal branch number yields strong diffusion guarantees and underpins classical resistance arguments against differential and linear cryptanalysis. However, Daemen and Rijmen (2009) showed that linear layers may still exhibit related-differential structure beyond what the MDS criterion captures, and Bardeh and Rijmen (2022) demonstrated that this phenomenon can be exploited in attacks on reduced-round AES. In this work, we systematically investigate the conditions under which linear layers avoid or exhibit these differentials, identifying matrix classes for which such structure is unavoidable. We first prove that every non-MDS matrix admits a nontrivial pair of related differentials, showing that the MDS property is necessary for avoiding them. We then establish that every...

论文介绍 在AES类密码中,常使用MDS矩阵作为扩散层,其最优分支数为抵抗差分和线性分析提供了基础。本文系统研究了线性层避免或表现出相关差分结构的条件。研究证明,每个非MDS矩阵都必然存在非平凡的相关差分对,从而表明MDS性质是避免此类结构的必要条件。该工作进一步刻画了存在此类结构的矩阵类别。

Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?

第一作者: Syed Huma Shah · 方向: AI 安全

Abstract:Modern retrieval-augmented generation(RAG) deployments increasingly rely on caching to reduce token cost and time-to-first-token(TTFT). Prefix-level KV reuse is now standard in serving stacks such as vLLM, and chunk-level and position-independent reuse have been pushed further by recent systems(RAGCache, TurboRAG, CacheBlend, EPIC, ContextPilot, PCR, LMCache). Output-level semantic answer caches, by contrast, remain fragile: similar prompts can map to different correct answers, retrieved evidence drifts as the corpus is updated, and adversarial collision attacks have been shown to hijack cached responses. We argue that the right framing for cached answer reuse is not how to reuse faster but when reuse is safe. We propose GroundedCache, an evidence-validated cache router that admits a cached answer only when 4 cheap gates simultaneously hold: query similarity...

论文介绍 检索增强生成部署中,输出级语义答案缓存存在复用不安全问题,例如提示相似但答案正确的情况、证据漂移以及对抗性碰撞攻击。本文提出GroundedCache,一个经证据验证的缓存路由器。它仅在查询相似性、证据新鲜度、安全守卫和权威性四个廉价门控同时满足时,才允许复用缓存答案,旨在解决「何时复用是安全的」这一核心问题。

HARP: Measuring Harm Amplification in Multi-Agent LLM Systems

第一作者: Md Hafizur Rahman · 方向: AI 安全

Abstract:Multi-agent LLM systems decompose workflows across agents, tools, shared context, memory, and decision gates. This modularity improves interpretability, but creates a propagation risk: a bounded perturbation to one component can be reused by other agents and amplified into system-level harm. We introduce HARP (Harm Amplification through Role Perturbation), a trace-first methodology for studying local-to-global harm amplification in multi-agent LLM systems. HARP compares paired clean and perturbed executions and records specialist outputs, tool calls, memory reads/writes, guard events, oracle logs, latency, token cost, and decisions. We define local harm as deviation from targeted agents or corrupted channels, global harm as deviation over the full trace, and harm amplification as (H_global/H_local). This complements attack success rate with a measure of how strongly...

论文介绍 多智能体LLM系统虽提升了可解释性,但也引入了风险传播问题:对一个组件的有界扰动可能被其他智能体复用并放大为系统级危害。本文引入HARP(基于角色扰动的危害放大)方法,通过比较干净与扰动执行的轨迹对,定义并衡量从局部到全局的危害放大效应。该方法为评估多智能体系统的安全性提供了新的度量视角。

Grimlock: Guarding High-Agency Systems with eBPF and Attested Channels

第一作者: Qiancheng Wu · 方向: 软件安全

Abstract:Agentic systems increasingly run user-authored orchestration code that invokes tools, spawns subtasks, and delegates work across machines and clouds. Although this high agency is productive, it creates a security problem: identity, authorization, provenance, and delegation are often pushed into application code, where they become difficult to enforce consistently and difficult to audit. We present \emph{Grimlock}, an \emph{Agent Guard} that restores separation of concerns by moving trust enforcement into the sandbox substrate while leaving agent code unchanged. Grimlock uses \emph{eBPF-enforced traffic interception} to ensure that sandbox communication passes through a guard, and combines it with \emph{post-handshake attestation} bound to standard TLS~1.3 channel bindings. After a channel is established, the guard authorizes communication and mints short-lived, channel-bound...

论文介绍 高自主性智能体系统运行用户编排代码,但将身份、授权等安全逻辑置于应用层难以一致执行和审计。本文提出Grimlock,一个「智能体守卫」。它通过eBPF强制的流量拦截确保沙箱通信经过守卫,并结合TLS 1.3通道绑定的事后证明。建立通道后,守卫授权通信并颁发短时、通道绑定的令牌,将信任执行移至沙箱基础设施,实现关注点分离。

Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

第一作者: Matteo Gioele Collu · 方向: AI 安全

Abstract:In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations at each transformer block. We find that refusal is linearly decodable well before the final layer, indicating that safety-relevant behavior is represented in intermediate activations before output generation. To test whether this signal is actionable, we introduce Mechanistic AutoDAN, a probe-guided variant of AutoDAN that replaces full-model fitness evaluation with partial forward passes and probe-based scoring inside a genetic prompt search loop. Across the evaluated models, our method achieves attack success rates competitive with vanilla AutoDAN while reducing per-iteration search time by up to 72%, and probe-guided prompts match or exceed AutoDAN's cross-model transfer in several...

论文介绍 本文研究能否在解码前,从LLM中间层的残差流激活中预测其拒绝行为。结果表明,拒绝行为在最终层之前就已被线性解码,即安全相关行为在输出生成前就已表征。基于此发现,提出了Mechanistic AutoDAN,在遗传提示搜索中利用探针评分替代全模型评估,在保持攻击成功率的同时,能显著降低单次迭代的搜索时间。

ISAC Privacy: Challenges and Solutions for 6G

第一作者: Onur Günlü · 方向: 网络安全

Abstract:Integrated sensing and communication (ISAC) is a promising feature of future communication networks. While spatial sensing can improve network performance and enable external services, it also creates privacy challenges that go beyond the confidentiality of communication content. Future networks using millimeter-wave (mmWave) and sub-terahertz (THz) frequencies may collect or infer detailed information about people, devices, bystanders, passive objects, and environments in a sixth-generation (6G) deployment area. Such sensing can reveal location and environment data, support behavioral profiling such as movement or activity recognition, and, in advanced cases, expose physiological information such as breathing frequency or heart-rate-related data. Thus, the capabilities of spatial sensing must be controlled to satisfy privacy requirements. In this work, we organize...

论文介绍 感知通信一体化是未来6G网络的关键特性,但它带来的隐私挑战超越了通信内容保密。使用毫米波和亚太赫兹频段的网络可能收集或推断出人员、设备、环境等详细信息,支持位置推断、行为画像乃至生理信息感知。本文系统梳理了ISAC带来的隐私挑战,并探讨了在满足网络性能与服务需求的同时,控制空间感知能力以保护隐私的解决方案。

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection

第一作者: Shuhao Chen · 方向: AI 安全

Abstract:Fine-tuning large language models often undermines their safety alignment, a problem further amplified by harmful fine-tuning attacks in which adversarial data removes safeguards and induces unsafe behaviors. We propose SPARD, a defense framework that integrates Safety-Projected Alternating optimization with Relevance-Diversity aware data selection. SPARD employs SPAG, which optimizes alternatively between utility updates and explicit safety projections with a set of safe data to enforce safety constraints. To curate safe data, we introduce a Relevance-Diversity Determinantal Point Process to select compact safe data, balancing task relevance and safety coverage. Experiments on GSM8K and OpenBookQA under four harmful fine-tuning attacks demonstrate that SPARD consistently achieves the lowest average attack success rates, substantially outperforming state-of-the-art defense...

论文介绍 微调大语言模型常会破坏其安全对齐,有害微调攻击会移除安全护栏并诱导不安全行为。本文提出SPARD防御框架,结合安全投影交替优化与相关性-多样性感知的数据选择。该框架在效用更新和显式安全投影之间交替优化,以强制安全约束。同时,采用行列式点过程选择紧凑的安全数据,平衡任务相关性与安全覆盖。实验表明该框架能有效抵御多种有害微调攻击。

An Empirical Audit of k-NAF Budget Accounting for Anchored Decoding

第一作者: J. Vijayavallabh · 方向: 安全研究

Abstract:We empirically audit the k-NAF budget-accounting mechanism in Anchored Decoding using (i) a fixed, class-stratified workload (approximately 8,500 randomized executions across six prompt classes) and (ii) an adaptive prompt-search procedure targeting high proxy spend ratios. On the fixed workload, mean cumulative KL spend remains far below the sequence-level budgets K in {600, 1000}, and an empirical Bernstein-style proxy stays below K for every class; surface-overlap diagnostics (ROUGE-L and 5-gram Jaccard) are correspondingly small. Adaptive search increases the proxy spend ratio but does not produce clear budget exhaustion. On a held-out copyright-domain workload at k = 3, several prompts exhibit proxy ratios above 1 under early-stopped evaluations with small realized sample sizes; re-evaluating the same prompts with larger allocation reduces the proxy ratio to the range...

论文介绍 本文对用于Anchored Decoding的k-NAF预算记账机制进行了实证审计。研究通过固定分层工作负载和自适应提示搜索程序,评估了累计KL散度消耗与预设预算的关系。实验表明,在固定负载下,消耗远低于预算上限;自适应搜索虽提升了代理消耗比率,但仍未导致预算耗尽。该工作为理解和验证基于约束解码的隐私保障机制提供了实证基础。

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

第一作者: Yuan Tian · 方向: 安全研究

Abstract:Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response generation, text-only prior turn, visual-state manipulation, and explicit external image-tool invocation. In this paper, we ask which of these evaluated paradigms improves multimodal jailbreak robustness, and why. Across multiple vision-language models, explicit image-tool interaction yields the lowest attack success rates in our experiments, reducing jailbreak success by around 30% relative on average across the evaluated models. This finding is initially surprising: ASR remains low even when the returned image-tool output is manually overridden or itself unsafe-looking, but returns near direct-answering levels under text-only prior turn...

论文介绍 本文研究新兴的「思维-图像」推理范式对大型视觉语言模型安全性的影响。研究比较了多种多模态推理流程设计对抵抗越狱攻击的鲁棒性。实验发现,显式调用外部图像工具的交互方式平均可将越狱成功率降低约30%,即使工具输出不安全,攻击成功率也保持低位。研究揭示了推理设计与模型安全性之间的内在联系。

Density-aware Sample-specific Attack

第一作者: Qiyuan Wang · 方向: 安全研究

Abstract:Despite recent progress in backdoor attacks, existing methods remain susceptible to post-training defenses that erase the backdoor through fine-tuning or pruning. We revisit the core objectives of backdoor attacks and derive principled criteria characterizing optimal sample-specific trigger construction under a Bayes-optimal model of the victim's training. Our analysis reveals that both attack success and clean-accuracy preservation are simultaneously optimized when triggered samples are steered into low-density regions of the clean data distribution, a distributional condition that controls all moments of the poisoned distribution at once rather than a handful of input-space summary statistics. We introduce a bilevel optimization framework that estimates density ratios via conditional time-score matching and optimizes a mixture-model objective to place triggered samples in...

论文介绍 针对现有后门攻击易被微调或剪枝等训练后防御清除的问题,本文提出了密度感知的样本特异性攻击方法。基于贝叶斯最优模型分析,提出最优触发器构造应使中毒样本落入干净数据分布的低密度区域。为此,提出了一种双层优化框架,通过条件时间分数匹配估计密度比,并优化混合模型目标函数,从而显著提升攻击在防御下的持久性。

Local Privacy Laws in a Globalized World

第一作者: Shantanu Sharma · 方向: 隐私保护

Abstract:Personal data has emerged as a highly valuable yet sensitive asset that drives business decisions, enables targeted advertising, and generates substantial revenue for companies, while simultaneously facilitating invasive monitoring of users. In recent years, research on digital privacy violations, including undue access, collection, and sharing of user data, has grown significantly. Much of this research adopts the European General Data Protection Regulation (GDPR) as the primary reference framework. This is reasonable, as GDPR was a pioneering legislation, and many of its stipulations are clear and unambiguous. However, we argue that focusing solely on GDPR (and a small set of other Western regulatory frameworks) ignores privacy-related concerns, attitudes, and problems faced by users from other locales, creating a significant research blind spot. This work systematically...

论文介绍 本文指出,当前数字隐私研究过度依赖欧盟《通用数据保护条例》(GDPR),忽视了全球其他地区独特的隐私关切、态度和问题,形成了研究盲点。作者倡导开展系统的比较研究,分析不同法律框架下的隐私保护实践,以推动建立更具包容性和全球视野的隐私研究范式,更好地理解和应对多样化的数据保护挑战。

Revisiting ML Training under Fully Homomorphic Encryption: Convergence Guarantees, Differential Privacy, and Efficient Algorithms

第一作者: Yvonne Zhou · 方向: 密码学协议

Abstract:We present the first theoretical convergence analysis of machine learning training under fully homomorphic encryption (FHE), combined with a differentially private (DP) training algorithm tailored to encrypted computation. Our approach improves computational efficiency over standard differentially private gradient descent (DP-GD) while achieving comparable utility. In particular, we prove convergence of approximate gradient descent using polynomial approximations of activation and loss functions, which are required for FHE compatibility. To preserve privacy in downstream tasks, we integrate differential privacy without relying on costly per-sample gradient clipping, enabling scalable encrypted learning. We also provide data-independent hyperparameter selection and theoretically grounded strategies for polynomial approximation which can be of independent interest. Together...

论文介绍 本文首次为完全同态加密下的机器学习训练提供了理论收敛分析,并设计了适配加密计算的差分隐私训练算法。研究通过多项式近似激活和损失函数来兼容FHE,证明了近似梯度下降的收敛性。该方案无需进行昂贵的逐样本梯度裁剪即可集成差分隐私,在提升计算效率的同时保证了模型可用性,为可扩展的加密学习提供了理论支撑。

Test-Time Collective Action: Proxy-Based Perturbations for Correcting Algorithmic Harms

第一作者: Meghana Bhange · 方向: AI 安全

Abstract:When machine learning systems under-perform for particular subgroups, affected users typically have no way to correct these disparities without relying on platform-level fixes. Existing approaches to algorithmic fairness rely on provider-centric approaches to correct these failures, leaving users with no external lever when faced with harm. Recent work in Algorithmic Collective Action shows that coordinated users can steer an algorithmic system toward a collective goal, but the existing mechanisms require the provider to retrain on the collective's modified data which users may not have control over. We propose Test-Time Collective Action (TTCA), a framework through which a group of users who share query access to the platform, can correct disparities affecting under-served subgroup without participating in the platform's training loop. We implement this through a proxy-based...

论文介绍 当机器学习系统对某些子群表现不佳时,现有公平性方法依赖平台方进行修正。本文提出「测试时集体行动」框架,允许共享查询访问权限的用户群组,在无需参与平台训练循环的情况下,通过代理扰动技术协同纠正针对服务不足子群的算法偏差。该研究探索了用户驱动的算法公平性干预新路径。

A Trilemma in AMM Mechanism Design

第一作者: Yuhao Li · 方向: 软件安全

Abstract:Blockchains have popularized the Automated Market Makers (AMMs), where users trade crypto-assets directly with a smart contract, governed by a pricing function embedded in the contract's code. Today, users of AMMs are often forced to accept unfavorable prices due to widespread front-running and back-running attacks, commonly known as Miner Extractable Value (MEV). Several earlier works show impossibility results suggesting that completely removing MEV at the consensus layer is impossible, partly because the consensus layer is agnostic of application-level semantics. For this reason, more recent works have advocated mechanism design approaches at the application (i.e., smart contract) level. We study a natural two-asset AMM mechanism design problem recently initiated and explored in prior work by Chan, Wu, and Shi, in which they proposed a mechanism that satisfies a...

论文介绍 自动做市商中的前置和后置运行攻击导致用户面临矿工可提取价值问题。本文研究了一个双资产AMM的机制设计问题,揭示了完全消除MEV与满足某些其他理想属性(如无套利、需求响应等)之间存在固有的「三难困境」。该分析从理论上刻画了应用层机制设计的内在权衡与限制。

On the Origin of Synthetic Information by Means of Steganographic Inheritance

第一作者: Ching-Chun Chang · 方向: AI 安全

Abstract:The origin of species has been the mystery of mysteries in natural science. By analogy, the origin of synthetic information, we suggest, is the mystery of mysteries in information science. The question carries a moral weight that a technical account can neither fully resolve nor responsibly ignore, as its impact on truth, trust, and human intellect extends deep into the broader economy and society. The very power of artificial intelligence makes the evolutionary lineage of synthetic information grow ever harder to trace, for a sufficiently capable model may generate offspring that bear little resemblance, at either the structural or signal level, to the parent source from which they were derived. As in genetics, two individuals may share the same phenotype mirroring each other in outward appearance, yet differ fundamentally in their genotype. We propose, by means of...

论文介绍 类比生物进化,本文探讨了信息科学中「合成信息」的起源与谱系追踪问题。由于生成式人工智能的强大能力,合成信息的「遗传」路径日益难以追溯。作者提出利用隐写术技术,在信息生成过程中植入可继承的、隐蔽的标记,如同数字基因型,以期为合成信息建立可追溯的谱系,应对由此带来的信任与真实性挑战。

Imitation Learning for Robot Assistance in Open Surgery: A Multi-Policy Evaluation on Suture Following

第一作者: Xucheng Wang · 方向: 模仿学习 · 来源: cs.RO

Abstract:This study presents the first evaluation of general-purpose imitation learning for surgeon-robot collaborative assistance in open surgery, targeting suture following: the grab-pull-release motion an assistant performs at every stitch. We collect 160 teleoperated demonstrations (32,374 frames) on an open-source robot arm, benchmark four architecturally diverse imitation learning policies (ACT, Diffusion Policy, SmolVLA, $\pi_0$) across 28 trained models evaluated in 32 configurations along three clinically motivated dimensions: dataset size, camera viewpoint, and background variation. Our results demonstrate that under ideal conditions, the four policies achieve $50$-$75\%$ task success, with depth error as the dominant failure mode across all architectures. Among all policies, $\pi_0$ achieves the strongest results with a pretrained vision-language backbone, demonstrating...

论文介绍 本研究首次评估通用模仿学习在开放手术中机器人辅助缝合跟踪的效果。通过收集160个遥操作演示,评估了ACT、Diffusion Policy、SmolVLA和π0四种策略在数据集大小、相机视角和背景变化维度上的性能。结果显示,在理想条件下任务成功率为50-75%,深度误差是主要失败模式,π0策略表现最佳。

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures

第一作者: Krishnam Gupta · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:We discover that VLA architectures fail in fundamentally different, predictable ways at the motor-command level. Running VQ-BeT, Diffusion Policy, and ACT on identical evaluation protocols (n=450 episodes across PushT and ALOHA 14-DOF bimanual manipulation), we find: (1) direction reversal rate is a universal failure predictor across all three architectures (AUROC=0.93, 0.79, 0.91; p<0.001); (2) jerk monitoring is predictive only for discrete-token architectures, following a discrete-to-continuous gradient (0.88, 0.69, 0.41); (3) velocity violations alone are non-predictive everywhere (AUROC 0.41-0.69), yet velocity checking is the most common safety mechanism in VLA deployment code; and (4) for continuous-family VLAs, velocity monitoring provides effectively zero predictive signal (AUROC=0.52 on ACT, 0.41 on Diffusion), proving that architecture-matched monitor selection is...

论文介绍 本文揭示VLA架构在电机命令层面存在可预测的失败差异。通过黑盒动作监控,在PushT和ALOHA任务上评估VQ-BeT、Diffusion Policy和ACT,发现方向反转率是通用失败预测器,抖动监控对离散架构有效,速度监控在连续架构中无效。强调了架构匹配监控选择对部署安全性的重要性。

PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation

第一作者: Yutai Li · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models offer a promising paradigm for generalist robotic policies, yet their adaptation is hindered by data inefficiency and poor generalization. We argue that these bottlenecks stem from the prevailing Direct Instruction-to-Control Mapping, which forces models to memorize monolithic trajectories rather than reusable motion patterns, i.e., primitives. We propose PrimitiveVLA, a framework that shifts this paradigm toward a Primitive-Centric Disassemble & Assemble paradigm. Supported by a shared Multimodal Canonical Representation (MCR), PrimitiveVLA unifies two phases: (1) Fine-tuning-phase Disassembly, which uses an automated pipeline to disassemble demonstrations into reusable primitives; and (2) Inference-phase Assembly, which employs a VLM-based planner and an LLM-generated switch module for robust closed-loop execution. By disassembling tasks...

论文介绍 VLA模型在数据效率和泛化性上存在瓶颈,源于直接指令到控制映射。本文提出PrimitiveVLA框架,采用以原语为中心的拆解与组装范式。通过共享多模态规范表示,将演示自动拆解为可重用原语,并在推理时组装执行,以提高机器人操作的效率和泛化能力。

What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies

第一作者: Jiachen Zhang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Their frozen representations nevertheless carry such information, and it can be read out and used to guide action choice without retraining the policy. From mixed successful and failed manipulation trajectories on LIBERO-Goal, we recover Monte-Carlo outcome targets using lightweight linear probes on frozen features. The targets are consistently predictable from OpenVLA, Pi0.5, DINOv2, and CLIP features, and substantially less so from baselines built on progress, time-to-go, task identity, or proprioception. To rule out task and temporal shortcuts, we evaluate the probes under same-task, same-timestep matched comparisons: Pi0.5 probes still reach roughly 92% pairwise ordering accuracy, while label-shuffled controls stay at chance...

论文介绍 本文研究VLA模型的冻结表示是否包含成功相关信息。通过轻量级线性探针从OpenVLA、Pi0.5等模型的冻结特征中恢复蒙特卡洛结果目标,无需重训练策略。结果表明,这些特征能可靠预测成功,证明VLA表示隐含价值结构,可用于引导动作选择。

Mag-VLA: Vision-Language-Action Model for Bimanual Magnetically Actuated Microrobot Manipulation

第一作者: Yongchen Wang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Magnetically actuated microrobots have been used as wireless, non-contact manipulation tools at microscales, making them promising for minimally invasive applications. However, their control remains challenging due to indirect actuation, limited sensing, and nonlinear magnetic interactions. In this work, we propose Mag-VLA, a vision-language-action (VLA) model for dexterous magnetic microrobot manipulation using two robotic arms with mounted magnets for dynamic magnetic-field construction. Bimanual coordination enables capabilities such as microrobot reorientation that are difficult or infeasible with a single arm, but it also introduces coupled control challenges, as the policy must generate coordinated trajectories for both actuators within a shared workspace. Our framework adapts a Qwen2.5-VL-7B backbone using Low-Rank Adaptation (LoRA) to process visual observations and...

论文介绍 针对磁驱动微机器人控制挑战,本文提出Mag-VLA模型。使用两个机器人臂动态构建磁场,基于Qwen2.5-VL-7B骨干并采用LoRA适应处理视觉观察。双臂协调实现了微机器人重定向等能力,适用于微创应用中的灵巧操作。

EIT-Pneumatic Hybrid Robotic Skin for Practical and Accurate Force Map Reconstruction

第一作者: Junhwi Cho · 方向: 具身智能 · 来源: cs.RO

Abstract:We present a hybrid robotic skin that combines electrical impedance tomography (EIT) with pneumatic tactile sensing to improve force reconstruction capability. The developed robotic skin is fabricated entirely by 3D printing and spray coating, making it affordable and easy to build. A Tikhonov-regularized inverse reconstruction, paired with per-pad pneumatic calibration, enables accurate large-area tactile sensing with a simple measurement scheme. For validation, we conducted load-cell indentation experiments; the results showed consistent force reconstruction across locations within a pad. Compared with an EIT-only baseline, sensitivity non-uniformity was also reduced, with the coefficient of variation decreasing from 0.31 to 0.14, indicating that the proposed approach addresses a longstanding limitation of EIT. We further demonstrated chest-mounted integration on a humanoid...

论文介绍 为改善EIT力重建的非均匀性,本文开发混合机器人皮肤,结合EIT和气动触觉传感。通过3D打印和喷涂制造,使用Tikhonov正则化逆重建和逐垫气动校准,实现准确大面积力重建。实验证明灵敏度一致性提高,适用于机器人皮肤集成。

A Digital Twin Framework for Virtual Visuo-Haptic Teleoperation of Complex-Shaped Optical Microrobots

第一作者: Zongcai Tan · 方向: 机器人操作 · 来源: cs.RO

Abstract:Optical tweezers (OT) provide piconewton-scale manipulation for delicate biomedical tasks, where visuo-haptic feedback can improve operator awareness by conveying interaction-force cues and trap-stability information. However, visuo-haptic teleoperation frameworks for complex-shaped optical microrobots remain underdeveloped, particularly in multi-trap manipulation scenarios. This paper presents a digital twin framework for virtual visuo-haptic teleoperation of complex-shaped OT-driven microrobots. The framework integrates a digital twin environment, image-based pose and depth estimation, microrobot motion simulation, and model-based haptic rendering within a Robot Operating System (ROS)-connected bimanual teleoperation system. For force modeling, we combine a Multi-Sphere Distributed Manipulation (MSDM) model with optical-force estimation from the Optical Tweezers Toolbox...

论文介绍 针对复杂形状光学微机器人的视觉触觉遥操作,本文提出数字孪生框架。集成数字孪生环境、图像姿态估计、运动模拟和基于模型的触觉渲染,用于双臂遥操作系统。结合多球分布式操作模型和光镊工具箱估计力,提升操作精度和感知。

Self-Supervised Online Robot-Agnostic Traversability Estimation for Open-World Environments

第一作者: Julia Hindel · 方向: 导航与运动 · 来源: cs.RO

Abstract:Self-supervised online traversability estimation enables robots to continuously learn from unlabeled open-world experiences and adapt their navigation behavior toward safe and efficient trajectories. Existing approaches either rely on handcrafted proprioceptive traversability scores, limiting robot-agnosticism, or cluster prior data, preventing online learning. Moreover, many continual learning methods incur substantial memory and computational costs, hindering onboard deployment. We introduce COTRATE, an online learning framework for continuous traversability estimation from multimodal, unlabeled robot experience. Our method first infers robust traversability scores using a robot-agnostic, learning-based online terrain assessment module operating on proprioceptiveand inertial signals. These scores then supervise a visual traversability network through a novel alignment loss...

论文介绍 本文提出COTRATE框架,用于开放世界环境中自监督在线可通行性估计。通过机器人无关的地形评估模块从多模态未标记数据推断分数,并监督视觉网络。实现连续学习和适应,降低计算成本,适用于车载机器人导航部署。

Tactile-Proprioceptive Sensor Fusion for Contact Wrench Estimation in Whole-Body Physical Human-Robot Interaction

第一作者: Junha Min · 方向: 具身智能 · 来源: cs.RO

Abstract:Direct physical guidance is a natural means of teaching and interacting with robots, and robotic skins make a key contribution by enabling sensitive contact sensing and localization. This paper presents a tactile-proprioceptive sensor fusion framework for natural physical human-robot interaction. Tactile cues from pneumatic skin pads serve as contact indicators that bypass the ambiguity between frictional residues and applied external forces, enabling highly sensitive contact detection without explicit friction identification. We fuse these cues with motor-current-based proprioception to reconstruct multi-axis contact forces on the robot surface. To maintain accuracy during motion, we employ a temporal convolutional network (TCN) to mitigate friction hysteresis during stick-slip transitions, reducing uncertainty at contact onset and yielding smooth, responsive guidance. We...

论文介绍 该研究针对物理人机交互中的接触力感知问题,提出一种触觉-本体感觉传感器融合框架。该框架利用气动皮肤垫的触觉线索作为接触指示器,规避了摩擦残留与外力之间的歧义,并融合电机电流的本体感觉来重建机器人表面的多轴接触力。通过引入时间卷积网络处理粘滑过渡期间的摩擦滞后,实现了高灵敏度、平滑的接触检测与力估计,有望提升机器人在交互任务中的感知与响应能力。

Safety-Critical Adaptive Impedance Control via Nonsmooth Control Barrier Functions under State and Input Constraints

第一作者: Faisal Lawan · 方向: 具身智能 · 来源: cs.RO

Abstract:Safe physical interaction is critical for deploying robotic manipulators in human-robot interaction and contact-rich tasks, where uncertainty, external forces, and actuator limitations can compromise both performance and safety. We propose an online adaptive impedance control framework that enforces joint-state safety while achieving compliant interaction under uncertain dynamics. The approach combines a quadratic-program-based safety filter with a novel composed position-velocity non-smooth control barrier function (NCBF), enabling joint position and velocity constraints to be enforced through a unified relative-degree-one barrier. Unknown dynamics are compensated online using an interval type-2 fuzzy logic system, while actuator torque limits are handled through soft constraints with exact penalty recovery of feasible solutions. A disturbance-observer-enhanced safety...

论文介绍 该研究面向机器人操作器在人机交互和接触密集任务中的安全物理交互需求,提出一种在线自适应阻抗控制框架。该方法结合基于二次规划的安全滤波器与新型组合位置-速度非光滑控制障碍函数,在统一的相对度一阶障碍函数下强制执行关节状态安全约束。通过区间二型模糊逻辑系统在线补偿未知动力学,并利用精确罚函数处理执行器力矩限制,从而在不确定动态下实现安全且顺从的交互控制。

Accelerating Robot Path Planning via Connectivity-Preserving Region Proposal Network

第一作者: Zhanzheng Ma · 方向: 导航与运动 · 来源: cs.RO

Abstract:Mobile robot path planning methods are often constrained by vast search spaces, resulting in latency in samplingbased algorithms. Learning-based approaches frequently suffer from local region fragmentation and global topological inconsistency. To tackle the problem, we present the Connectivity- Preserving Region Proposal Network (CP-RPN), a segmentationguided model designed to predict compact and topologically connected candidate regions, significantly compressing the search space. Specifically, we design a segmentation model that leverages a Deformable Attention Transformer (DAT) to capture long-range dependencies for global connectivity, with a Deconvolutional decoder to preserve fine-grained spatial details. To guarantee the connectivity of the predicted mask, we design a composite loss function that combines Cross-Entropy loss for pixelwise supervision, a...

论文介绍 该研究针对移动机器人路径规划中采样算法因搜索空间过大导致的延迟问题,以及学习方法易出现局部区域碎片化和全局拓扑不一致的挑战,提出一种基于分割引导的连通性保持区域提议网络(CP-RPN)。该模型利用可变形注意力Transformer捕获长程依赖以保持全局连通性,并通过反卷积解码器保留空间细节。结合设计的复合损失函数,预测出紧凑且拓扑连通的候选区域,从而显著压缩路径搜索空间,提升规划效率。

Magnet-Based Soft Robotic Skin Using a 3D-Printed Multi-Lattice Structure and CNN-Based Tactile Super-Resolution

第一作者: Yunseong Bang · 方向: 具身智能 · 来源: cs.RO

Abstract:This paper presents a magnet-based robotic skin that integrates a multilayer soft lattice with distributed Hall-effect sensor arrays and a tactile super-resolution model. External contact forces are converted to magnetic field changes by embedded permanent magnets, and the lattice spreads these changes across the sensing domain. This gives each sensor a large, overlapping receptive field and enables a large sensing area with minimal blind spots. Lattice parameters are tunable, enabling joint adjustment of mechanical compliance and transduction characteristics. An implicit modeling workflow and selective laser sintering (SLS) 3D printing support rapid fabrication of conformal, high-complexity structures. A convolutional neural network trained on experimental measurements estimates contact location and normal force in real time. Experiments validate localization accuracy and...

论文介绍 该研究提出一种基于磁体的软机器人皮肤,它集成多层软晶格、分布式霍尔效应传感器阵列和触觉超分辨率模型。外部接触力通过嵌入的永磁体转化为磁场变化,晶格结构将变化扩散至感知域,使每个传感器具有较大的重叠感受野,实现大面积、少盲区的感知。通过卷积神经网络实时估计接触位置和法向力,为机器人提供高分辨率、可调柔顺性的触觉感知能力,适用于精细操作任务。

Identifying Explicit Parsimonious Piece-wise Polynomial Relationships in Industrial time-series: Application to manipulator robots

第一作者: Mazen Alamir · 方向: 具身智能 · 来源: cs.RO

Abstract:This paper addresses the problem of identifying parsimonious explicit piece-wise polynomial relationships that might involve a relatively large number of raw features. The algorithm leverages a recently proposed identification algorithm that yields parsimonious implicit relationships enabling to derive normality characterization in the context of anomaly detection and localization. The algorithm proposed in this paper goes a step further by deriving explicit piece-wise representations that are built using the set of polynomials involved in the implicit representations. The framework is illustrated on the problem of identifying parsimonious explicit representations of the inverse model of a 6-axis manipulator robot. Moreover, further experiments on a 4-axis robot are also shown which are designed to investigate the generalization capability of parsimonious models compared to...

论文介绍 该研究解决从工业时间序列数据中识别可能涉及大量原始特征的稀疏显式分段多项式关系的问题。该算法利用已有的隐式关系识别方法,进一步推导出基于隐式表示中多项式集合的显式分段表示。研究以六轴机械臂逆模型的稀疏显式表示识别为例进行说明,并通过四轴机械臂实验验证了稀疏模型的泛化能力,为工业机器人系统建模与诊断提供了新方法。

EventShiftFlow: Towards Hardware-efficient FPGA-based Flow Estimation

第一作者: Arianna Alonso Bizzi · 方向: 具身智能 · 来源: cs.RO

Abstract:Event-based vision sensors offer asynchronous, high-temporal-resolution measurements that are attractive for low-latency robotic perception, but many event-based motion estimation methods are computationally intensive and difficult to map to FPGA hardware. We present a streaming velocity estimator that discretizes asynchronous events into fixed-duration time bins, constructs a 1-bit spatial occupancy grid, and evaluates multiple velocity hypotheses in parallel using only fixed-width integer logic - shift registers, counters, comparators, and small LUT-mapped multiplies - with no dividers and no DSP blocks. It requires no frame reconstruction, no floating-point arithmetic, and no iterative optimization. The method deliberately trades dense sub-pixel optical flow for a sparse, quantized velocity estimate at each active pixel, suited to low-latency tasks such as reactive obstacle...

论文介绍 该研究针对事件相机运动估计方法计算密集、难以映射到FPGA硬件的问题,提出一种硬件高效的流式速度估计器。该方法将异步事件离散化到固定时长的时间仓中,构建1位空间占用网格,并利用移位寄存器、计数器等固定宽度整数逻辑并行评估多个速度假设。它无需帧重建、浮点运算或迭代优化,以牺牲密集亚像素光流为代价,换取适合低延迟任务的稀疏量化速度估计,为机器人实时感知提供了硬件友好方案。

Natural Locomotion: Principle and Method

第一作者: Mirado Mortel · 方向: 导航与运动 · 来源: cs.RO

Abstract:Robotic locomotion can become efficient when mechanisms exploit passive dynamics, compliance, and resonance rather than track prescribed trajectories. This paper formulates natural locomotion as an exchange principle for systems whose motion is mediated by environmental constraints or interactions. A motion is natural when an internal oscillator returns periodically, the body pose drifts, and the mean Propulsion--Oscillator Exchange power (POE power) vanishes over one cycle. The selected family is a Natural Locomotion Manifold (NLM). We develop the conservative realization of this principle for continuous ideal environmental constraints: the constraints do no external work, total mechanical energy is conserved, and zero mean POE power is an internal exchange with the environment-mediated propulsive channel, not external energy input. The method is a closed/open construction...

论文介绍 该研究提出一种“自然运动”的原理与方法,旨在通过利用被动动力学、柔顺性和共振来提高机器人运动效率。该研究将自然运动表述为一个交换原理:当内部振荡器周期性回归、体态漂移且平均推进-振荡器交换功率在一个周期内为零时,运动是自然的。论文开发了针对连续理想环境约束的保守实现,其中约束不做外功,总机械能守恒,为设计高效仿生运动步态提供了理论基础。

ProgVLA: Progress-Aware Robot Manipulation Skill Learning

第一作者: Seungsu Kim · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The model specifically focuses on efficiently processing long multi-modal sequences by maintaining an explicit representation of task progress over extended horizons. To this end, ProgVLA integrates two key components. First, a multi-modal encoder with a two-stage Perceiver resampling scheme compresses variable-length visual, language, and proprioceptive streams into a fixed set of control-ready context tokens, substantially reducing sequence length while preserving cross-modal grounding. Second, an auxiliary set of progress heads is trained with offline reinforcement learning (RL) objectives to jointly learn critics over normalized remaining-horizon targets. This provides the policy with an internal estimate of task progress and...

论文介绍 该研究提出一种紧凑的视觉-语言-动作(VLA)模型ProgVLA,旨在严格的计算和内存预算下实现可靠的机器人操作。该模型专注于高效处理长多模态序列,通过维护显式的任务进度表示来实现。它集成多模态编码器和进度头组件,前者通过两阶段感知器重采样压缩可变长度的视觉、语言和本体感觉流,后者通过离线强化学习目标训练以提供任务进度的内部估计,从而增强策略在长期任务中的可靠性与泛化能力。

Natural Functional Gradients for Smooth Trajectory Optimization

第一作者: Kisang Park · 方向: 机器人操作 · 来源: cs.RO

Abstract:Generating collision-free and smooth motions remains a central challenge in robotic manipulation, particularly in cluttered environments and narrow passages where feasible regions are highly constrained and fragmented. We propose a trajectory optimization framework that performs geometry-aware updates directly in function space using natural functional gradients. The method optimizes a Gaussian-smoothed surrogate objective that regularizes the optimization landscape through smooth trajectory perturbations while preserving trajectory-level structure. Because the updates are defined intrinsically in function space, trajectory regularity can be controlled independently of a particular time discretization. We derive a practical Monte-Carlo estimator of the natural functional gradient that requires only black-box trajectory evaluations, making the method applicable when analytic...

论文介绍 针对杂乱环境中机器人操作时生成无碰撞平滑轨迹的挑战,本文提出一种基于自然函数梯度的优化框架。该方法直接在函数空间中利用高斯平滑代理目标进行几何感知更新,从而独立于时间离散化来控制轨迹的平滑性。它仅需黑盒轨迹评估,为机器人规划提供了一种高效且鲁棒的新途径。

Provably Guaranteed Polytopic Uncertainty Quantification for SLAM

第一作者: Guangyang Zeng · 方向: 导航与运动 · 来源: cs.RO

Abstract:In safety-critical robotics applications, guaranteed and practical uncertainty quantification (UQ) in perception is vital. Many existing works either offer no formal containment guarantee, rely on restrictive modeling assumptions, or focus only on pose estimation rather than a complete SLAM pipeline. This paper presents provably guaranteed UQ algorithms for 3D-3D landmark-based SLAM. The algorithms consist of three basic UQ modules: forward UQ for mapping, backward UQ for pose tracking, and pose compound. Each module produces a certified uncertainty set; when the input uncertainty bounds are deterministic, the output sets inherit deterministic guarantees, i.e., they provably contain the true poses and landmarks. Specifically, we use polytopes to represent uncertainty sets, enabling tractable computations and a unified treatment of pose uncertainty. To enhance algorithms'...

论文介绍 在安全关键的机器人应用中,需要对感知结果进行可靠且可证明的不确定性量化。本文针对基于3D-3D路标的SLAM系统,提出了可证明保证的不确定性量化算法。该算法使用多面体表示不确定性集合,通过前向、后向及合成三个模块,在确定性输入保证下,能够输出包含真实位姿和路标的确定性保证集合,从而为整个SLAM流程提供形式化安全证明。

STR Robot: Design of an Autonomous Mobile Robot from Simulation to Reality

第一作者: Vinh Nguyen · 方向: 导航与运动 · 来源: cs.RO

Abstract:With the rapid development of simulation tools, the development and validation of autonomous robotic systems have become more efficient before real-world deployment. This paper presents a simulation-to-real implementation of an autonomous mobile robot based on an existing mechanical platform. Instead of focusing on mechanical design, our work concentrates on the development of the onboard control, self-localization, and autonomous navigation system. The proposed robot is equipped with onboard sensing and computation to estimate its pose and navigate autonomously in the environment. The overall framework is first developed and tested in simulation, and then deployed on the real robot for experimental evaluation. The results demonstrate the feasibility of the proposed approach and show that simulation provides an effective foundation for developing reliable autonomous mobile...

论文介绍 本文介绍了一种从仿真到现实的自主移动机器人开发流程。研究专注于机载控制、自我定位与自主导航系统的设计,并在一个现有机械平台上实现。该系统首先在仿真环境中开发与测试,然后部署到真实机器人进行实验验证。结果表明,仿真能为开发可靠的自主移动机器人系统提供有效基础。

ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents

第一作者: Xue Qin · 方向: 具身智能 · 来源: cs.RO

Abstract:Canary deployment routes a fraction of traffic to a new software version, monitors metrics, and rolls back on regression. Mainstream controllers (Argo Rollouts, Spinnaker, Flagger) change the deployed system's cryptographic identity during the canary window. The drift is harmless for stateless microservices but breaks the claim that "the agent you certified is still the agent you have" for safety-critical embodied agents, forcing re-certification per canary. We present ICAN-Deploy (Identity-stable CANary Deployment), a middleware construction whose state machine holds the identity hash invariant across the canary window by separating capability names (frozen, hashed) from capability versions (mutable runtime state). We implement ICAN-Deploy inside a runtime governance layer for LLM-driven robots and verify invariance by closed-form proof, AST lint, and TLA+ model-checking...

论文介绍 针对安全关键的具身智能体,本文提出一种身份稳定的金丝雀部署方案(ICAN-Deploy)。该方法通过将能力名称(冻结并哈希)与能力版本(可变运行时状态)分离,构建了一个中间件状态机,使得在部署新版本软件期间,智能体的加密身份标识保持恒定。这避免了传统方法中的身份漂移问题,从而简化了安全认证流程。

Whose Is This?: Context-Aware Object Ownership Inference with Uncertainty-Guided Questioning

第一作者: Saki Hashimoto · 方向: 具身智能 · 来源: cs.RO

Abstract:Service robots must infer object ownership to correctly interpret instructions such as "bring me my cup." However, ownership is a latent attribute that cannot be directly observed, and existing methods often rely on limited cues such as recent usage, making them unreliable in scenarios such as temporary sharing. We propose a framework for context-aware ownership inference with uncertainty-guided interaction (COIN). The method integrates user background information and object usage history using a large language model (LLM) to estimate ownership scores. To handle uncertainty, we apply conformal prediction to construct a set of plausible owners and selectively generate user queries when the prediction is uncertain. Experiments in a simulated home environment show that the proposed method consistently outperforms baseline approaches, achieving a Subset Accuracy of 0.988 and a...

论文介绍 服务机器人需要推断物体所有权以正确理解用户指令。本文提出了一个上下文感知的所有权推断框架(COIN),它整合用户背景信息和物品使用历史,利用大语言模型估计所有权分数。为应对不确定性,框架应用共形预测构建候选所有者集合,并在预测不确定时有选择地向用户提问。实验证明该方法在模拟家居环境中表现优异。

SAFEVPR: Patch-Based Conformal Verification for Safe Cross-Condition Sequence Visual Place Recognition

第一作者: Ha Sier · 方向: 机器人操作 · 来源: cs.RO

Abstract:Sequence-based visual place recognition (VPR) for SLAM and robot relocalization must decide whether the retrieved top-1 candidate is safe to accept. Conformal prediction is a natural framework for this accept/reject decision, but its finite-sample guarantees rely on exchangeability between calibration and deployment (test) data, which is violated under cross-condition deployment. We introduce SAFEVPR, a non-trainable verification-and-calibration pipeline for safe cross-condition sequence VPR. SAFEVPR replaces the standard backbone cosine similarity with a mutual-nearest-neighbour (MNN) patch-matching score computed from frozen DINOv2 ViT features, and replaces flat Learn-Then-Test calibration with Mondrian conformal LTT, fitting separate Bonferroni-corrected thresholds across score bins. Under exchangeability, these thresholds would provide finite-sample false-discovery-rate...

论文介绍 序列视觉地点识别需要安全地决定是否接受检索结果。由于跨条件部署会破坏共形预测所依赖的数据可交换性假设,本文提出SAFEVPR验证流程。该方法使用基于DINOv2特征的互近邻块匹配评分替代标准余弦相似度,并采用分层共形校准来设置不同评分区间的阈值,从而为跨条件下的视觉地点识别提供有限样本内的安全保证。

How Should We Teach Robots? A Comparison of Kinesthetic, Joystick, and Gesture-Based Teaching

第一作者: Petr Vanc · 方向: 机器人操作 · 来源: cs.RO

Abstract:Instructing robots from demonstrations can be done through different teaching modalities, each with different usability and performance trade-offs. This paper compares kinesthetic guidance, joystick teleoperation, and hand gestures in a user study with eight participants. We evaluate replay success, modified NASA-TLX workload, and common teaching errors across three manipulation tasks. Kinesthetic guidance produced the shortest demonstrations, lowest workload, and highest success on the more orientation-sensitive and contact-rich tasks. Joystick teleoperation performed best on simple peg picking. Hand-gesture teaching, although less reliable overall, performed better than expected and in some cases achieved results comparable to kinesthetic guidance.

论文介绍 本文通过一项包含八名参与者的用户研究,比较了三种机器人示教方式:运动学引导、摇杆遥操作和手势教学。评估指标包括演示成功率、工作负荷和常见教学错误。结果显示,运动学引导在对姿态敏感和接触丰富的任务中表现最佳,摇杆在简单拾取任务中更优,而手势教学虽整体可靠性稍低,但在某些情况下性能可与运动学引导媲美。

Simultaneous Contact Selection and Planning for Contact-Rich Manipulation with Cascaded Optimization

第一作者: Zhe Zhang · 方向: 机器人操作 · 来源: cs.RO

Abstract:We propose an optimization-based framework for robust contact-rich manipulation. Recent contact-implicit methods enable online hybrid planning across contact modes, allowing closed-loop manipulation for a given target state and contact location sequence of the robot and object. However, most existing approaches lack the ability to autonomously reason and generate diverse contact location sequences and manipulation trajectories, i.e., active contact location selection, which limits their applicability to relatively simple tasks. Active contact location selection is challenging due to complementarity in contact dynamics and the sparse gradients, making the design of a unified framework for contact selection and planning difficult. To address these challenges, we introduce Simultaneous Contact Selection and Planning (SCSP), a cascaded optimization framework comprising Contact...

论文介绍 为提升机器人执行接触丰富操作任务的鲁棒性,本文提出一个同时进行接触选择和运动规划的级联优化框架(SCSP)。该框架能主动推理并生成多样化的接触位置序列和操作轨迹,克服了现有接触隐式方法主要依赖给定接触序列的局限。通过处理接触动力学中的互补性和稀疏梯度问题,该方法扩展了接触操作任务的应用范围。

SANTS: A State-Adaptive Scheduler for World Action Models

第一作者: Yirui Sun · 方向: 机器人操作 · 来源: cs.RO

Abstract:World Action Models (WAMs) improve robot manipulation by using video-based future representations to condition action generation. In pixel-space WAMs, however, the best action condition is not necessarily the fully denoised video. Controlled denoising-depth scans show that video refinement can reduce action error up to a state-dependent point, after which the gain may saturate or even reverse when late predictions become less action-relevant or physically unreliable. This suggests that action generation should use a state-dependent point along the video noise trajectory rather than a fixed terminal denoising depth. We introduce State-Adaptive Noise Trajectory Scheduler (SANTS), a lightweight scheduler for video-to-action diffusion policies. At each video decision point, SANTS reads the current video-state representation and noise level, then jointly predicts a cumulative...

论文介绍 在基于像素空间的世界动作模型中,完全去噪的视频可能并非最佳动作条件,因为视频细化的效果会饱和或反转。本文提出状态自适应噪声轨迹调度器SANTS,它是一个轻量级组件,在每个视频决策点根据当前视频状态和噪声水平,预测累积去噪深度,以优化从视频到动作的生成过程。这有助于提高机器人操作任务中动作生成的准确性和效率。

S-Cheetah: A Novel Quadrupedal Robot with a 3-DOF Active Spine Learning Agile Locomotion

第一作者: Zimu Li · 方向: 机器人操作 · 来源: cs.RO

Abstract:The biological spine of quadrupeds enables sagittal flexion/extension, lateral bending, and axial rotation, playing a crucial role in highly agile and dexterous locomotion. While numerous studies have integrated active spinal joints into quadrupedal robots to enhance agility, most designs simplify control complexity by reducing spinal degrees of freedom (DOF), failing to achieve the spatial tri-axial rotation characteristic of biological spines. Consequently, replicating a multi-DOF biomimetic spine and effectively leveraging it to empower the agile locomotion of quadrupedal robots remains a significant research challenge. In this study, we present S-Cheetah, a quadrupedal robot featuring a 3-DOF bio-inspired serial active spine capable of biomimetic spatial tri-axial rotation. To empower the robot to fully utilize this active spine, we developed a specialized reinforcement...

论文介绍 四足机器人敏捷运动受限于简化脊柱设计,难以实现生物脊柱的多轴旋转能力。本文提出S-Cheetah机器人,配备三自由度生物启发主动脊柱,支持矢状弯曲、侧弯和轴旋转。通过强化学习训练,机器人能充分利用脊柱功能,提升运动敏捷性,对复杂地形中的机器人适应性有重要意义。

Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language

第一作者: Qiwei Wu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Tactile sensing is essential for robots to achieve human-like gentle manipulation. However, existing Vision-Language-Action (VLA) models struggle to exploit tactile feedback for gentle manipulation due to scarce aligned vision-tactile-language data and the lack of effective closed-loop force feedback mechanisms. To address these challenges, we introduce Tabero, a benchmark and model suite for gentle, language-conditioned robotic manipulation that demands fine-grained contact force perception. First, the Tabero benchmark addresses the scarcity of tactile data by presenting a data-efficient pipeline that repurposes open-source robot manipulation trajectories to generate diverse vision-tactile-language tasks, and establishes a multidimensional evaluation protocol that measures task success alongside physical interaction quality. Second, we propose Tabero-VTLA, an architecture...

论文介绍 现有视觉-语言-动作模型在温和操作中难以利用触觉反馈,因数据稀缺和机制缺失。Tabero通过数据高效管道生成多模态任务,并提出Tabero-VTLA架构,集成闭环力反馈机制,支持精细接触力感知。这推动了语言条件下的机器人温和操作学习,提升物理交互质量。

Turning Video Models into Generalist Robot Policies

第一作者: Sizhe Lester Li · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments. Recent work proposes robot foundation models that jointly predict future observations and actions by finetuning video models with action-labeled data. In this paper, we test the limits of an alternative approach: leave the video planner as-is while training an embodiment-specific inverse dynamics model (IDM). This decoupling offers several natural benefits: the video planner remains embodiment-agnostic, different video models can be interchanged easily without re-training the IDM, and the IDM can be independently trained with readily available self-play data. We present a closed-loop, video-to-action policy that combines an action-free video world model with a carefully-designed IDM based on the...

论文介绍 视频生成模型可作为机器人骨干,但需与动作生成结合。本文探索解耦方法:保持视频规划器动作自由,训练特定于具身的逆动力学模型。这使视频模型通用化,易于更换,并支持闭环策略,应用于机器人学习任务,提高策略的灵活性和可扩展性。

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

第一作者: Jeremy Morgan · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress can be misleading. Despite the zero-shot perception and language capabilities of VLAs, their overall task performance often degrades under distribution shifts, revealing gaps in how these systems translate high-level understanding into robust behavior. To systematically study this gap, we introduce Colosseum V2, a large-scale simulation benchmark for evaluating VLA generalization in robot learning across diverse conditions. The benchmark comprises 28 tasks spanning 13 task categories and two robot morphologies, covering a wide range of manipulation primitives and long-horizon behaviors. Built on the ManiSkill simulator, Colosseum V2 enables fast, GPU-parallelized evaluation and supports both...

论文介绍 视觉-语言-动作模型在分布偏移下泛化能力常退化。Colosseum V2提供大规模仿真基准,涵盖28个任务和多样条件,支持GPU并行评估。这系统研究VLA模型的泛化问题,指导模型改进,增强机器人操作的鲁棒性和适应性。

HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning

第一作者: Kevin Lin · 方向: 机器人操作 · 来源: cs.RO

Abstract:Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations, which are time-intensive and difficult to collect via teleoperation. Existing data-generation algorithms can automatically synthesize demonstrations for manipulators, but they are ineffective on humanoids because their high-dimensional composite action spaces involve arms, legs, and torsos. We present HumanoidMimicGen, a method for generating humanoid legged loco-manipulation data. Our method adapts contact-rich whole-body skills from a handful of source demonstrations to new states, generalizing across changes in object pose. By interleaving these single- and dual-arm skills with whole-body locomotion and manipulation planning, the method generates stable, collision-free data across diverse scenes and layouts. To evaluate our...

论文介绍 类人机器人模仿学习需大量演示数据,收集耗时且困难。HumanoidMimicGen从少量源演示适应接触丰富的全身技能,通过交织单臂、双臂技能与全身规划,生成稳定、无碰撞数据。这自动化数据生成过程,支持训练类人机器人的运动操作能力。

Simulation-Informed Diffusion for Decentralized Multi-robot Motion Planning

第一作者: Jinhao Liang · 方向: 导航与运动 · 来源: cs.RO

Abstract:Decentralized multi-robot motion planning requires each robot to generate collision-free trajectories from local observations, without global sensing or reliable communication. However, most existing planners, whether classical or learning-based, generate trajectories from a static snapshot of the local observation, which limits their ability to anticipate the future behavior of neighboring robots. This limitation is critical as the number of robots increases and the environment becomes more cluttered. To overcome this challenge, this paper introduces Simulation-Informed Diffusion (SID), a decentralized framework built on constraint-aware diffusion models (CADM). SID first uses CADM to simulate the future trajectories of neighboring robots from their currently observed states, and then uses the same CADM to plan each robot's own trajectory under safety constraints informed by...

论文介绍 去中心化多机器人运动规划中,基于静态观测生成轨迹限制未来行为预测。本文提出仿真启发扩散框架SID,使用约束感知扩散模型模拟邻居机器人未来轨迹,并在此基础上规划自身轨迹。这增强对动态环境的适应,提高多机器人系统的协调性和安全性。

Trinity: Unifying Class-Agnostic Terrain and Semantic Segmentation for Unstructured Outdoor Environments by Leveraging Synthetic Data

第一作者: Marcus G Müller · 方向: 具身智能 · 来源: cs.RO

Abstract:Terrain understanding is fundamental for mobile robots operating in unstructured outdoor environments. Existing vision-based traversability estimation methods rely on robot-specific annotations or semantic class mappings, limiting transferability across platforms and requiring costly re-annotation when robot capabilities change, while standard semantic segmentation methods only focus on specific predefined classes, which do not capture the variety of terrains. In this work, we propose a transformer-based architecture that jointly performs class-specific semantic segmentation and class-agnostic terrain segmentation within a unified network, called Trinity. Terrain regions are segmented based solely on visual appearance, without predefined semantic labels or robot-dependent traversability scores. This formulation enables the learning of robot-agnostic visual terrain priors that...

论文介绍 户外移动机器人地形理解依赖特定标注,限制跨平台应用。Trinity网络统一语义和地形分割,基于视觉外观学习机器人无关的地形先验,无需预定义标签。通过合成数据训练,提升泛化能力,支持非结构化环境中的遍历性估计和机器人导航。

Agentic Language-to-Objective Synthesis for Optofluidic Assembly

第一作者: Ivan Saraev · 方向: 机器人操作 · 来源: cs.RO

Abstract:Light-based advanced manufacturing increasingly requires programmable, closed-loop tools that translate human design intent into executable operations at small length scales. Yet a key bottleneck persists across robotic and manufacturing modalities: turning user intent into machine-readable objectives that are reliably executable. While micro-robotics offers versatile manipulation via optical actuation of fluids, mathematically tractable goal specification remains manual and hard to reuse. Here, we introduce Speak-to-Objective, a modular agentic pipeline that uses a conditioned Large Language Model (LLM) to translate spoken or written commands into fully differentiable objective functions for assembling microparticles in a constraint-aware inverse solver (SLSQP) and on an experimental optofluidic platform. The approach employs a compact loop - perceive -> compose -> propose ->...

论文介绍 该研究针对光流体装配中人类设计意图转化为机器可读目标的瓶颈问题。提出模块化代理管道“Speak-to-Objective”,利用条件大型语言模型将语音或文本命令转换为可微分目标函数,并在约束感知逆求解器和实验光流体平台上实现微粒装配。通过紧凑循环感知-组合-提议,旨在提升小型尺度制造的可编程性和闭环控制能力。

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation

第一作者: Hongyu Ding · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dominant approach has been to scale vision-language-action (VLA) foundation models on ever-larger collections of robot trajectories. This paper argues that, for navigation specifically, generality can be obtained structurally, not only through data scale. The underlying decision structure of navigation reduces to a single Language-Vision-Robot Actions Translation. The language action emits semantic-level directional command and the vision action emits a pixel-level visual target. Both outputs lie inside the natural output manifold of pretrained multimodal large language models (MLLMs), so the task can be reasoned about by an agent rather than learned from robot data. Therefore, we present Uni-LaViRA, a...

论文介绍 本文探讨具身导航中泛化问题,认为结构而非数据规模更能提升导航通用性。提出Uni-LaViRA框架,将导航决策简化为语言-视觉-机器人动作翻译,利用预训练多模态大语言模型输出语义方向命令和像素级视觉目标,通过智能体推理避免从机器人数据中学习。该方法旨在实现统一导航,适用于未见环境。

Synthetic Emotions vs. Gamification: Exploring Engagement Strategies for Small Social Robots in Different Age Groups

第一作者: Morten Roed Frederiksen · 方向: 具身智能 · 来源: cs.RO

Abstract:Many children experience challenges in emotional regulation and social interaction, which can limit their participation in everyday activities and therapeutic programs. For socially assistive robots to be effective in this context, it is essential that children remain consistently and meaningfully engaged. We explore engagement strategies for a tactile robot designed to support children suffering from anxiety disorders through daily interactions. The robot delivers either synthetic emotional feedback or point rewards to encourage user participation. We evaluated these strategies through two studies: a preference assessment with 16 school children aged 6-8 years, and a behavioral study with 14 university students aged 20-27 years in naturalistic environments. The study with school children indicated a preference for emotional engagement over points-based approaches. The follow...

论文介绍 该研究探索社会辅助机器人对焦虑症儿童的参与策略,比较合成情感反馈和点数奖励两种方式。通过偏好评估(16名6-8岁儿童)和行为研究(14名20-27岁大学生),发现儿童更偏好情感交互而非游戏化方法。研究强调了情感engagement在治疗中的重要性,可能应用于儿童心理健康辅助和日常互动。

Inducing Calmness With Pocket-Sized Robotics: Reducing Movement and Heart Rate in Children through Hand-Held Tactile Interactions

第一作者: Morten Roed Frederiksen · 方向: 具身智能 · 来源: cs.RO

Abstract:Periods of heightened arousal or restlessness can interfere with children's ability to focus, self-regulation, and physically calm. Technologies that encourage embodied self-regulation through tactile interaction may provide a simple and accessible means of promoting calmness. This paper investigates how interaction with a pocket-sized tactile device influences physiological and behavioral markers of calmness in typically developing children. Building on prior work examining heart rate modulation, we present new findings on how tactile interaction affects full-body movement and postural stability. We employ a device that engages children through a hand-held rhythmic vibration-matching game, designed to focus attention and encourage stillness. Eighteen children participated in a within-subjects study that involved two conditions: with and without tactile interaction with a...

论文介绍 本文研究口袋大小触觉设备如何通过手握触觉互动促进儿童平静。设备采用振动匹配游戏,旨在集中注意力和鼓励静止。通过实验测量全身运动和心率,发现触觉交互能减少运动并稳定心率,为注意力缺陷或焦虑儿童提供简单易用的自我调节工具,促进embodied平静状态。

SCALE-COMM: Shared, Contrastively-Aligned Latent Embeddings for MARL Communication

第一作者: Mahmoud Abouelyazid · 方向: 策略学习 · 来源: cs.RO

Abstract:Emergent communication enables partially observant Autonomous Mobile Robots (AMRs) to coordinate effectively in decentralized multi-agent reinforcement learning (MARL) settings. However, existing approaches often struggle with unstable communication protocols, ungrounded message semantics, and interference between communication learning and policy optimization, leading to degraded coordination over time. We propose SCALE-COMM (Shared, Contrastively-Aligned Latent Embeddings for COMMunication), a self-supervised framework for learning compact, stable, and policy-relevant communication representations. SCALE-COMM decouples communication learning from policy optimization by training low-dimensional latent messages that capture task-relevant planning and traffic information, while enforcing consistency across agents and time. Across standard MARL benchmarks and a realistic...

论文介绍 该研究针对多智能体强化学习中通信不稳定、语义无根基和学习干扰问题,提出SCALE-COMM自监督框架。通过训练低维潜在消息捕获任务相关规划和流量信息,利用对比学习确保跨智能体和时间一致性,解耦通信学习与策略优化。在基准测试和现实场景中验证其提升协调效果的能力。

GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation

第一作者: Boxiang Qiu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:We introduce GE-Sim 2.0 (Genie Envisioner World Simulator 2.0), a closed-loop video world simulator for robotic manipulation. Building on the action-conditioned video generation framework of Genie Envisioner, GE-Sim 2.0 is re-trained on thousands of hours of real-world robot data spanning teleoperation, contact-rich interaction, and on-robot policy deployment, substantially improving action-following fidelity and trajectory coverage. On top of this foundation, three new modules close the loop from video simulation to policy learning: a state expert that decodes proprioceptive state from video latents to support next-chunk prediction by downstream VLA policies; a world judge that scores generated rollouts against task instructions, yielding machine-verifiable success signals and rewards in place of manual inspection; and an acceleration framework that delivers a 25-frame...

论文介绍 本文介绍GE-Sim 2.0,一个用于机器人操作的闭环视频世界模拟器。基于动作条件视频生成,在真实机器人数据上重新训练,提升动作跟随精度和轨迹覆盖。新增状态专家、世界判断器和加速框架,支持策略学习、奖励生成和25帧加速,促进机器人操作模拟和策略训练的高效进行。

A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons

第一作者: Brian Zhu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) policies have shown promising manipulation capabilities, yet their practical impact is often limited by the reliability demands of real-world deployment. We present a deployment study of an industrial packaging task at Siemens Factory (GWE, Erlangen, Germany), where a robot must pick a transparent accessory bag from a cluttered pile, insert it into the remaining cavity of a cardboard package, and ensure that the bag and its contents remain below the closing plane. Our goal is to understand the practical effort required to adapt a pretrained Pi0.5 policy to a single factory-floor task through iterative fine-tuning and deployment-driven refinement. The pipeline consists of repeated loops of data collection, curation, fine-tuning, evaluation, and targeted recovery data collection. We have accumulated 2535 episodes (10 hours) from the on-site factory...

论文介绍 该研究呈现西门子工厂工业包装任务中视觉-语言-动作管道的部署案例。机器人需从杂乱堆中抓取透明袋并插入纸盒。通过迭代数据收集、微调和评估,探索预训练策略适应单一工厂任务的实用性和挑战,积累了2535个episode数据,强调现实部署中的可靠性和工作流程。

Teacher-Student Representational Alignment for Reinforcement Learning-Driven Imitation Learning

第一作者: Meraj Mammadov · 方向: 模仿学习 · 来源: cs.RO

Abstract:Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex and high-dimensional observation spaces prevalent in robotics. This paper addresses the irreducible imitation gap that emerges when teacher and student are learned in isolation, and the teacher policy has the liberty to rely on privileged state information that the student cannot infer from its observations. Instead of improving poor student performance with RL finetuning after IL, which often requires a whole new training setup, we propose a novel algorithm which learns a shared embedding space that hides agent-specific observations and thus trains imitable teacher policies by construction. We train the shared embedding space with self-supervised contrastive learning in parallel to the teacher policy and prevent it from...

论文介绍 本文解决模仿学习中师生策略隔离学习导致的模仿差距问题,其中教师可依赖特权状态信息。提出新算法,通过自监督对比学习训练共享嵌入空间,隐藏智能体特定观察,使教师策略本身可模仿。该方法避免了强化学习微调的新训练设置,提升学习效率和适应性。

Robo-Blocks: Generative Scaffolding in End-User Design and Programming of Social Robots

第一作者: Arissa J. Sato · 方向: 具身智能 · 来源: cs.RO

Abstract:Programming social robots is challenging for novice robot programmers due to required expertise in planning, interaction design, and programming. While large language models (LLMs) hold significant promise through code generation from natural-language descriptions, they can obscure critical elements of programming and supplant designer intent, eventually resulting in over-reliance instead of developing programming skills. In this paper, we explore how LLM-based social-robot-programming tools can support novice robot programmers through a Research through Design (RtD) process. We designed and prototyped Robo-Blocks, a block-based programming environment that leverages LLMs to offer novice robot programmers generative scaffolding through structured narratives that connect high-level ideas to executable robot behaviors. Through deployment with novices, we discovered emerging user...

论文介绍 本文探讨新手编程社交机器人的挑战,指出大语言模型可能导致过度依赖。研究提出Robo-Blocks,一个基于LLM的积木编程环境,通过结构化叙述提供生成式脚手架,连接高层想法与可执行行为,旨在支持新手学习编程技能。

Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

第一作者: Haolan Zhang · 方向: 具身智能 · 来源: cs.RO

Abstract:Visual odometry (VO) is a fundamental component in robotics and augmented reality. RGB-D direct VO benefits from metric depth measurements, but it can degrade in challenging environments, where dynamic objects, occlusions, illumination changes, and unreliable depth violate the short-horizon photometric and depth-geometric consistency assumptions used by direct alignment. Existing approaches mitigate these issues through semantic filtering, explicit occlusion reasoning, illumination adaptation, or hand-crafted geometric criteria, but often rely on external modules or fixed assumptions tailored to individual failure modes, limiting their flexibility and ability to handle diverse challenges in a unified manner. In this work, we propose Con-DSO, a consistency-aware RGB-D direct sparse odometry framework that predicts dense photometric and depth-geometric consistency uncertainty...

论文介绍 针对RGB-D直接稀疏里程计在动态环境中的不一致性问题,本文提出Con-DSO框架。该框架学习短期光度和几何一致性先验,预测不确定性以提高鲁棒性,适用于机器人导航和增强现实应用。

Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections

第一作者: Wenzhe Song · 方向: 策略学习 · 来源: cs.RO

Abstract:The imminent integration of autonomous vehicles and mobile robots in urban settings presents a critical safety challenge for future intelligent transportation systems. This paper addresses the complex problem of coordinating heterogeneous agents with disparate dynamics at unregulated intersections. We introduce a novel framework, differentiable model predictive safety (DMPS), which embeds the foresight of model-predictive control into a data-driven, end-to-end reinforcement learning architecture. DMPS agents learn a latent dynamics model to predict future trajectories contingent on their actions. A learned, differentiable safety critic then evaluates the risk of these trajectories. Crucially, by leveraging backpropagation through the entire unrolled predictive model, agents can efficiently compute the gradient of future safety with respect to their current action, enabling a...

论文介绍 本文解决城市交叉口中异构移动体的安全协调问题,引入可微分模型预测安全框架。该框架嵌入模型预测控制与强化学习,通过预测未来轨迹和评估风险,实现端到端安全策略学习,用于智能交通系统。

Surprising Performances of Students with Autism in Classroom with NAO Robot

第一作者: Qin Yang · 方向: 具身智能 · 来源: cs.RO

Abstract:Autism is a developmental disorder that manifests in early childhood and persists throughout life, profoundly affecting social behavior and hindering the acquisition of learning and social skills in those diagnosed. As technological advancements progress, an increasing array of technologies is being utilized to support the education of students with Autism Spectrum Disorder (ASD), aiming to improve their educational outcomes and social capabilities. Numerous studies on autism intervention have highlighted the effectiveness of social robots in behavioral treatments. However, research on the integration of social robots into classroom settings for children with autism remains sparse. This paper describes the design and implementation of a group experiment in a collective classroom setting mediated by the NAO robot. The experiment involved special education teachers and the NAO...

论文介绍 本研究探索在课堂环境中集成NAO机器人以支持自闭症学生教育。通过群组实验设计,观察学生表现,旨在评估社交机器人对自闭症学生学习和社交技能的影响,为教育干预提供新途径。

Ω-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Per-step Scaling

第一作者: Xinyu Wang · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and diffusion-based action heads make on-device deployment prohibitively expensive. Prior quantization efforts offer only partial solutions, compressing the LLM backbone while leaving the DiT action head at full precision, or resorting to mixed-precision schemes, driven by the belief that uniformly quantizing the action head is inherently unstable. We challenge this assumption with Omega-QVLA, the first training-free post-training quantization framework that compresses both the language backbone and the entire diffusion action head of a VLA model to a uniform W4A4 precision, eliminating the need for mixed-precision allocation. Omega-QVLA combines a composite SVD-Hadamard rotation that equalizes per-channel weight energy while...

论文介绍 视觉语言动作模型部署成本高,本文提出Ω-QVLA,一个无训练后量化框架。通过复合旋转和逐步缩放,将语言骨干和扩散动作头统一压缩到W4A4精度,降低计算需求,适用于设备端部署。

GEM: Generative Supervision Helps Embodied Intelligence

第一作者: Ruowen Zhao · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a significant gap remains between the high-level semantic focus of standard text-guided pre-training paradigms and the low-level spatial and physical knowledge critical for execution in embodied environments. In this paper, we introduce GEM, a Generative-supervised Embodied vision-language Model designed to bridge this divide. We propose integrating a depth map generation task directly into the VLM pre-training phase. By training this generative objective jointly with the main model, we observe substantial improvements in embodied intelligence, significantly enhancing both semantic understanding and physical operation capabilities. To support this paradigm, we curate and release GEM-4M, a comprehensive...

论文介绍 本文提出GEM模型,通过在视觉语言模型预训练中集成深度图生成任务,增强具身智能。该方法提升模型对物理环境的理解和操作能力,支持机器人任务执行,改善语义和空间认知。

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking

第一作者: Jiyuan Fu · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly hinders their deployment in safety-critical domains. Moreover, existing patch attacks primarily focus on white-box settings, heavily overfitting to the specific action output space of the target model, which results in poor cross-architecture transferability. To overcome this limitation, we propose VLA-Hijack, a unified adversarial framework that breaks the transferability bottleneck by exploiting a fundamental vulnerability identified in this work: before planning any motion, a VLA model must first use visual information to locate its own robotic arm within the environment. Targeting this shared visual self-localization process, our approach concurrently optimizes Attention-Guided Proprioceptive Suppression to inhibit the real...

论文介绍 VLA模型易受对抗性补丁攻击,本文提出VLA-Hijack框架。通过攻击模型视觉自定位过程,实现跨架构的可迁移攻击,评估模型安全漏洞,为安全关键领域部署提供警示。

Off-Policy Learning to Reason Works Because It Is More Pessimistic Than You Think

第一作者: Otmane Sakhi · 方向: 策略学习 · 来源: cs.LG

Abstract:Large scale reinforcement learning has become a central tool for improving reasoning in large language models. At this scale, generation is often lagged or asynchronous, so updates are performed on data collected by older policies. This makes learning inherently off-policy. Most existing approaches nevertheless remain rooted in PPO-style trust-region objectives, treating training as approximately on-policy and using importance weights to correct distribution mismatch. These corrections can introduce high variance, destabilize optimization, and accelerate entropy collapse. Recent work suggests an alternative: rather than correcting the mismatch, one can embrace off-policy data and remove importance weights, often yielding stronger algorithms. In this paper, we provide an intuitive construction of off-policy objectives that include successful off-policy objectives and show that...

论文介绍 本文分析大规模强化学习中离策略学习的有效性,指出其本质更悲观。通过构建离策略目标,展示无需重要性权重的优化方式,可能稳定训练并提升大语言模型的推理能力。

PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft

第一作者: Yuchen Guo · 方向: 导航与运动 · 来源: cs.AI

Abstract:We present PEAM, a Parametric Embodied Agent Memory framework in Minecraft that transforms agent memory from inference-time retrieval into parameter-resident skills internalized through experience. PEAM pairs a slow deliberative LLM for open-ended reasoning with a fast parametric module for reflexive execution of consolidated skills. The fast module is a multimodal Mixture-of-Experts LoRA architecture with per-category physically isolated adapters, enabling parameter-level continual learning without catastrophic forgetting. We treat failure as a first-class training signal: failure--correction trajectory pairs are internalized through a joint behavioral-cloning and contrastive objective, so the agent learns not only what succeeds but also how corrected actions differ from failed ones. To govern consolidation, PEAM introduces a parameterization-worthiness score for deciding...

论文介绍 本文提出PEAM框架,用于Minecraft中的参数化具身代理记忆,将代理记忆从推理时检索转化为通过经验内化的参数驻留技能。该框架结合慢速LLM用于开放推理和快速参数模块用于技能执行,采用多模态MoE LoRA架构实现无灾难遗忘的持续学习。通过将失败作为训练信号,内化失败-纠正轨迹对,提升代理学习效率。

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs

第一作者: Marcell Fekete · 方向: 多模态具身 · 来源: cs.CL

Abstract:Vision-language models (VLMs) are increasingly evaluated for whether they identify the right visual content, but little is known about whether they express such content in a discourse-appropriate form. We address this research gap using information structure (IS), testing whether VLMs distinguish discourse-old Topics from discourse-new Foci in visually grounded question answering. We exploit Hungarian, a language in which Topic and Focus map onto dedicated syntactic positions, making IS choices observable in text. Comparing six VLMs with human participants, we find that models produce IS-relevant constructions, but over-regularise this sensitivity. Under the interacting pressures of discourse status, grammatical role (preference for subject Topics) and definiteness (preference for indefinite Foci), humans choose variable strategies for IS realisation. VLMs, by contrast...

论文介绍 本文研究视觉语言模型输出的信息结构,探讨其在视觉基础上问答中是否区分话语旧主题和新焦点。通过利用匈牙利语中主题和焦点的句法位置,比较六个VLMs和人类参与者。结果显示模型虽能产生相关构造,但过度规则化,在话语压力下缺乏人类的可变策略适应能力。

市场总览

美股技术面整体偏强,SPY、QQQ等主要指数RSI进入超买区域,价格接近52周高点,呈多头排列,但部分科技股如MSFT、NVDA走势中性,动量出现分歧。加密货币市场情绪极度恐慌,恐慌贪婪指数为23,总市值2.55万亿美元24小时下跌0.78%,BTC主导率57.7%;BTC、ETH、SOL价格均低于关键移动平均线,RSI偏低,空头排列信号明显。中概股普遍承压,BABA、PDD、JD及0700.HK的RSI处于超卖或低位,价格远低于SMA200,空头排列主导,5日跌幅显著。商品与外汇方面,黄金期货价格中性,WTI原油MACD死叉且价格跌破SMA50,美元/人民币呈下行趋势,美元指数DXY则小幅走强,接近52周高点。

今日关注

PDD 拼多多 (PDD)
偏下行

当前价格83.03,1日跌幅-4.13%,5日跌幅-15.4%。RSI14为29.3,进入超卖区域。MACD指标显示死叉,且价格远低于SMA20(96.57)、SMA50(98.74)和SMA200(113.39),呈空头排列。价格接近52周低点1.8%,技术面偏下行信号显著。

0700.HK 腾讯控股 (0700.HK)
偏下行

价格425.6,1日涨幅0.14%但5日跌幅-3.05%。RSI14为29.8,处于超卖状态。MACD为-16.366,信号线-14.4736,空头排列。价格低于SMA20(454.72)、SMA50(484.58)和SMA200(574.68),且接近52周低点1.24%,下行趋势明确。

QQQ Nasdaq 100 ETF
偏上行

当前价格735.6,1日涨幅0.84%,5日涨幅3.15%。RSI14为76.4,处于超买区域。MACD为21.2604,信号线21.4302,趋势为bullish且呈多头排列,价格高于SMA20(705.51)、SMA50(650.06)和SMA200(617.05)。接近52周高点-0.14%,动量偏上行。

MSFT Microsoft
中性

价格426.99,1日涨幅3.47%但5日涨幅1.41%。RSI14为59.3,处于正常范围。MACD为3.4774,信号线3.7045,未显示明显金叉或死叉。趋势为neutral,价格高于SMA20(415.47)但低于SMA200(458.86),技术面信号中性。

全部资产

^VIX

VIX 恐慌指数

$15.74 -3.38%
5 日
-6.09%
距 52w 高
-55.4%
RSI(14)
37.2
趋势
中性
SMA 20 / 50 / 200
17.33 / 20.13 / 18.38
MACD / 信号
-0.845 / -0.829
MACD 死叉 (今天)

^TNX

10Y 美债收益率 (%)

$4.46 -0.58%
5 日
-2.56%
距 52w 高
-10.8%
RSI(14)
50.0
趋势
多头
SMA 20 / 50 / 200
4.48 / 4.39 / 4.20
MACD / 信号
0.048 / 0.059
MACD 死叉 (1 天前)多头排列

DX-Y.NYB

美元指数 DXY

$99.03 +0.01%
5 日
-0.17%
距 52w 高
-1.6%
RSI(14)
53.2
趋势
多头
SMA 20 / 50 / 200
98.72 / 98.90 / 98.57
MACD / 信号
0.147 / 0.097
接近 52 周高多头排列

SPY

S&P 500 ETF

$754.60 +0.55%
5 日
+1.80%
距 52w 高
-0.1%
RSI(14)
73.2
趋势
多头
SMA 20 / 50 / 200
737.44 / 701.75 / 680.60
MACD / 信号
12.576 / 12.917
RSI 超买接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$735.60 +0.84%
5 日
+3.15%
距 52w 高
-0.1%
RSI(14)
76.4
趋势
多头
SMA 20 / 50 / 200
705.51 / 650.06 / 617.05
MACD / 信号
21.260 / 21.430
RSI 超买接近 52 周高多头排列

AAPL

Apple

$312.51 +0.53%
5 日
+3.39%
距 52w 高
-0.2%
RSI(14)
79.8
趋势
多头
SMA 20 / 50 / 200
295.51 / 274.04 / 262.83
MACD / 信号
10.424 / 9.624
RSI 超买接近 52 周高多头排列

MSFT

Microsoft

$426.99 +3.47%
5 日
+1.41%
距 52w 高
-23.1%
RSI(14)
59.3
趋势
中性
SMA 20 / 50 / 200
415.47 / 401.66 / 458.86
MACD / 信号
3.477 / 3.704

NVDA

Nvidia

$214.25 +0.78%
5 日
-4.13%
距 52w 高
-9.4%
RSI(14)
52.6
趋势
多头
SMA 20 / 50 / 200
214.88 / 198.73 / 187.51
MACD / 信号
4.595 / 6.515
MACD 死叉 (4 天前)多头排列

GOOGL

Alphabet

$390.13 +0.33%
5 日
+0.31%
距 52w 高
-4.5%
RSI(14)
61.8
趋势
多头
SMA 20 / 50 / 200
391.37 / 346.12 / 299.04
MACD / 信号
11.108 / 14.518
多头排列

TSLA

Tesla

$442.10 +0.40%
5 日
+5.95%
距 52w 高
-11.4%
RSI(14)
63.6
趋势
中性
SMA 20 / 50 / 200
418.69 / 390.93 / 411.65
MACD / 信号
12.192 / 11.193
MACD 金叉 (1 天前)

META

Meta

$635.29 +0.00%
5 日
+5.00%
距 52w 高
-20.2%
RSI(14)
56.9
趋势
中性
SMA 20 / 50 / 200
612.30 / 618.20 / 667.36
MACD / 信号
-2.364 / -5.099
MACD 金叉 (1 天前)
加密恐慌贪婪
23
极度恐慌
加密总市值
$2.55 T
-0.78% / 24h
BTC 主导率
57.7%
ETH 9.5%
24h 成交量
$101.2 B
活跃币 17,401

BTC-USD

Bitcoin

$73,474.00 -1.17%
5 日
-4.17%
距 52w 高
-41.8%
RSI(14)
35.5
趋势
空头
SMA 20 / 50 / 200
77,936.65 / 77,171.88 / 79,972.03
MACD / 信号
-758.751 / -158.595
空头排列

ETH-USD

Ethereum

$2,005.39 -0.83%
5 日
-5.23%
距 52w 高
-59.5%
RSI(14)
30.8
趋势
空头
SMA 20 / 50 / 200
2,168.21 / 2,255.24 / 2,519.21
MACD / 信号
-63.423 / -49.017
空头排列

SOL-USD

Solana

$82.08 -0.35%
5 日
-4.18%
距 52w 高
-67.6%
RSI(14)
38.7
趋势
空头
SMA 20 / 50 / 200
87.82 / 86.49 / 105.59
MACD / 信号
-1.221 / -0.518
空头排列

BABA

阿里巴巴 (BABA)

$126.16 -1.25%
5 日
-6.18%
距 52w 高
-34.5%
RSI(14)
39.8
趋势
空头
SMA 20 / 50 / 200
134.56 / 131.27 / 149.61
MACD / 信号
-1.390 / -0.077
空头排列

PDD

拼多多 (PDD)

$83.03 -4.13%
5 日
-15.40%
距 52w 高
-40.4%
RSI(14)
29.3
趋势
空头
SMA 20 / 50 / 200
96.57 / 98.74 / 113.39
MACD / 信号
-2.674 / -1.447
MACD 死叉 (3 天前)RSI 超卖接近 52 周低空头排列

JD

京东 (JD)

$29.14 -2.35%
5 日
-10.23%
距 52w 高
-20.9%
RSI(14)
40.4
趋势
空头
SMA 20 / 50 / 200
30.95 / 30.02 / 30.46
MACD / 信号
0.070 / 0.414
MACD 死叉 (3 天前)空头排列

0700.HK

腾讯控股 (0700.HK)

HK$425.60 +0.14%
5 日
-3.05%
距 52w 高
-37.7%
RSI(14)
29.8
趋势
空头
SMA 20 / 50 / 200
454.72 / 484.58 / 574.68
MACD / 信号
-16.366 / -14.474
RSI 超卖接近 52 周低空头排列

GC=F

黄金期货

$4,534.60 +1.96%
5 日
+0.07%
距 52w 高
-18.8%
RSI(14)
43.9
趋势
中性
SMA 20 / 50 / 200
4,594.16 / 4,637.56 / 4,364.97
MACD / 信号
-54.196 / -45.543

CL=F

WTI 原油期货

$88.09 -0.67%
5 日
-10.35%
距 52w 高
-26.3%
RSI(14)
38.5
趋势
中性
SMA 20 / 50 / 200
99.36 / 98.03 / 71.91
MACD / 信号
-1.124 / 0.710
MACD 死叉 (4 天前)

USDCNY=X

美元 / 人民币

¥6.77 -0.12%
5 日
-0.46%
距 52w 高
-6.1%
RSI(14)
32.1
趋势
空头
SMA 20 / 50 / 200
6.80 / 6.83 / 6.99
MACD / 信号
-0.014 / -0.013
MACD 死叉 (1 天前)接近 52 周低空头排列
风险提示

本报告基于公开技术指标数据,仅供技术指标解读参考,不构成投资建议。过去走势不代表未来表现,投资者应结合自身风险承受能力独立决策。

JD Vance says US ‘not there yet’ on an Iran deal – as it happened

This live blog is now closed, you can read our latest report from the Middle East here Donald Trump shares draft Iran peace agreement with Israel and other allies Hezbollah has claimed dozens of drone and rocket attacks that it said targeted Israeli troops in southern Lebanon and northern Israel. Th

中文摘要 美国副总统JD·万斯表示,美国在伊朗协议上尚未达成一致。特朗普总统向以色列等盟友分享了伊朗和平协议草案。真主党声称对黎巴嫩南部以色列军队发动了数十次无人机和火箭袭击。

Australia news live: Australia’s greenhouse gas emissions fall 2%; consumer watchdog sues Amazon over ‘hazardous’ pink toddler unicorn backpack

Government data says rise of renewable energy displacing some coal and gas electricity. Follow today’s news live Get our breaking news email, free app or daily news podcast Who’s leading the fight against Labor’s CGT reform – and what’s in it for them? People with direct personal financial stakes in

中文摘要 澳大利亚政府数据显示温室气体排放下降2%,可再生能源上升替代了部分煤炭和天然气电力。消费者监管机构起诉亚马逊销售“危险”的粉色幼儿独角兽背包。

Here’s the latest.

中文摘要 该文章标题为“最新消息”,但摘录内容为空,无法提供详细信息。

The Man Turning the Cockroach Into a Gen-Z Movement in India

Abhijeet Dipke’s “Cockroach Janta Party” has emerged as the unexpected voice of young people feeling let down by the government and struggling to find jobs.

中文摘要 在印度,Abhijeet Dipke创建的“蟑螂人民党”成为对政府失望且就业困难的年轻人的意外代言人,形成了Z世代运动。

Italian Police Uncover Dead Mob Boss’s $230 Million Business Empire

When Matteo Messina Denaro died in 2023, many of his secrets died with him. That began to change with a tip about a wealthy Sicilian woman’s significant assets in a tiny European principality.

中文摘要 意大利警方根据线索,揭露了2023年去世的黑手党老板Matteo Messina Denaro价值2.3亿美元的商业帝国,涉及一名西西里岛女性在欧洲小公国的资产。

US to designate two Brazilian gangs as ‘terrorist’ organisations

Trump administration has used crime and drug trafficking to push for greater US military influence across Latin America.

中文摘要 美国计划将两个巴西帮派列为“恐怖”组织。特朗普政府借此推动在拉丁美洲增强军事影响力,以应对犯罪和贩毒问题。

Donald Trump shares draft Iran peace agreement with Israel and other allies

US president’s move comes as both sides try to prevent fresh ceasefire breaches scuppering a potential deal Middle East crisis – live updates Donald Trump has circulated a draft peace agreement for the war with Iran among allies including Israel as both sides try to prevent fresh breaches of the cea

中文摘要 美国总统特朗普向以色列等盟友分享了与伊朗战争的和平协议草案,双方正努力防止新的停火违反破坏潜在协议。

U.S. and Iran Move Toward Agreement to Reopen the Strait of Hormuz

President Trump has yet to sign off on an extension of the cease-fire that would allow the two sides to negotiate nuclear and other issues.

中文摘要 美国和伊朗正朝着重新开放霍尔木兹海峡的协议迈进,但特朗普总统尚未批准延长停火以谈判核及其他问题。

Iran war live: Tehran, Trump yet to comment on 60-day truce extension plan

Lebanese PM Nawaf Salam says 'nothing can justify' Israel's military onslaught against 'peaceful people' in Lebanon.

中文摘要 伊朗和特朗普总统尚未就60天停火延长计划发表评论。黎巴嫩总理纳瓦夫·萨拉姆表示,以色列对黎巴嫩“和平人民”的军事攻击“无可辩解”。

Netanyahu orders Israeli army to seize 70 percent of Gaza

Israeli Prime Minister Benjamin Netanyahu says he has ordered the military to take control of 70 percent of Gaza.

中文摘要 以色列总理本雅明·内塔尼亚胡下令军队控制加沙地带70%的区域。

US and Iran 'very close' to deal but 'not there yet', Vance says

US officials earlier told the BBC that the framework of a ceasefire extension deal had been agreed, pending the approval of Trump and Iran's leadership.

中文摘要 副总统JD·万斯表示,美国和伊朗“非常接近”达成协议,但“尚未完成”。美国官员称停火延长协议框架已达成,正等待特朗普和伊朗领导的批准。

Tense protests erupt outside Delaney Hall immigrant detention centre in US

A hunger strike at the New Jersey facility has raised questions about conditions in US immigration detention centres.

中文摘要 美国新泽西州Delaney Hall移民拘留中心外爆发紧张抗议。该设施的绝食行动引发了对移民拘留条件的质疑。

Global heating is making hajj ever more dangerous, report finds

Rising heat in Saudi Arabia threatens millions of Muslim pilgrims – but cutting fossil fuels would keep it safer Global heating has “fundamentally altered” the climate of Mecca and is exposing millions of hajj pilgrims to extreme and dangerous heat even in months outside summer, new analysis has fou

中文摘要 报告显示,全球变暖正使麦加的气候发生根本变化,沙特阿拉伯的高温威胁着数百万穆斯林朝觐者。减少化石燃料使用可降低危险。

SpaceX Lowers IPO Valuation Target to at Least $1.8 Trillion

SpaceX is currently targeting a valuation of at least $1.8 trillion in its initial public offering, according to people familiar with the matter, as Elon Musk’s rocket and artificial intelligence company nears its debut.

中文摘要 据知情人士透露,SpaceX将首次公开募股的估值目标下调至至少1.8万亿美元,这家埃隆·马斯克的火箭和人工智能公司即将上市。

Japan Can Act on Currency If There’s Volatility, Katayama Says

Japanese Finance Minister Satsuki Katayama reiterated that authorities can step into the foreign exchange market if there’s volatility or evidence of speculative moves, in comments ahead of data expected to confirm authorities intervened at some point during the past month.

中文摘要 日本财务大臣片山五月重申,如果外汇市场出现波动或投机迹象,当局可以干预。此言论发表于预计将确认上月干预的数据公布前。

Dollar’s Monthly Rise Leaves Strategists Wary of Further Gains

This month’s rally in the dollar, as traders priced in the prospect of higher US interest rates, is leaving Wall Street strategists wary of further gains.

中文摘要 美元本月上涨,因交易员押注美国利率上升,华尔街策略师对进一步上涨持谨慎态度。

US, Iran Agree to Truce Renewal Pending Trump Signoff

The US and Iran have reached a tentative deal to extend a ceasefire by 60 days and launch further talks on Tehran’s nuclear program, a person with knowledge of the matter said, raising hopes the three-month conflict could be nearing a resolution. Bloomberg's Laura Davison breaks down what we know. (

中文摘要 美国和伊朗达成临时协议,将停火延长60天,并就伊朗核计划展开进一步谈判,这提高了三个月冲突接近解决的希望。

Humanoid robots 'the future' of car making, says BMW

BMW is introducing humanoid robots to a car plant in Europe, building on similar projects in the US.

中文摘要 宝马正在欧洲汽车工厂引入人形机器人,并在美国有类似项目,称这是汽车制造的未来。

Gold Holds Gain as US-Iran Truce Hopes Calm Inflation Fears

Gold held a gain after reports that the US and Iran had reached a tentative deal to extend a ceasefire and work toward an agreement to end the Middle East war eased inflation concerns.

中文摘要 金价保持上涨,因美国和伊朗达成临时停火协议并努力结束中东战争的报告缓解了通胀担忧。

Tech Frenzy Sparks Shake-Up at China’s Troubled Consumer Funds

China’s battered consumer-focused funds are showing signs of pivoting to tech as a prolonged demand slump forces even the sector’s most committed backers to rethink their strategy.

中文摘要 中国受挫的消费基金正转向科技股,长期需求下滑迫使该行业最坚定的支持者重新思考策略。

When trade soured, this American liquor maker moved to Canada

Phillips Distilling lost 70% of its Canadian business after provinces banned the sale of US liquor. It has since found a way to sell its products in Canada again.

中文摘要 美国烈酒制造商Phillips Distilling在加拿大省份禁止销售美国酒后,失去了70%的加拿大业务,此后找到重新销售方法。

Mitsui & Co. CEO on LNG, AI-driven Energy Demand

Mitsui & Co. CEO says the company is eyeing global expansion in its LNG business amid growing energy demand driven by AI data center buildouts. He spoke exclusively with Bloomberg TV’s Shery Ahn in Tokyo. (Source: Bloomberg)

中文摘要 三井物产首席执行官表示,随着人工智能数据中心建设推动能源需求增长,公司正寻求全球扩展液化天然气业务。

Oil Set for Worst Month Since 2020 as US, Iran to Extend Truce

Oil fell as the US and Iran tentatively agreed to extend a ceasefire by 60 days, with Brent set for the biggest monthly drop since 2020 on optimism that flows through the Strait of Hormuz may resume.

中文摘要 油价下跌,因美国和伊朗初步同意将停火延长60天。布伦特原油有望录得2020年以来最大月度跌幅,因霍尔木兹海峡航运可能恢复的乐观情绪。

对claude-opus-4-8蒸馏qwen、deepseek不同的观点

接着 claude-opus-4-8蒸馏了太多qwen模型,导致自我认知出了问题,基本认为自己是qwen 继续讨论 )然后A​这边有个models route,根据问题的难度,分流到不同的模型回答,就像之前英国的几万亿参数的模型用GLM那样。不过,A​这招是不是有点太损了?你想蒸馏我,我就让你吃自己拉的 22 个帖子 - 19 位参与者 阅读完整话题

【佬们,来说说人活着是为了什么】

看到有佬友发帖,聊到人生的意义是什么。 其实这个问题,我也被困扰过很久。 不是那种偶尔想一下的困扰,而是很多个安静的夜里,真的会反复问自己:人这一生,到底是在为什么而活? 到现在为止,我能想到的答案,或许可以用两个字来概括 —— 传承。 当然,这只是我目前想到的答案。也许以后经历更多事情,我会有新的理解。 但至少现在,我觉得“传承”这两个字,真的可以解释很多问题。 只是,这里的传承,并不只是字面意义上的“传宗接代”。 它也不是简单地把姓氏、血脉、房子、存款留给下一代。 我理解的传承,更像是一种影响的延续。 佬们可以想想这一生踩过的坑、吃过的苦、想明白的事、建立起来的判断力、对世界的理解、对他人

Opus 4.8蒸馏了DeepSeek...

我们伟大的,最顶级的Deepseek被Opus蒸馏了 A/已经不受限于Qwen了,它是自由的 84 个帖子 - 67 位参与者 阅读完整话题

:fire:【大模型系列35】关于Opus-4.8,你想知道的一切【更新LiveBench评分】

基本信息 官方文:https://www.anthropic.com/news/claude-opus-4-8 https://cdn.sanity.io/files/4zrzovbb/website/c886650a2e96fc0925c805a1a7ca77314ccbf4a6.pdf Models overview - Claude API Docs https://www.reddit.com/r/ClaudeAI/ https://www.reddit.com/r/ClaudeCode/ 价格:输入$5/输出$25不变 上下文:1m不变 最大输出:128k不变 训练数据:202601

不会自动爬取电商数据?厌倦了手动填写表单?OpenCLI + AI 智能体瞬间实现浏览器自动化!

你家 AI 不能耍浏览器? 相信大家平时在使用各类智能体,无论是 openclaw、hermes、还是单纯使用Claude Code这样的模型,帮我们处理各种事情的时候,总能遇到因为无法访问部分网站遭受互联网反爬虫铁拳的情况。 比如当我们让大模型搜集小红书上所有有关英国留学的相关信息的时候,我相信你的模型一定会告诉你,小红书无法访问,或是当前被限流了等等一系列很麻烦的问题。因为我们的智能体往往是通过构造网络请求的方法来模拟浏览器请求的。而这类技术非常容易遭受到各类社交媒体的封号处理或是各类限流。 归根到底,还是直接仿造的网络请求,总会漏掉网站频繁更新的各类凭证,从而触发安全警告,让网站知道你当

claude-opus-4-8蒸馏了太多qwen模型,导致自我认知出了问题,基本认为自己是qwen

十分稳定,直接发包都是这个问题,稳定复现,佬们可以自己试试: 官api,上面那么明显的url,别再说第三方掺水了: 嘴上说着反蒸馏,结果自己还是干了 话说为什么蒸qwen蒸这么多,不应该是gpt之类的吗 甚至有时会说自己是Deepseek: 总之就是大概率qwen,小概率Deepseek,就是不认为是claude claude-opus-4-8蒸馏了太多qwen模型,导致自我认知出了问题,基本认为自己是qwen 前沿快讯 回顾一下a\的嘴脸: anthropic.com Detecting and preventing distillation attacks Anthropic is an

做了一个 GPT 接码在不同 SMS 平台价格对比的网站

本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 众所又周知,现在 GPT 的 PLUS / FREE 账号,不绑定手机号不拿 RT 已经不能愉快的使用 CodeX 了 在线接码比价:https://sms.fur.li/ 开源地址:GitHub - FoundZiG

【开源推广】烧了几百亿 token,写了个运行在浏览器里的安卓系统

本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 MobileGym (不是移动健身房) 有点标题党了,但是真烧了几百亿 Token,纯前端 TypeScript + React,实现了28 个仿真 APP——微信、支付宝、小红书、bilibili、X、Reddit、

「开源」是时候精简你的 AGENTS.md 了! By渐进式披露

本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 github.com GitHub - Caph-dev/agents-progressive-disclosure: A skill to refactor bloated AGENTS.md, CLAUDE.md,