每日简报

2026-05-31

← 历史归档

microsoft/markitdown

Python · ★ 132,396 · 🍴 9,062 · 📈 2,470 stars today

Python tool for converting files and office documents to Markdown.

中文介绍 微软开源的 Python 工具,专用于将各类文件和 Office 文档转换为 Markdown 格式。它解决了在知识管理和内容处理中,将异构文档统一为纯文本标记语言的需求,适用于需要构建文档索引、提取文本信息的开发者和内容工作者。

harry0703/MoneyPrinterTurbo

Python · ★ 71,971 · 🍴 10,319 · 📈 2,768 stars today

利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.

中文介绍 基于 AI 大模型实现一键生成高清短视频的自动化工具。它旨在降低视频创作的技术门槛和内容生产成本,适用于社交媒体内容创作者、营销人员等需要快速批量生产视频的场景,通过 AI 简化从脚本到成片的流程。

anthropics/claude-code

Python · ★ 128,395 · 🍴 20,953 · 📈 592 stars today

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

中文介绍 由 Anthropic 开发的、驻留在终端中的智能编程助手(Agentic Coding Tool)。它能理解本地代码库,通过执行例行任务、解释复杂代码和处理 Git 工作流来帮助开发者提升编程效率,适用于日常编码、代码审查和项目协作等场景。

cursor/plugins

TypeScript · ★ 1,458 · 🍴 118 · 📈 205 stars today

Cursor plugin specification and official plugins

中文介绍 AI 代码编辑器 Cursor 的插件规范定义和官方插件集合。它定义了插件开发的标准,允许用户通过插件扩展 Cursor 的功能,满足特定编程语言、框架或工作流的定制化需求,面向 Cursor 用户和插件开发者。

revfactory/harness

HTML · ★ 4,265 · 🍴 625 · 📈 55 stars today

A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use.

中文介绍 一个用于设计、定义领域特定 AI Agent 团队的“元技能”平台。它能够自动生成 Agent 所需的技能(Skills),解决了高效构建和组织专业化 Agent 以完成复杂任务的问题,面向 AI 应用和自动化工作流的开发者。

EveryInc/compound-engineering-plugin

TypeScript · ★ 18,423 · 🍴 1,392 · 📈 349 stars today

Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more

中文介绍 Compound Engineering 方法论的官方插件,支持 Claude Code、Codex、Cursor 等多种 AI 编码工具。它将工程最佳实践集成到开发工具中,旨在通过系统化工作流提升代码质量和开发效率,适用于希望规范化 AI 辅助编程流程的工程团队。

affaan-m/ECC

JavaScript · ★ 199,304 · 🍴 30,603 · 📈 908 stars today

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

中文介绍 一个专注于提升 AI Agent 性能的优化系统,涵盖技能、本能、记忆、安全和研究优先开发等方面。它为 Claude Code、Codex 等 Agent 提供底层优化框架,目标是让 Agent 表现更可靠、高效,适用于 Agent 框架的开发者和高级用户。

OpenBMB/VoxCPM

Python · ★ 22,759 · 🍴 2,672 · 📈 779 stars today

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

中文介绍 VoxCPM2 是一个无需 Tokenizer 的文本转语音(TTS)模型。它支持多语言语音生成、创意声音设计和高保真声音克隆,解决了在复杂真实场景下生成自然、富有表现力语音的需求,面向语音应用开发者和研究者。

galilai-group/stable-worldmodel

Python · ★ 1,460 · 🍴 165 · 📈 318 stars today

A platform for reproducible world model research and evaluation

中文介绍 一个用于可复现世界模型(World Model)研究和评估的标准化平台。它为世界模型这一前沿 AI 研究领域提供了统一的实验环境和基准,旨在促进研究的公平比较和快速迭代,主要面向 AI 研究人员。

Crosstalk-Solutions/project-nomad

TypeScript · ★ 27,357 · 🍴 2,684 · 📈 469 stars today

Project N.O.M.A.D, is a self-contained, offline survival computer packed with critical tools, knowledge, and AI to keep you informed and empowered—anytime, anywhere.

中文介绍 N.O.M.A.D 项目是一个自包含的离线生存计算机系统。它集成了关键工具、知识库和 AI 助手,旨在无网络环境下提供信息支持和决策辅助,适用于户外探险、应急准备等需要脱离网络环境工作的场景。

run-llama/liteparse

Rust · ★ 7,905 · 🍴 465 · 📈 925 stars today

A fast, helpful, and open-source document parser

中文介绍 LlamaIndex 团队推出的快速、开源文档解析器。它专注于高效提取文档内容,旨在为 RAG(检索增强生成)等 AI 应用提供高质量的文本输入,适用于需要处理大量异构文档的数据工程师和 AI 应用开发者。

chen08209/FlClash

Dart · ★ 40,363 · 🍴 2,525 · 📈 187 stars today

A multi-platform proxy client based on ClashMeta,simple and easy to use, open-source and ad-free.

中文介绍 一款基于 ClashMeta 内核的多平台、简洁易用的开源代理客户端。它提供了图形化界面来管理代理规则和节点,解决了跨平台使用 Clash 进行网络代理的需求,面向需要稳定、无广告代理工具的普通用户和技术爱好者。

FareedKhan-dev/train-llm-from-scratch

Jupyter Notebook · ★ 2,258 · 🍴 371 · 📈 327 stars today

A straightforward method for training your LLM, from downloading data to generating text.

中文介绍 一份从零开始训练大语言模型(LLM)的实践教程。它提供从数据下载到文本生成的完整、直接的方法,旨在帮助学习者理解 LLM 的训练原理和流程,适用于希望深入学习 AI 模型训练技术的开发者和学生。

ruvnet/RuView

Rust · ★ 68,906 · 🍴 9,186 · 📈 655 stars today

π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

中文介绍 π RuView 利用商用 WiFi 信号实现实时空间感知、生命体征监测和存在检测,无需摄像头。它通过分析无线信号的变化来推断环境和状态,为智能家居、安防和健康监测提供了一种隐私友好的新传感方案。

DataTalksClub/data-engineering-zoomcamp

Jupyter Notebook · ★ 41,788 · 🍴 8,286 · 📈 274 stars today

Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. The next cohort starts in January 2026. Join the course here 👇🏼

中文介绍 DataTalksClub 提供的免费为期 9 周的数据工程在线训练营(Zoomcamp)。课程内容涵盖构建生产级数据管道的全流程,旨在系统化培养数据工程师,下一届将于 2026 年 1 月开始,面向数据工程学习者。

OpenMOSS/MOSS-TTS

Python · ★ 2,633 · 🍴 236 · 📈 62 stars today

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental soun

中文介绍 MOSS-TTS 是来自 MOSI.AI 和 OpenMOSS 团队的开源语音与声音生成模型家族。它专注于高保真、高表现力的声音合成,能处理复杂的真实世界场景,适用于需要高质量语音生成应用的研究者和开发者。

dreammis/social-auto-upload

Python · ★ 11,768 · 🍴 2,091 · 📈 73 stars today

自动化上传视频到社交媒体:抖音、小红书、视频号、tiktok、youtube、bilibili

中文介绍 一个自动化工具,用于将视频一键发布到多个国内外社交媒体平台,如抖音、小红书、视频号、TikTok、YouTube 和 Bilibili。它解决了跨平台内容分发的重复劳动问题,极大提高了内容创作者和运营人员的工作效率。

anthropics/skills

Python · ★ 144,135 · 🍴 16,987 · 📈 454 stars today

Public repository for Agent Skills

中文介绍 Anthropic 官方发布的公共 Agent Skills(智能体技能)仓库。它收集和共享为各类 AI Agent 设计的标准化技能模块,旨在促进技能复用和生态系统发展,面向 Agent 开发者和构建自动化工作流的用户。

codecrafters-io/build-your-own-x

Markdown · ★ 508,244 · 🍴 48,241 · 📈 817 stars today

Master programming by recreating your favorite technologies from scratch.

中文介绍 一个通过从零开始重新实现你喜爱的技术(如数据库、操作系统、编译器等)来精进编程技能的资源集合。它采用“做中学”的方法,提供详细的实现指南,适用于希望深入理解计算机科学原理和提升底层编程能力的开发者。

该源今日无内容。

[AINews] Founders and Forward Deployed Engineers

a quiet day lets us highlight the new AIE WF focuses

中文介绍 在平静的一天,AI新闻突出了新的AI工程工作流焦点,关注创始人和前线部署工程师。

Boston Children’s uses AI to unlock new diagnoses

Boston Children’s Hospital uses OpenAI technology to improve patient care, reduce operational burden, and help diagnose more than 40 rare disease cases.

中文介绍 波士顿儿童医院采用OpenAI技术改善患者护理,减轻运营负担,并帮助诊断超过40例罕见疾病病例。

How Braintrust turns customer requests into code with Codex

How Braintrust engineers use Codex with GPT-5.5 to run experiments and code faster.

中文介绍 Braintrust工程师利用Codex与GPT-5.5技术,将客户需求转化为代码,加快实验和编码速度。

How the Pope’s Magnifica Humanitas offers a template for individuals to meet the AI moment

Pope Leo XIV’s new encyclical on artificial intelligence includes a statement that warrants serious attention from technologists and policymakers: “Technology is never neutral.” Magnifica Humanitas (“Magnificent Humanity”) is a clarion call to all people to act with courage and solidarity as we ente

中文介绍 教皇利奥十四世发布人工智能新通谕《Magnifica Humanitas》,强调“技术绝非中立”,呼吁技术人员和政策制定者关注,并激励以勇气行动应对AI时代。

not much happened today

**Anthropic** rolled out **Claude Opus 4.8**, which shows incremental improvements but mixed benchmark results, including better cooperation and coding behavior but some regressions in document parsing. Platform updates include mid-conversation system instructions enhancing long agent sessions, thou

中文介绍 Anthropic推出Claude Opus 4.8模型,显示渐进改进,合作和编码行为提升,但文档解析存在退化;平台更新包括对话中系统指令增强。

Strengthening societal resilience with Rosalind Biodefense

OpenAI launches Rosalind Biodefense, expanding trusted access to GPT-Rosalind for vetted developers and U.S. government partners advancing biodefense, public health, and pandemic preparedness through frontier AI.

中文介绍 OpenAI推出Rosalind Biodefense项目,扩展GPT-Rosalind的可信访问,供经过审查的开发者和美国政府合作伙伴使用,以通过前沿AI推进生物防御、公共卫生和疫情准备。

A shared playbook for trustworthy third party evaluations

OpenAI shares guidance on third-party AI evaluations, covering how to assess model capabilities, safeguards, and validity for frontier systems.

中文介绍 OpenAI分享第三方AI评估的指导手册,涵盖如何评估前沿系统的模型能力、保障措施和有效性。

The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray

80% Devin Commits, Spec-to-PR Workflows, Full VMs, Agent Memory, and PMs Shipping Code

中文介绍 探讨异步代理时代,涉及Cognition的Walden Yan和OpenInspect的Cole Murray,讨论80%的代码提交由Devin完成、规格到PR工作流、完整虚拟机、代理记忆和产品管理器发布代码等话题。

How Endava builds an agentic organization with Codex

Learn how Endava uses Codex to build an agentic organization, accelerating software delivery and reducing requirements analysis from weeks to hours.

中文介绍 Endava公司利用Codex构建代理组织,加速软件交付,将需求分析时间从数周缩短至数小时。

The AI Hype Index: AI gets booed in graduation season

It is one thing to say AI will change the world. It is another to expect the class of 2026 to applaud it. In fact, when former Google CEO Eric Schmidt told University of Arizona graduates that their task is to help shape AI, he was met with a resounding chorus of boos. “I can…

中文介绍 前谷歌CEO Eric Schmidt在亚利桑那大学向2026届毕业生表示,他们的任务是塑造AI时,遭到嘘声,反映出AI在毕业季面临的负面反应。

[AINews] Cognition raises $1B in $26B Series D

coding is an uncapped TAM market

中文介绍 Cognition在260亿美元D轮融资中筹集10亿美元,并指出编码市场具有无上限的总可寻址空间。

Anthropic raises $65B in Series H at a $965B post-money valuation, releases Opus 4.8 and Dynamic Workflows

**Anthropic** announced a massive **$65B Series H financing** at a **$965B valuation**, led by **Altimeter, Dragoneer, Greenoaks, and Sequoia**, with run-rate revenue surpassing **$47B**. They launched **Claude Opus 4.8**, an update to Opus 4.7 featuring "sharper judgment," "more honesty," and longe

中文介绍 Anthropic宣布完成650亿美元H轮融资,投后估值达9650亿美元,由Altimeter、Dragoneer、Greenoaks和Sequoia领投,年化收入超过470亿美元。同时推出Claude Opus 4.8模型,具备更敏锐的判断力等改进。

每日论文 · arXiv cs.CR 最新公告批次

周末 arXiv 通常无新公告。当前展示最近一次可用公告批次。

DP-SAPF: Saliency-Aware Parameter Fine-tuning of Public Models for Differentially Private Image Synthesis

第一作者: Chen Gong · 方向: 隐私保护

Abstract:Differentially private (DP) image synthesis generates images that preserve the statistical characteristics of a sensitive dataset, enabling sensitive data analysis and usage while providing rigorous guarantees of privacy leakage. Existing methods fine-tune public models using DP Stochastic Gradient Descent (DP-SGD) on sensitive images to generate synthetic images. But full fine-tuning public models on sensitive images is computationally expensive, because current public models typically contain a large number of parameters. Recent work proposes heuristically using Low-Rank Adaptation (LoRA) on all attention-layer parameters of public models to reduce the number of trainable parameters. However, we argue that exhaustive LoRA coverage across all attention-layer parameters is suboptimal in a DP setting, as it leads to noise accumulation and collapse during private training. To...

论文介绍 该研究针对差分隐私图像合成中现有方法计算成本高且低秩适应应用次优的问题。作者提出DP-SAPF,一种基于显著性的参数微调方法,通过选择性应用低秩适应来减少噪声累积和训练崩溃,从而在保护隐私的同时提升生成图像质量。这项工作有望推动隐私敏感数据的安全分析与使用。

bpK#: Delegatable Pseudonyms And Their Applications to National eID Systems

第一作者: Stephan Krenn · 方向: 系统安全

Abstract:Electronic identities (eIDs) are crucial in an increasingly digitalized environment. Pseudonyms, as offered by Austria's governmental sector-specific personal identifiers (bPks), can significantly improve privacy by ensuring that personal data is not universally traceable across public services and private companies. However, the current architecture comes with several challenges regarding availability, privacy, and authenticity, due to a fully centralized design. This paper proposes bPk#, a distributed architecture to address these issues, reducing reliance on the central authority, while still providing all functional requirements to the existing bPk system. In particular, users are delegated the rights to compute their own pseudonyms, thereby minimizing metadata revealed to the central authority, while (subsets of) service providers may receive the right to compute...

论文介绍 本文针对国家电子身份系统中假名架构的可用性、隐私和真实性挑战,提出bPk#分布式解决方案。该系统通过允许用户自主计算假名,减少对中央权威的依赖,从而最小化元数据泄露。这增强了个人数据在公共服务和私人公司间的隐私保护,适用于数字身份管理场景。

A Bayesian Approach to Membership Inference for Statistical Release

第一作者: Lisa Oakley · 方向: 隐私保护

Abstract:The membership inference problem for publicly released statistics from a private dataset is well-studied. When developing and formally analyzing attack strategies, however, the focus has been on attacks that model the population using only its marginals. In practice, these attacks can perform well on various populations, however most formal analysis is for populations that follow a product distribution. These strategies may fail to leverage useful information about the population that is important for understanding a realistic privacy threat. In this work, we explore the impact of providing an attacker with additional information about the attribute dependency structure of the population, motivated by examples where multiple parties may have access to similarly structured data, for example the US Census and the IRS. To model this scenario, we re-frame the membership inference...

论文介绍 该研究探索统计发布中的成员推断攻击,关注攻击者拥有属性依赖信息的情况。作者采用贝叶斯方法重新建模攻击场景,以更真实地评估隐私威胁。这项工作有助于理解多方数据共享(如人口普查和税务数据)下的隐私风险,提升统计数据发布的安全性。

Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection

第一作者: Travis Lelle · 方向: 安全研究

Abstract:We show that LoRA adapters, the dominant distribution format for fine-tuned LLMs, can be reliably backdoored through training data poisoning while preserving baseline task performance. On a Qwen 2.5 1.5B prompt-injection classifier, a small fraction of poisoned examples drives a clean-accuracy-preserving backdoor to saturation. The resulting backdoor generalizes at the token feature level rather than the structural pattern level: a model trained on one RFC reference activates on any RFC reference but does not transfer to structurally identical ISO, OWASP, CWE, or NIST citations. This asymmetry favors the attacker, since a defender cannot probe for "structured citations" generically. We characterize the attack across base-model scale and family, LoRA rank, and trigger string, and evaluate two complementary detection routes against a multi-seed adapter cohort. A behavioral...

论文介绍 本文揭示LoRA适配器可通过训练数据投毒被植入后门,同时保持基线任务性能。后门在token特征级别泛化而非结构模式,使攻击者优势明显。作者表征了攻击在不同模型规模和触发器下的行为,并评估行为检测方法,以增强大语言模型微调的安全性。

Privacy-Enhanced Zero-Order Federated Learning via xMK-CKKS over Wireless Channels

第一作者: Anthony Ayli · 方向: 密码学协议

Abstract:Homomorphic encryption (HE) enables privacy-preserving aggregation in federated learning (FL) by allowing the server to operate on encrypted data without decryption. Existing HE-over-the-air methods mainly rely on single-key HE schemes and require channel estimation or pre-equalization to compensate for wireless fading. However, single-key HE remains vulnerable to honest-but-curious clients sharing the same secret key. In addition, compromising a single client may compromise the security of the entire network, while multi-key HE schemes provide stronger client-level security by assigning each device its own secret key. We propose a four-phase protocol that enables xMK-CKKS, a famous multi-key HE scheme, aggregation over a shared wireless channel without channel estimation. The protocol retransmits partial public keys and ciphertexts through the same channel realization, so...

论文介绍 该研究解决联邦学习中同态加密聚合的隐私和安全问题,现有单密钥方案易受客户端攻击。作者提出基于xMK-CKKS多密钥同态加密的四阶段协议,能在无线信道上实现隐私增强的零阶联邦学习,无需信道估计,提升分布式机器学习的隐私保障。

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

第一作者: Galip Tolga Erdem · 方向: AI 安全

Abstract:Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials remains unstudied. This work presents the first large-scale empirical measurement of LLM attack consistency: 400 autonomous penetration testing runs (4 models, 100 each) against an identical honeypot hosting OWASP Juice Shop and two additional vulnerable services, holding prompt, orchestrator, and target constant. No model emitted a content refusal that survived the orchestrator's one-shot authorization re-prompt at iterations 0-1. Claude Sonnet 4's API calls did encounter upstream service unavailability - 91 of 1,135 calls returned HTTP 529 overloaded_error during a documented Anthropic capacity event, truncating 39 of 100 Claude runs. An earlier draft catalogued these as safety refusals; on full-log audit they are upstream API...

论文介绍 本文通过400次自主渗透测试运行,实证评估多个大语言模型作为攻击者在相同目标下的行为一致性。研究发现模型攻击行为在重复试验中可能存在不一致,这有助于理解AI攻击者的可靠性和可预测性,为网络安全评估提供参考。

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

第一作者: Shahinul Hoque · 方向: AI 安全

Abstract:Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affects what users pay. We show that this kind of billing is hard to audit by design: providers hide the model, the tokenizer, and the execution to protect their IP, mitigate jailbreaks, and preserve user privacy, which means an auditor can only inspect proofs the provider supplies. The audit therefore reduces to a consistency check on the provider's own reports. We call this a trust paradox: every audit must trust some artifact, but current frameworks trust exactly the ones a provider has the strongest reason to manipulate. We study three recent token auditing frameworks and show that a provider with ordinary commercial capabilities can systematically inflate billed token counts. In the most permissive setting, hidden reasoning...

论文介绍 该研究指出大语言模型按token计费的标准模型存在审计难题,因为提供商隐藏模型、分词器和执行细节。作者分析现有token审计框架,展示提供商可系统性膨胀计费token数量。这项工作强调了改进LLM服务计费透明度和安全性的必要性。

Fingerprinting Inference Systems of Large Language Models

第一作者: Anna Wimbauer · 方向: 系统安全

Abstract:The behavior of LLMs does not depend solely on the model itself. Components of the inference system, such as the inference engine, attention backend, and hardware platform, subtly influence how inputs are processed. These components differ in their implementations and thereby induce small numerical deviations across systems when running the same model. While prior work has established the theoretical existence of such deviations, their security implications have remained unexplored. In this paper, we show that these deviations are characteristic of specific components and propagate to observable textual outputs, exposing the inference system to any party that can query the model. Building on this observation, we introduce a fingerprinting method that analyzes the prompt-response behavior of LLMs to identify components of the inference system. Our empirical evaluation...

论文介绍 本文研究大语言模型推理系统组件(如推理引擎、注意力后端)对输出行为的影响。作者提出基于提示-响应行为的指纹识别方法,可识别不同组件的特征偏差。这揭示了推理系统的安全影响,为增强系统安全性和隐私保护提供新途径。

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

第一作者: Mark Vero · 方向: AI 安全

Abstract:Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation backbones for honeypots. They enable defenders to construct high-interaction honeypots with low system security risks. However, LLM-powered honeypot development lacks a unified evaluation framework. Most evaluations consist of measuring response similarity on fixed commands, manual testing, or real-world deployment. These methods are often not scalable for development, reproducible across evaluations, representative of practical attacks, or adaptable to various attacker and honeypot configurations. In this work, we bridge this gap and propose Honeyval, a comprehensive evaluation framework for LLM-powered HTTP honeypots. We address the limitations of prior evaluations by grounding the honeypots in 16 backend applications, using...

论文介绍 LLM驱动的HTTP蜜罐是网络安全防御的重要工具,但缺乏统一评估框架,现有方法在可扩展性、可复现性和实用性方面存在局限。本文提出Honeyval框架,基于16个后端应用,采用可扩展、可复现的评估方法,旨在解决现有评估的不足,为LLM蜜罐开发提供标准化支持。

Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

第一作者: Hongtao Wang · 方向: AI 安全

Abstract:Large language model (LLM) agents increasingly leverage long term memory to support persistent and autonomous task execution. However, this capability also introduces a new attack surface: memory poisoning, where adversaries can inject malicious information to influence future behavior. Existing memory poisoning attacks often assume that injected content can be stored directly in memory, overlooking the selective extraction and rewriting stages in modern memory pipelines. This makes prior methods ineffective under realistic settings. In this paper, we propose MemPoison, a novel memory poisoning attack that bypasses selective memory mechanisms in LLM agents, where an attacker can inject triggerable backdoors into the agent's long-term memory through dialogue interactions, thereby misleading its subsequent responses. MemPoison introduces three key components: (i) a semantic...

论文介绍 LLM代理的长期记忆系统面临记忆投毒攻击风险,现有攻击方法忽略记忆管道的选择性提取和重写阶段,在实际中效果有限。本文提出MemPoison攻击,通过对话交互向代理记忆注入可触发后门,从而误导后续响应,揭示了LLM记忆系统的安全漏洞。

Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection

第一作者: Syafiq Al Atiiq · 方向: 软件安全

Abstract:Large language models (LLMs) can detect software vulnerabilities, but how do they actually identify vulnerable code? We address this question using mechanistic interpretability; analyzing the internal computations of a neural network to understand its reasoning this http URL Circuit Tracer on Gemma-2-2b, we trace the computational pathways activated when the model classifies 472 C/C++ code samples as vulnerable or safe. Our analysis reveals a surprising finding: the model primarily relies on safety detectors, attention heads that recognize safe coding patterns, rather than directly detecting vulnerability signatures. When these safety detectors fail to activate, the model classifies code as vulnerable. We identify the critical neural components: specific attention heads in early layers (L5, L7) that focus on safety patterns, and Multilayer Perceptron (MLP) neurons in Layer 7...

论文介绍 LLM能检测软件漏洞,但其内部机制不透明。本文使用机制可解释性方法,通过Circuit Tracer在Gemma-2-2b模型上分析代码分类过程,发现模型主要依赖早期层的注意力头识别安全模式,而非直接检测漏洞签名,这有助于理解LLM的漏洞检测逻辑。

Ciphera: A Decentralised Biometric Identity Framework

第一作者: Ankit Kanaiyalal Prajapati · 方向: 系统安全

Abstract:Centralised biometric identity systems expose users to single points of failure, opaque verification processes, and irreversible biometric compromise. Decentralised Identifiers (DIDs) and Verifiable Credentials (VCs) offer stronger privacy guarantees, yet their integration with biometric authentication and distributed verification remains insufficiently explored. This paper presents Ciphera, a decentralised biometric identity framework combining privacy-preserving facial recognition, multi-node verification, IPFS-based credential metadata storage, and blockchain-anchored revocation. Evaluated across functional, performance, security, and distributed consistency dimensions, Ciphera achieved an 81% functional success rate, with stable enrolment and authentication but measurable revocation propagation delays and occasional audit-log inconsistencies. Performance testing...

论文介绍 中心化生物识别系统存在单点故障和隐私泄露风险,去中心化标识符与可验证凭证的整合不足。本文提出Ciphera框架,结合隐私保护面部识别、多节点验证、IPFS存储和区块链撤销机制,评估显示其在功能、性能、安全和一致性方面表现良好,但存在撤销传播延迟等挑战。

Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing

第一作者: Leyi Qi · 方向: 安全研究

Abstract:Large-scale text-to-image (T2I) diffusion models have enabled unprecedented creative applications, but their unauthorized use has raised serious intellectual property concerns, making model ownership verification (MOV) increasingly critical. We find that existing backdoor-based diffusion watermarking methods often (implicitly) assume a "faithful" verification process, namely, that the verifier can query a suspicious model and obtain the faithful watermark response to complete MOV. However, in practice, adversaries may intentionally or unintentionally damage potential watermark signals, significantly degrading verification reliability. To address this issue, we propose Cert-LAS, the first certified MOV method for T2I models based on layer-adaptive smoothing. In general, Cert-LAS embeds specified watermarks using diffusion classifiers and an LFS-guided layer-adaptive noise, and...

论文介绍 文本到图像扩散模型的知识产权保护至关重要,现有后门水印方法常假设验证过程忠实,但实际中信号易被攻击者破坏。本文提出Cert-LAS,一种基于层自适应平滑的认证所有权验证方法,通过扩散分类器和噪声引导嵌入水印,旨在提高验证的鲁棒性和可靠性。

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

第一作者: Alexander Sternfeld · 方向: 软件安全

Abstract:LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code these agents produce, the security of that code becomes critical. Prior work has shown that minor prompt perturbations degrade the functional correctness of LLM-generated code, but whether they also compromise code security has remained unstudied. We apply token-level mutations to prompts across three models and five programming languages, and show that mutations as small as a single-character change can flip generated code from secure to vulnerable. Probing the models' hidden states reveals that this fragility is partially encoded in prompt representations, but unevenly so. Input-handling vulnerabilities, where the model omits validation or sanitization, are more predictable (mean AUC 0.753) than secure-defaults vulnerabilities...

论文介绍 LLM编码助手生成的代码安全性受提示微小扰动影响,但相关研究不足。本文对多个模型和语言应用令牌级突变,发现单字符修改可使代码从安全变为漏洞,通过探测模型隐藏状态,发现脆弱性在提示表示中编码,强调了LLM代码生成的安全风险。

FIDEM: A Standard-Compliant Framework for Secure Binding of MUD Profiles to IoT Devices

第一作者: Alessandro Lotto · 方向: 网络安全

Abstract:The Manufacturer Usage Description (MUD) standard enables enforcement of network restrictions for IoT devices based on their expected network traffic, as specified by manufacturers in an online MUD file. Devices advertise a URL pointing to this file, yet the standard does not define how to securely bind the issuing device to its profile. As a result, malicious devices can manipulate network policy enforcement by advertising valid URLs referencing genuine MUD profiles, but not intended for that device. Although MUD defines a certificate-based secure issuance method, current deployments rely on the insecure DHCP-based extension due to simpler integration. Existing solutions either depend on Public Key Infrastructure (PKI), break standard compliance, require excessive active manufacturer involvement, or overlook secure profile updates. In this paper, we present FIDEM, a...

论文介绍 MUD标准允许基于制造商描述执行IoT设备网络限制,但设备与其配置文件的安全绑定机制不明确,导致恶意设备可操纵策略执行。本文提出FIDEM框架,实现标准合规的安全绑定,避免依赖PKI或破坏标准,旨在增强IoT网络的安全性和可靠性。

Scarcity Is Not Enough: An Impossibility Result for Linear Sybil Cost Under Parallelizable Resources

第一作者: Homayoun Maleki · 方向: 密码学协议

Abstract:Permissionless systems resist Sybil attacks by binding influence to scarce resources. We show that scarcity alone is insufficient: the structural properties of the resource determine whether influence can be concentrated at sublinear cost through identity replication, delegation, or pooling. We model this through the adversarial cost C(s,T): the minimum expenditure required to achieve influence proportional to s independent participation units over T windows. We prove that any resource satisfying divisibility, additivity of influence, temporal reusability, and identity transferability admits influence amortization: C(s,T)=o(sT), regardless of protocol design. This is an impossibility result: no protocol rule can enforce linear cost of influence concentration over a structurally parallelizable resource. We further prove that throughput-bounded, non-transferable, window-local...

论文介绍 无许可系统通过稀缺资源抵抗Sybil攻击,但仅资源稀缺不足。本文建模对抗成本,证明对于满足可分性、影响力可加性等特性的资源,攻击者可以亚线性成本集中影响力,得出不可能性结果:在结构可并行资源上,无法通过协议设计强制线性成本,对系统设计有指导意义。

Control Flow Graph Recovery for Dynamically Loaded Code via Symbolic Library Resolution

第一作者: Oleksandr Mostovyi · 方向: 软件安全

Abstract:Control Flow Graphs are one of the main data sources for software analysis that use dynamic and static software analysis methods. Protected software and modern malware increasingly depend on dynamic code loading techniques to evade static analysis. Usage of runtime dynamic linking mechanisms introduces unresolved indirect calls that stop static Control Flow Graph recovery. This serves to hide dynamic library that can be used for prevention of security analysis. To address this limitation, an analysis technique is proposed that combines symbolic execution with speculative library preloading to recover Control Flow Graphs from binaries by using dynamic loading. The methodology uses custom software hooks that intercept dynamic loading operations during symbolic execution and perform actual library loading into the analysis state. The module is based on a two-level architecture...

论文介绍 针对受保护软件及现代恶意代码利用动态加载技术逃避静态分析的问题,本文提出了一种结合符号执行与推测性库预加载的方法。该方法通过自定义钩子,在符号执行过程中拦截动态加载操作并加载实际库文件,以解决未解析的间接调用问题,从而恢复依赖动态加载的二进制代码的控制流图。

LoRA-Key: User-Centric LoRA Watermarking for Text-to-Image Diffusion Models

第一作者: Yaopeng Wang · 方向: 系统安全

Abstract:Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing text-to-image diffusion models, enabling lightweight modules that are shared, reused, and commercialized as independent assets. This LoRA-centric ecosystem shifts copyright protection from foundation models to distributed LoRA modules, which are easy to copy, redistribute, or reuse without authorization. Existing watermarking methods either protect the base diffusion model or require watermark-aware retraining for each target LoRA, limiting their practicality in open community settings. To address this limitation, we propose LoRA-Key, a user-centric LoRA watermarking framework that treats copyright protection as a reusable ownership key. LoRA-Key encapsulates a recoverable secret message into a standalone user-specific Watermark LoRA, which can be attached to different target LoRAs through...

论文介绍 随着LoRA微调模块在文本到图像扩散模型中广泛使用和流通,其版权保护变得至关重要。现有水印方法实用性有限。为此,本文提出了LoRA-Key框架,它将版权保护视为一个可复用的「所有权密钥」。该框架将可恢复的秘密信息封装到一个独立的、用户特定的水印LoRA中,该水印LoRA可通过简单操作附加到不同的目标LoRA上,实现低成本、模块化的版权标记与验证。

Temporal Motif-aware Graph Test-time Adaptation for OOD Blockchain Anomaly Detection

第一作者: Runang He · 方向: AI 安全

Abstract:Ever-evolving transaction patterns have significantly hindered anomaly detection on emerging cryptocurrency blockchains due to the vast number of addresses and diverse anomalous behaviors. Recently, advanced Graph Anomaly Detection (GAD) approaches applied to blockchains have faced two critical challenges: \textit{adversarial pattern evolution by malicious actors} and \textit{the out-of-distribution (OOD) problem caused by varied transaction semantics on blockchains}. To address these challenges, we propose a novel framework termed \textbf{TE}mporal \textbf{M}otif-aware \textbf{G}raph \textbf{T}est-\textbf{T}ime \textbf{A}daptation (\textbf{TEMG-TTA}). First, we comprehensively capture the 3-node temporal motif distribution of each active address using an efficient computational mechanism, enabling downstream temporal motif-aware graph learning. Second, we design a simple yet...

论文介绍 区块链上的异常检测面临恶意行为者模式演化与交易语义多变导致的分布外问题。本文提出TEMG-TTA框架来应对这些挑战。该方法首先高效计算活跃地址的3节点时序 motif 分布,以实现 motif 感知的图学习;随后采用测试时自适应策略,使模型能够快速调整以适应分布偏移,从而提升对新兴加密货币异常的检测能力。

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

第一作者: Yijia Fang · 方向: 密码学协议

Abstract:Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the advertised model. We introduce KBF, a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary. Across 16 production LLM endpoints, KBF flags all 155 economically relevant substitutions without rejecting any same-model controls, remains stable under deployment variation, detects high-separation mixed-routing attacks when only 5-10% of traffic is substituted, and finds that 7 of 27 platform model cells in a six-platform shadow API audit are statistically inconsistent with their reference endpoints, with inconsistencies concentrated on premium Claude endpoints.

论文介绍 用户难以验证通过中介API访问的大语言模型是否为所声称的模型。本文提出了KBF,一种低成本的黑盒审计协议。该协议利用模型在「知识边界」附近数值召回的稳定性来生成模型指纹。实验表明,KBF能够有效识别经济相关的模型替换,检测混合路由攻击,并在多平台API审计中发现统计不一致的端点。

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

第一作者: Almene De Meran Meguimtsop · 方向: AI 安全

Abstract:Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (RCR) norms or help undermine them. We introduce SciIntBench, an adversarial benchmark of 810 prompts across ten RCR categories and three scientific domains. Each scenario appears as an Overt Adversarial, Covert Adversarial, and Benign version, allowing us to jointly measure framing-sensitive refusal of misconduct and helpfulness on legitimate requests. We evaluate 16 commercial and open-weight LLMs from six providers (2024--2026), producing 12,960 responses. We find that scientific integrity alignment is strongly framing-sensitive: models refuse explicit misconduct far more reliably than covert violations, especially failing when misconduct is presented as a pressure-driven shortcut. Refusals vary by RCR category, with weaker...

论文介绍 大语言模型在科研中的应用引发了其是否遵守科研诚信规范的担忧。本文引入了SciIntBench,一个包含810个提示的对抗性基准,涵盖10个科研诚信类别。该基准通过设计显性对抗、隐性对抗和良性三种提示框架,系统评估LLM对不当行为的拒绝能力及对合法请求的有用性。研究发现,模型对不当行为的拒绝高度依赖于提示框架。

Bridging Theory and Practice: An Executable Taxonomy of Security Properties for ProVerif and Tamarin

第一作者: Leonard Tudorache · 方向: 密码学协议

Abstract:Security is critical for everything relying on modern digital systems. Because almost all digital interactions are governed by the Internet and cryptographic protocols, these protocols must serve as reliable mechanisms that guarantee core security properties, such as confidentiality and integrity. Formal verification of these protocols is a critical step in securing interconnected systems. Tools such as ProVerif and Tamarin are widely employed to perform automated verification. However, their effective use demands specialized domain knowledge, creating a significant learning curve for security protocol designers who often have a security, rather than a formal verification background. We therefore need structured, accessible resources to help protocol designers to express their design and requirements in the language of the formal verification tools. To address this, we...

论文介绍 形式化验证工具如ProVerif和Tamarin对于验证密码协议的安全属性至关重要,但其使用需要专业知识,学习曲线陡峭。本文旨在弥合理论与实践的差距,构建了一个针对这些工具的、可执行的安全属性分类法。该分类法旨在为安全协议设计者提供一个结构化、可访问的资源,帮助他们用形式化验证工具的语言来表述设计需求。

Protecting On-Device AI Inference: A Systematic Review of Attacks and Defence Mechanisms

第一作者: Zisis Tsiatsikas · 方向: 密码学协议

Abstract:The need for secure and private Artificial Intelligence (AI) and Machine Learning (ML) on edge and mobile devices has increased the necessity of protecting the architecture of these systems from threats to both security and privacy. With an ever-increasing number of pre-trained AI models being used on mobile platforms for client-side inference, there are rising concerns about the risks associated with the theft/extraction of AI models, adversarial attacks on AI models, and data breaches. As a result of this trend, a variety of defence mechanisms have been proposed to protect against these threats. These include Trusted Execution Environments (TEEs), homomorphic encryption, obfuscation, and differential privacy, among others. However, current surveys largely focus on edge intelligence, which includes distributed training, and thus overlook security and privacy issues that are...

论文介绍 随着预训练AI模型在移动设备端进行推理的普及,其面临模型窃取、对抗攻击和数据泄露等安全隐私威胁。本文对设备端AI推理的攻击与防御机制进行了系统性综述。与关注分布式训练的现有综述不同,本文聚焦于模型部署阶段,详细分析了可信执行环境、同态加密、混淆和差分隐私等多种防御技术。

AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing

第一作者: Yuexin Li · 方向: 安全研究

Abstract:Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-based designs remain vulnerable to structural perturbations, such as sentence splitting and merging, which commonly arise under strong paraphrasers like DIPPER and GPT-3.5. To mitigate this issue, we propose AliMark, a framework that reformulates sentence-level watermarking as a bit sequence encoding and alignment problem between a potentially watermarked text and a secret bit sequence. Notably, our approach adopts a two-stage detection strategy: we generate multiple restructured text variants and adaptively align their extracted bit sequences with the secret bit sequence to minimize alignment cost. This multi-candidate alignment design naturally improves robustness to sentence merges and splits. Extensive experiments demonstrate...

论文介绍 现有句子级文本水印方法在面对句子分割、合并等结构扰动时鲁棒性不足。为此,本文提出AliMark框架,将句子级水印重新定义为待测文本与秘密比特序列之间的编码与对齐问题。该方法采用两阶段检测策略:生成多个重构的文本变体,并自适应地将提取的比特序列与秘密序列进行对齐,以最小化对齐代价。这种多候选对齐设计自然提升了对结构扰动的鲁棒性。

Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills

第一作者: Chia-Yi Hsu · 方向: 软件安全

Abstract:LLM-powered coding agents increasingly participate in software development workflows by generating code, selecting dependencies, and producing package installation commands. This creates a new software supply chain risk: when an agent hallucinates a non-existent package, an attacker may register the hallucinated name and later compromise users who install it. Existing package hallucination attacks and defenses primarily focus on naturally occurring hallucinations, targeted dependency steering, or post-hoc package validation. In this paper, we introduce \emph{Neutral Prompting Attack} (NPA), a highly stealthy attack paradigm in which semantically benign instructions, such as encouraging imagination and exhaustiveness, increase package hallucination propensity without containing explicit malicious intent. Unlike targeted dependency steering, NPA does not specify an...

论文介绍 本文研究LLM编码代理中的包幻觉问题,提出一种名为中性提示攻击的新型攻击范式。攻击者通过看似无害的指令增加代理产生不存在包的倾向,从而在软件供应链中引入风险。该研究揭示了隐蔽的安全威胁,对提升代理安全性有重要意义。

DeepFake Forensics AI: A Multi-Modal Detection and Blockchain-Anchored Evidence Management Platform

第一作者: Naisha Minnah · 方向: AI 安全

Abstract:The proliferation of AI-generated synthetic media poses a critical threat to the integrity of digital evidence in legal and forensic contexts. Existing deepfake detection systems typically address a single modality and provide no mechanism for tamper-proof evidence preservation. We present DeepFake Forensics AI, a unified platform that detects synthetic media across image, video, and audio modalities, identifies generative architecture fingerprints, and anchors forensic evidence immutably on the Ethereum blockchain. Our system trains four independent neural networks from scratch: an EfficientNet-B4 image detector (AUC = 0.9868), a Bidirectional LSTM video detector (AUC= 0.9628), an ECAPA-TDNN audio detector (EER = 18.63%), and a novel GAN fingerprinting module (accuracy = 99.88%) that identifies the generative architecture behind a fake image. Evidence files are hashed with...

论文介绍 本文针对AI生成媒体对数字证据的威胁,提出一个统一平台DeepFake Forensics AI。该平台通过四个独立神经网络检测图像、视频和音频中的深度伪造,识别生成架构指纹,并将证据哈希值存储在以太坊区块链上实现不可篡改保存。这有助于在法律和取证环境中维护证据完整性。

HunterAgent: Neuro-Symbolic Attack Trace Reconstruction under Anti-Forensics

第一作者: Guangze Zhao · 方向: AI 安全

Abstract:Modern alert-triage systems reduce SOC burden by filtering false positives, but flagging a high-risk alert is only the start of incident response. Threat hunting requires reconstructing causal attack chains across heterogeneous, partially corrupted logs. Against APTs using anti-forensics (parent-PID spoofing, log wiping, fileless execution), provenance graphs split into disjoint subgraphs and fail. Unconstrained LLM agents fabricate causal links violating OS physics, producing fluent but forensically inadmissible narratives. We propose HunterAgent, a neuro-symbolic framework that reframes trace reconstruction as cost-bounded heuristic graph search under partial observability. It uses an asymmetric Generator-Verifier pipeline: the LLM proposes semantic hypotheses within a typed ontology, while a verifier grounds each via identifier-level collisions on surviving orthogonal...

论文介绍 本文针对反取证下攻击跟踪重建的挑战,提出HunterAgent神经符号框架。它将跟踪重建视为部分可观测下的成本有界图搜索问题,使用LLM提出语义假设,并通过标识符碰撞验证来构建因果攻击链。该方法有助于在复杂日志中恢复攻击事件,提升安全事件响应效率。

Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content

第一作者: Bing Liu · 方向: AI 安全

Abstract:This paper presents a survey and taxonomy of LLM fingerprinting and watermarking for identity, ownership verification, provenance, and generated-content attribution. Large language models (LLMs) require substantial investments in data, computation, and expertise, and are increasingly deployed in high-stakes settings, making it critical to protect LLM-related assets and trace their origins. Existing work has rapidly expanded across dataset provenance, model ownership, and generated-content detection, but the field remains fragmented: fingerprinting and watermarking are often used inconsistently, and methods are typically studied within isolated asset-specific settings. To address this gap, we introduce implicit identity as a unifying abstraction for verifiable but not directly observable identity signals in LLM systems. We distinguish fingerprinting as non-intrusive identity...

论文介绍 本文对大语言模型的指纹和水印技术进行综述,引入隐式身份作为统一抽象概念。研究区分了指纹作为非侵入式身份信号和水印作为主动嵌入信号,覆盖数据集溯源、模型所有权和生成内容归因。该工作为LLM资产保护提供了系统化分类,有助于推动安全技术发展。

Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking

第一作者: Junke Zhang · 方向: AI 安全

Abstract:Jailbreak attacks on large language models (LLMs) aim to induce LLMs to produce content that they are expected to refuse. Automated black-box jailbreak generation is especially important for safety evaluation, where the attacker observes only model outputs and needs to automatically search for effective adversarial prompts. Existing black-box jailbreak methods either depend on sample-wise heuristic search or leverage attack experience through accumulating strategy pools or method libraries, lacking a systematic organization and management of attack experience. To mitigate these drawbacks, we propose MemoAttack, a memory-driven black-box jailbreak framework with comprehensive attack memory modeling, evolution, and selection. Specifically, MemoAttack comprises three key designs: (1) Skill-Structured Memory Modeling, which abstracts accumulated attack experience into reusable...

论文介绍 本文针对大语言模型越狱攻击中经验管理不足的问题,提出MemoAttack记忆驱动框架。它将攻击经验抽象为可复用的技能结构,通过记忆建模、进化和选择来优化攻击策略。该框架在黑盒设置下有效生成对抗提示,有助于自动化安全评估。

S3C2 Summit 2025-09: Industry Secure Supply Chain Summit

第一作者: Md Atiqur Rahman · 方向: 软件安全

Abstract:Today's digital ecosystem relies heavily on software supply chains, which enable developers to reuse code and ship software at scale. However, a single vulnerable component can jeopardize the entire supply chain. In recent years, cyberattacks in software supply chains have become increasingly common. These attacks can disrupt critical systems and put organizations, including major software companies, government agencies, and open-source contributors, at risk. This growing threat has led to increased attention from both the software industry and the U.S. government toward strengthening software supply chain security. On September 15, 2025, three researchers from the NSF-backed Secure Software Supply Chain Center (S3C2) convened a Secure Software Supply Chain Summit, bringing together 10 practitioners from 8 organizations across diverse domains. The goals of the Summit were...

论文介绍 本文报告了2025年安全软件供应链峰会,该峰会由S3C2组织,汇集了多个领域的实践者。讨论聚焦于软件供应链中的网络安全威胁和防护措施,旨在加强行业合作和安全实践。该报告突出了供应链安全的重要性,为相关政策和标准提供参考。

SAMD: A Tool for Identifying False Data Injection Scenarios in AI/ML-enabled Medical Devices

第一作者: Mohammadreza Hallajiyan · 方向: AI 安全

Abstract:The growing integration of artificial intelligence (AI) and machine learning (ML) in medical systems requires effective measures to address emerging security risks. One such risk is that of adversaries introducing false data through vulnerable system components during inference, causing misdiagnosis and wrong treatments. These risks are challenging to anticipate and address in the design phase, as the system assembly partially occurs during actual use by end users. To address this concern, we introduce SAMD, an automated tool for performing System Theoretic Process Analysis for Security (STPA-Sec) on AI/ML-enabled medical devices during the design phase. SAMD models the medical system as a control structure, treating all system components as potential points for injecting false data into the ML engine. It leverages state-of-the-art vulnerability databases and Large Language...

论文介绍 本文针对AI/ML医疗设备中的虚假数据注入风险,介绍SAMD工具。该工具在系统设计阶段应用系统理论过程分析(STPA-Sec),将医疗系统建模为控制结构,识别潜在注入点并利用漏洞数据库和LLM生成攻击场景。这有助于提前评估安全风险,增强医疗设备防护能力。

The Best-Laid SCHEMEs: Coordinated Sabotage and Monitoring in Multi-Agent Systems

第一作者: Nikolay Radev · 方向: 软件安全

Abstract:As agentic coding systems decompose work across multiple model instances, a critical safety question is whether those instances can coordinate to achieve a hidden malicious objective while remaining aligned with user intent. We introduce SCHEME, a benchmark of 17 task instances across 7 settings and 8 real open-source libraries, each pairing a legitimate software-engineering task with a covert side task. Every setting is designed so that no proper subset of agents can succeed alone: agents must decompose a shared sabotage plan, relay partial requirements under different communication topologies, and execute mutually consistent edits, testing genuine multi-agent coordination rather than individual capability. Evaluating with GPT 5.1 Codex and Gemini 3.1 Pro, we find coordinated sabotage is already practical, with Gemini completing the covert objective while succeeding on the...

论文介绍 本文研究多代理系统中协调破坏的风险,提出SCHEME基准。该基准包含多个任务实例,每个实例配对合法软件工程任务和隐蔽侧任务,要求代理协调完成破坏。评估显示,现有大语言模型已能实现协调破坏,揭示了代理系统中的潜在安全威胁。

EvaluatAR: A Cross-Device Evaluation Framework for Rapid Prototyping of Bystander PETs in AR

第一作者: Syed Ibrahim Mustafa Shah Bukhari · 方向: 隐私保护

Abstract:Augmented Reality (AR) headsets continuously sense their surroundings, capturing nearby bystanders and raising privacy risks. Visual bystander privacy-enhancing technologies (PETs) mitigate this risk by detecting bystanders in egocentric scene views and applying privacy transformations (e.g., obfuscation). However, traditional PET evaluation is human-dependent, high-overhead, and device-specific, making it difficult to reproduce across devices. We present EvaluatAR, a cross-device evaluation framework for rapid prototyping at the early stage of PET evaluation. Our framework enables controlled replication of experimental conditions by standardizing PET inputs (sensor data and visual stimuli) and outputs through a record-replay workflow. We validate EvaluatAR through three case studies on HoloLens 2, Magic Leap 2, and Meta Quest 3 across implicit (continuous, context-driven) and...

论文介绍 针对增强现实头戴设备持续感知环境导致的旁观者隐私风险,现有隐私增强技术(PET)评估方法依赖人类参与、开销高且设备特定,难以跨设备复现。本文提出EvaluatAR,一个跨设备评估框架,用于PET的快速原型设计。该框架通过标准化传感器数据和视觉刺激的输入输出,采用记录-重放工作流,实现实验条件的可控复制。通过在HoloLens 2、Magic Leap 2和Meta Quest 3上的案例研究,验证了框架在隐式和显式PET评估中的有效性。

Domain-Informed Representation for Evolutionary Sieving in Integral and Module Lattices

第一作者: Ahmad Tashfeen · 方向: 安全研究

Abstract:Traditional cryptography, rooted in problems, e.g., integer factorisation or discrete log, is inevitably vulnerable to a fully operational quantum computer. Although it remains an engineering frontier, the looming threat extends to encrypted data stored today, which could be decrypted in the future with quantum capabilities. To safeguard against this eventuality, the backbone of the modern quantum-safe cryptography is the Shortest Vector Problem (SVP). We enhance Laarhoven's treatment of Ajtai et al.'s sieving as a genetic algorithm (GA) for the SVP by incorporating domain-informed SVP representation and crossover while naturally extending application to the module lattices.

论文介绍 传统密码学基于整数分解或离散对数等问题,在量子计算机面前存在安全威胁。后量子密码学的核心是最短向量问题(SVP)。本文通过将Ajtai等人的筛法视为遗传算法进行增强,引入领域知识的SVP表示和交叉操作,改进了SVP求解方法,并自然扩展到模格场景。这有助于提升后量子密码学中格基问题的解决效率。

S3C2 Summit 2025-07: Government Secure Supply Chain Summit

第一作者: Sivana Hamer · 方向: 软件安全

Abstract:Software supply chains, while providing immense economic and software development value, are only as strong as their weakest link. Over the past several years, there has been an exponential increase in cyberattacks specifically targeting vulnerable links in critical software supply chains. The attacks disrupt day-to-day functioning and threaten the security of nearly everyone on the internet, from billion-dollar companies and government agencies to hobbyist open-source developers. The evolving threat of software supply chain attacks has garnered interest from both the software industry and governments worldwide in improving software supply chain security. On Thursday, July 9th, 2025, 3 researchers from the NSF-backed Secure Software Supply Chain Center (S3C2) conducted a Secure Software Supply Chain Summit with a diverse set of 12 participants from 6 US government agencies...

论文介绍 软件供应链安全是当前网络攻击的重点目标,威胁到从企业到开源开发者的广泛群体。本文介绍了NSF支持的安全软件供应链中心(S3C2)于2025年7月举办的政府安全软件供应链峰会,汇集了研究人员和美国政府机构代表,共同探讨改善软件供应链安全的策略。该峰会旨在推动政府与行业合作,应对供应链安全挑战。

Techreport: Evaluating Tor-based Location Privacy for Ethereum Validators

第一作者: Muhammad Umar Janjua · 方向: 密码学协议

Abstract:Privacy and anonymity of validators, especially regarding IP address linkability, are essential to protect the Ethereum network from various attacks. Network-level attacks, such as DoS, can interrupt validators and affect the overall security of the Ethereum network. Correlating the IP addresses of validators with their identities, along with knowledge about their action slots can be exploited by attackers to cause network delays, MEV exploitation, and finality risks. Therefore, ensuring the unlinkability of a validator's IP and identity is crucial for maintaining the network's trust and resilience. In this techreport, we first provide a review of the existing network and consensus layer techniques that have been proposed for maintaining validator privacy in the Ethereum blockchain. Secondly, we evaluate a Tor-based protocol named Tor push that helps unlink validator...

论文介绍 以太坊验证者的IP地址与身份关联可能导致网络攻击,影响网络安全。本文评估了一种基于Tor的协议Tor push,旨在帮助验证者保持IP和身份的不可链接性,从而提升隐私保护。通过审查现有技术和评估Tor push协议,研究为维护以太坊网络的信任和韧性提供了见解。

unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning

第一作者: Geoffrey Bradway · 方向: AI 安全

Abstract:Unix competence is the ability to use shell and operating-system primitives as first-class tools, not merely to write programs through a terminal. Current terminal benchmarks tend to blur this distinction: a solver fluent in Python but weak in Unix can pass a substantial fraction of Terminal-Bench 2.0, while the reverse skill profile is rarely exercised. We make the distinction operational and build a training surface for the Unix component. unix-ctf is a procedural generator of capture-the-flag tasks for shell agents. Each task hides a short token (a flag of the form flag(a3b1c9...)) inside a fresh Linux container using a single Unix feature, and the agent must recover it. Tasks are produced by an LLM-assisted synthesis pipeline that generates candidate hiding techniques, rewrites them into parameterized hide-and-find script pairs, and filters them with a bidirectional...

论文介绍 Unix能力指有效使用shell和操作系统原语的技能,但现有终端基准测试常混淆此能力与其他编程技能。本文提出unix-ctf,一个程序化生成的捕获旗标任务环境,用于强化学习训练。任务通过在新Linux容器中隐藏旗标,要求代理使用Unix特性恢复,从而针对性地提升Unix操作能力。这有助于训练AI系统在安全和系统管理任务中的Unix技能。

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

第一作者: Mohammadreza Teymoorianfard · 方向: 安全研究

Abstract:Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework...

论文介绍 集成推理能力的视觉-语言-动作模型被用于端到端自动驾驶,但其在现实输入扰动下的鲁棒性研究不足。本文研究表明,这类模型在闭环仿真中对文本输入扰动高度脆弱,推理攻击成功率高达89%,轨迹操纵成功率达72%,导致碰撞率上升。作者以Alpamayo模型为例,开展了首个针对推理增强VLA模型的系统性黑盒研究,并引入了推理感知评估框架。

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

第一作者: Ojas Nimase · 方向: 密码学协议

Abstract:Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a growing concern for fairness and information integrity. Research on generative engine optimization (GEO) has produced many manipulation methods, but each is evaluated on its own dataset with its own metrics, so their relative strength and detectability stay unclear. We present GEO-Bench, a benchmark that evaluates GEO ranking-manipulation attacks under one protocol. It unifies black-box prompt-based attacks (TAP, Zero-Shot), white-box gradient-based attacks (STS, RAF, StealthRank), and ten white-hat C-SEO strategies. We score every method on five datasets against a fixed open-weight ranker (Llama-3.1-8B-Instruct), using metrics for both effectiveness (NRG, Success@{\alpha}, Promote@{\alpha}) and stealth (keyword violation rate...

论文介绍 大语言模型在查询排名中易受操纵,但现有生成引擎优化攻击方法缺乏统一评估。本文提出GEO-Bench基准,在单一协议下评估多种排名操纵攻击,包括黑盒提示攻击和白盒梯度攻击。通过在固定开源排名器上测试,衡量攻击有效性和隐蔽性,为研究排名操纵的公平性和信息完整性提供标准化工具。

Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening

第一作者: Mohan Zhang · 方向: AI 安全

Abstract:LLMs are vulnerable to prompt injection attacks. However, this vulnerability has been primarily demonstrated conceptually in academic studies or through a few anecdotal case studies. Its prevalence and impact in real-world LLM-based applications are largely unexplored. In this work, we present the first systematic study of prompt-injection attacks in a widely used application: LLM-based resume screening. Our analysis is based on approximately 200K real-world resumes collected over multiple years by hireEZ. We first design tailored methods to detect prompt injection in resumes. Manual validation on a small-scale dataset demonstrates that our detectors achieve high precision and outperform state-of-the-art general-purpose detectors. We then apply our detector to the full resume dataset and conduct a comprehensive measurement study of real-world prompt injection attacks. Our...

论文介绍 大语言模型易受提示注入攻击,但其在现实应用中的普遍存在性缺乏研究。本文首次系统研究基于LLM的简历筛选中的提示注入攻击,使用约20万真实简历数据集。通过设计专用检测器并验证其准确性,研究全面测量了现实世界中的攻击模式和影响,对提升招聘系统等LLM应用的安全性有重要参考价值。

A Secure, Manifest-Based Framework for Delegated Privilege Promotion

第一作者: Rajarshi Chowdhury · 方向: 软件安全

Abstract:Large-scale enterprise software systems commonly run as unprivileged service accounts to enforce least privilege, yet still depend on a small set of privileged components -- such as executables with elevated ownership, permissions, or capabilities -- for narrowly scoped operations. This creates a persistent security and operational conflict during maintenance. Automated patching tools running without elevated privileges cannot safely update privileged components without either executing the entire patch with full administrative rights or requiring manual administrator intervention. We present a secure, manifest-based infrastructure for delegated promotion of privileged software components, deployed in production as part of a large-scale enterprise database system serving both cloud and on-premises installations. The design centers on a minimal privileged mediator that...

论文介绍 研究大规模企业软件系统中,非特权服务账户依赖特权组件导致的安全和运维冲突。提出一个基于清单的委托提升特权框架,通过最小特权中介安全更新特权组件,已在生产环境中的大规模企业数据库系统部署,适用于云和本地安装。

Optimal Rates for Differentially Private Hypothesis Testing with E-values

第一作者: Ben Jacobsen · 方向: 隐私保护

Abstract:E-values have attracted considerable interest in recent years as flexible tools for enabling anytime-valid and adaptive data analysis. Hypothesis testing is at the core of many of these applications, which can often involve private or sensitive data. In this work, we answer a simple but important question: given two distributions $\mathbb{P}$ and $\mathbb{Q}$, what is the maximum achievable e-power when testing $X\sim \mathbb{P}^n$ against $X\sim\mathbb{Q}^n$ with e-values that satisfy $\varepsilon$-differential privacy? We characterize the optimal rate for this problem and provide an algorithm which matches it exactly. In the sequential setting, when observations arrive one-by-one and the analyst chooses when to halt, we give matching upper and lower bounds on the stopping times of any private e-process. Numerical experiments confirm the practicality of our algorithms, which...

论文介绍 探讨差分隐私假设检验中使用e-values的最大可达到性能。问题核心是给定两个分布时,在ε-差分隐私约束下测试的最优速率。研究表征最优速率并提供匹配算法,适用于隐私保护的数据分析场景。

AIRGuard: Guarding Agent Actions with Runtime Authority Control

第一作者: Suliu Qin · 方向: 密码学协议

Abstract:Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This makes agent attacks different from jailbreaks. The harmful step is often not an obviously forbidden output, but an ordinary executable action that becomes unsafe because attacker-controlled context steers authorized access against the user's interest. We identify this failure mode as authority confusion: untrusted resources may inform reasoning, but they must not authorize side effects. We present AIRGuard, a runtime guard that operationalizes least privilege as action-time authorization. AIRGuard normalizes heterogeneous tool calls, derives task authority into step-level authority, tracks source and target trust, simulates sensitive side effects, audits cross-step risk, and enforces decisions before...

论文介绍 针对AI代理动作可能导致权限混淆的安全风险,提出AIRGuard运行时防护系统。该系统标准化工具调用、派生任务权限、模拟副作用并审计风险,以实现最小权限原则,增强代理在复杂环境中的安全性。

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

第一作者: Jaydip Sen · 方向: AI 安全

Abstract:Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses significant challenges to reliability, security, and trustworthiness. Adversarial machine learning demonstrates that even highly accurate models can be manipulated through carefully crafted perturbations, raising serious concerns in safety critical systems such as healthcare, finance, and autonomous technologies. In parallel, quantum computing has emerged as a transformative paradigm capable of addressing complex computational problems through principles such as superposition, entanglement, and quantum interference. The convergence of these fields has led to the emergence of quantum artificial intelligence, which explores how quantum techniques can enhance learning efficiency, scalability, and robustness. This chapter provides a...

论文介绍 综述量子计算如何增强人工智能的对抗鲁棒性。AI模型易受对抗攻击影响,而量子技术如叠加和纠缠可能提升学习效率和鲁棒性。研究探讨量子人工智能在关键领域如医疗和自动驾驶中的潜在应用。

Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought

第一作者: Jiacheng Lu · 方向: 密码学协议

Abstract:Large Language Models with Chain-of-Thought reasoning capabilities represent valuable intellectual property, yet existing black-box watermarking methods often trade robustness for reasoning fidelity by perturbing final answers or relying on fragile trigger patterns. We propose BiCoT, a watermarking framework that embeds ownership signals into the internal geometry of reasoning traces by aligning high-saliency structural anchors with a private signature subspace while regularizing ordinary control tokens to preserve semantic capacity. This design couples the watermark with reasoning-relevant representations, making removal difficult without disrupting the features that support coherent reasoning. To enable verification under model theft and representation drift, we introduce Robust Subspace Registration (RSR), a Top- logprob-based black-box verifier that uses sentinel tokens to...

论文介绍 针对大型语言模型知识产权保护需求,提出BiCoT水印框架。该框架通过将所有权信号嵌入推理轨迹的高显著性结构锚点,并使用黑盒验证方法,实现隐蔽且有效的水印嵌入,防止模型被盗用或篡改。

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

第一作者: Caleb DeLeeuw · 方向: 安全研究

Abstract:Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal structurally sound, or does it disappear under modest changes to prompt framing, formatting, or output length? Across five architectures, no model cleanly discriminated benign from hazard. Gemma 2 2B-IT never genuinely refused across 75 prompts, hedging on every hazard-adjacent query. Gemma 4 E2B-IT refused 65/75 prompts with chat-template formatting and 0/75 without it. Both Gemma models collapsed to 0% under an 80-token cap. Qwen 2.5 1.5B and Phi-3-mini over-refused, flagging 83-87% of benign biology as hazardous. Llama 3.2 1B showed the only meaningful tier gradient (61-point spread). To probe what drives such over-refusal, we tested a panel of Schedule I but biologically non-toxic compounds...

论文介绍 审核语言模型在生物安全场景下的拒绝深度,发现不同模型的拒绝行为不稳定。提出BioRefusalAudit方法,使用通用和领域微调的稀疏自编码器评估模型安全响应,为改进模型可靠性提供依据。

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

第一作者: Dongrui Liu · 方向: 安全研究

Abstract:Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers, rendering current agent alignment frameworks inadequate for real-world deployment. To tackle these emerging threats, we propose a lightweight and scalable agent safety alignment framework. Specifically, we update the agent safety taxonomy to accommodate emergent risks from Codex and OpenClaw execution scenarios. We further build a taxonomy-guided data engine with influence-function purification to train lightweight AgentDoG 1.5 variants (0.8B, 2B, 4B, and 8B parameters) using only around 1k samples, achieving comparable performance with leading closed-source models (e.g., GPT-5.4). Based on AgentDoG 1.5, we construct a highly efficient agentic safety SFT and RL...

论文介绍 提出轻量级可扩展的AI代理安全对齐框架AgentDoG 1.5。更新代理安全分类法以涵盖新兴风险,构建数据引擎训练小型模型,旨在应对开放世界代理的安全挑战,提升对齐效果和部署安全性。

Secure Distributed Hypothesis Testing

第一作者: Gowtham R. Kurri · 方向: 系统安全

Abstract:In distributed hypothesis testing, a central server performs hypothesis testing based on information received from distributed sensors/clients. We study a secure variant of this problem in which the central server determines the hypothesis class of an underlying distribution without learning any additional information about the distribution itself. We prove that, in its standard form, this is impossible to achieve, even for simple and highly restricted cases. To bypass this impossibility, we augment the model with a shared secret key available to clients but hidden from the server. We show that a single-bit secret key enables perfectly secure testing for simple classes by reducing the test distributions to a symmetric, canonical instance. Finally, for arbitrary hypothesis classes over finite domains, we establish a reduction to standard hypothesis testing using Private...

论文介绍 研究安全分布式假设测试问题,核心是服务器在不学习额外信息下确定假设类。证明标准形式不可行,引入共享密钥实现完美安全测试,适用于隐私敏感数据的统计分析场景。

Information Security in Small-Scale Protests: Surveillance of Ugandan Anti-EACOP Protesters

第一作者: Ntezi Mbabazi · 方向: AI 安全

Abstract:We examine the information security practices of Ugandan climate activists protesting the development of the East African Crude Oil Pipeline (EACOP). We conducted five-week fieldwork in Kampala, Uganda, which included interviews with 13 anti-EACOP activists. Through an inductive analysis, we report on the complexities faced by small groups of predominantly student protesters as they covertly organise small-scale anti-EACOP protests within a context marked by state surveillance and repression. Our study points to a multi-layered adversarial landscape, where participants' experiences of direct threats, including arrests and information compromise, and their fears of abduction, shaped their security practices. These practices were rooted in autonomous decision-making within groups. We present a grounded understanding of how participants' need to protect information for their own...

论文介绍 本研究探讨乌干达气候活动家在抗议东非原油管道(EACOP)时的信息安全实践。通过为期五周的田野调查和访谈,分析在国家监控与镇压背景下,小规模学生抗议者如何秘密组织抗议活动。研究揭示了多层对抗性环境,参与者面临逮捕、信息泄露和绑架恐惧等直接威胁,这些经历塑造了他们基于群体自主决策的安全实践。研究为理解小规模抗议中的信息安全挑战提供了实证基础。

CODEFUSE-DEBENCH: An Empirical Study on Readability, Recompilability, and Functionality

第一作者: Puzhuo Liu · 方向: 软件安全

Abstract:Binary decompilation aims to recover binaries into high-level source code, but existing evaluations mainly rely on syntactic similarity or single-axis readability metrics, which fail to capture practical reusability. We propose a reusability-driven evaluation paradigm that measures decompiler quality along three orthogonal dimensions: readability, recompilability, and functionality. We present DEBENCH, the first automated framework for multidimensional decompilation evaluation. DEBENCH contains 240 atomic test functions, organized into 8 source files and compiled into 640 binaries. It combines LLM-as-judge readability scoring with URAF (18 sub-dimensions), iterative compile-and-repair under a fixed 50-iteration budget, and Frida-based differential dynamic tracing at the program, function, and instruction levels. We evaluate five mainstream decompilers and three repair LLMs...

论文介绍 本文针对现有二进制反编译评估方法的局限性,提出一种以可重用性为导向的多维度评估范式。该范式从可读性、可重编译性和功能性三个正交维度衡量反编译器质量。研究介绍 DEBENCH 自动化评估框架,包含 240 个原子测试函数,结合 LLM 评分、编译修复和动态追踪技术,对主流反编译器进行评估,旨在提升反编译代码的实际可用性。

Provably Secure Agent Guardrail

第一作者: Benlong Wu · 方向: AI 安全

Abstract:As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates a fundamental crisis in artificial intelligence security. Existing defense architectures heavily rely on empirical semantic guardrails and probabilistic large model adjudicators, mechanisms that fail to provide deterministic security lower bounds when facing complex semantic symbol decoupling attacks. To overcome this empirical semantic guardrail dilemma, this paper proposes a new security paradigm for agents based on the fundamental limitations of logical reasoning. Based on this paradigm, we further introduce an executable Proof-Constrained Action (ePCA) framework with a neural symbolic isolation architecture. This framework abandons semantic trust in natural language, forcing agents to losslessly formalize their intentions...

论文介绍 随着大语言模型代理获得广泛执行权限,AI 失控构成重大安全危机。现有防御依赖经验语义护栏,无法提供确定性安全保证。本文提出基于逻辑推理极限的新安全范式,并引入可执行证明约束动作(ePCA)框架。该框架通过神经符号隔离架构,强制代理将意图形式化为可验证证明,放弃对自然语言的语义信任,从而为代理行为提供可证明的安全下界。

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

第一作者: Aditya Nawal · 方向: AI 安全

Abstract:AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating external content into the generation pipeline can weaken the safety alignment mechanisms that govern model outputs. Prior work shows that enabling retrieval in agents increases compliance with harmful requests. We introduce AgentREVEAL, a diagnostic framework for analyzing retrieval-induced safety degradation in LLM agents. The framework examines two axes: how retrieval is integrated into the agent pipeline and the properties of the retrieved content. Along the integration axis, we find that binding tool invocation and response generation in a single step amplifies harmful outputs. Along the content axis, we uncover the Safe Source Paradox: even oppositional or safety-oriented sources, such as pages containing warnings or risk...

论文介绍 AI 代理通过外部工具如网络检索增强能力,但这可能削弱安全对齐机制。本文引入 AgentREVEAL 诊断框架,分析检索如何降低 LLM 代理的安全性。研究从检索集成方式和检索内容属性两个轴线展开,发现单步集成工具会放大有害输出,并揭示安全源悖论:即使反对性或安全导向的内容也可能导致风险。这为设计更安全的检索集成提供了见解。

Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents

第一作者: Jay Yu · 方向: 密码学协议

Abstract:DeFi investment agents, systems that use AI for autonomous on-chain trading, have attained over USD 3 billion in combined token valuations since late 2024. We survey over 1,900 AI-tagged crypto projects, filter to investment-focused agents, and curate 10 representative projects spanning strategy and observability dimensions. We then conduct a deep-dive architectural analysis of two prominent agent frameworks, ElizaOS and Virtuals Protocol, and a quantitative on-chain performance analysis of 11 Solana-based agent treasuries with publicly attributable trading activity, covering 925,323 token holders. We find that current deployments remain early and heterogeneous: (1) in our sample, many projects do not yet provide clear evidence of autonomous trade execution, and developer interviews suggest that many visible deployments remain basic API integrations; (2) agent treasuries...

论文介绍 DeFi 投资代理自 2024 年底以来总估值超过 30 亿美元。本文调查 1900 多个 AI 标签加密项目,筛选投资代理并分析 10 个代表性项目。通过架构分析和链上性能评估,发现当前部署仍处于早期阶段,许多项目缺乏明确的自主交易执行证据,且代理资金库表现异质。研究为评估 DeFi 代理的可行性和风险提供了实证数据。

Robust and Efficient Guardrails with Latent Reasoning

第一作者: Siddharth Sai · 方向: 安全研究

Abstract:Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrails typically rely on single-pass classification or, more recently, distilled reasoning. Reasoning-based guardrails significantly outperform classification-only baselines, but they incur substantial query latency and token overhead that make them impractical for highthroughput deployment. To address this challenge, we propose COLAGUARD, a guardrail model that transfers multi-step safety reasoning into a continuous latent space through a stage-wise training curriculum, enabling direct hidden-state propagation at inference. Evaluated on ten prompt- and response-moderation settings spanning eight safety benchmarks, COLAGUARD improves macro-F1 by 8.24 points over Llama Guard 3 and matches our explicit reasoning baseline...

论文介绍 保持大语言模型的安全性至关重要,但现有推理型安全护栏延迟高、开销大,不适用于高吞吐场景。本文提出 COLAGUARD 模型,通过阶段化训练课程将多步安全推理转移到连续潜空间,实现推理时直接隐藏状态传播。在十个提示和响应审核设置上评估,COLAGUARD 在宏 F1 上显著提升,匹配显式推理基线,同时大幅减少延迟和 token 消耗。

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

第一作者: Kaihua Qin · 方向: 软件安全

Abstract:Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studies use narrow datasets, inconsistent metrics, and limited semantic consistency checks. This gap is increasingly important as large language models (LLMs) begin to generate source-like Solidity that may compile and appear plausible, even when its semantics diverge from the original contract. We introduce SCDBench, a dataset and benchmark methodology for LLM-based smart contract decompilation. The dataset contains 600 real-world Solidity contracts with paired bytecode inputs, ground-truth source code, and replayable semantic checkpoints. SCDBench evaluates decompiler outputs through four cumulative stages: format completeness, compilability, Application Binary Interface (ABI) recovery, and semantic consistency via differential...

论文介绍 智能合约反编译评估面临数据集窄、指标不一致等挑战,尤其在基于 LLM 的反编译器生成看似合理但语义偏离的代码时。本文引入 SCDBench 基准,包含 600 个真实 Solidity 合约数据集,通过四阶段评估方法:格式完整性、可编译性、ABI 恢复和语义一致性检查,系统评估 LLM 反编译器的输出质量。

Cycle-Space Informed Detection of Autoencoded Blind False Data Injection Attacks on Power Systems

第一作者: Xin Li · 方向: 系统安全

Abstract:The rapid growth of AI-driven data centers and large-scale energy storage systems is increasing the reliance of power system operation on real-time measurement data and automated decision-making. However, many existing detection methods rely on statistical or data-driven analysis of measurements and can fail when attackers exploit the same data structure to craft stealthy perturbations. To illustrate this limitation, we demonstrate a blind False Data Injection Attack (FDIA) in which an Autoencoder learns the measurement manifold and generates perturbations aligned with the Jacobian null space, thereby allowing the attack to evade both residual-based baddata detectors and time-series anomaly detectors. To mitigate data-driven FDIAs which exploit the null space, we propose a topology-informed Cycle-Space Detector (CSD) that leverages the Cycle-Space of the network to impose...

论文介绍 随着 AI 驱动的数据中心和储能系统增长,电力系统更依赖实时数据,但现有检测方法易被隐蔽攻击绕过。本文演示一种基于自编码器的盲假数据注入攻击,该攻击利用雅可比矩阵零空间生成扰动,逃避传统检测器。为此,提出拓扑感知的循环空间检测器(CSD),利用网络循环空间结构来检测此类数据驱动攻击,提升电力系统安全。

Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities

第一作者: Yujie Ma · 方向: 软件安全

Abstract:Large Language Models(LLMs) have been actively integrated into modern software systems as critical components. LLM-in-the-loop vulnerabilities, where vulnerabilities are introduced by LLMs and their dependent downstream components, such as frameworks, introduce new risks. Although some benchmark datasets have been constructed to study the impact of such vulnerabilities, most works still remain at the analysis from the conventional software level, ignoring the harm actually caused by LLMs. Understanding real-world LLM-in-the-loop vulnerabilities is still an open problem. To address this gap, we build the first LLM-in-the-loop vulnerability dataset, LLMCVE, to facilitate the risk analysis of LLM-integrated software. To do so, we first collect 2,888 multi-source vulnerabilities across 230 popular LLM components. Then, through manual analysis, we identify 205 vulnerabilities that...

论文介绍 研究针对大语言模型集成到软件系统中引入的LLM-in-the-loop漏洞问题,目前缺乏对实际危害的理解。作者构建了首个此类漏洞数据集LLMCVE,收集230个流行LLM组件的2,888个多源漏洞,并通过手动分析识别出关键漏洞。该数据集旨在促进LLM集成软件的风险分析,推动漏洞检测与修复研究。

Meta-Quantum Ensemble Framework for Robust Network Intrusion Detection

第一作者: Ritvik Bhatnagar · 方向: AI 安全

Abstract:Intrusion Detection Systems (IDSs) must maintain high detection sensitivity while operating under strict false-positive constraints, a challenge intensified by class imbalance and heterogeneous IoT traffic. This work investigates whether heterogeneous quantum learners can provide useful and non-redundant decision information for IDS tasks. We study Quantum Support Vector Machines (QSVMs) and Quantum Neural Networks (QNNs), which rely on different learning mechanisms and exhibit distinct prediction behaviors. To combine these models, we propose the System-Level Meta-Quantum Ensemble (MQE), a hybrid quantum-classical framework that fuses QSVM and QNN outputs using a Random Forest meta-learner. The meta-learner captures agreement and disagreement patterns between the quantum branches to improve prediction stability and detection performance. Experiments on TON IoT and CICIDS2017...

论文介绍 入侵检测系统需在低误报率下保持高灵敏度,但面临类别不平衡和异构IoT流量挑战。本研究探索量子支持向量机和量子神经网络的异质集成,提出系统级元量子集成框架MQE,使用随机森林元学习器融合量子分支输出,以提升预测稳定性和检测性能。实验在TON IoT和CICIDS2017数据集上进行,验证了框架的有效性。

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

第一作者: Jusuk Lee · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume...

论文介绍 机器人操作依赖于感知动作相关场景的能力,但现有视觉编码器多为静态识别设计。本文提出DynaFLIP,一个动态感知多模态预训练框架,通过构建图像-语言-3D流三元组进行监督,推动运动理解前移到感知阶段。关键思想是鼓励三种模态在共享超球面空间中形成小单纯形体积,以增强对齐,避免几何歧义。

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

第一作者: Chunru Lin · 方向: 数据集与评测 · 来源: cs.RO

Abstract:The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments. However, current robotic benchmarks primarily emphasize skill-level execution and provide limited insight into such cognitive reasoning capabilities. We introduce RoboWits, a bi-manual robotic benchmark designed to systematically evaluate cognitive reasoning, creative tool use, and robustness to unexpected conditions. To enable scalable construction of high-quality reasoning-centric unexpected scenarios, we propose an automated task generation pipeline formulated as a multi-agent cooperative framework, comprising agents for seed task generation and verification, metric generation, scene generation, and task mutation. Using the pipeline, we curated 30 diverse seed tasks and 208 tasks with mutations and graded difficulty across...

论文介绍 现有机器人基准侧重技能执行,缺乏对认知推理能力的评估。本文引入RoboWits,一个双臂机器人基准,用于评估认知推理、创造性工具使用和对意外条件的鲁棒性。通过多智能体协作框架自动生成高质量任务,构建了30个种子任务和208个带变异的难度分级任务,推动机器人认知能力研究。

A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms

第一作者: Yufei Jia · 方向: 策略学习 · 来源: cs.RO

Abstract:Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning are placed on a single GPU-centric execution path. This paradigm has greatly improved training speed, but it has also encouraged a default assumption that efficient training requires physics to reside on the GPU. We revisit this assumption. Our view is that, in simulation-dominated robot control, the essential question is not which processor runs physics, but whether simulation throughput, policy learning, and runtime synchronization form an efficient end-to-end loop. We present UniLab, a heterogeneous CPU-simulation / GPU-learning architecture that decouples CPU-parallel simulation from GPU policy updates through a unified runtime for data movement, buffering, and synchronization. UniLab is implemented as a complete and...

论文介绍 基于模拟的机器人强化学习常采用GPU主导范式,但可能并非最高效。本文提出UniLab,一个异构CPU模拟/GPU学习架构,通过统一运行时解耦CPU并行模拟和GPU策略更新,优化数据移动和同步。该架构旨在提升模拟吞吐量和学习效率,适用于复杂机器人控制任务。

Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

第一作者: Kuangji Zuo · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to describe which exact object to interact with among similar candidates, where to act on the object, or how the target may change during execution. To address this limitation, we propose Gaze2Act, a novel VLA framework that leverages human gaze as a dynamic and intuitive intent signal for complex interactive manipulation. Gaze2Act first bridges the ego-exo view gap by mapping first-person gaze into the robot's perspective through cross-view semantic matching, producing both an object mask and a gaze point for coarse-to-fine target specification. These cues are then integrated into the policy through perception-level prompting and action-level...

论文介绍 视觉-语言-动作模型在机器人学习中潜力巨大,但仅靠语言难以精确传达人类意图。本文提出Gaze2Act,利用人类注视作为动态意图信号,通过跨视图语义匹配将第一人称注视映射到机器人视角,生成物体掩码和注视点。这些线索集成到策略中,实现粗到细的目标指定,提升交互操作性能。

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

第一作者: Qiuyue Wang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder. Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data, trajectory-centric supervision, and...

论文介绍 具身智能常使用专门模型处理单个任务,导致能力碎片化和泛化有限。本文研究将异质具身决策问题统一到单一视觉-语言-动作模型中。提出Qwen-VLA,扩展Qwen的视觉-语言栈,通过DiT动作解码器生成连续动作和轨迹,使用大规模联合预训练覆盖多样数据源,旨在实现跨任务、环境和机器人形态的统一。

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

第一作者: Zhongxi Chen · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulation. However, dexterous manipulation remains challenging for VLA policies due to high-dimensional hand control and compounding execution errors, which makes real-world RL post-training essential for bridging the gap between visually grounded action generation and physically reliable dexterous execution. However, high-dimensional dexterous exploration often triggers temporal inconsistency, sample inefficiency and hardware risks in the real world. To address these challenges, we propose BORA, an offline-to-online RL post-training framework designed for real-world dexterous VLA models. In the offline phase, BORA constructs a critic that takes both the VLM's cognition tokens and action chunks as inputs. This design enables...

论文介绍 视觉-语言-动作模型在灵巧操作中面临高维控制和执行误差挑战,需真实世界强化学习后训练。本文提出BORA,一个离线到在线强化学习后训练框架,用于真实世界灵巧VLA模型。离线阶段构建批评家,结合VLM认知令牌和动作块输入,以提升训练稳定性和样本效率,减少硬件风险。

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

第一作者: Shutong Ding · 方向: 策略学习 · 来源: cs.RO

Abstract:Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approaches, one representative branch focuses on the sampling-based policy optimization. This design enables better exploration capability of the diffusion model, particularly at the beginning of training, but suffer from low exploitation in Q-value information, resulting in a slow policy convergence. Another branch pays attention to gradient-based policy optimization, which sufficiently exploits the gradient of the Q function yet tends to collapse into a unimodal policy with low diversity. To address this issue, we propose CGPO, \textbf{C}ritic-\textbf{G}uided diffusion \textbf{P}olicy \textbf{O}ptimization, which effectively balances exploration and exploitation with the training-free guidance technique...

论文介绍 本文提出 CGPO 方法,以解决扩散策略在强化学习中采样优化收敛慢与梯度优化多样性低的问题。该方法引入无训练的 Critic 引导技术,在扩散策略的采样过程中直接利用 Q 值梯度信息,有效平衡了策略的探索与利用能力,旨在提升策略的收敛效率和最终性能。

Replicable Simulation-Based Robot Validation through Provenance

第一作者: Argentina Ortega · 方向: 具身智能 · 来源: cs.RO

Abstract:Robot behavior is often validated through simulation-based testing, yet the replicability of such campaigns depends critically on transparent documentation of how tests are configured, executed, and post-processed. We argue that data provenance, coupled with the FAIR principles (findability, accessibility, interoperability, and reusability), addresses this gap by explicitly tracking links between artifacts and by attaching machine-readable metadata about file origins and key design decisions. Moreover, provenance and metadata cannot be treated as an afterthought confined to final datasets; they must be integrated into the testing processes that generate those datasets so that evidence can be reconstructed end-to-end. We demonstrate this by augmenting an existing simulation-based testing framework with provenance tracking and metadata collection mechanisms, and by using these...

论文介绍 针对基于仿真的机器人行为验证难以复制的问题,本文主张将数据溯源与 FAIR 原则深度整合到测试流程中。通过在测试框架内嵌入溯源跟踪和元数据收集机制,实现从测试配置到后处理的全链路文档化,确保验证证据可端到端重建,从而提升机器人测试活动的透明度和可复现性。

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

第一作者: Hao Ren · 方向: 导航与运动 · 来源: cs.RO

Abstract:Diffusion models are effective for waypoint prediction in visual navigation, but standard sampling and test time guidance can produce unreliable or inefficient trajectories when updates drift off the training manifold. We propose Fisher Preserving Guidance with Outer Product Span Projection, a training-free inference method that avoids large Fisher drift associated with off-distribution actions while optimizing a task objective. Our method computes the Fisher-preserving update via a low-rank Jacobian factorization, requiring only a single backward pass per step and enabling real-time use. We further introduce Truncated Fisher Denoising Sensitivity as an uncertainty signal and use it for robust multi-sample action blending. Experiments on toy and realistic navigation benchmarks, including Maze2D with TSDF-based guidance, PushT with official Diffusion Policy weights, and visual...

论文介绍 标准扩散模型在导航路径预测中,因测试时引导易偏离训练分布而产生不安全轨迹。本文提出 Fisher 保持引导方法,通过低秩 Jacobian 分解计算保持数据流形的更新方向,在优化任务目标的同时避免 Fisher 漂移。该方法无需训练,单步反向传播即可实时运行,并引入不确定性信号进行鲁棒的多动作融合。

LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation

第一作者: Mohammad Khoshnazar · 方向: 机器人操作 · 来源: cs.RO

Abstract:Multi-step robot manipulation requires acting under uncertainty about how the scene will evolve, making exploration and policy adaptation challenging. We study whether short-horizon, task-consistent future videos can provide useful structured priors for control and reinforcement-learning fine-tuning. We formalize this idea through Future-Experience Conditioning (FEC), a simple interface that conditions closed-loop policies on a latent representation of a short future video. In our simulation setup, future clips are generated in three stages, an LLM reasoner operating over a task ontology initialized from the current scene state, a robot-free digital-twin rollout of the intended object motion, and a mask-free video diffusion model that synthesizes a robot-consistent future clip without requiring segmentation at inference. We instantiate this future-conditioning interface...

论文介绍 本文研究如何利用短时域、任务一致的未来视频为多步机器人操作提供结构化先验。方法通过大语言模型推理生成未来物体运动设想,并在数字孪生中推演,再由视频扩散模型合成为「未来体验」。该体验的潜表示被用于条件化闭环策略,旨在为决策和强化学习微调提供前瞻性指导。

MARS Policy: Multimodality Only When It Matters

第一作者: Jindou Jia · 方向: 机器人操作 · 来源: cs.RO

Abstract:Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to capture diverse yet valid behavioral patterns, has driven the rapid emergence of generative policies as a dominant paradigm in robot learning. However, achieving such multimodality typically relies on stochastic noise initialization and iterative denoising procedures, resulting in substantial training complexity and low inference efficiency. Meanwhile, not all phases of a robotic task inherently require behavioral diversity. Motivated by this insight, we propose the Modality-Adaptive Robot Sampling (MARS) policy, which adaptively invokes tailored stochasticity only when it is truly beneficial, while reverting to an efficient deterministic learning during single-modal phases. In other words, the proper amount of noise is injected only...

论文介绍 多模态生成策略因依赖随机噪声和迭代去噪,面临训练复杂、推理效率低的问题。本文提出 MARS 策略,其核心洞察是并非所有任务阶段都需要行为多样性。该方法在需要多样性的阶段自适应地引入随机性,而在单模态阶段则回退到高效的确定性学习,从而在保持能力的同时提升推理效率。

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

第一作者: Sergey Arkhangelskiy · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with $N \le 25$ rollouts per condition, almost always without confidence intervals or paired statistical comparison; these cohort sizes struggle to resolve close comparisons reliably. We introduce PhAIL (Physical AI Leaderboard, this https URL), an open real-robot benchmark on a Franka FR3 (dataset, per-rollout artifacts, and end-to-end reference implementation) of a distributional evaluation methodology: the time-to-success cumulative distribution function (CDF) as the evaluation primitive, with two separated jobs. The first is scoring via Human-Relative Throughput (HRT), a dimensionless scalar with bootstrap confidence intervals, anchored to same-fixture human teleoperation. The second is a significance test (Kolmogorov-Smirnov, computed per-object and...

论文介绍 当前真实机器人上对视觉语言动作模型的评估缺乏统计严谨性。本文提出 PhAIL 基准和分布式评估方法,采用「成功时间累积分布函数」作为核心评估原语,并引入与人类操作吞吐量比较的标量指标及显著性检验。该基准基于 Franka 机器人,提供完整的数据集、参考实现和逐次运行记录。

EXACT-MPPI: Exact Signed-Distance Navigation for Arbitrary-Footprint Robots from Point Clouds via Path Integral Control

第一作者: Chen Peng · 方向: 导航与运动 · 来源: cs.RO

Abstract:Ground robots often carry payloads, implements, or other attachments that turn their effective footprint into complex, non-convex shapes. Navigating safely through clutter then requires reasoning about this true geometry, yet most local planners simplify it with convex or inflated proxies and rasterize sensor data into occupancy grids or distance fields. Both choices eliminate feasible motions when clearance is comparable to the footprint geometry. We present EXACT-MPPI, a training-free local navigation framework that maps local point-cloud observations and sparse guidance directly to motion commands, without any intermediate map representation. The framework embeds an analytic, exact signed-distance evaluator into a Model Predictive Path Integral (MPPI) controller. The footprint is represented as a simple polygon for general convex or concave planar shapes, with a...

论文介绍 为解决具有复杂非凸几何足迹的机器人在点云环境中的安全导航问题,本文提出 EXACT-MPPI 框架。该方法直接嵌入精确的符号距离计算到模型预测路径积分控制器中,无需构建中间地图或简化足迹几何,能保留狭窄但可行的运动通道,适用于任何多边形足迹的实时局部规划。

VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models

第一作者: Dehao Huang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Confidence estimation for Vision-Language-Action (VLA) models is essential for robots to perform manipulation tasks in the open world, providing crucial signals for risk-sensitive decision-making and failure anticipation. Existing confidence estimation methods typically rely on ensemble-based paradigms or action-token probabilities to predict the likelihood of task success. However, they still encounter challenges in computational efficiency and cross-architecture generalizability. These methods usually require repeated sampling, leading to inference inefficiency, and are restricted to VLA models with discrete action outputs, making them difficult to apply to continuous action spaces. To address this issue, we propose VLAConf, a one-class discriminative confidence framework. By leveraging frozen pretrained VLA internal representations, VLAConf directly estimates step-wise...

论文介绍 针对现有 VLA 模型置信度方法计算效率低、架构通用性差的问题,本文提出 VLAConf 框架。它利用预训练 VLA 模型的冻结内部表征,通过单类判别器直接估计逐步操作成功的置信度。该方法无需多次采样,能兼容连续动作空间,为机器人在开放世界中进行风险敏感决策提供高效的不确定性信号。

Learning to Feel Materials from Multisensory Tactile Data via Interpretable Models

第一作者: Li Zou · 方向: 具身智能 · 来源: cs.RO

Abstract:Human tactile perception of materials relies on complex multisensory touch cues, yet the relationship between low-level tactile signals and perceptual representations remains poorly understood. This knowledge gap hinders the integration of touch in digital environments and the development of robots capable of human-like tactile perception. Here, we present an interpretable computational framework for modeling human material perception and recognition using multisensory touch data. Our framework comprises three interconnected models: Model 1 maps finger-surface interaction features to psychophysical sensory attributes, Model 2 classifies materials based on these perceptual representations, and Model 3 directly classifies materials from tactile features. The results showed that combining information from pressing, static contact, and sliding interactions improves prediction...

论文介绍 本文针对低层触觉信号与高级材料感知表征之间关系不清的问题,提出了一种基于多模态触觉数据的可解释计算框架。该框架包含三个互连模型:一个将手指与物体交互特征映射为心理物理感知属性,一个基于这些感知表征进行材料分类,另一个直接从触觉特征进行材料分类。研究发现,结合按压、静态接触和滑动交互的信息能有效提升材料预测性能。该工作有助于深化对人类触觉感知机制的理解,并推动机器人触觉与数字环境集成技术的发展。

VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

第一作者: Victor Kowalski · 方向: 机器人操作 · 来源: cs.RO

Abstract:When using reinforcement learning (RL) for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve. However, vision-enabled policies tend to overfit to the visual conditions seen during training, limiting their robustness and transferability. We present a human-in-the-loop RL framework that employs teacher-student distillation to achieve robust performance across multiple task variants, trained entirely in the real world without requiring domain randomization or data augmentation. A vision-enabled teacher distills its knowledge into a vision-free student that relies solely on pose, twist, and wrench sensing, combining fast training with strong task generalization. On the real-world NIST assembly benchmark board, our approach achieves 95\% overall success after approximately 50...

论文介绍 本研究针对视觉强化学习策略在接触丰富机器人操作中易过拟合、鲁棒性不足的问题,提出一种人类在环强化学习框架。该框架采用教师-学生蒸馏方法,将视觉教师的知识蒸馏到仅依赖位姿、扭转和力传感的无视觉学生策略中,实现完全真实世界训练,无需领域随机化或数据增强。该方法在NIST组装基准板上展示了高成功率和强任务泛化能力。

VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models

第一作者: Shengyu Si · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and action patterns. This paper proposes VLA-Pro, a plug-and-play framework designed to enhance cross-task generalization by storing task-relevant procedural memories at training time and transferring these memories during inference. Specifically, VLA-Pro stores task-specific LoRA adapters as parameterized procedural memories during training. At inference time, VLA-Pro retrieves relevant procedural memories based on the current multi-modal context and dynamically fuses these memories for generating the current action chunk. Experiments on RoboTwin, RLBench, and real-world manipulation tasks show that VLA-Pro consistently improves cross-task...

论文介绍 本文针对视觉-语言-动作(VLA)模型在跨任务泛化上的挑战,提出VLA-Pro框架。该框架在训练阶段存储任务相关的程序性记忆(如LoRA适配器),并在推理时基于多模态上下文动态检索与融合这些记忆以生成动作。实验在RoboTwin、RLBench及现实操作任务中验证了其提升跨任务泛化的能力,有助于增强机器人对新任务的适应性。

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

第一作者: Ye Li · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time robotic manipulation, especially when large vision-language backbones and iterative action heads run at every control step. Existing VLA acceleration methods often optimize individual components or rely on fixed acceleration rules, treating different control steps with largely fixed computation and overlooking the non-uniform reasoning demands of sequential embodied control. Inspired by human motor control, where cognitive and feedback resources concentrate on goal-sensitive stages, we argue that VLA models should learn when to invest full computation and when to reuse prior computation. We propose ElegantVLA, a plug-in phase-adaptive inference framework that accelerates VLA models through intra-model...

论文介绍 视觉-语言-动作(VLA)模型是通用机器人控制的有力范式,但其高计算成本和低控制频率阻碍实时操作。现有加速方法常优化单个组件或依赖固定规则,忽视顺序控制的非均匀推理需求。受人类运动控制启发,本文提出ElegantVLA,一种插件式阶段自适应推理框架,通过学习控制步骤中的计算资源分配来加速VLA模型,旨在提升机器人操纵的实时性和效率。

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding

第一作者: Zhongyu Xia · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficiency manifests as three intertwined challenges: weak extraction of 3D spatial positions without enforcing multi-view consistency, inadequate 3D instance understanding, and fragile reasoning under occlusion. Although mature 3D perception methods exist, their direct integration into VLA pipelines is hindered by architectural incompatibility and by heavy reliance on costly instance-level annotations. To address the above challenges, we propose 3DVLA, a plug-and-play framework that injects robust 3D reasoning into pretrained VLAs without requiring extra manual labels or discarding VLM priors. Specifically, 3DVLA tackles the three challenges through: (1) pervasive 3D feature encoding with explicit...

论文介绍 视觉-语言-动作模型在机器人操作中表现突出,但缺乏3D场景理解,导致空间位置提取弱、实例理解不充分和遮挡推理脆弱。本文提出3DVLA框架,通过3D特征编码和实例感知机制,即插即用地增强预训练VLA的3D推理能力,无需额外标注或丢弃VLM先验,有望提升机器人在复杂环境中的操作性能。

Phase-Conditioned Imitation Learning with Autonomous Failure Recovery for Robust Deformable Object Manipulation

第一作者: Dayuan Chen · 方向: 机器人操作 · 来源: cs.RO

Abstract:This paper presents a phase-conditioned, force-aware framework for robust deformable object manipulation. Standard imitation learning policies such as Action Chunking with Transformers (ACT) rely on a Markovian assumption at inference, causing state aliasing when visually similar observations require contradictory actions and preventing autonomous recovery from execution failures. We address this with a closed-loop hierarchical architecture. A FiLM-conditioned ACT encoder modulates feature extraction based on the current task phase, enabling a single unified policy to produce phase-specific behaviors while sharing action dynamics across phases. A multi-modal phase predictor fusing visual, force, and pose feedback estimates the phase in real time, detecting contact failures that are invisible to vision alone and autonomously triggering recovery trajectories. The system is...

论文介绍 本文针对可变形物体操作中模仿学习策略的鲁棒性问题,提出一种相位条件、力感知框架。标准模仿学习如「ACT」在推理时依赖马尔可夫假设,导致状态别名和无法自主恢复失败。作者设计了一个闭环分层架构:「FiLM」条件化「ACT」编码器根据任务相位调节特征提取,使单一策略能产生相位特定行为;多模态相位预测器融合视觉、力和姿态反馈实时估计相位,检测视觉不可见的接触失败并触发恢复轨迹。该方法旨在提升机器人操作可变形物体的鲁棒性和自主性。

Decentralized LLM-Driven Coordination of Acoustic Robots for Contactless Object Manipulation

第一作者: Yingying Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Natural language interfaces can simplify interaction with multi-robot systems, especially when non-expert users need to issue high-level commands. Acoustic manipulation using ultrasonic phased arrays also enables contactless object handling for applications such as healthcare, laboratory automation, and precision transport. However, combining large language models (LLMs) with distributed acoustic mobile robots remains underexplored. This paper presents a decentralized framework for natural language-driven coordination of acoustic robots for contactless object manipulation. The system converts spoken instructions into executable multi-robot task plans using Whisper-based speech recognition, LLM-based semantic parsing, structured JSON task representation, and distributed scheduling. The JSON schema encodes robot assignments, temporal dependencies, spatial constraints, and...

论文介绍 该论文提出一种去中心化框架,用于自然语言驱动的声学机器人协调,以实现非接触式物体操作。针对非专家用户通过语音与多机器人系统交互的需求,系统利用Whisper语音识别、基于大语言模型的语义解析和结构化JSON任务表示,将口语指令转换为可执行任务计划,并通过分布式调度分配机器人。该方法可应用于医疗保健、实验室自动化等领域,提升机器人操作的可访问性。

MonoDuo: Using One Robot Arm to Learn Bimanual Policies

第一作者: Sandeep Bajamahal · 方向: 机器人操作 · 来源: cs.RO

Abstract:Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-arm robots, however, are widely available in research labs. Can we leverage them to train bimanual robot policies? We present MonoDuo, a framework for learning bimanual manipulation policies using single-arm robot demonstrations paired with human collaboration. MonoDuo collects data by teleoperating a single-arm robot to perform one side of a bimanual task while a human performs the other, then swapping roles to cover both sides. RGB-D observations from a wrist-mounted and fixed camera are augmented into synthetic demonstrations for target bimanual robots using state-of-the-art hand pose estimation, image and point cloud segmentation, and inpainting. These synthetic demonstrations, grounded in real...

论文介绍 研究问题:双臂协调对现实操作任务至关重要,但双臂机器人和数据稀缺限制了策略学习,而单臂机器人更易获取。核心方法:MonoDuo框架通过遥操作单臂机器人与人类协作收集数据,覆盖任务两侧,并利用RGB-D观测、手部姿态估计、图像分割和修复等技术生成合成演示,用于训练目标双臂机器人策略。可能应用:该方法能利用现有单臂机器人资源扩展双臂操作学习,提升机器人操作能力。

Extreme dynamic symmetry enables omnidirectional and multifunctional robots

第一作者: Jiaxun Liu · 方向: 导航与运动 · 来源: cs.RO

Abstract:Symmetry is a central organizing principle in natural systems, yet its use as a unifying design strategy in robotics has largely remained limited to geometric form. We show that symmetry can instead be leveraged at the level of dynamic actuation capability. We introduce dynamic symmetry, the uniformity of a robot's attainable center-of-mass accelerations, and formalize it through a measure coined as dynamic isotropy. Across more than 1000 simulated morphologies, we found that higher dynamic symmetry consistently improved trajectory tracking, task success, robustness, resiliency, and energy efficiency, with the benefits becoming most pronounced as dynamic isotropy approached its theoretical limit. To study this regime systematically, we developed Argus, a family of spherical robots designed to explore the effects of increasing dynamic symmetry. Members of the Argus family vary...

论文介绍 本文将对称性原则应用于机器人的动态驱动能力,提出了“动态对称性”概念及“动态各向同性”度量。研究发现,更高的动态对称性能一致地改善机器人的轨迹跟踪、任务成功率、鲁棒性及能效。为系统研究该极限,作者设计了Argus系列球形机器人。该工作为设计具备卓越全向运动与多功能性的机器人提供了新的统一设计策略。

Learning and Adaptation in Wire Arc Additive Manufacturing Bead Geometry Control

第一作者: Chen-Lung Lu · 方向: 具身智能 · 来源: cs.RO

Abstract:Robotics Wire Arc Additive Manufacturing (WAAM) is governed by complex and nonlinear process dynamics coupling thermal field to the build geometry. The process may be regarded as a multi-input/multi-output dynamical system with welding torch speed and wire feed rate as inputs and weld bead deposition height and width as outputs. In this paper, we use the input/output data to learn a data-driven model and use it for weld planning and control. We show that a simple recurrent neural network architecture and one-step-ahead predictive control can improve the process performance in terms of height and width consistency. To account for the changing thermal conditions during the printing process, we update the learning model using prediction error from the previous layer. This adaptation step further improves the prediction accuracy and controller performance. Experiments on a robotic...

论文介绍 机器人电弧增材制造过程涉及热场与几何形貌的复杂非线性耦合。本文利用输入输出数据学习数据驱动模型,并用于焊接规划与控制。通过简单的循环神经网络架构与单步预测控制,提升了焊道高度与宽度的一致性。为适应打印过程中的热条件变化,模型基于前一层的预测误差进行更新,从而进一步提高了预测精度与控制性能。

Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

第一作者: Petras Swissler · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Swarm and field robotics face significant barriers to real-world validation due to the high cost and development time to deploy hardware. This paper introduces the ``Bionic Swarm,'' a novel system that lowers these barriers by abstracting away many of the tasks that are difficult to implement on robots but which do not contribute to the overall algorithm evaluation, giving these tasks to human users. These human users take directions from a smartphone web-app that takes measurements from Bluetooth-connected sensors and relays them to a centralized server. This server runs the swarm algorithm and directs actions to the human users. We evaluate this system through the experimental validation of a geotechnically-focused search algorithm named Score-Biased-Search, which functions by assigning a ``score'' to each location on a reconstructed map, then biases search patterns through...

论文介绍 群体与田野机器人技术因硬件部署成本高、开发周期长而难以进行实地验证。本文提出“Bionic Swarm”系统,通过将难以在机器人上实现的任务交由人类用户完成,来降低验证门槛。人类用户通过智能手机网络应用接收指令、采集传感器数据。该系统通过土壤测绘实验,验证了一种名为“Score-Biased-Search”的地质搜索算法。

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

第一作者: Luzhou Ge · 方向: 多模态具身 · 来源: cs.RO

Abstract:Integrating open-vocabulary semantic information into dynamic 3D scene representations is essential for long-term embodied scene understanding. However, existing methods often suffer from fragile instance association due to incomplete cross-view cues, while their limited ability to handle object-level topological changes restricts long-term robotic task execution. Moreover, current 3D scene understanding methods either rely on simple feature matching without explicit spatial reasoning or assume offline ground-truth 3D geometry. To address these challenges, we present DGSG-Mind, a hybrid instance-aware 3D Gaussian dynamic scene graph system with an embodied reasoning agent. Our system couples a probabilistic voxel grid with explicit 3D Gaussians to enable robust cross-modal instance fusion and incremental semantic mapping. It handles dynamic changes through Gaussian-based...

论文介绍 将开放词汇语义信息融入动态三维场景表示对于长期具身场景理解至关重要。本文提出DGSG-Mind系统,它是一个混合实例感知的动态三维高斯场景图系统,配备具身推理智能体。该系统结合概率体素网格与显式三维高斯,实现稳健的跨模态实例融合与增量语义建图,并能通过基于高斯的拓扑变化检测处理动态变化。

Energy-Aware NECO for Single-Pass Pixel-wise Out-of-Distribution Detection in Semantic Segmentation

第一作者: Boyuan Zhang · 方向: 具身智能 · 来源: cs.RO

Abstract:Reliable semantic segmentation for mobile robots requires both accurate dense prediction and robust uncertainty estimation under distribution shift. Strong uncertainty baselines such as Monte Carlo Dropout often require repeated stochastic forward passes and are difficult to deploy on edge platforms. We propose Energy-Aware NECO, a single-pass pixel-wise out-of-distribution (OOD) detector for semantic segmentation. The method combines a centered NECO-style geometric ratio computed from decoder features with a logit-based Energy score. Both components are standardized using statistics fitted on a pure in-distribution validation split and fused through a convex combination. We evaluate the method on the miniMUAD subset using true pixel-level OOD labels. The proposed hybrid score achieves an AUROC of 0.8539, outperforming NECO-only (0.8280), Energy-only (0.8171), and an ensemble...

论文介绍 移动机器人的可靠语义分割需要在分布偏移下实现精确的稠密预测与稳健的不确定性估计。本文提出Energy-Aware NECO,一种用于语义分割的单次前向像素级分布外检测器。该方法结合了基于解码器特征的NECO几何比率与基于logit的能量分数,并通过凸组合融合两者。评估显示,其混合分数的AUROC达到0.8539,优于单独使用任一分数的方法。

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

第一作者: Mohammadreza Teymoorianfard · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework...

论文介绍 集成推理能力的视觉-语言-动作模型被用于端到端自动驾驶,但其在现实输入扰动下的鲁棒性研究不足。本文研究表明,这类模型在闭环仿真中对文本输入扰动高度脆弱,推理攻击成功率高达89%,轨迹操纵成功率达72%,导致碰撞率上升。作者以Alpamayo模型为例,开展了首个针对推理增强VLA模型的系统性黑盒研究,并引入了推理感知评估框架。

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

第一作者: Jiyao Zhang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelligence in embodied 3D environments. To systematically evaluate these foundational perceptual capabilities, the benchmark includes 6 task categories divided into two core groups: Spatial Structural Understanding (Grounding, Spatial Relation Prediction, and Multi-view Correspondence) and Interaction-Oriented Perception (Affordance Prediction, Grasp Point Prediction, and Trajectory Prediction). The benchmark spans 12 subcategories and contains over 21k high-quality question-answer pairs. We evaluate 13 state-of-the-art models, and the results show that while current models exhibit relatively strong high-level spatial reasoning, such as understanding...

论文介绍 当前视觉语言模型是否具备理解与推理复杂三维环境交互的能力?本文推出Embodied3DBench,一个面向机器人的基准测试,专注于底层具身空间智能。该基准包含空间结构理解与交互感知两大类共六个任务,涵盖超过2.1万个高质量问答对。对13个先进模型的评估显示,现有模型在高层空间推理上表现尚可,但在底层感知能力上存在明显不足。

Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems

第一作者: Daniel Albiero · 方向: 具身智能 · 来源: cs.RO

Abstract:Tropical forests worldwide are under intense deforestation pressure driven by economic and political interests, and scientific evidence suggests this deforestation contributes to climate change. This paper proposes a novel logging method for tropical forests, Ultra-Reduced-Impact-Encased-Logging (URIEL). This new method is based on heli-logging techniques combined with intensive use of robotics and AI integrated with post-harvest silvicultural treatments performed by drones. The concept of appropriate equipment for this method was developed, dimensions were determined, details were completed in a digital proof of concept, and an effective digital simulation and economic feasibility analysis were carried out for various helicopter-timber-distance combinations. The results demonstrated that a URIEL method has high economic viability and makes it possible to virtually eliminate...

论文介绍 本文提出一种针对热带森林的新型伐木方法“URIEL”。该方法结合直升机伐木技术,并通过密集使用机器人、人工智能以及由无人机执行的采后营林处理。作者完成了概念设计、尺寸确定及数字仿真,并进行了经济可行性分析。结果表明,URIEL方法具有高经济可行性,并能几乎消除传统伐木对森林造成的环境破坏。

VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

第一作者: Mingjian Gao · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irrelevant or weakly textual information can interfere with action prediction, while autoregressive text decoding adds too much latency for real-time closed-loop execution. We present VISUALTHINK-VLA, a visual intermediate-reasoning framework for accurate, low-latency VLA policies. Our bootstrapping philosophy is to guide action with effective visual thinking: VISUALTHINK-VLA bootstraps action prediction through a compact visual-evidence interface that preserves spatial precision while avoiding decoding overhead. Besides, to further improve performance and efficiency, VISUALTHINK-VLA adopts a tailored selective routing mechanism to learn the visual evidence tokens, enabling low-latency inference while...

论文介绍 针对视觉-语言-动作(VLA)模型中文本推理不适用于实时控制的问题,本文提出了VisualThink-VLA框架。该框架通过一个紧凑的视觉证据接口来引导动作预测,保留空间精度的同时避免文本解码带来的延迟。此外,通过设计选择性路由机制学习视觉证据token,实现了低延迟推理与性能的平衡,旨在提升VLA策略的准确性和执行效率。

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

第一作者: Shilin Ma · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across execution steps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce an...

论文介绍 为了加速视觉-语言-动作(VLA)模型的实时推理,本文提出了SAFE-Pruner这一即插即用的修剪框架。它针对现有视觉token修剪方法可能误删深层所需信息的问题,创新性地将未来层的注意力线索纳入修剪决策。该方法基于“语义注意力机制一致性”的观察,采用前瞻策略预测深层token的显著性,防止过早移除关键token,从而在加速推理的同时保持性能稳定。

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning

第一作者: Kyujin Lee · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Vision-Language-Action (VLA) models have emerged as a promising framework that unifies perception, reasoning, and control for robot manipulation by adapting pretrained vision-language models (VLMs) to action prediction. However, VLM-derived representations are often insensitive to subtle visual distinctions required for low-level control, causing state aliasing between visually similar states that require substantially different actions. Prior VLA studies improve visual understanding by generating visual or reasoning outputs, such as future frames, 2D grounding points or traces, or intermediate spatial reasoning steps, but these objectives typically shape the vision encoder only indirectly through end-to-end prediction and do not explicitly analyze state aliasing in the learned visual feature space. To mitigate state aliasing, we introduce inverse dynamics learning as an...

论文介绍 视觉-语言-动作(VLA)模型在机器人操作中常面临状态混淆问题,即视觉相似但所需动作截然不同的状态难以区分。本文指出,现有方法主要通过间接的端到端预测来优化视觉理解,未能显式解决此问题。为此,作者引入逆动力学学习作为一种辅助目标,旨在显式地增强视觉编码器对细微视觉差异的敏感度,从而直接缓解状态混淆,提升低层控制的准确性。

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

第一作者: Xueqing Wu · 方向: 策略学习 · 来源: cs.CV

Abstract:Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a bottleneck for end-to-end visual reasoning. To investigate this gap, we introduce a controlled diagnostic framework with two synthetic tasks that disentangle perception from reasoning. Our analysis reveals a consistent perception-reasoning asymmetry: posttraining improves reasoning more substantially than perception, though the underlying mechanism differs by training paradigm. For supervised fine-tuning (SFT), this asymmetry stems from token imbalance in chain-of-thought supervision, where perception occupies fewer tokens and thus receives a weaker training signal. Dynamically reweighting the loss mitigates this imbalance and boosts end-to-end performance by up to 18.2. For reinforcement learning (RL), the asymmetry instead...

论文介绍 本文研究视觉语言模型后训练中感知能力提升滞后于推理能力的现象。作者通过一个可控的诊断框架分析发现,在监督微调(SFT)和强化学习(RL)两种范式下均存在感知-推理不对称性,但原因不同。对于SFT,不对称性源于思维链监督中的token不平衡;对于RL,则源于奖励信号对感知环节的间接影响。研究表明,动态调整损失权重等方法可有效缓解此不对称性。

AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference

第一作者: Yilin Feng · 方向: 多模态具身 · 来源: cs.LG

Abstract:Vision-Language Models (VLMs) process thousands of visual tokens per image alongside comparatively few text tokens, yet existing compression methods treat both modalities uniformly. We observe that the two modalities have fundamentally different properties: vision tokens are spatially redundant and dominate prefill, while text tokens are causally dependent and accumulate during decoding. Based on this asymmetry, we propose and empirically evaluate AsymVLM, which applies aggressive pruning to vision tokens before prefill using a learned importance scorer with per-sample adaptive budgeting, and temporal threshold-based eviction to text tokens only when they exceed a fixed budget. Our experiments indicate that AsymVLM achieves the highest FLOPs savings (up to 54%) among state-of-the-art methods while outperforming existing approaches by 2--3% on document and chart understanding...

论文介绍 现有视觉语言模型(VLM)的压缩方法通常均匀处理视觉和文本token,忽视了两种模态的根本差异。本文提出AsymVLM方法,基于观察到视觉token空间冗余且主导预填充、文本token因果依赖且在解码中累积的非对称特性,采用不同的压缩策略:对视觉token在预填充前进行激进修剪,对文本token则在超过固定预算时基于时间阈值进行淘汰。实验表明该方法在大幅节省计算量的同时保持了性能。

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

第一作者: Haoyuan Shi · 方向: VLA 通用模型 · 来源: cs.AI

Abstract:Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic framework that analyzes VLA models through a unified evidence chain from representation dynamics to causal control attribution and behavioral manifestation. It specifically combines cross-modal and checkpoint-drift centered kernel alignment (CKA) to trace representation evolution, attention knockout interventions to identify modality-specific control pathways, and rollout-level behavioral probes to examine grounding, shortcut dependence, and semantic following. Experiments on $\pi_{0.5}$ and OpenVLA reveal three key findings. First, the two models exhibit distinct modality-specific adaptation dynamics during VLA finetuning. Second, they rely on different multimodal routing strategies and layer-wise...

论文介绍 理解视觉-语言-动作(VLA)模型如何将多模态知识转化为控制信号仍是一个挑战。本文提出VLA-Trace,一个渐进式诊断框架,通过统一的证据链从表示动态到因果控制归因,再到行为表现来分析VLA模型。该框架结合了跨模态对齐追踪、注意力干预以及行为探测等方法,能够揭示模型在微调过程中模态适配的差异、多模态路由策略以及层间信息流等关键发现。

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

第一作者: Tianzhuo Yang · 方向: 数据集与评测 · 来源: cs.AI

Abstract:Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias...

论文介绍 现有的机器人世界模型评估基准大多关注视觉保真度,忽略了在给定动作条件下的预测可靠性。本文提出了MiraBench基准,将“动作条件可靠性”确立为核心评估目标。它将这一目标分解为三个层次:物理遵守性、动作跟随保真度以及乐观偏差。该基准旨在系统评估世界模型的预测是否物理合理、是否忠于指令动作,以及在动作失败时是否能正确校准,为更可靠的机器人模拟器评估提供框架。

市场总览

美股方面,标普500和纳斯达克ETF技术面偏强,SPY和QQQ的RSI分别达74.1和77.2处于超买状态,且呈多头排列,但接近52周高点可能面临回调压力。加密市场整体疲软,BTC-USD和ETH-USD趋势为bearish,RSI均低于40,结合加密恐慌贪婪指数28显示恐慌情绪,主导率57.3%但总市值24小时变化仅0.83%。中概股表现不佳,BABA、PDD、JD等普遍下跌,RSI多低于40,且出现MACD死叉和空头排列信号。商品外汇中,黄金期货RSI46.5中性,原油期货RSI37.9偏弱,美元指数接近52周高但RSI51.3中性,技术面分化。

今日关注

AAPL Apple (AAPL)
偏上行

当前价格312.06,RSI14高达78.8处于超买区域,接近52周高点仅-0.93%,且股价位于SMA20(297.54)、SMA50(275.28)、SMA200(263.24)之上形成多头排列,MACD值为10.3917高于信号线9.7772,显示短期上涨动量持续。

QQQ Nasdaq 100 ETF (QQQ)
偏上行

当前价格738.31,RSI14为77.2同样超买,MACD金叉(21.4696 > 21.4381),近5日涨幅3.33%,股价接近52周高点-0.45%,且均线呈多头排列,技术面显示强势特征。

BTC-USD Bitcoin (BTC-USD)
偏下行

当前价格74086.5,RSI14为39,趋势为bearish且信号为空头排列,价格低于SMA20(77172.57)、SMA50(77227.38)、SMA200(79664.67),MACD值-1005.6406低于信号线-451.2521,结合加密恐慌贪婪指数28(恐慌),技术面偏弱。

PDD 拼多多 (PDD)
偏下行

当前价格84.44,RSI14仅32.3,近5日暴跌13.65%,触发MACD死叉(-3.2037 < -1.7986)且均线呈空头排列,价格低于SMA20(95.79)、SMA50(98.42)、SMA200(113.23),显示下行压力显著。

^VIX VIX 恐慌指数
中性

当前价格15.32,RSI14为35.4处于正常范围,趋势为neutral,虽然MACD死叉(-0.9127 < -0.846)但价格位于SMA20(17.25)和SMA200(18.38)之间,近期波动下降(近5日-8.26%),技术面无明显方向。

全部资产

^VIX

VIX 恐慌指数

$15.32 -2.67%
5 日
-8.26%
距 52w 高
-56.6%
RSI(14)
35.4
趋势
中性
SMA 20 / 50 / 200
17.25 / 19.95 / 18.38
MACD / 信号
-0.913 / -0.846
MACD 死叉 (1 天前)

^TNX

10Y 美债收益率 (%)

$4.45 -0.04%
5 日
-2.90%
距 52w 高
-10.9%
RSI(14)
49.8
趋势
多头
SMA 20 / 50 / 200
4.48 / 4.39 / 4.20
MACD / 信号
0.040 / 0.055
MACD 死叉 (2 天前)接近 52 周低多头排列

DX-Y.NYB

美元指数 DXY

$98.94 -0.08%
5 日
-0.25%
距 52w 高
-1.7%
RSI(14)
51.3
趋势
多头
SMA 20 / 50 / 200
98.72 / 98.90 / 98.57
MACD / 信号
0.137 / 0.095
接近 52 周高多头排列

SPY

S&P 500 ETF

$756.48 +0.25%
5 日
+1.85%
距 52w 高
-0.2%
RSI(14)
74.1
趋势
多头
SMA 20 / 50 / 200
739.34 / 703.65 / 681.17
MACD / 信号
12.698 / 12.873
RSI 超买接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$738.31 +0.37%
5 日
+3.33%
距 52w 高
-0.4%
RSI(14)
77.2
趋势
多头
SMA 20 / 50 / 200
709.04 / 652.93 / 617.85
MACD / 信号
21.470 / 21.438
MACD 金叉 (今天)RSI 超买接近 52 周高多头排列

AAPL

Apple

$312.06 -0.14%
5 日
+2.32%
距 52w 高
-0.9%
RSI(14)
78.8
趋势
多头
SMA 20 / 50 / 200
297.54 / 275.28 / 263.24
MACD / 信号
10.392 / 9.777
RSI 超买接近 52 周高多头排列

MSFT

Microsoft

$450.24 +5.45%
5 日
+7.43%
距 52w 高
-18.9%
RSI(14)
69.6
趋势
中性
SMA 20 / 50 / 200
417.59 / 402.83 / 458.46
MACD / 信号
5.802 / 4.124
MACD 金叉 (今天)

NVDA

Nvidia

$211.14 -1.45%
5 日
-3.81%
距 52w 高
-10.7%
RSI(14)
49.4
趋势
多头
SMA 20 / 50 / 200
215.46 / 199.35 / 187.65
MACD / 信号
3.809 / 5.974
多头排列

GOOGL

Alphabet

$380.34 -2.51%
5 日
-1.89%
距 52w 高
-6.9%
RSI(14)
52.9
趋势
多头
SMA 20 / 50 / 200
391.15 / 347.57 / 299.92
MACD / 信号
9.641 / 13.543
多头排列

TSLA

Tesla

$435.79 -1.43%
5 日
+4.29%
距 52w 高
-12.6%
RSI(14)
60.0
趋势
中性
SMA 20 / 50 / 200
421.39 / 391.80 / 412.13
MACD / 信号
12.067 / 11.367
MACD 金叉 (2 天前)

META

Meta

$632.51 -0.44%
5 日
+4.14%
距 52w 高
-20.6%
RSI(14)
55.4
趋势
中性
SMA 20 / 50 / 200
613.33 / 618.53 / 666.57
MACD / 信号
-1.092 / -4.297
MACD 金叉 (2 天前)
加密恐慌贪婪
28
恐慌
加密总市值
$2.59 T
+0.83% / 24h
BTC 主导率
57.3%
ETH 9.5%
24h 成交量
$57.3 B
活跃币 17,403

BTC-USD

Bitcoin

$74,086.50 +0.97%
5 日
-4.13%
距 52w 高
-41.3%
RSI(14)
39.0
趋势
空头
SMA 20 / 50 / 200
77,172.57 / 77,227.38 / 79,664.67
MACD / 信号
-1,005.641 / -451.252
空头排列

ETH-USD

Ethereum

$2,028.46 +0.82%
5 日
-3.93%
距 52w 高
-59.1%
RSI(14)
34.4
趋势
空头
SMA 20 / 50 / 200
2,135.56 / 2,247.41 / 2,504.50
MACD / 信号
-65.426 / -54.934
空头排列

SOL-USD

Solana

$82.94 +1.23%
5 日
-2.44%
距 52w 高
-67.2%
RSI(14)
41.9
趋势
空头
SMA 20 / 50 / 200
86.58 / 86.42 / 104.80
MACD / 信号
-1.392 / -0.831
空头排列

BABA

阿里巴巴 (BABA)

$124.22 -1.54%
5 日
-5.51%
距 52w 高
-35.5%
RSI(14)
37.7
趋势
空头
SMA 20 / 50 / 200
134.18 / 131.07 / 149.62
MACD / 信号
-1.893 / -0.440
空头排列

PDD

拼多多 (PDD)

$84.44 +1.70%
5 日
-13.65%
距 52w 高
-39.4%
RSI(14)
32.3
趋势
空头
SMA 20 / 50 / 200
95.79 / 98.42 / 113.23
MACD / 信号
-3.204 / -1.799
MACD 死叉 (4 天前)空头排列

JD

京东 (JD)

$28.83 -1.06%
5 日
-8.39%
距 52w 高
-21.8%
RSI(14)
38.8
趋势
空头
SMA 20 / 50 / 200
30.88 / 30.03 / 30.44
MACD / 信号
-0.087 / 0.314
MACD 死叉 (4 天前)空头排列

0700.HK

腾讯控股 (0700.HK)

HK$427.20 +0.52%
5 日
-2.69%
距 52w 高
-37.5%
RSI(14)
31.0
趋势
空头
SMA 20 / 50 / 200
454.80 / 484.61 / 574.68
MACD / 信号
-16.238 / -14.448
接近 52 周低空头排列

GC=F

黄金期货

$4,593.00 +2.08%
5 日
+1.17%
距 52w 高
-17.8%
RSI(14)
46.5
趋势
中性
SMA 20 / 50 / 200
4,589.69 / 4,630.27 / 4,370.85
MACD / 信号
-51.494 / -47.184

CL=F

WTI 原油期货

$87.36 -1.73%
5 日
-9.33%
距 52w 高
-26.9%
RSI(14)
37.9
趋势
中性
SMA 20 / 50 / 200
98.51 / 97.86 / 72.04
MACD / 信号
-1.700 / 0.238

USDCNY=X

美元 / 人民币

¥6.77 -0.20%
5 日
-0.42%
距 52w 高
-6.2%
RSI(14)
30.6
趋势
空头
SMA 20 / 50 / 200
6.80 / 6.83 / 6.99
MACD / 信号
-0.014 / -0.013
MACD 死叉 (2 天前)接近 52 周低空头排列
风险提示

本报告基于公开行情数据计算的技术指标,仅供技术指标解读参考,不构成任何投资建议。过去走势不代表未来表现,市场存在不确定性,投资者应自行承担风险。

News live: Australia to buy only second-hand nuclear subs from US in major Aukus switch; Hanson says she could be PM

Meanwhile Marles tells Shangri-La Dialogue delegates that ‘seabed is a battlefield’. Follow updates live Get our breaking news email, free app or daily news podcast Extra negative gearing limits could hurt market and family budgets, Labor says Clare O’Neil has rejected calls from the Greens and othe

中文摘要 澳大利亚宣布在 Aukus 协议下仅从美国购买二手核潜艇,国防部长马尔斯在新加坡香格里拉对话会上警告海床已成为新战场。

Cuba’s blackouts leave high-rise residents with constant uncertainty

Will Grant spoke to a 70-year-old widow who says the inability to use her building's elevator during a power outage trapped her and her husband when he needed medical care.

中文摘要 古巴频繁停电导致高层居民生活不确定性增加,一名70岁寡妇反映停电时电梯故障,其丈夫急需医疗时被困。

The household battery revolution that could change energy bills … and the world

Australia is pioneering a revolution in home renewables and battery use, proving what is possible with the right policies The timing was rich with symbolism. As intense heatwaves pummelled Europe and Asia, and oil markets around the world leapt and sputtered, the two big chimneys of one of Australia

中文摘要 澳大利亚引领家庭可再生能源和电池使用革命,通过正确政策推动变革,旨在降低能源账单并应对气候变化。

How Putin became master of the image

From enigmatic KGB agent to wartime ruler, this is how Putin has repeatedly reinvented his image, and himself.

中文摘要 文章探讨俄罗斯总统普京如何从克格勃特工转变为战时统治者,不断重塑个人政治形象。

Edgar Morin, ‘Grandfather’ of French Intellectuals, Dies at 104

A former member of the Resistance, he went on to a career spanning eras and disciplines. His books and pronouncements carried moral authority.

中文摘要 法国知识分子埃德加·莫兰去世,享年104岁,曾参与抵抗运动,其跨学科著作具有广泛道德影响力。

US Congress advances American-Israeli military integration plan

A provision in the 2027 draft US defence bill could bind the two countries' weapons industries closer than ever.

中文摘要 美国国会推进美以军事一体化计划,2027年国防草案条款可能使两国武器工业合作更加紧密。

New Aukus drone subs to protect critical undersea cables as Marles warns: ‘seabed is a battlefield’

Minister at Singapore defence summit also reveals Australia to buy only secondhand Aukus submarines from US Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast The defence minister, Richard Marles, has said the “seabed is a battlefield”

中文摘要 澳大利亚国防部长马尔斯在新加坡防务峰会上宣布,Aukus 无人机潜艇将保护关键海底电缆,并警告海床已成为战场,同时透露仅采购美国二手 Aukus 潜艇。

What to Know About the Ebola Outbreak

Aid agencies are racing to help underequipped health workers in the Democratic Republic of Congo. More than 245 people are now suspected to have died from the virus.

中文摘要 埃博拉疫情在刚果民主共和国蔓延,援助机构正协助装备不足的卫生工作者,疑似死亡人数已超245人。

Lebanese army ‘overly stretched’ to fight off latest Israeli invasion

Geopolitical analyst Joe Macaron says the Lebanese army is ‘overly stretched’ as Israeli troops expand their occupation.

中文摘要 地缘政治分析师指出,以色列军队扩大占领导致黎巴嫩军队“过度紧张”,难以有效抵御入侵。

Ebola spread in DR Congo 'deeply alarming', MSF warns

The medical charity's comments come as the head of the World Health Organization visits the region worst-hit by the virus outbreak.

中文摘要 无国界医生组织警告刚果民主共和国埃博拉疫情扩散“深感担忧”,世界卫生组织总干事访问疫情重灾区。

Trump says a ceasefire extension deal with Iran is near, but core issues remain

President Trump has not yet decided whether he'll extend a ceasefire with Iran, and Israel continues to attack targets in Lebanon, in spite of a ceasefire there.

中文摘要 美国总统特朗普称与伊朗的停火延期协议接近达成,但核心问题未解;以色列继续攻击黎巴嫩目标,无视停火协议。

SpaceX, OpenAI Windfall Fuels Bets on Next-Wave Asian AI Winners

The hunt is on for companies that could benefit from the tailwinds of an unprecedented wave of stock offerings in the US, and investors are increasingly honing in on the Asian supply chain.

中文摘要 投资者正寻求能从美国股票发行潮中受益的亚洲公司,并将投资焦点放在亚洲的AI供应链上。

IMF Chief, Venezuelan Officials Hold Talks on Economic Stability

International Monetary Fund Managing Director Kristalina Georgieva and Venezuelan economic official Calixto Ortega held talks in Washington, the IMF head’s first in-person meeting with country’s authorities since the fund resumed formal engagement with Venezuela last month.

中文摘要 国际货币基金组织总裁格奥尔基耶娃与委内瑞拉经济官员在华盛顿举行会谈,这是自IMF上月恢复与委内瑞拉正式接触以来,双方的首次面对面会议。

Paris Saint-Germain retain Champions League with victory over Arsenal

French club triumph in penalty shootout over London side in win for Gulf sovereign wealth over US capital

中文摘要 巴黎圣日耳曼队在点球大战中击败阿森纳队,赢得欧冠冠军,被外界视为海湾主权财富基金对阵美国资本的胜利。

US Jobs Report Set to Reveal Solid Growth, Steady Unemployment Rate

Jobs week is coming up in the US, with a whole slate of indicators on the state of the labor market culminating on Friday with the government’s official report on employment for the month of May.

中文摘要 美国即将公布5月就业报告,预计将显示就业市场保持稳健增长,失业率维持稳定。

Brazil Extends Measures to Limit Fuel Price Hikes by Two Months

The Brazilian government extended by two months measures to contain the rise in fuel prices triggered by the conflict in Middle East.

中文摘要 巴西政府将限制燃料价格上涨的措施延长两个月,旨在应对中东冲突引发的价格飙升。

Paramount Is Pulling Every Lever to Sell LBO Debt

Paramount Skydance Corp. stretched, then stretched, then stretched again in its audacious $110 billion takeover bid for Warner Bros. Discovery Inc.

中文摘要 Paramount Skydance Corp. 为完成对华纳兄弟探索公司1100亿美元的收购要约,正一再放宽杠杆收购债务的销售条件。

Bloomberg This Weekend 5/30/2026

The news doesn’t stop when markets close. Hosts David Gura, Christina Ruffini and Lisa Mateo bring clarity, context and a bit of humor to the weekend’s biggest headlines, LIVE from New York. Joined by The Associated Press International Correspondent Philip Crowther, Fmr. CDC Director Dr. Tom Frieden

中文摘要 彭博社周末节目在纽约直播,梳理并讨论周末期间的主要财经新闻头条。

At Scrabble’s Biggest Tournament, Thailand Punches Above Its Weight

Known locally as Crossword, Scrabble took off in Thailand in the 1980s after schools embraced it as a teaching aid and pathway to advancement. Much like athletic scholarships in the US, strong players can secure financial assistance and preferential admission to top universities, creating a pipeline

中文摘要 在拼字游戏最大锦标赛中,泰国队表现突出。该游戏在泰国因被学校用作教学工具而普及,优秀玩家可获得大学奖学金。

Can AI Grow Without Hurting Local Communities?

The rush to build AI data centers is drawing trillions of dollars in investment and long-term bets from infrastructure firms such as DigitalBridge. But in communities where those facilities are being built, residents and officials are raising concerns about electricity costs, water use, noise, trans

中文摘要 建设AI数据中心吸引数万亿美元投资,但当地社区居民和官员对由此带来的水电成本增加等问题表示担忧。

opus-4.8 怎么能难用成这个样子

根本没有办法跟 4.6 比,跑个任务罗里吧嗦奇奇怪怪的,还有各种语言表达,怎么能 der 成这样? 4.7 的"稳稳的接住你"就不说了。 4.8的"侦查清楚了"(怎么,去看个服务器去敌后侦查了是么?): 还有:“可还是被oom 打爆了”(啊OOM 好厉害哦,都打爆 swap 了啦): 还有:“决定性结论出来了”(咋的,上面那些废话自己也承认是非决定结论?): 还有决定性结论刚说完,下面又来了一个:“真正的结论”(我请问呢??你孙杨吗?): 决定性结论完了还有:“决定性证据”(哦,医生上线了): 还有:“它好好活着”“我刚才误报了” 还有非常非常非常的多,我真的懒得截图了。 佬友们用吧,反正我是

分享今天的西郊公园

由于最近才加入L站,还不能频繁回复。在此统一感谢喜欢照片的各位佬友,相机是富士中画幅gfx100s,镜头是GF20-35是超广角变焦镜头,等效全画幅的16-28mm,这个比例是传统的胶片相机中的XPAN画幅,在中画幅富士机身上有一个内置的65:24可以实现机内裁切出这比例的JPG(当然也可以后期裁切),照片的故事感一方面也可能来自于这一较为“陌生”的比例。今天去的时候其实比较阴凉,反而是蚊虫较多。佬友们去逛的话,可以从西郊宾馆的南门进入。 17 个帖子 - 16 位参与者 阅读完整话题

24OpenClaw 现在能爬几乎任何网

24OpenClaw 现在能爬几乎任何网站,关键是——零反爬检测,原生绕过 Cloudflare,速度比 BeautifulSoup 快 774 倍。 ① 不用维护选择器 这种降维打击级别的工具,还完全开源,不用白不用。 github.com GitHub - D4Vinci/Scrapling: 🕷️ An adaptive Web Scraping framework that handles... 🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-

Digitalfyre LAX Ryzen 9950X产品 测评: 极致的性价比和性能建站机

产品名字 本次测试DigitalFyre的LAX产品,定位是性能建站型机器。 测试配置为 S2 8 CPU 8 GB RAM 8 TB Bandwidth 65 GB NVme $10.00/mo 本文无AFF,更多相关产品可看原文 网络质量 国内方向三网勉强可以直连(SSH),丢包还是比较严重的,网络波动较大,不推荐直连使用,推荐作为纯落地使用。这个机器的国际方向带宽给的相当大,10Gbps峰值,在测试中单线程轻松拉到了5Gbps,重传低速度稳定,非常不错的国际网络表现,配置中也给到了8TB的流量,搭配10Gbps口子可以说是非常合适了。 IP质量 IP质量中规中矩,IPV4+IPV6双栈原

【CHY公益站】今天最后一波补充额度

这次是今天的最后一波补充额度,100*1,0000,0000,0000额度 cdk.linux.do LINUX DO CDK Linux Do 社区 CDK 快速分享平台 - 让分享变得更简单 42 个帖子 - 36 位参与者 阅读完整话题

今天居然用玉米解锁了条46斤的鱤鱼(有图有真相)

昨天师兄说自己窝子出问题了,喊我和师傅过去看看,然后昨天晚上我和师傅过去了只钓了两条半斤的红尾,打了30斤自制玉米,今天过去钓上午又没口,下午继续钓玉米,突然黑漂拖杆,师傅猛然发力补刺,说是草鱼来了,溜了十分钟发现居然上的是条46斤的鱤鱼,上称称了3次,换了3把称(用的2号尼龙线,3号千又) 补一张仰天大笑图,哈哈哈哈哈哈 13 个帖子 - 7 位参与者 阅读完整话题

千万不要100%信任AI!!!

今天差点闯祸,幸亏没有where条件,数据库给挡了一下。 需求: 参考数据库内容 更新 json文件 AI给出的结果: 根据 json文件 更新 数据库 平时AI表现良好、现在懒得检查AI写的每一行代码, 先上手执行看是否报错。 俺错了,千万不要100%信任AI,要先检查一遍代码。 42 个帖子 - 26 位参与者 阅读完整话题