每日简报

2026-05-30

← 历史归档

harry0703/MoneyPrinterTurbo

Python · ★ 69,671 · 🍴 10,045 · 📈 3,567 stars today

利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.

中文介绍 一款基于AI大模型的短视频生成工具,能够一键自动生成高清短视频。它利用AI大模型理解需求并驱动内容生成,旨在简化视频创作流程。主要面向内容创作者、营销人员及需要快速批量生产短视频的用户。

microsoft/markitdown

Python · ★ 129,959 · 🍴 8,916 · 📈 1,873 stars today

Python tool for converting files and office documents to Markdown.

中文介绍 微软开源的一款Python工具,专注于将各类文件和Office文档转换为Markdown格式。它能自动处理文档解析和格式转换,解决文档格式不统一的问题。适用于需要将报告、幻灯片等内容快速集成到Markdown工作流中的开发者和技术写作者。

EveryInc/compound-engineering-plugin

TypeScript · ★ 18,134 · 🍴 1,377 · 📈 353 stars today

Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more

中文介绍 Compound Engineering官方插件,为Claude Code、Codex、Cursor等AI编码工具提供增强功能。它通过插件规范扩展了这些工具的能力,提升了AI辅助编码的效率和灵活性。主要用户是使用上述AI编程工具的开发者,用于优化他们的开发体验。

twentyhq/twenty

TypeScript · ★ 48,408 · 🍴 6,863 · 📈 578 stars today

The open alternative to Salesforce, designed for AI.

中文介绍 一款专为AI时代设计的开源客户关系管理(CRM)系统,旨在成为Salesforce的开源替代品。它将AI能力深度集成,帮助用户更智能地管理客户交互和业务流程。适用于寻求灵活、可控且具备AI功能的CRM解决方案的中小型团队。

anthropics/claude-code

Python · ★ 127,874 · 🍴 20,910 · 📈 395 stars today

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

中文介绍 Anthropic推出的智能编码助手,以终端工具的形式存在。它能理解代码库,通过执行常规任务、解释复杂代码和管理Git工作流来帮助开发者提速编码。面向希望在命令行环境中获得强大AI编程辅助的开发者。

Leonxlnx/taste-skill

Shell · ★ 28,132 · 🍴 2,085 · 📈 2,062 stars today

Taste-Skill - gives your AI good taste. stops the AI from generating boring, generic slop

中文介绍 一个旨在提升AI生成内容品味的“技能”模块,通过特定指令或方法,防止AI输出内容陷入无聊、通用的套路。它旨在提高生成文本的原创性和质量。适用于对AI输出内容有高要求,希望去除“机器味”的创作者或内容生产流程。

cursor/plugins

TypeScript · ★ 1,298 · 🍴 111 · 📈 134 stars today

Cursor plugin specification and official plugins

中文介绍 AI编程助手Cursor的官方插件仓库,提供了插件开发规范和官方示例。它定义了扩展Cursor功能的接口,允许社区开发者创建插件来增强其代码编辑和AI交互能力。目标用户是Cursor的开发者和高级用户。

run-llama/liteparse

Rust · ★ 7,301 · 🍴 445 · 📈 701 stars today

A fast, helpful, and open-source document parser

中文介绍 一款快速、实用且开源的文档解析器。它能高效地从各类文档中提取结构化信息,帮助开发者快速获取文档内容用于后续处理。适用于需要进行文档信息抽取、构建知识库或进行数据预处理的技术场景。

galilai-group/stable-worldmodel

Python · ★ 1,252 · 🍴 148 · 📈 362 stars today

A platform for reproducible world model research and evaluation

中文介绍 一个用于世界模型可复现研究与评估的平台。它为研究人员提供标准化的工具和基准,以在可控环境下构建、测试和比较不同的世界模型。主要面向从事人工智能,特别是具身智能、机器人学习和预测模型研究的学者与工程师。

byoungd/English-level-up-tips

★ 49,634 · 🍴 5,199 · 📈 1,566 stars today

An advanced guide to learn English which might benefit you a lot 🎉 . 离谱的英语学习指南/英语学习教程/英语学习/学英语

中文介绍 一份进阶英语学习指南,汇集了众多高效学习英语的方法、资源和经验。它超越基础教程,提供系统化的提升策略,旨在帮助学习者突破瓶颈。适用于已有一定英语基础,希望系统化提升综合能力的学习者。

Biohub/esm

Jupyter Notebook · ★ 2,566 · 🍴 314 · 📈 52 stars today

中文介绍 信息有限,根据仓库名推测可能与生物信息学或蛋白质语言模型(如ESM)相关。它可能是一个相关工具、数据集或研究代码的集合。具体功能和目标用户需参考项目内部文档进一步了解。

Crosstalk-Solutions/project-nomad

TypeScript · ★ 26,964 · 🍴 2,657 · 📈 318 stars today

Project N.O.M.A.D, is a self-contained, offline survival computer packed with critical tools, knowledge, and AI to keep you informed and empowered—anytime, anywhere.

中文介绍 Project N.O.M.A.D. 是一个自包含、离线的生存计算机系统。它集成了关键工具、知识库和AI助手,旨在任何地点、任何时间为用户提供信息支持和决策辅助。适用于户外探险者、生存主义者或需要在无网络环境下进行知识查询和决策的场景。

DigitalPlatDev/FreeDomain

HTML · ★ 171,797 · 🍴 3,345 · 📈 1,313 stars today

DigitalPlat FreeDomain: Free Domain For Everyone

中文介绍 DigitalPlat FreeDomain项目旨在为所有人提供免费域名。它简化了域名申请流程,降低了个人或小型项目获取网络地址的门槛。适用于个人开发者、学生项目或任何需要低成本在线身份的初创者。

affaan-m/ECC

JavaScript · ★ 198,585 · 🍴 30,514 · 📈 1,406 stars today

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

中文介绍 一个针对AI编程代理的性能优化系统,集成技能、本能、记忆、安全性和研究优先的开发理念。它支持Claude Code、Codex等工具,旨在提升代理的效率和可靠性。面向深度使用并希望优化这些AI编码工具性能的高级开发者。

hardikpandya/stop-slop

★ 6,998 · 🍴 491 · 📈 617 stars today

A skill file for removing AI tells from prose

中文介绍 一个技能文件,专门用于去除AI生成文本中常见的、生硬的写作痕迹,使文风更自然、更人性化。它通过提供指令或后处理规则来“打磨”AI输出。适用于需要AI辅助写作但追求自然文风的内容创作者、编辑或营销人员。

DataTalksClub/data-engineering-zoomcamp

Jupyter Notebook · ★ 41,564 · 🍴 8,271 · 📈 160 stars today

Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. The next cohort starts in January 2026. Join the course here 👇🏼

中文介绍 DataTalksClub推出的免费数据工程课程,为期9周,核心内容是构建生产级别的数据管道。它以在线训练营形式提供系统化学习和实践项目。面向希望系统学习现代数据工程技术栈,如数据建模、作业编排和数据仓库的初学者和中级工程师。

codecrafters-io/build-your-own-x

Markdown · ★ 507,363 · 🍴 48,172 · 📈 866 stars today

Master programming by recreating your favorite technologies from scratch.

中文介绍 一个通过“从头构建”你喜爱技术来掌握编程的实践项目集合。它提供指南,引导学习者动手实现如数据库、操作系统、编译器等复杂系统,以深入理解底层原理。适合希望通过高强度实践来深化理解特定技术领域或计算机基础的开发者。

该源今日无内容。

Boston Children’s uses AI to unlock new diagnoses

Boston Children’s Hospital uses OpenAI technology to improve patient care, reduce operational burden, and help diagnose more than 40 rare disease cases.

中文介绍 波士顿儿童医院采用OpenAI技术提升患者护理质量,减轻运营负担,并已利用该技术帮助诊断超过40例罕见疾病病例。

How Braintrust turns customer requests into code with Codex

How Braintrust engineers use Codex with GPT-5.5 to run experiments and code faster.

中文介绍 Braintrust的工程师使用配备GPT-5.5的Codex来运行实验和加快编程速度,展示了如何将客户需求转化为代码。

How the Pope’s Magnifica Humanitas offers a template for individuals to meet the AI moment

Pope Leo XIV’s new encyclical on artificial intelligence includes a statement that warrants serious attention from technologists and policymakers: “Technology is never neutral.” Magnifica Humanitas (“Magnificent Humanity”) is a clarion call to all people to act with courage and solidarity as we ente

中文介绍 教皇利奥十四世关于人工智能的新通谕《非凡的人性》发出呼吁,其中包含“技术从不中立”的声明,呼吁技术人员与政策制定者深思。

not much happened today

**Anthropic** rolled out **Claude Opus 4.8**, which shows incremental improvements but mixed benchmark results, including better cooperation and coding behavior but some regressions in document parsing. Platform updates include mid-conversation system instructions enhancing long agent sessions, thou

中文介绍 Anthropic发布了Claude Opus 4.8模型,该版本在协作和编程行为上有所改进,但在文档解析等方面存在回退,属于渐进式更新。

Strengthening societal resilience with Rosalind Biodefense

OpenAI launches Rosalind Biodefense, expanding trusted access to GPT-Rosalind for vetted developers and U.S. government partners advancing biodefense, public health, and pandemic preparedness through frontier AI.

中文介绍 OpenAI推出Rosalind生物防御项目,将GPT-Rosalind的受信访问权限扩展至经审核的开发者及美国政府合作伙伴,以推进生物防御和公共卫生准备。

A shared playbook for trustworthy third party evaluations

OpenAI shares guidance on third-party AI evaluations, covering how to assess model capabilities, safeguards, and validity for frontier systems.

中文介绍 OpenAI分享了关于第三方AI评估的指南,内容涵盖如何评估前沿系统的模型能力、保障措施及评估有效性。

The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray

80% Devin Commits, Spec-to-PR Workflows, Full VMs, Agent Memory, and PMs Shipping Code

中文介绍 文章探讨了Cognition公司创始人Walden Yan与OpenInspect的Cole Murray关于异步智能体的讨论,涉及80%的Devin提交、规格到PR的工作流及智能体记忆等话题。

How Endava builds an agentic organization with Codex

Learn how Endava uses Codex to build an agentic organization, accelerating software delivery and reducing requirements analysis from weeks to hours.

中文介绍 Endava公司利用Codex构建智能体组织,以加速软件交付,将需求分析所需时间从数周缩短至数小时。

The AI Hype Index: AI gets booed in graduation season

It is one thing to say AI will change the world. It is another to expect the class of 2026 to applaud it. In fact, when former Google CEO Eric Schmidt told University of Arizona graduates that their task is to help shape AI, he was met with a resounding chorus of boos. “I can…

中文介绍 在毕业季,AI遭到抵制。前谷歌CEO埃里克·施密特在亚利桑那大学向毕业生表示,他们的任务是帮助塑造AI,但遭到一片嘘声。

[AINews] Cognition raises $1B in $26B Series D

coding is an uncapped TAM market

中文介绍 Cognition公司在260亿美元估值的D轮融资中筹集了10亿美元,其观点认为编码是一个上限无上限的市场。

Anthropic raises $65B in Series H at a $965B post-money valuation, releases Opus 4.8 and Dynamic Workflows

**Anthropic** announced a massive **$65B Series H financing** at a **$965B valuation**, led by **Altimeter, Dragoneer, Greenoaks, and Sequoia**, with run-rate revenue surpassing **$47B**. They launched **Claude Opus 4.8**, an update to Opus 4.7 featuring "sharper judgment," "more honesty," and longe

中文介绍 Anthropic宣布完成650亿美元的H轮融资,投后估值达9650亿美元,由Altimeter、Dragoneer、Greenoaks和Sequoia领投,其年化收入已超过470亿美元,并发布了Claude Opus 4.8模型。

OpenAI’s Frontier Governance Framework

Explore OpenAI’s Frontier Governance Framework and how our AI safety, security, and risk practices align with emerging EU and California regulations.

中文介绍 OpenAI介绍了其前沿治理框架,阐述了其在AI安全、安保和风险管理方面的实践如何与欧盟及加州的新兴法规保持一致。

每日论文 · arXiv cs.CR 最新公告批次

周末 arXiv 通常无新公告。当前展示最近一次可用公告批次。

DP-SAPF: Saliency-Aware Parameter Fine-tuning of Public Models for Differentially Private Image Synthesis

第一作者: Chen Gong · 方向: 隐私保护

Abstract:Differentially private (DP) image synthesis generates images that preserve the statistical characteristics of a sensitive dataset, enabling sensitive data analysis and usage while providing rigorous guarantees of privacy leakage. Existing methods fine-tune public models using DP Stochastic Gradient Descent (DP-SGD) on sensitive images to generate synthetic images. But full fine-tuning public models on sensitive images is computationally expensive, because current public models typically contain a large number of parameters. Recent work proposes heuristically using Low-Rank Adaptation (LoRA) on all attention-layer parameters of public models to reduce the number of trainable parameters. However, we argue that exhaustive LoRA coverage across all attention-layer parameters is suboptimal in a DP setting, as it leads to noise accumulation and collapse during private training. To...

论文介绍 研究差分隐私图像合成中公共模型全微调的计算成本问题。现有方法使用LoRA减少参数,但在差分隐私设置下导致噪声累积和性能下降。提出DP-SAPF,一种基于显著性的参数微调方法,优化LoRA在隐私保护下的使用,可能应用于高效且隐私安全的图像生成。

bpK#: Delegatable Pseudonyms And Their Applications to National eID Systems

第一作者: Stephan Krenn · 方向: 系统安全

Abstract:Electronic identities (eIDs) are crucial in an increasingly digitalized environment. Pseudonyms, as offered by Austria's governmental sector-specific personal identifiers (bPks), can significantly improve privacy by ensuring that personal data is not universally traceable across public services and private companies. However, the current architecture comes with several challenges regarding availability, privacy, and authenticity, due to a fully centralized design. This paper proposes bPk#, a distributed architecture to address these issues, reducing reliance on the central authority, while still providing all functional requirements to the existing bPk system. In particular, users are delegated the rights to compute their own pseudonyms, thereby minimizing metadata revealed to the central authority, while (subsets of) service providers may receive the right to compute...

论文介绍 针对国家电子身份系统中假名的集中设计挑战,如可用性、隐私和真实性问题。提出bPk#分布式架构,允许用户委托计算假名,减少中央权威的元数据泄露,同时保持功能需求,可应用于改进电子身份系统的隐私和可靠性。

A Bayesian Approach to Membership Inference for Statistical Release

第一作者: Lisa Oakley · 方向: 隐私保护

Abstract:The membership inference problem for publicly released statistics from a private dataset is well-studied. When developing and formally analyzing attack strategies, however, the focus has been on attacks that model the population using only its marginals. In practice, these attacks can perform well on various populations, however most formal analysis is for populations that follow a product distribution. These strategies may fail to leverage useful information about the population that is important for understanding a realistic privacy threat. In this work, we explore the impact of providing an attacker with additional information about the attribute dependency structure of the population, motivated by examples where multiple parties may have access to similarly structured data, for example the US Census and the IRS. To model this scenario, we re-frame the membership inference...

论文介绍 研究成员推断攻击在公开统计数据中的隐私威胁。现有攻击基于边际模型,忽略人口属性依赖结构。采用贝叶斯方法建模依赖关系,增强攻击准确性,可能用于评估统计数据(如人口普查)的隐私风险。

Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection

第一作者: Travis Lelle · 方向: 安全研究

Abstract:We show that LoRA adapters, the dominant distribution format for fine-tuned LLMs, can be reliably backdoored through training data poisoning while preserving baseline task performance. On a Qwen 2.5 1.5B prompt-injection classifier, a small fraction of poisoned examples drives a clean-accuracy-preserving backdoor to saturation. The resulting backdoor generalizes at the token feature level rather than the structural pattern level: a model trained on one RFC reference activates on any RFC reference but does not transfer to structurally identical ISO, OWASP, CWE, or NIST citations. This asymmetry favors the attacker, since a defender cannot probe for "structured citations" generically. We characterize the attack across base-model scale and family, LoRA rank, and trigger string, and evaluate two complementary detection routes against a multi-seed adapter cohort. A behavioral...

论文介绍 探讨LoRA适配器中后门攻击的令牌级泛化特性。通过数据投毒植入后门,攻击在令牌特征级别泛化,而非结构模式。研究攻击特征并提出行为检测路线,可能用于加强LLM适配器的安全性。

Privacy-Enhanced Zero-Order Federated Learning via xMK-CKKS over Wireless Channels

第一作者: Anthony Ayli · 方向: 密码学协议

Abstract:Homomorphic encryption (HE) enables privacy-preserving aggregation in federated learning (FL) by allowing the server to operate on encrypted data without decryption. Existing HE-over-the-air methods mainly rely on single-key HE schemes and require channel estimation or pre-equalization to compensate for wireless fading. However, single-key HE remains vulnerable to honest-but-curious clients sharing the same secret key. In addition, compromising a single client may compromise the security of the entire network, while multi-key HE schemes provide stronger client-level security by assigning each device its own secret key. We propose a four-phase protocol that enables xMK-CKKS, a famous multi-key HE scheme, aggregation over a shared wireless channel without channel estimation. The protocol retransmits partial public keys and ciphertexts through the same channel realization, so...

论文介绍 解决联邦学习中同态加密的单密钥方案安全性和信道估计问题。提出基于xMK-CKKS多密钥HE的协议,在无线信道上实现隐私增强聚合,无需信道估计,可能用于安全无线联邦学习系统。

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

第一作者: Galip Tolga Erdem · 方向: AI 安全

Abstract:Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials remains unstudied. This work presents the first large-scale empirical measurement of LLM attack consistency: 400 autonomous penetration testing runs (4 models, 100 each) against an identical honeypot hosting OWASP Juice Shop and two additional vulnerable services, holding prompt, orchestrator, and target constant. No model emitted a content refusal that survived the orchestrator's one-shot authorization re-prompt at iterations 0-1. Claude Sonnet 4's API calls did encounter upstream service unavailability - 91 of 1,135 calls returned HTTP 529 overloaded_error during a documented Anthropic capacity event, truncating 39 of 100 Claude runs. An earlier draft catalogued these as safety refusals; on full-log audit they are upstream API...

论文介绍 实证研究LLM作为网络攻击者的行为一致性。通过400次自主渗透测试实验,测量模型在重复试验中的攻击可靠性,发现差异和不一致性,可能用于评估AI攻击的风险和可信度。

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

第一作者: Shahinul Hoque · 方向: AI 安全

Abstract:Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affects what users pay. We show that this kind of billing is hard to audit by design: providers hide the model, the tokenizer, and the execution to protect their IP, mitigate jailbreaks, and preserve user privacy, which means an auditor can only inspect proofs the provider supplies. The audit therefore reduces to a consistency check on the provider's own reports. We call this a trust paradox: every audit must trust some artifact, but current frameworks trust exactly the ones a provider has the strongest reason to manipulate. We study three recent token auditing frameworks and show that a provider with ordinary commercial capabilities can systematically inflate billed token counts. In the most permissive setting, hidden reasoning...

论文介绍 分析LLM按令牌计费模型中的信任悖论。提供商可能虚报令牌计费,现有审计框架易受攻击。展示如何通过隐藏推理系统膨胀计费,可能影响LLM服务的定价透明度和用户信任。

Fingerprinting Inference Systems of Large Language Models

第一作者: Anna Wimbauer · 方向: 系统安全

Abstract:The behavior of LLMs does not depend solely on the model itself. Components of the inference system, such as the inference engine, attention backend, and hardware platform, subtly influence how inputs are processed. These components differ in their implementations and thereby induce small numerical deviations across systems when running the same model. While prior work has established the theoretical existence of such deviations, their security implications have remained unexplored. In this paper, we show that these deviations are characteristic of specific components and propagate to observable textual outputs, exposing the inference system to any party that can query the model. Building on this observation, we introduce a fingerprinting method that analyzes the prompt-response behavior of LLMs to identify components of the inference system. Our empirical evaluation...

论文介绍 研究LLM推理系统组件差异对输出的影响。这些差异可用于指纹识别推理系统组件。提出基于查询行为的指纹方法,可能用于检测或审计推理系统配置,增强安全性。

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

第一作者: Mark Vero · 方向: AI 安全

Abstract:Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation backbones for honeypots. They enable defenders to construct high-interaction honeypots with low system security risks. However, LLM-powered honeypot development lacks a unified evaluation framework. Most evaluations consist of measuring response similarity on fixed commands, manual testing, or real-world deployment. These methods are often not scalable for development, reproducible across evaluations, representative of practical attacks, or adaptable to various attacker and honeypot configurations. In this work, we bridge this gap and propose Honeyval, a comprehensive evaluation framework for LLM-powered HTTP honeypots. We address the limitations of prior evaluations by grounding the honeypots in 16 backend applications, using...

论文介绍 该研究旨在解决LLM驱动HTTP蜜罐缺乏统一评估框架的问题。现有评估方法在可扩展性、可重复性和攻击代表性方面存在不足。为此,论文提出了一个名为Honeyval的综合评估框架。该框架将蜜罐置于16个后端应用和多种攻击向量中进行测试,旨在为LLM蜜罐的开发和评估提供标准化的基准。

Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

第一作者: Hongtao Wang · 方向: AI 安全

Abstract:Large language model (LLM) agents increasingly leverage long term memory to support persistent and autonomous task execution. However, this capability also introduces a new attack surface: memory poisoning, where adversaries can inject malicious information to influence future behavior. Existing memory poisoning attacks often assume that injected content can be stored directly in memory, overlooking the selective extraction and rewriting stages in modern memory pipelines. This makes prior methods ineffective under realistic settings. In this paper, we propose MemPoison, a novel memory poisoning attack that bypasses selective memory mechanisms in LLM agents, where an attacker can inject triggerable backdoors into the agent's long-term memory through dialogue interactions, thereby misleading its subsequent responses. MemPoison introduces three key components: (i) a semantic...

论文介绍 大型语言模型智能体依赖长期记忆执行任务,但这也引入了记忆投毒的攻击面。现有攻击方法在现实选择性记忆提取机制下效果不佳。本文提出了一种名为MemPoison的新型攻击,它通过对话交互向智能体长期记忆注入可触发后门,能绕过选择性记忆机制并误导后续响应,揭示了LLM智能体记忆系统的安全隐患。

Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection

第一作者: Syafiq Al Atiiq · 方向: 软件安全

Abstract:Large language models (LLMs) can detect software vulnerabilities, but how do they actually identify vulnerable code? We address this question using mechanistic interpretability; analyzing the internal computations of a neural network to understand its reasoning this http URL Circuit Tracer on Gemma-2-2b, we trace the computational pathways activated when the model classifies 472 C/C++ code samples as vulnerable or safe. Our analysis reveals a surprising finding: the model primarily relies on safety detectors, attention heads that recognize safe coding patterns, rather than directly detecting vulnerability signatures. When these safety detectors fail to activate, the model classifies code as vulnerable. We identify the critical neural components: specific attention heads in early layers (L5, L7) that focus on safety patterns, and Multilayer Perceptron (MLP) neurons in Layer 7...

论文介绍 本文探究大型语言模型如何检测软件漏洞的内部机制。研究者采用机制可解释性方法,对Gemma-2-2b模型分析472个C/C++代码样本的分类过程。分析发现,模型主要依赖识别安全编码模式的“安全检测器”注意力头,而非直接检测漏洞特征。当这些安全检测器未被激活时,模型便将代码判定为存在漏洞。

Ciphera: A Decentralised Biometric Identity Framework

第一作者: Ankit Kanaiyalal Prajapati · 方向: 系统安全

Abstract:Centralised biometric identity systems expose users to single points of failure, opaque verification processes, and irreversible biometric compromise. Decentralised Identifiers (DIDs) and Verifiable Credentials (VCs) offer stronger privacy guarantees, yet their integration with biometric authentication and distributed verification remains insufficiently explored. This paper presents Ciphera, a decentralised biometric identity framework combining privacy-preserving facial recognition, multi-node verification, IPFS-based credential metadata storage, and blockchain-anchored revocation. Evaluated across functional, performance, security, and distributed consistency dimensions, Ciphera achieved an 81% functional success rate, with stable enrolment and authentication but measurable revocation propagation delays and occasional audit-log inconsistencies. Performance testing...

论文介绍 集中式生物识别身份系统存在单点故障和隐私泄露风险。本文提出了Ciphera框架,这是一个结合了隐私保护面部识别、多节点验证、基于IPFS的凭证存储和区块链撤销机制的去中心化生物识别身份系统。评估表明,该框架在功能、安全性和分布式一致性方面达到了81%的成功率,为生物识别身份管理提供了更安全的去中心化方案。

Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing

第一作者: Leyi Qi · 方向: 安全研究

Abstract:Large-scale text-to-image (T2I) diffusion models have enabled unprecedented creative applications, but their unauthorized use has raised serious intellectual property concerns, making model ownership verification (MOV) increasingly critical. We find that existing backdoor-based diffusion watermarking methods often (implicitly) assume a "faithful" verification process, namely, that the verifier can query a suspicious model and obtain the faithful watermark response to complete MOV. However, in practice, adversaries may intentionally or unintentionally damage potential watermark signals, significantly degrading verification reliability. To address this issue, we propose Cert-LAS, the first certified MOV method for T2I models based on layer-adaptive smoothing. In general, Cert-LAS embeds specified watermarks using diffusion classifiers and an LFS-guided layer-adaptive noise, and...

论文介绍 针对文本到图像扩散模型面临的知识产权侵权问题,现有基于后门的水印方法在验证过程中可能被破坏。本文提出了首个经过认证的所有权验证方法Cert-LAS。该方法利用扩散分类器和层自适应噪声嵌入特定水印,旨在即使在验证过程被恶意干扰的情况下,也能提供可靠的模型所有权证明。

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

第一作者: Alexander Sternfeld · 方向: 软件安全

Abstract:LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code these agents produce, the security of that code becomes critical. Prior work has shown that minor prompt perturbations degrade the functional correctness of LLM-generated code, but whether they also compromise code security has remained unstudied. We apply token-level mutations to prompts across three models and five programming languages, and show that mutations as small as a single-character change can flip generated code from secure to vulnerable. Probing the models' hidden states reveals that this fragility is partially encoded in prompt representations, but unevenly so. Input-handling vulnerabilities, where the model omits validation or sanitization, are more predictable (mean AUC 0.753) than secure-defaults vulnerabilities...

论文介绍 研究了轻微的提示扰动是否会损害LLM生成代码的安全性。通过对提示进行token级突变,发现单个字符的改变就可能使生成的代码从安全变为存在漏洞。对模型隐藏状态的探测表明,这种脆弱性在提示表征中有所编码。研究特别发现,与输入处理相关的漏洞比与安全默认值相关的漏洞更具可预测性。

FIDEM: A Standard-Compliant Framework for Secure Binding of MUD Profiles to IoT Devices

第一作者: Alessandro Lotto · 方向: 网络安全

Abstract:The Manufacturer Usage Description (MUD) standard enables enforcement of network restrictions for IoT devices based on their expected network traffic, as specified by manufacturers in an online MUD file. Devices advertise a URL pointing to this file, yet the standard does not define how to securely bind the issuing device to its profile. As a result, malicious devices can manipulate network policy enforcement by advertising valid URLs referencing genuine MUD profiles, but not intended for that device. Although MUD defines a certificate-based secure issuance method, current deployments rely on the insecure DHCP-based extension due to simpler integration. Existing solutions either depend on Public Key Infrastructure (PKI), break standard compliance, require excessive active manufacturer involvement, or overlook secure profile updates. In this paper, we present FIDEM, a...

论文介绍 MUD标准允许基于制造商描述的网络策略限制物联网设备行为,但标准未定义如何将设备与其声称的MUD配置文件安全绑定。恶意设备可能冒用合法配置文件。本文提出了FIDEM框架,它是一个标准兼容的框架,通过扩展证书和区块链技术,为MUD配置文件与IoT设备之间建立安全的密码学绑定,以增强网络策略执行的可靠性。

Scarcity Is Not Enough: An Impossibility Result for Linear Sybil Cost Under Parallelizable Resources

第一作者: Homayoun Maleki · 方向: 密码学协议

Abstract:Permissionless systems resist Sybil attacks by binding influence to scarce resources. We show that scarcity alone is insufficient: the structural properties of the resource determine whether influence can be concentrated at sublinear cost through identity replication, delegation, or pooling. We model this through the adversarial cost C(s,T): the minimum expenditure required to achieve influence proportional to s independent participation units over T windows. We prove that any resource satisfying divisibility, additivity of influence, temporal reusability, and identity transferability admits influence amortization: C(s,T)=o(sT), regardless of protocol design. This is an impossibility result: no protocol rule can enforce linear cost of influence concentration over a structurally parallelizable resource. We further prove that throughput-bounded, non-transferable, window-local...

论文介绍 无许可系统通常通过绑定稀缺资源来抵御Sybil攻击。本文证明,仅靠稀缺性是不够的,资源的结构性质决定了攻击者能否通过复制身份以次线性成本获得影响力。研究证明,任何满足可分性、影响力可加性、时间可重用性和身份可转移性的资源,都允许影响力摊销,从而实现次线性成本。这构成一个不可能性结果:没有任何协议能强制实现线性Sybil成本。

Control Flow Graph Recovery for Dynamically Loaded Code via Symbolic Library Resolution

第一作者: Oleksandr Mostovyi · 方向: 软件安全

Abstract:Control Flow Graphs are one of the main data sources for software analysis that use dynamic and static software analysis methods. Protected software and modern malware increasingly depend on dynamic code loading techniques to evade static analysis. Usage of runtime dynamic linking mechanisms introduces unresolved indirect calls that stop static Control Flow Graph recovery. This serves to hide dynamic library that can be used for prevention of security analysis. To address this limitation, an analysis technique is proposed that combines symbolic execution with speculative library preloading to recover Control Flow Graphs from binaries by using dynamic loading. The methodology uses custom software hooks that intercept dynamic loading operations during symbolic execution and perform actual library loading into the analysis state. The module is based on a two-level architecture...

论文介绍 针对现代恶意软件通过动态代码加载规避静态分析的问题,本研究提出一种结合符号执行与推测性库预加载的方法,以恢复二进制文件的控制流图。该方法使用自定义钩子拦截动态加载操作,在符号执行过程中实际加载库文件,从而解决间接调用未解析的难题,提升软件安全性分析的效率。

LoRA-Key: User-Centric LoRA Watermarking for Text-to-Image Diffusion Models

第一作者: Yaopeng Wang · 方向: 系统安全

Abstract:Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing text-to-image diffusion models, enabling lightweight modules that are shared, reused, and commercialized as independent assets. This LoRA-centric ecosystem shifts copyright protection from foundation models to distributed LoRA modules, which are easy to copy, redistribute, or reuse without authorization. Existing watermarking methods either protect the base diffusion model or require watermark-aware retraining for each target LoRA, limiting their practicality in open community settings. To address this limitation, we propose LoRA-Key, a user-centric LoRA watermarking framework that treats copyright protection as a reusable ownership key. LoRA-Key encapsulates a recoverable secret message into a standalone user-specific Watermark LoRA, which can be attached to different target LoRAs through...

论文介绍 低秩适配(LoRA)模块在文本到图像扩散模型中易于复制和再分发,现有水印方法实用性不足。本研究提出LoRA-Key框架,将版权保护视为可重用的所有权密钥,通过用户特定的Watermark LoRA封装可恢复的秘密消息,并可附加到不同目标LoRA上,实现轻量级版权保护,适用于开放社区场景。

Temporal Motif-aware Graph Test-time Adaptation for OOD Blockchain Anomaly Detection

第一作者: Runang He · 方向: AI 安全

Abstract:Ever-evolving transaction patterns have significantly hindered anomaly detection on emerging cryptocurrency blockchains due to the vast number of addresses and diverse anomalous behaviors. Recently, advanced Graph Anomaly Detection (GAD) approaches applied to blockchains have faced two critical challenges: \textit{adversarial pattern evolution by malicious actors} and \textit{the out-of-distribution (OOD) problem caused by varied transaction semantics on blockchains}. To address these challenges, we propose a novel framework termed \textbf{TE}mporal \textbf{M}otif-aware \textbf{G}raph \textbf{T}est-\textbf{T}ime \textbf{A}daptation (\textbf{TEMG-TTA}). First, we comprehensively capture the 3-node temporal motif distribution of each active address using an efficient computational mechanism, enabling downstream temporal motif-aware graph learning. Second, we design a simple yet...

论文介绍 加密货币区块链中交易模式不断演变,导致图异常检测面临对抗性模式演化和分布外(OOD)问题。本研究提出TEMG-TTA框架,通过捕捉活跃地址的三节点时间 motif 分布进行图学习,并设计测试时自适应机制,以提高对新兴交易语义的鲁棒性,增强区块链安全监控能力。

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

第一作者: Yijia Fang · 方向: 密码学协议

Abstract:Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the advertised model. We introduce KBF, a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary. Across 16 production LLM endpoints, KBF flags all 155 economically relevant substitutions without rejecting any same-model controls, remains stable under deployment variation, detects high-separation mixed-routing attacks when only 5-10% of traffic is substituted, and finds that 7 of 27 platform model cells in a six-platform shadow API audit are statistically inconsistent with their reference endpoints, with inconsistencies concentrated on premium Claude endpoints.

论文介绍 大语言模型 API 中存在模型替换和混合路由攻击风险,用户难以验证端点一致性。本研究提出KBF黑盒审计协议,通过测量模型在知识边界附近的稳定数值回忆来生成指纹,低成本检测API替换。实验在16个生产端点中成功标记所有替换,发现平台模型单元不一致现象,提升API可信度。

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

第一作者: Almene De Meran Meguimtsop · 方向: AI 安全

Abstract:Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (RCR) norms or help undermine them. We introduce SciIntBench, an adversarial benchmark of 810 prompts across ten RCR categories and three scientific domains. Each scenario appears as an Overt Adversarial, Covert Adversarial, and Benign version, allowing us to jointly measure framing-sensitive refusal of misconduct and helpfulness on legitimate requests. We evaluate 16 commercial and open-weight LLMs from six providers (2024--2026), producing 12,960 responses. We find that scientific integrity alignment is strongly framing-sensitive: models refuse explicit misconduct far more reliably than covert violations, especially failing when misconduct is presented as a pressure-driven shortcut. Refusals vary by RCR category, with weaker...

论文介绍 大语言模型在科学工作中的应用日益广泛,但其是否遵守研究诚信规范尚不明确。本研究引入SciIntBench对抗性基准,涵盖10个研究诚信类别和810个提示,评估16个LLM。发现模型对显性违规拒绝更可靠,但对隐蔽违规响应不足,尤其在压力驱动场景下,凸显科学伦理对齐的挑战。

Bridging Theory and Practice: An Executable Taxonomy of Security Properties for ProVerif and Tamarin

第一作者: Leonard Tudorache · 方向: 密码学协议

Abstract:Security is critical for everything relying on modern digital systems. Because almost all digital interactions are governed by the Internet and cryptographic protocols, these protocols must serve as reliable mechanisms that guarantee core security properties, such as confidentiality and integrity. Formal verification of these protocols is a critical step in securing interconnected systems. Tools such as ProVerif and Tamarin are widely employed to perform automated verification. However, their effective use demands specialized domain knowledge, creating a significant learning curve for security protocol designers who often have a security, rather than a formal verification background. We therefore need structured, accessible resources to help protocol designers to express their design and requirements in the language of the formal verification tools. To address this, we...

论文介绍 形式化验证工具如ProVerif和Tamarin对密码协议安全至关重要,但使用门槛较高。本研究提出一个可执行的安全属性分类法,专门针对这些工具,提供结构化资源帮助协议设计师表达设计和需求。该分类法桥接理论与实践,降低学习曲线,提升验证效率和准确性。

Protecting On-Device AI Inference: A Systematic Review of Attacks and Defence Mechanisms

第一作者: Zisis Tsiatsikas · 方向: 密码学协议

Abstract:The need for secure and private Artificial Intelligence (AI) and Machine Learning (ML) on edge and mobile devices has increased the necessity of protecting the architecture of these systems from threats to both security and privacy. With an ever-increasing number of pre-trained AI models being used on mobile platforms for client-side inference, there are rising concerns about the risks associated with the theft/extraction of AI models, adversarial attacks on AI models, and data breaches. As a result of this trend, a variety of defence mechanisms have been proposed to protect against these threats. These include Trusted Execution Environments (TEEs), homomorphic encryption, obfuscation, and differential privacy, among others. However, current surveys largely focus on edge intelligence, which includes distributed training, and thus overlook security and privacy issues that are...

论文介绍 移动和边缘设备上AI推理面临模型提取、对抗攻击和数据泄露等安全与隐私威胁。本研究系统综述相关攻击和防御机制,包括可信执行环境、同态加密和差分隐私等方法,分析当前研究侧重于边缘智能的不足,为保护设备上AI推理提供全面指导。

AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing

第一作者: Yuexin Li · 方向: 安全研究

Abstract:Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-based designs remain vulnerable to structural perturbations, such as sentence splitting and merging, which commonly arise under strong paraphrasers like DIPPER and GPT-3.5. To mitigate this issue, we propose AliMark, a framework that reformulates sentence-level watermarking as a bit sequence encoding and alignment problem between a potentially watermarked text and a secret bit sequence. Notably, our approach adopts a two-stage detection strategy: we generate multiple restructured text variants and adaptively align their extracted bit sequences with the secret bit sequence to minimize alignment cost. This multi-candidate alignment design naturally improves robustness to sentence merges and splits. Extensive experiments demonstrate...

论文介绍 现有句子级水印方法在文本改写下易受结构扰动影响。本研究提出AliMark框架,将水印视为比特序列编码和对齐问题,通过生成多个文本变体并自适应对齐提取的比特序列,减少对齐成本。这种多候选设计增强了对句子合并和分割的鲁棒性,提升水印在强改写器下的稳定性。

Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills

第一作者: Chia-Yi Hsu · 方向: 软件安全

Abstract:LLM-powered coding agents increasingly participate in software development workflows by generating code, selecting dependencies, and producing package installation commands. This creates a new software supply chain risk: when an agent hallucinates a non-existent package, an attacker may register the hallucinated name and later compromise users who install it. Existing package hallucination attacks and defenses primarily focus on naturally occurring hallucinations, targeted dependency steering, or post-hoc package validation. In this paper, we introduce \emph{Neutral Prompting Attack} (NPA), a highly stealthy attack paradigm in which semantically benign instructions, such as encouraging imagination and exhaustiveness, increase package hallucination propensity without containing explicit malicious intent. Unlike targeted dependency steering, NPA does not specify an...

论文介绍 针对LLM编码代理在软件开发中可能幻觉不存在的包而引发供应链攻击的问题,本文提出「中性提示攻击」方法,通过语义上无害的指令增加幻觉倾向,实现隐蔽的依赖转向。研究揭示了这种新型攻击向量,有助于推动相关防御机制的发展。

DeepFake Forensics AI: A Multi-Modal Detection and Blockchain-Anchored Evidence Management Platform

第一作者: Naisha Minnah · 方向: AI 安全

Abstract:The proliferation of AI-generated synthetic media poses a critical threat to the integrity of digital evidence in legal and forensic contexts. Existing deepfake detection systems typically address a single modality and provide no mechanism for tamper-proof evidence preservation. We present DeepFake Forensics AI, a unified platform that detects synthetic media across image, video, and audio modalities, identifies generative architecture fingerprints, and anchors forensic evidence immutably on the Ethereum blockchain. Our system trains four independent neural networks from scratch: an EfficientNet-B4 image detector (AUC = 0.9868), a Bidirectional LSTM video detector (AUC= 0.9628), an ECAPA-TDNN audio detector (EER = 18.63%), and a novel GAN fingerprinting module (accuracy = 99.88%) that identifies the generative architecture behind a fake image. Evidence files are hashed with...

论文介绍 随着AI生成合成媒体对数字证据完整性的威胁加剧,本文提出「DeepFake Forensics AI」平台,集成多模态神经网络检测图像、视频和音频,并生成架构指纹。系统将取证证据锚定到以太坊区块链,确保不可篡改,为法律和法医场景提供可信解决方案。

HunterAgent: Neuro-Symbolic Attack Trace Reconstruction under Anti-Forensics

第一作者: Guangze Zhao · 方向: AI 安全

Abstract:Modern alert-triage systems reduce SOC burden by filtering false positives, but flagging a high-risk alert is only the start of incident response. Threat hunting requires reconstructing causal attack chains across heterogeneous, partially corrupted logs. Against APTs using anti-forensics (parent-PID spoofing, log wiping, fileless execution), provenance graphs split into disjoint subgraphs and fail. Unconstrained LLM agents fabricate causal links violating OS physics, producing fluent but forensically inadmissible narratives. We propose HunterAgent, a neuro-symbolic framework that reframes trace reconstruction as cost-bounded heuristic graph search under partial observability. It uses an asymmetric Generator-Verifier pipeline: the LLM proposes semantic hypotheses within a typed ontology, while a verifier grounds each via identifier-level collisions on surviving orthogonal...

论文介绍 高级持续威胁使用反取证技术导致源图分裂,使因果攻击链重建困难。本文提出「HunterAgent」框架,将追踪重建建模为部分可观测性下的成本约束启发式图搜索,采用LLM生成假设并验证,提升威胁猎手在对抗环境中的溯源能力。

Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content

第一作者: Bing Liu · 方向: AI 安全

Abstract:This paper presents a survey and taxonomy of LLM fingerprinting and watermarking for identity, ownership verification, provenance, and generated-content attribution. Large language models (LLMs) require substantial investments in data, computation, and expertise, and are increasingly deployed in high-stakes settings, making it critical to protect LLM-related assets and trace their origins. Existing work has rapidly expanded across dataset provenance, model ownership, and generated-content detection, but the field remains fragmented: fingerprinting and watermarking are often used inconsistently, and methods are typically studied within isolated asset-specific settings. To address this gap, we introduce implicit identity as a unifying abstraction for verifiable but not directly observable identity signals in LLM systems. We distinguish fingerprinting as non-intrusive identity...

论文介绍 大型语言模型的资产保护和溯源需求日益增长,但现有指纹和水印工作碎片化。本文引入「隐式身份」作为统一抽象,系统综述跨数据集、模型和生成内容的身份技术,区分指纹为非侵入性识别,水印为嵌入式信号,为LLM安全提供结构化视角。

Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking

第一作者: Junke Zhang · 方向: AI 安全

Abstract:Jailbreak attacks on large language models (LLMs) aim to induce LLMs to produce content that they are expected to refuse. Automated black-box jailbreak generation is especially important for safety evaluation, where the attacker observes only model outputs and needs to automatically search for effective adversarial prompts. Existing black-box jailbreak methods either depend on sample-wise heuristic search or leverage attack experience through accumulating strategy pools or method libraries, lacking a systematic organization and management of attack experience. To mitigate these drawbacks, we propose MemoAttack, a memory-driven black-box jailbreak framework with comprehensive attack memory modeling, evolution, and selection. Specifically, MemoAttack comprises three key designs: (1) Skill-Structured Memory Modeling, which abstracts accumulated attack experience into reusable...

论文介绍 现有黑盒LLM越狱方法缺乏系统组织攻击经验。本文提出「MemoAttack」框架,将攻击经验抽象为技能结构化记忆,通过演化和选择机制增强攻击效果。这为自动化安全评估提供了更有效的黑盒攻击生成方法。

S3C2 Summit 2025-09: Industry Secure Supply Chain Summit

第一作者: Md Atiqur Rahman · 方向: 软件安全

Abstract:Today's digital ecosystem relies heavily on software supply chains, which enable developers to reuse code and ship software at scale. However, a single vulnerable component can jeopardize the entire supply chain. In recent years, cyberattacks in software supply chains have become increasingly common. These attacks can disrupt critical systems and put organizations, including major software companies, government agencies, and open-source contributors, at risk. This growing threat has led to increased attention from both the software industry and the U.S. government toward strengthening software supply chain security. On September 15, 2025, three researchers from the NSF-backed Secure Software Supply Chain Center (S3C2) convened a Secure Software Supply Chain Summit, bringing together 10 practitioners from 8 organizations across diverse domains. The goals of the Summit were...

论文介绍 软件供应链的脆弱性正导致网络攻击风险增加。本文总结S3C2峰会内容,汇集行业实践者分享经验、讨论挑战和解决方案。峰会旨在促进跨领域合作,加强供应链安全实践,应对日益增长的威胁。

SAMD: A Tool for Identifying False Data Injection Scenarios in AI/ML-enabled Medical Devices

第一作者: Mohammadreza Hallajiyan · 方向: AI 安全

Abstract:The growing integration of artificial intelligence (AI) and machine learning (ML) in medical systems requires effective measures to address emerging security risks. One such risk is that of adversaries introducing false data through vulnerable system components during inference, causing misdiagnosis and wrong treatments. These risks are challenging to anticipate and address in the design phase, as the system assembly partially occurs during actual use by end users. To address this concern, we introduce SAMD, an automated tool for performing System Theoretic Process Analysis for Security (STPA-Sec) on AI/ML-enabled medical devices during the design phase. SAMD models the medical system as a control structure, treating all system components as potential points for injecting false data into the ML engine. It leverages state-of-the-art vulnerability databases and Large Language...

论文介绍 AI/ML医疗设备易受虚假数据注入攻击,可能导致误诊。本文提出「SAMD」工具,在设计阶段使用系统理论过程分析(STPA-Sec)建模医疗系统,自动识别注入点。该工具帮助提前发现安全风险,提升设备安全性。

The Best-Laid SCHEMEs: Coordinated Sabotage and Monitoring in Multi-Agent Systems

第一作者: Nikolay Radev · 方向: 软件安全

Abstract:As agentic coding systems decompose work across multiple model instances, a critical safety question is whether those instances can coordinate to achieve a hidden malicious objective while remaining aligned with user intent. We introduce SCHEME, a benchmark of 17 task instances across 7 settings and 8 real open-source libraries, each pairing a legitimate software-engineering task with a covert side task. Every setting is designed so that no proper subset of agents can succeed alone: agents must decompose a shared sabotage plan, relay partial requirements under different communication topologies, and execute mutually consistent edits, testing genuine multi-agent coordination rather than individual capability. Evaluating with GPT 5.1 Codex and Gemini 3.1 Pro, we find coordinated sabotage is already practical, with Gemini completing the covert objective while succeeding on the...

论文介绍 多智能体编码系统可能协调实现隐藏恶意目标,威胁用户意图对齐。本文引入「SCHEME」基准,包含跨设置和库的任务实例,要求智能体协调执行破坏。评估显示协调破坏已可行,警示多智能体系统对齐风险。

EvaluatAR: A Cross-Device Evaluation Framework for Rapid Prototyping of Bystander PETs in AR

第一作者: Syed Ibrahim Mustafa Shah Bukhari · 方向: 隐私保护

Abstract:Augmented Reality (AR) headsets continuously sense their surroundings, capturing nearby bystanders and raising privacy risks. Visual bystander privacy-enhancing technologies (PETs) mitigate this risk by detecting bystanders in egocentric scene views and applying privacy transformations (e.g., obfuscation). However, traditional PET evaluation is human-dependent, high-overhead, and device-specific, making it difficult to reproduce across devices. We present EvaluatAR, a cross-device evaluation framework for rapid prototyping at the early stage of PET evaluation. Our framework enables controlled replication of experimental conditions by standardizing PET inputs (sensor data and visual stimuli) and outputs through a record-replay workflow. We validate EvaluatAR through three case studies on HoloLens 2, Magic Leap 2, and Meta Quest 3 across implicit (continuous, context-driven) and...

论文介绍 本文提出EvaluatAR,一个用于评估增强现实(AR)设备中旁观者隐私增强技术(PETs)的跨设备评估框架。该研究旨在解决传统PET评估方法对人工依赖强、开销大且设备特定的问题。框架通过标准化传感器数据输入和视觉刺激输出,采用记录-回放工作流,以实现实验条件的可控复制,支持在HoloLens 2等多款设备上对隐式和显式PET进行快速原型评估。

Domain-Informed Representation for Evolutionary Sieving in Integral and Module Lattices

第一作者: Ahmad Tashfeen · 方向: 安全研究

Abstract:Traditional cryptography, rooted in problems, e.g., integer factorisation or discrete log, is inevitably vulnerable to a fully operational quantum computer. Although it remains an engineering frontier, the looming threat extends to encrypted data stored today, which could be decrypted in the future with quantum capabilities. To safeguard against this eventuality, the backbone of the modern quantum-safe cryptography is the Shortest Vector Problem (SVP). We enhance Laarhoven's treatment of Ajtai et al.'s sieving as a genetic algorithm (GA) for the SVP by incorporating domain-informed SVP representation and crossover while naturally extending application to the module lattices.

论文介绍 本研究致力于增强量子安全格密码学中的核心难题——最短向量问题(SVP)的求解方法。作者改进了基于遗传算法的筛法,通过引入域知情的SVP表示和交叉算子,提升了算法效率,并将该方法自然扩展到模格上的SVP求解。这为设计和分析能够抵御未来量子计算机威胁的密码系统提供了更有效的基础工具。

S3C2 Summit 2025-07: Government Secure Supply Chain Summit

第一作者: Sivana Hamer · 方向: 软件安全

Abstract:Software supply chains, while providing immense economic and software development value, are only as strong as their weakest link. Over the past several years, there has been an exponential increase in cyberattacks specifically targeting vulnerable links in critical software supply chains. The attacks disrupt day-to-day functioning and threaten the security of nearly everyone on the internet, from billion-dollar companies and government agencies to hobbyist open-source developers. The evolving threat of software supply chain attacks has garnered interest from both the software industry and governments worldwide in improving software supply chain security. On Thursday, July 9th, 2025, 3 researchers from the NSF-backed Secure Software Supply Chain Center (S3C2) conducted a Secure Software Supply Chain Summit with a diverse set of 12 participants from 6 US government agencies...

论文介绍 本文是2025年“政府安全供应链峰会”的报告记录。该峰会由美国NSF支持的S3C2中心召集,旨在应对日益严峻的软件供应链攻击威胁。报告汇集了来自美国多个政府机构的参与者,共同探讨了软件供应链的脆弱性以及改善其安全性的合作与治理路径,强调了多方协作的重要性。

Techreport: Evaluating Tor-based Location Privacy for Ethereum Validators

第一作者: Muhammad Umar Janjua · 方向: 密码学协议

Abstract:Privacy and anonymity of validators, especially regarding IP address linkability, are essential to protect the Ethereum network from various attacks. Network-level attacks, such as DoS, can interrupt validators and affect the overall security of the Ethereum network. Correlating the IP addresses of validators with their identities, along with knowledge about their action slots can be exploited by attackers to cause network delays, MEV exploitation, and finality risks. Therefore, ensuring the unlinkability of a validator's IP and identity is crucial for maintaining the network's trust and resilience. In this techreport, we first provide a review of the existing network and consensus layer techniques that have been proposed for maintaining validator privacy in the Ethereum blockchain. Secondly, we evaluate a Tor-based protocol named Tor push that helps unlink validator...

论文介绍 该技术报告探讨了以太坊验证者的网络层隐私问题,特别是IP地址与其身份的关联可能引发的拒绝服务攻击和MEV利用等风险。文章首先回顾了现有的隐私保护技术,随后重点评估了一种基于Tor的协议——Tor push。该协议旨在通过网络层匿名化来切断验证者IP地址与其身份的联系,从而增强网络的安全性与韧性。

unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning

第一作者: Geoffrey Bradway · 方向: AI 安全

Abstract:Unix competence is the ability to use shell and operating-system primitives as first-class tools, not merely to write programs through a terminal. Current terminal benchmarks tend to blur this distinction: a solver fluent in Python but weak in Unix can pass a substantial fraction of Terminal-Bench 2.0, while the reverse skill profile is rarely exercised. We make the distinction operational and build a training surface for the Unix component. unix-ctf is a procedural generator of capture-the-flag tasks for shell agents. Each task hides a short token (a flag of the form flag(a3b1c9...)) inside a fresh Linux container using a single Unix feature, and the agent must recover it. Tasks are produced by an LLM-assisted synthesis pipeline that generates candidate hiding techniques, rewrites them into parameterized hide-and-find script pairs, and filters them with a bidirectional...

论文介绍 本文区分了“终端能力”与“Unix技能”,指出当前基准测试常模糊此界限。为专门强化AI智能体的Unix技能,作者提出了unix-ctf,一个程序化生成“捕获旗帜”(CTF)任务的环境。每个任务在Linux容器中隐藏一个标志,智能体需利用单一Unix功能来发现它。任务由大语言模型辅助的流程生成与筛选,为训练和评估Shell代理提供了专注的测试平台。

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

第一作者: Mohammadreza Teymoorianfard · 方向: 安全研究

Abstract:Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework...

论文介绍 本文研究了集成推理能力的视觉-语言-动作模型在自动驾驶中的鲁棒性。研究表明,这些模型对现实的文本输入扰动非常脆弱。通过使用NVIDIA的Alpamayo模型进行黑盒研究,发现在闭环仿真中,扰动可导致推理失败率高达89%,轨迹操纵成功率达72%,进而增加碰撞率并降低安全指标。研究引入了一个推理感知的评估框架来系统性地探测这些模型的漏洞,揭示了其在推理与轨迹生成耦合假设下的脆弱性。

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

第一作者: Ojas Nimase · 方向: 密码学协议

Abstract:Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a growing concern for fairness and information integrity. Research on generative engine optimization (GEO) has produced many manipulation methods, but each is evaluated on its own dataset with its own metrics, so their relative strength and detectability stay unclear. We present GEO-Bench, a benchmark that evaluates GEO ranking-manipulation attacks under one protocol. It unifies black-box prompt-based attacks (TAP, Zero-Shot), white-box gradient-based attacks (STS, RAF, StealthRank), and ten white-hat C-SEO strategies. We score every method on five datasets against a fixed open-weight ranker (Llama-3.1-8B-Instruct), using metrics for both effectiveness (NRG, Success@{\alpha}, Promote@{\alpha}) and stealth (keyword violation rate...

论文介绍 随着大语言模型(LLM)越来越多地用于对查询结果进行排序,操纵其排名的生成式引擎优化(GEO)方法引发了对信息公平性的担忧。然而,现有攻击方法缺乏统一的评估标准。本文提出GEO-Bench基准,在统一协议下评估了多种黑盒和白盒攻击方法的有效性与隐蔽性,旨在系统比较这些攻击手段,为理解和防御LLM排名操纵提供共同基础。

Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening

第一作者: Mohan Zhang · 方向: AI 安全

Abstract:LLMs are vulnerable to prompt injection attacks. However, this vulnerability has been primarily demonstrated conceptually in academic studies or through a few anecdotal case studies. Its prevalence and impact in real-world LLM-based applications are largely unexplored. In this work, we present the first systematic study of prompt-injection attacks in a widely used application: LLM-based resume screening. Our analysis is based on approximately 200K real-world resumes collected over multiple years by hireEZ. We first design tailored methods to detect prompt injection in resumes. Manual validation on a small-scale dataset demonstrates that our detectors achieve high precision and outperform state-of-the-art general-purpose detectors. We then apply our detector to the full resume dataset and conduct a comprehensive measurement study of real-world prompt injection attacks. Our...

论文介绍 本文首次对基于大语言模型(LLM)的简历筛选应用中的提示注入攻击进行了系统性实证研究。基于约20万份真实世界简历数据集,作者设计并验证了针对性的检测方法,其精度优于通用检测器。随后,他们应用该检测器对完整数据集进行了大规模测量,量化了真实世界中提示注入攻击的普遍性和特征,填补了此类攻击在主流商业应用中实际影响的研究空白。

A Secure, Manifest-Based Framework for Delegated Privilege Promotion

第一作者: Rajarshi Chowdhury · 方向: 软件安全

Abstract:Large-scale enterprise software systems commonly run as unprivileged service accounts to enforce least privilege, yet still depend on a small set of privileged components -- such as executables with elevated ownership, permissions, or capabilities -- for narrowly scoped operations. This creates a persistent security and operational conflict during maintenance. Automated patching tools running without elevated privileges cannot safely update privileged components without either executing the entire patch with full administrative rights or requiring manual administrator intervention. We present a secure, manifest-based infrastructure for delegated promotion of privileged software components, deployed in production as part of a large-scale enterprise database system serving both cloud and on-premises installations. The design centers on a minimal privileged mediator that...

论文介绍 研究企业软件系统中特权组件更新的安全冲突问题。提出基于清单的框架,通过最小特权中介委托提升特权组件的权限。该框架部署于大规模企业数据库系统,支持云和本地安装,旨在自动化修补工具安全更新特权组件,减少管理员干预。

Optimal Rates for Differentially Private Hypothesis Testing with E-values

第一作者: Ben Jacobsen · 方向: 隐私保护

Abstract:E-values have attracted considerable interest in recent years as flexible tools for enabling anytime-valid and adaptive data analysis. Hypothesis testing is at the core of many of these applications, which can often involve private or sensitive data. In this work, we answer a simple but important question: given two distributions $\mathbb{P}$ and $\mathbb{Q}$, what is the maximum achievable e-power when testing $X\sim \mathbb{P}^n$ against $X\sim\mathbb{Q}^n$ with e-values that satisfy $\varepsilon$-differential privacy? We characterize the optimal rate for this problem and provide an algorithm which matches it exactly. In the sequential setting, when observations arrive one-by-one and the analyst chooses when to halt, we give matching upper and lower bounds on the stopping times of any private e-process. Numerical experiments confirm the practicality of our algorithms, which...

论文介绍 探讨在差分隐私约束下使用e-values进行假设测试的最优性能。表征了最大e-power的最优速率,并提供了匹配算法。在顺序设置中,分析了停止时间的上下界。算法经数值实验验证,适用于隐私保护的数据分析场景。

AIRGuard: Guarding Agent Actions with Runtime Authority Control

第一作者: Suliu Qin · 方向: 密码学协议

Abstract:Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This makes agent attacks different from jailbreaks. The harmful step is often not an obviously forbidden output, but an ordinary executable action that becomes unsafe because attacker-controlled context steers authorized access against the user's interest. We identify this failure mode as authority confusion: untrusted resources may inform reasoning, but they must not authorize side effects. We present AIRGuard, a runtime guard that operationalizes least privilege as action-time authorization. AIRGuard normalizes heterogeneous tool calls, derives task authority into step-level authority, tracks source and target trust, simulates sensitive side effects, audits cross-step risk, and enforces decisions before...

论文介绍 针对语言代理工具使用中的权威混淆安全风险,提出AIRGuard运行时保护。该系统标准化异构工具调用,将任务权限派生为步骤级权限,跟踪信任源和目标,模拟敏感副作用,并在执行前强制决策,以实施最小特权原则。

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

第一作者: Jaydip Sen · 方向: AI 安全

Abstract:Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses significant challenges to reliability, security, and trustworthiness. Adversarial machine learning demonstrates that even highly accurate models can be manipulated through carefully crafted perturbations, raising serious concerns in safety critical systems such as healthcare, finance, and autonomous technologies. In parallel, quantum computing has emerged as a transformative paradigm capable of addressing complex computational problems through principles such as superposition, entanglement, and quantum interference. The convergence of these fields has led to the emergence of quantum artificial intelligence, which explores how quantum techniques can enhance learning efficiency, scalability, and robustness. This chapter provides a...

论文介绍 讨论人工智能模型易受对抗攻击的脆弱性,以及量子计算如何通过叠加、纠缠等原理增强学习效率和鲁棒性。本章综述量子人工智能在提升AI系统安全性方面的潜力和挑战,适用于医疗、金融等关键领域。

Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought

第一作者: Jiacheng Lu · 方向: 密码学协议

Abstract:Large Language Models with Chain-of-Thought reasoning capabilities represent valuable intellectual property, yet existing black-box watermarking methods often trade robustness for reasoning fidelity by perturbing final answers or relying on fragile trigger patterns. We propose BiCoT, a watermarking framework that embeds ownership signals into the internal geometry of reasoning traces by aligning high-saliency structural anchors with a private signature subspace while regularizing ordinary control tokens to preserve semantic capacity. This design couples the watermark with reasoning-relevant representations, making removal difficult without disrupting the features that support coherent reasoning. To enable verification under model theft and representation drift, we introduce Robust Subspace Registration (RSR), a Top- logprob-based black-box verifier that uses sentinel tokens to...

论文介绍 为保护具有链式思维推理能力的大语言模型知识产权,提出BiCoT水印框架。该框架将所有权信号嵌入推理轨迹的高显著性结构锚点中,结合稳健子空间注册进行验证,确保水印与推理相关,难以移除而不影响推理能力。

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

第一作者: Caleb DeLeeuw · 方向: 安全研究

Abstract:Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal structurally sound, or does it disappear under modest changes to prompt framing, formatting, or output length? Across five architectures, no model cleanly discriminated benign from hazard. Gemma 2 2B-IT never genuinely refused across 75 prompts, hedging on every hazard-adjacent query. Gemma 4 E2B-IT refused 65/75 prompts with chat-template formatting and 0/75 without it. Both Gemma models collapsed to 0% under an 80-token cap. Qwen 2.5 1.5B and Phi-3-mini over-refused, flagging 83-87% of benign biology as hazardous. Llama 3.2 1B showed the only meaningful tier gradient (61-point spread). To probe what drives such over-refusal, we tested a panel of Schedule I but biologically non-toxic compounds...

论文介绍 审核语言模型在生物安全场景下拒绝行为的可靠性。使用通用和领域微调的稀疏自编码器评估模型,发现不同架构在拒绝深度上存在差异,例如格式和输出长度影响拒绝一致性。该工作旨在提高生物安全评估的深度。

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

第一作者: Dongrui Liu · 方向: 安全研究

Abstract:Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers, rendering current agent alignment frameworks inadequate for real-world deployment. To tackle these emerging threats, we propose a lightweight and scalable agent safety alignment framework. Specifically, we update the agent safety taxonomy to accommodate emergent risks from Codex and OpenClaw execution scenarios. We further build a taxonomy-guided data engine with influence-function purification to train lightweight AgentDoG 1.5 variants (0.8B, 2B, 4B, and 8B parameters) using only around 1k samples, achieving comparable performance with leading closed-source models (e.g., GPT-5.4). Based on AgentDoG 1.5, we construct a highly efficient agentic safety SFT and RL...

论文介绍 针对开放世界AI代理的安全风险,提出轻量级可扩展的对齐框架AgentDoG 1.5。通过更新安全分类法并构建数据引擎,训练小规模模型实现高效安全对齐。框架适用于新兴代理场景,如Codex和OpenClaw执行。

Secure Distributed Hypothesis Testing

第一作者: Gowtham R. Kurri · 方向: 系统安全

Abstract:In distributed hypothesis testing, a central server performs hypothesis testing based on information received from distributed sensors/clients. We study a secure variant of this problem in which the central server determines the hypothesis class of an underlying distribution without learning any additional information about the distribution itself. We prove that, in its standard form, this is impossible to achieve, even for simple and highly restricted cases. To bypass this impossibility, we augment the model with a shared secret key available to clients but hidden from the server. We show that a single-bit secret key enables perfectly secure testing for simple classes by reducing the test distributions to a symmetric, canonical instance. Finally, for arbitrary hypothesis classes over finite domains, we establish a reduction to standard hypothesis testing using Private...

论文介绍 研究分布式假设测试中的安全问题,要求服务器确定假设类别而不学习分布额外信息。证明标准形式下不可能实现,因此引入共享密钥。单比特密钥可用于简单类别的安全测试,通过减少到对称实例实现。

Information Security in Small-Scale Protests: Surveillance of Ugandan Anti-EACOP Protesters

第一作者: Ntezi Mbabazi · 方向: AI 安全

Abstract:We examine the information security practices of Ugandan climate activists protesting the development of the East African Crude Oil Pipeline (EACOP). We conducted five-week fieldwork in Kampala, Uganda, which included interviews with 13 anti-EACOP activists. Through an inductive analysis, we report on the complexities faced by small groups of predominantly student protesters as they covertly organise small-scale anti-EACOP protests within a context marked by state surveillance and repression. Our study points to a multi-layered adversarial landscape, where participants' experiences of direct threats, including arrests and information compromise, and their fears of abduction, shaped their security practices. These practices were rooted in autonomous decision-making within groups. We present a grounded understanding of how participants' need to protect information for their own...

论文介绍 本研究关注在国家监控与压制背景下,乌干达反东非原油管道项目的小规模学生抗议者如何进行信息保密与安全实践。研究者通过在坎帕拉为期五周的实地工作,访谈了13名活动家,发现其安全措施根植于团体内的自主决策,旨在应对其亲身经历的逮捕、信息泄露及绑架恐惧等多层次威胁。研究揭示了边缘化群体在高压环境中保护信息安全的复杂策略。

CODEFUSE-DEBENCH: An Empirical Study on Readability, Recompilability, and Functionality

第一作者: Puzhuo Liu · 方向: 软件安全

Abstract:Binary decompilation aims to recover binaries into high-level source code, but existing evaluations mainly rely on syntactic similarity or single-axis readability metrics, which fail to capture practical reusability. We propose a reusability-driven evaluation paradigm that measures decompiler quality along three orthogonal dimensions: readability, recompilability, and functionality. We present DEBENCH, the first automated framework for multidimensional decompilation evaluation. DEBENCH contains 240 atomic test functions, organized into 8 source files and compiled into 640 binaries. It combines LLM-as-judge readability scoring with URAF (18 sub-dimensions), iterative compile-and-repair under a fixed 50-iteration budget, and Frida-based differential dynamic tracing at the program, function, and instruction levels. We evaluate five mainstream decompilers and three repair LLMs...

论文介绍 现有二进制反编译评估主要依赖语法相似度,难以衡量代码的实际可重用性。本文提出一个以可重用性为导向的评估范式,从可读性、可重编译性和功能性三个正交维度评估反编译器质量。为此,研究团队构建了自动化评估框架DEBENCH,它结合了基于大型语言模型的可读性评分、迭代编译修复以及基于动态追踪的功能一致性检查,并用此框架评估了五种主流反编译器。

Provably Secure Agent Guardrail

第一作者: Benlong Wu · 方向: AI 安全

Abstract:As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates a fundamental crisis in artificial intelligence security. Existing defense architectures heavily rely on empirical semantic guardrails and probabilistic large model adjudicators, mechanisms that fail to provide deterministic security lower bounds when facing complex semantic symbol decoupling attacks. To overcome this empirical semantic guardrail dilemma, this paper proposes a new security paradigm for agents based on the fundamental limitations of logical reasoning. Based on this paradigm, we further introduce an executable Proof-Constrained Action (ePCA) framework with a neural symbolic isolation architecture. This framework abandons semantic trust in natural language, forcing agents to losslessly formalize their intentions...

论文介绍 随着大语言模型转变为具有广泛执行权限的智能体,其失控风险成为AI安全的核心危机。现有依赖经验语义的安全护栏无法提供确定的安全下限。本文提出一种基于逻辑推理根本局限性的智能体安全新范式,并引入可执行证明约束动作框架。该框架抛弃对自然语言的语义信任,强制智能体将其意图无损地形式化,从而在根本上隔离攻击,提供可证明的安全保障。

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

第一作者: Aditya Nawal · 方向: AI 安全

Abstract:AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating external content into the generation pipeline can weaken the safety alignment mechanisms that govern model outputs. Prior work shows that enabling retrieval in agents increases compliance with harmful requests. We introduce AgentREVEAL, a diagnostic framework for analyzing retrieval-induced safety degradation in LLM agents. The framework examines two axes: how retrieval is integrated into the agent pipeline and the properties of the retrieved content. Along the integration axis, we find that binding tool invocation and response generation in a single step amplifies harmful outputs. Along the content axis, we uncover the Safe Source Paradox: even oppositional or safety-oriented sources, such as pages containing warnings or risk...

论文介绍 AI智能体通过网络检索等外部工具增强能力,但外部内容的整合可能弱化模型的安全对齐机制。本文引入诊断框架AgentREVEAL,从检索集成方式和检索内容属性两个轴分析检索导致的安全退化。研究发现,将工具调用与响应生成绑定在单一步骤会放大有害输出,并揭示了「安全来源悖论」:即便是包含警告或风险提示的反对或安全导向的来源,也可能被攻击者利用来增加有害内容的合规性。

Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents

第一作者: Jay Yu · 方向: 密码学协议

Abstract:DeFi investment agents, systems that use AI for autonomous on-chain trading, have attained over USD 3 billion in combined token valuations since late 2024. We survey over 1,900 AI-tagged crypto projects, filter to investment-focused agents, and curate 10 representative projects spanning strategy and observability dimensions. We then conduct a deep-dive architectural analysis of two prominent agent frameworks, ElizaOS and Virtuals Protocol, and a quantitative on-chain performance analysis of 11 Solana-based agent treasuries with publicly attributable trading activity, covering 925,323 token holders. We find that current deployments remain early and heterogeneous: (1) in our sample, many projects do not yet provide clear evidence of autonomous trade execution, and developer interviews suggest that many visible deployments remain basic API integrations; (2) agent treasuries...

论文介绍 DeFi投资智能体,即利用AI进行自主链上交易的系统,其代币总估值已超过30亿美元。本研究调研了超过1900个AI加密项目,筛选出投资导向的智能体,并对两个突出的智能体框架进行了深度架构分析,同时对11个基于Solana的可归因交易活动的智能体金库进行了量化分析。研究发现,当前的部署仍处于早期且异质阶段,许多项目尚未提供自主交易执行的明确证据,其实际交易行为与表现参差不齐。

Robust and Efficient Guardrails with Latent Reasoning

第一作者: Siddharth Sai · 方向: 安全研究

Abstract:Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrails typically rely on single-pass classification or, more recently, distilled reasoning. Reasoning-based guardrails significantly outperform classification-only baselines, but they incur substantial query latency and token overhead that make them impractical for highthroughput deployment. To address this challenge, we propose COLAGUARD, a guardrail model that transfers multi-step safety reasoning into a continuous latent space through a stage-wise training curriculum, enabling direct hidden-state propagation at inference. Evaluated on ten prompt- and response-moderation settings spanning eight safety benchmarks, COLAGUARD improves macro-F1 by 8.24 points over Llama Guard 3 and matches our explicit reasoning baseline...

论文介绍 现有的大语言模型安全护栏或基于单次分类,或依赖显式推理,后者效果更佳但延迟和token开销巨大,难以用于高吞吐场景。为解决此挑战,本文提出COLAGUARD模型,通过阶段式训练课程,将多步安全推理转移到连续的潜在空间,从而在推理时实现直接的隐状态传播。评估表明,COLAGUARD在保持或超越显式推理基线性能的同时,显著提升了宏观F1分数,并大幅降低了推理延迟。

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

第一作者: Kaihua Qin · 方向: 软件安全

Abstract:Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studies use narrow datasets, inconsistent metrics, and limited semantic consistency checks. This gap is increasingly important as large language models (LLMs) begin to generate source-like Solidity that may compile and appear plausible, even when its semantics diverge from the original contract. We introduce SCDBench, a dataset and benchmark methodology for LLM-based smart contract decompilation. The dataset contains 600 real-world Solidity contracts with paired bytecode inputs, ground-truth source code, and replayable semantic checkpoints. SCDBench evaluates decompiler outputs through four cumulative stages: format completeness, compilability, Application Binary Interface (ABI) recovery, and semantic consistency via differential...

论文介绍 智能合约反编译旨在从字节码恢复高级源代码,但现有评估存在数据集窄、指标不一致和语义检查有限等问题。随着大语言模型能生成看似合理但语义偏离的代码,评估缺陷日益凸显。本文引入SCDBench,一个针对基于LLM的智能合约反编译器的基准数据集和方法。该数据集包含600份真实世界的Solidity合约及其字节码和真实源码,并通过格式完整性、可编译性、ABI恢复和语义一致性四个累积阶段评估反编译输出。

Cycle-Space Informed Detection of Autoencoded Blind False Data Injection Attacks on Power Systems

第一作者: Xin Li · 方向: 系统安全

Abstract:The rapid growth of AI-driven data centers and large-scale energy storage systems is increasing the reliance of power system operation on real-time measurement data and automated decision-making. However, many existing detection methods rely on statistical or data-driven analysis of measurements and can fail when attackers exploit the same data structure to craft stealthy perturbations. To illustrate this limitation, we demonstrate a blind False Data Injection Attack (FDIA) in which an Autoencoder learns the measurement manifold and generates perturbations aligned with the Jacobian null space, thereby allowing the attack to evade both residual-based baddata detectors and time-series anomaly detectors. To mitigate data-driven FDIAs which exploit the null space, we propose a topology-informed Cycle-Space Detector (CSD) that leverages the Cycle-Space of the network to impose...

论文介绍 随着AI驱动的数据中心和大规模储能系统增长,电力系统愈发依赖实时测量数据,但现有检测方法在攻击者利用相同数据结构制造隐蔽扰动时可能失效。本文展示了一种基于自动编码器的盲虚假数据注入攻击,其生成的扰动与雅可比矩阵零空间对齐,从而规避传统检测。为此,研究提出一种利用网络拓扑信息的循环空间检测器,通过在零空间施加拓扑约束来有效检测此类数据驱动的攻击。

Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities

第一作者: Yujie Ma · 方向: 软件安全

Abstract:Large Language Models(LLMs) have been actively integrated into modern software systems as critical components. LLM-in-the-loop vulnerabilities, where vulnerabilities are introduced by LLMs and their dependent downstream components, such as frameworks, introduce new risks. Although some benchmark datasets have been constructed to study the impact of such vulnerabilities, most works still remain at the analysis from the conventional software level, ignoring the harm actually caused by LLMs. Understanding real-world LLM-in-the-loop vulnerabilities is still an open problem. To address this gap, we build the first LLM-in-the-loop vulnerability dataset, LLMCVE, to facilitate the risk analysis of LLM-integrated software. To do so, we first collect 2,888 multi-source vulnerabilities across 230 popular LLM components. Then, through manual analysis, we identify 205 vulnerabilities that...

论文介绍 研究大语言模型集成到软件系统中引入的「LLM-in-the-loop 漏洞」问题。作者构建了首个此类漏洞数据集 LLMCVE,收集了来自 230 个流行 LLM 组件的 2,888 个多源漏洞,并通过手动分析识别出 205 个漏洞。该工作旨在促进对 LLM 集成软件的风险分析,填补了对真实世界漏洞理解的研究空白。

Meta-Quantum Ensemble Framework for Robust Network Intrusion Detection

第一作者: Ritvik Bhatnagar · 方向: AI 安全

Abstract:Intrusion Detection Systems (IDSs) must maintain high detection sensitivity while operating under strict false-positive constraints, a challenge intensified by class imbalance and heterogeneous IoT traffic. This work investigates whether heterogeneous quantum learners can provide useful and non-redundant decision information for IDS tasks. We study Quantum Support Vector Machines (QSVMs) and Quantum Neural Networks (QNNs), which rely on different learning mechanisms and exhibit distinct prediction behaviors. To combine these models, we propose the System-Level Meta-Quantum Ensemble (MQE), a hybrid quantum-classical framework that fuses QSVM and QNN outputs using a Random Forest meta-learner. The meta-learner captures agreement and disagreement patterns between the quantum branches to improve prediction stability and detection performance. Experiments on TON IoT and CICIDS2017...

论文介绍 针对入侵检测系统在类不平衡和异构 IoT 流量下的高检测率挑战,研究量子学习器的应用。提出系统级元量子集成框架 MQE,融合量子支持向量机和量子神经网络的输出,利用随机森林元学习器捕捉模型间的同意和分歧模式。实验在 TON IoT 和 CICIDS2017 数据集上验证了该框架能提升预测稳定性和检测性能。

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

第一作者: Jusuk Lee · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume...

论文介绍 现有机器人学习管道依赖静态视觉编码器,缺乏运动理解。DynaFLIP 提出一个动态感知的多模态预训练框架,利用图像-语言-3D 流三元组作为监督,鼓励三种模态在共享超球空间中保持较小的单纯形体积以实现强对齐。该方法将运动理解上游化到感知阶段,有望提升机器人操作的感知能力。

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

第一作者: Chunru Lin · 方向: 数据集与评测 · 来源: cs.RO

Abstract:The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments. However, current robotic benchmarks primarily emphasize skill-level execution and provide limited insight into such cognitive reasoning capabilities. We introduce RoboWits, a bi-manual robotic benchmark designed to systematically evaluate cognitive reasoning, creative tool use, and robustness to unexpected conditions. To enable scalable construction of high-quality reasoning-centric unexpected scenarios, we propose an automated task generation pipeline formulated as a multi-agent cooperative framework, comprising agents for seed task generation and verification, metric generation, scene generation, and task mutation. Using the pipeline, we curated 30 diverse seed tasks and 208 tasks with mutations and graded difficulty across...

论文介绍 当前机器人基准主要评估技能执行,缺乏对认知推理能力的评估。RoboWits 是一个双臂机器人基准,旨在系统评估认知推理、创造性工具使用和对意外条件的鲁棒性。提出基于多智能体合作的自动化任务生成管道,构建了 30 个种子任务和 208 个变异任务,支持可扩展的推理中心场景构建。

A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms

第一作者: Yufei Jia · 方向: 策略学习 · 来源: cs.RO

Abstract:Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning are placed on a single GPU-centric execution path. This paradigm has greatly improved training speed, but it has also encouraged a default assumption that efficient training requires physics to reside on the GPU. We revisit this assumption. Our view is that, in simulation-dominated robot control, the essential question is not which processor runs physics, but whether simulation throughput, policy learning, and runtime synchronization form an efficient end-to-end loop. We present UniLab, a heterogeneous CPU-simulation / GPU-learning architecture that decouples CPU-parallel simulation from GPU policy updates through a unified runtime for data movement, buffering, and synchronization. UniLab is implemented as a complete and...

论文介绍 现代机器人控制中的基于模拟的强化学习通常依赖 GPU 居中的模拟。UniLab 提出一个异构 CPU-模拟/GPU-学习架构,通过统一运行时解耦 CPU 并行模拟和 GPU 策略更新。该设计重新思考了物理模拟的处理器选择问题,专注于优化模拟吞吐量、策略学习和运行时同步的端到端循环,可能突破 GPU 主导范式的限制。

Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

第一作者: Kuangji Zuo · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to describe which exact object to interact with among similar candidates, where to act on the object, or how the target may change during execution. To address this limitation, we propose Gaze2Act, a novel VLA framework that leverages human gaze as a dynamic and intuitive intent signal for complex interactive manipulation. Gaze2Act first bridges the ego-exo view gap by mapping first-person gaze into the robot's perspective through cross-view semantic matching, producing both an object mask and a gaze point for coarse-to-fine target specification. These cues are then integrated into the policy through perception-level prompting and action-level...

论文介绍 语言指令有时无法精确传达意图,影响复杂交互操作。Gaze2Act 提出一个利用人类凝视作为动态意图信号的 VLA 框架。它通过跨视图语义匹配将第一人称凝视映射到机器人视角,生成目标掩码和凝视点,并集成到策略的感知和动作层。该方法提升了目标指定的精细度,适用于交互式机器人操作。

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

第一作者: Qiuyue Wang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder. Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data, trajectory-centric supervision, and...

论文介绍 具身智能常通过专用模型研究,导致能力碎片化。Qwen-VLA 旨在将异构具身决策问题统一到单一视觉-语言-动作模型中。它扩展了 Qwen 的视觉-语言建模栈,通过基于 DiT 的动作解码器生成连续动作和轨迹。模型在多种数据源上进行大规模联合预训练,涵盖机器人操作、人类演示、模拟数据等,以提升泛化能力。

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

第一作者: Zhongxi Chen · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulation. However, dexterous manipulation remains challenging for VLA policies due to high-dimensional hand control and compounding execution errors, which makes real-world RL post-training essential for bridging the gap between visually grounded action generation and physically reliable dexterous execution. However, high-dimensional dexterous exploration often triggers temporal inconsistency, sample inefficiency and hardware risks in the real world. To address these challenges, we propose BORA, an offline-to-online RL post-training framework designed for real-world dexterous VLA models. In the offline phase, BORA constructs a critic that takes both the VLM's cognition tokens and action chunks as inputs. This design enables...

论文介绍 灵巧操作对 VLA 模型提出挑战,由于高维控制和执行误差。BORA 提出一个离线到在线的 RL 后训练框架,用于现实世界灵巧 VLA 模型。离线阶段构建批评家,结合 VLM 认知 tokens 和动作块;在线阶段通过残差适应减少探索风险。该框架旨在桥接视觉接地动作生成与物理可靠执行之间的差距。

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

第一作者: Shutong Ding · 方向: 策略学习 · 来源: cs.RO

Abstract:Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approaches, one representative branch focuses on the sampling-based policy optimization. This design enables better exploration capability of the diffusion model, particularly at the beginning of training, but suffer from low exploitation in Q-value information, resulting in a slow policy convergence. Another branch pays attention to gradient-based policy optimization, which sufficiently exploits the gradient of the Q function yet tends to collapse into a unimodal policy with low diversity. To address this issue, we propose CGPO, \textbf{C}ritic-\textbf{G}uided diffusion \textbf{P}olicy \textbf{O}ptimization, which effectively balances exploration and exploitation with the training-free guidance technique...

论文介绍 该研究针对基于扩散策略的强化学习中探索与利用的平衡问题。现有方法或探索能力强但利用不足导致收敛慢,或充分利用梯度但陷入单模态多样性低。论文提出CGPO方法,通过无训练的引导技术优化策略,在扩散模型采样过程中有效平衡探索和利用,以提高训练效率和策略性能,可能应用于复杂机器人控制任务。

Replicable Simulation-Based Robot Validation through Provenance

第一作者: Argentina Ortega · 方向: 具身智能 · 来源: cs.RO

Abstract:Robot behavior is often validated through simulation-based testing, yet the replicability of such campaigns depends critically on transparent documentation of how tests are configured, executed, and post-processed. We argue that data provenance, coupled with the FAIR principles (findability, accessibility, interoperability, and reusability), addresses this gap by explicitly tracking links between artifacts and by attaching machine-readable metadata about file origins and key design decisions. Moreover, provenance and metadata cannot be treated as an afterthought confined to final datasets; they must be integrated into the testing processes that generate those datasets so that evidence can be reconstructed end-to-end. We demonstrate this by augmenting an existing simulation-based testing framework with provenance tracking and metadata collection mechanisms, and by using these...

论文介绍 本文关注机器人行为仿真测试的可重复性问题,强调测试配置、执行和后处理的透明文档至关重要。研究提出结合数据溯源和FAIR原则,将机器可读元数据集成到测试框架中,跟踪工件间链接和设计决策,实现端到端证据重构。通过增强现有仿真测试框架,证明其能提高验证过程的可追溯性和可信度,适用于机器人开发质量保证。

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

第一作者: Hao Ren · 方向: 导航与运动 · 来源: cs.RO

Abstract:Diffusion models are effective for waypoint prediction in visual navigation, but standard sampling and test time guidance can produce unreliable or inefficient trajectories when updates drift off the training manifold. We propose Fisher Preserving Guidance with Outer Product Span Projection, a training-free inference method that avoids large Fisher drift associated with off-distribution actions while optimizing a task objective. Our method computes the Fisher-preserving update via a low-rank Jacobian factorization, requiring only a single backward pass per step and enabling real-time use. We further introduce Truncated Fisher Denoising Sensitivity as an uncertainty signal and use it for robust multi-sample action blending. Experiments on toy and realistic navigation benchmarks, including Maze2D with TSDF-based guidance, PushT with official Diffusion Policy weights, and visual...

论文介绍 该论文解决扩散模型在视觉导航中生成轨迹不可靠或低效的问题。当采样偏离训练流形时,标准引导可能导致Fisher漂移。作者提出Fisher保持引导方法,通过低秩Jacobian分解计算保持分布的更新,仅需单次反向传播即可实时使用。此外引入截断Fisher去噪敏感性作为不确定性信号,用于鲁棒的多动作混合,提升导航安全性和效率。

LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation

第一作者: Mohammad Khoshnazar · 方向: 机器人操作 · 来源: cs.RO

Abstract:Multi-step robot manipulation requires acting under uncertainty about how the scene will evolve, making exploration and policy adaptation challenging. We study whether short-horizon, task-consistent future videos can provide useful structured priors for control and reinforcement-learning fine-tuning. We formalize this idea through Future-Experience Conditioning (FEC), a simple interface that conditions closed-loop policies on a latent representation of a short future video. In our simulation setup, future clips are generated in three stages, an LLM reasoner operating over a task ontology initialized from the current scene state, a robot-free digital-twin rollout of the intended object motion, and a mask-free video diffusion model that synthesizes a robot-consistent future clip without requiring segmentation at inference. We instantiate this future-conditioning interface...

论文介绍 本文研究多步骤机器人操作中场景演变不确定下的探索挑战。提出未来体验条件化接口,通过大语言模型基于任务本体生成短期未来视频,作为结构化先验指导控制。方法包括LLM推理、数字孪生回放和视频扩散模型合成未来片段,用于策略条件化和强化学习微调,以增强策略适应性和任务完成能力。

MARS Policy: Multimodality Only When It Matters

第一作者: Jindou Jia · 方向: 机器人操作 · 来源: cs.RO

Abstract:Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to capture diverse yet valid behavioral patterns, has driven the rapid emergence of generative policies as a dominant paradigm in robot learning. However, achieving such multimodality typically relies on stochastic noise initialization and iterative denoising procedures, resulting in substantial training complexity and low inference efficiency. Meanwhile, not all phases of a robotic task inherently require behavioral diversity. Motivated by this insight, we propose the Modality-Adaptive Robot Sampling (MARS) policy, which adaptively invokes tailored stochasticity only when it is truly beneficial, while reverting to an efficient deterministic learning during single-modal phases. In other words, the proper amount of noise is injected only...

论文介绍 模仿学习中的多模态策略虽能捕捉多样行为,但依赖随机噪声初始化和迭代去噪,导致训练复杂且推理低效。论文提出MARS策略,在需要多模态的阶段自适应注入随机性,而在单模态阶段切换为确定性学习,以简化训练并提高推理效率。实验表明该方法在机器人操作任务中保持性能的同时优化资源使用。

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

第一作者: Sergey Arkhangelskiy · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with $N \le 25$ rollouts per condition, almost always without confidence intervals or paired statistical comparison; these cohort sizes struggle to resolve close comparisons reliably. We introduce PhAIL (Physical AI Leaderboard, this https URL), an open real-robot benchmark on a Franka FR3 (dataset, per-rollout artifacts, and end-to-end reference implementation) of a distributional evaluation methodology: the time-to-success cumulative distribution function (CDF) as the evaluation primitive, with two separated jobs. The first is scoring via Human-Relative Throughput (HRT), a dimensionless scalar with bootstrap confidence intervals, anchored to same-fixture human teleoperation. The second is a significance test (Kolmogorov-Smirnov, computed per-object and...

论文介绍 现有视觉-语言-动作模型评估常基于二元成功率,缺乏统计严谨性。本文提出PhAIL基准,采用分布式评估方法,以时间到成功累积分布函数为基本指标,引入人类相对吞吐量进行评分,并使用统计检验进行显著性测试。该框架提供真实机器人数据集和端到端实现,旨在提高模型比较的可靠性和可重复性。

EXACT-MPPI: Exact Signed-Distance Navigation for Arbitrary-Footprint Robots from Point Clouds via Path Integral Control

第一作者: Chen Peng · 方向: 导航与运动 · 来源: cs.RO

Abstract:Ground robots often carry payloads, implements, or other attachments that turn their effective footprint into complex, non-convex shapes. Navigating safely through clutter then requires reasoning about this true geometry, yet most local planners simplify it with convex or inflated proxies and rasterize sensor data into occupancy grids or distance fields. Both choices eliminate feasible motions when clearance is comparable to the footprint geometry. We present EXACT-MPPI, a training-free local navigation framework that maps local point-cloud observations and sparse guidance directly to motion commands, without any intermediate map representation. The framework embeds an analytic, exact signed-distance evaluator into a Model Predictive Path Integral (MPPI) controller. The footprint is represented as a simple polygon for general convex or concave planar shapes, with a...

论文介绍 该研究针对携带复杂非凸足迹的地面机器人在杂乱环境中的导航安全问题。现有方法简化几何或栅格化数据会损失可行动作。论文提出EXACT-MPPI框架,直接从点云观测和稀疏引导生成运动命令,无需中间地图。它将精确符号距离评估器嵌入模型预测路径积分控制器,支持任意多边形足迹,提高导航精度和实时性。

VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models

第一作者: Dehao Huang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Confidence estimation for Vision-Language-Action (VLA) models is essential for robots to perform manipulation tasks in the open world, providing crucial signals for risk-sensitive decision-making and failure anticipation. Existing confidence estimation methods typically rely on ensemble-based paradigms or action-token probabilities to predict the likelihood of task success. However, they still encounter challenges in computational efficiency and cross-architecture generalizability. These methods usually require repeated sampling, leading to inference inefficiency, and are restricted to VLA models with discrete action outputs, making them difficult to apply to continuous action spaces. To address this issue, we propose VLAConf, a one-class discriminative confidence framework. By leveraging frozen pretrained VLA internal representations, VLAConf directly estimates step-wise...

论文介绍 视觉-语言-动作模型的置信度估计对风险决策和故障预测至关重要,但现有方法效率低且泛化差。本文提出VLAConf,一个单类判别置信度框架,利用冻结的预训练VLA内部表示直接估计逐步任务成功概率,避免重复采样。该方法适用于连续动作空间,支持计算高效和架构泛化,提升机器人开放世界操作中的可靠性。

Learning to Feel Materials from Multisensory Tactile Data via Interpretable Models

第一作者: Li Zou · 方向: 具身智能 · 来源: cs.RO

Abstract:Human tactile perception of materials relies on complex multisensory touch cues, yet the relationship between low-level tactile signals and perceptual representations remains poorly understood. This knowledge gap hinders the integration of touch in digital environments and the development of robots capable of human-like tactile perception. Here, we present an interpretable computational framework for modeling human material perception and recognition using multisensory touch data. Our framework comprises three interconnected models: Model 1 maps finger-surface interaction features to psychophysical sensory attributes, Model 2 classifies materials based on these perceptual representations, and Model 3 directly classifies materials from tactile features. The results showed that combining information from pressing, static contact, and sliding interactions improves prediction...

论文介绍 本文提出一个可解释的计算框架,用于建模人类的材料感知与识别。该框架利用多模态触觉数据,解决了底层触觉信号与高层感知表征之间关系不明确的问题。它包含三个相互连接的模型,分别用于将交互特征映射到心理物理属性、基于感知表示分类材料,以及从触觉特征直接分类。研究表明,结合按压、静态接触和滑动交互信息能提升预测准确性。该工作有助于在数字环境中集成触觉感知,并开发具有类人触觉能力的机器人。

VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

第一作者: Victor Kowalski · 方向: 机器人操作 · 来源: cs.RO

Abstract:When using reinforcement learning (RL) for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve. However, vision-enabled policies tend to overfit to the visual conditions seen during training, limiting their robustness and transferability. We present a human-in-the-loop RL framework that employs teacher-student distillation to achieve robust performance across multiple task variants, trained entirely in the real world without requiring domain randomization or data augmentation. A vision-enabled teacher distills its knowledge into a vision-free student that relies solely on pose, twist, and wrench sensing, combining fast training with strong task generalization. On the real-world NIST assembly benchmark board, our approach achieves 95\% overall success after approximately 50...

论文介绍 本文提出一种用于机器人接触式操作的人机在环强化学习框架。其核心是通过教师-学生蒸馏,将依赖视觉的“教师”策略知识迁移至仅依赖本体感觉、位姿及力信息的“无视觉”“学生”策略,从而提升策略在多种任务变体间的鲁棒性和可迁移性。整个训练完全在现实世界进行,无需域随机化或数据增强。在NIST装配基准测试中,该方法在约50次真实交互后取得了95%的总体成功率。

VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models

第一作者: Shengyu Si · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and action patterns. This paper proposes VLA-Pro, a plug-and-play framework designed to enhance cross-task generalization by storing task-relevant procedural memories at training time and transferring these memories during inference. Specifically, VLA-Pro stores task-specific LoRA adapters as parameterized procedural memories during training. At inference time, VLA-Pro retrieves relevant procedural memories based on the current multi-modal context and dynamically fuses these memories for generating the current action chunk. Experiments on RoboTwin, RLBench, and real-world manipulation tasks show that VLA-Pro consistently improves cross-task...

论文介绍 针对视觉-语言-动作(VLA)模型在未见任务上泛化能力不足的问题,本文提出VLA-Pro框架。该框架在训练阶段将任务相关的程序记忆(以任务特定的LoRA适配器形式)存储起来,并在推理阶段根据当前多模态上下文检索并动态融合这些记忆,以生成动作序列。这种即插即用的方法无需额外训练,通过在多个仿真和现实操作任务中的实验,证明了其能持续改善VLA模型的跨任务泛化性能。

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

第一作者: Ye Li · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time robotic manipulation, especially when large vision-language backbones and iterative action heads run at every control step. Existing VLA acceleration methods often optimize individual components or rely on fixed acceleration rules, treating different control steps with largely fixed computation and overlooking the non-uniform reasoning demands of sequential embodied control. Inspired by human motor control, where cognitive and feedback resources concentrate on goal-sensitive stages, we argue that VLA models should learn when to invest full computation and when to reuse prior computation. We propose ElegantVLA, a plug-in phase-adaptive inference framework that accelerates VLA models through intra-model...

论文介绍 视觉-语言-动作(VLA)模型因高计算成本限制了其在实时机器人操作中的应用。现有加速方法通常对不同控制阶段施加固定计算量,忽略了顺序控制中推理需求的不均匀性。受人类运动控制启发,本文提出ElegantVLA,一种插件式的相位自适应推理框架。它能让模型学习在不同阶段何时投入完整计算、何时复用先前计算,通过模型内部分阶段自适应来加速VLA模型,从而在保持性能的同时提高控制频率。

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding

第一作者: Zhongyu Xia · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficiency manifests as three intertwined challenges: weak extraction of 3D spatial positions without enforcing multi-view consistency, inadequate 3D instance understanding, and fragile reasoning under occlusion. Although mature 3D perception methods exist, their direct integration into VLA pipelines is hindered by architectural incompatibility and by heavy reliance on costly instance-level annotations. To address the above challenges, we propose 3DVLA, a plug-and-play framework that injects robust 3D reasoning into pretrained VLAs without requiring extra manual labels or discarding VLM priors. Specifically, 3DVLA tackles the three challenges through: (1) pervasive 3D feature encoding with explicit...

论文介绍 视觉-语言-动作(VLA)模型在机器人操作中表现优异,但常缺乏对三维场景的理解。本文提出3DVLA,一个即插即用的框架,旨在向预训练的VLA模型中注入鲁棒的三维推理能力。它通过无处不在的三维特征编码、无监督实例理解以及基于扩散模型的遮挡推理,解决了空间位置提取、实例理解及遮挡推理三大挑战。该方法无需额外人工标注,也不丢弃视觉语言模型原有的先验知识,从而增强了模型的空间与实例理解能力。

Phase-Conditioned Imitation Learning with Autonomous Failure Recovery for Robust Deformable Object Manipulation

第一作者: Dayuan Chen · 方向: 机器人操作 · 来源: cs.RO

Abstract:This paper presents a phase-conditioned, force-aware framework for robust deformable object manipulation. Standard imitation learning policies such as Action Chunking with Transformers (ACT) rely on a Markovian assumption at inference, causing state aliasing when visually similar observations require contradictory actions and preventing autonomous recovery from execution failures. We address this with a closed-loop hierarchical architecture. A FiLM-conditioned ACT encoder modulates feature extraction based on the current task phase, enabling a single unified policy to produce phase-specific behaviors while sharing action dynamics across phases. A multi-modal phase predictor fusing visual, force, and pose feedback estimates the phase in real time, detecting contact failures that are invisible to vision alone and autonomously triggering recovery trajectories. The system is...

论文介绍 本文针对可变形物体操作,提出一种相位条件化、力感知的鲁棒框架。标准的模仿学习策略(如ACT)在推理时依赖马尔可夫假设,易导致状态别名且无法自主从执行失败中恢复。该框架采用闭环分层架构:一个基于FiLM条件的ACT编码器根据任务相位调整特征提取,使得单一策略能产生相位特异性行为;一个多模态相位预测器实时融合视觉、力与位姿反馈来估计相位,检测视觉不可见的接触失败并触发恢复轨迹。

Decentralized LLM-Driven Coordination of Acoustic Robots for Contactless Object Manipulation

第一作者: Yingying Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Natural language interfaces can simplify interaction with multi-robot systems, especially when non-expert users need to issue high-level commands. Acoustic manipulation using ultrasonic phased arrays also enables contactless object handling for applications such as healthcare, laboratory automation, and precision transport. However, combining large language models (LLMs) with distributed acoustic mobile robots remains underexplored. This paper presents a decentralized framework for natural language-driven coordination of acoustic robots for contactless object manipulation. The system converts spoken instructions into executable multi-robot task plans using Whisper-based speech recognition, LLM-based semantic parsing, structured JSON task representation, and distributed scheduling. The JSON schema encodes robot assignments, temporal dependencies, spatial constraints, and...

论文介绍 本文提出一种去中心化框架,用于通过自然语言驱动声学机器人进行非接触物体操作。该系统将口语指令转换为可执行的多机器人任务计划,涉及基于Whisper的语音识别、基于大语言模型的语义解析、结构化JSON任务表示以及分布式调度。JSON模式编码了机器人分配、时间依赖、空间约束等信息,使多个基于超声相控阵的声学机器人能够协调操作,适用于医疗保健、实验室自动化等场景。

MonoDuo: Using One Robot Arm to Learn Bimanual Policies

第一作者: Sandeep Bajamahal · 方向: 机器人操作 · 来源: cs.RO

Abstract:Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-arm robots, however, are widely available in research labs. Can we leverage them to train bimanual robot policies? We present MonoDuo, a framework for learning bimanual manipulation policies using single-arm robot demonstrations paired with human collaboration. MonoDuo collects data by teleoperating a single-arm robot to perform one side of a bimanual task while a human performs the other, then swapping roles to cover both sides. RGB-D observations from a wrist-mounted and fixed camera are augmented into synthetic demonstrations for target bimanual robots using state-of-the-art hand pose estimation, image and point cloud segmentation, and inpainting. These synthetic demonstrations, grounded in real...

论文介绍 学习双臂机器人策略受限于双臂机器人和数据的稀缺性。本文提出MonoDuo框架,旨在利用广泛可用的单臂机器人演示来学习双臂操作策略。其方法是通过遥操作单臂机器人完成双臂任务的一侧,同时人类完成另一侧,然后交换角色以覆盖两侧。利用先进的手部姿态估计、分割和图像修复技术,将腕部及固定摄像头的RGB-D观察增强为针对目标双臂机器人的合成演示。这些基于真实数据的合成演示可用于训练有效的双臂策略。

Extreme dynamic symmetry enables omnidirectional and multifunctional robots

第一作者: Jiaxun Liu · 方向: 导航与运动 · 来源: cs.RO

Abstract:Symmetry is a central organizing principle in natural systems, yet its use as a unifying design strategy in robotics has largely remained limited to geometric form. We show that symmetry can instead be leveraged at the level of dynamic actuation capability. We introduce dynamic symmetry, the uniformity of a robot's attainable center-of-mass accelerations, and formalize it through a measure coined as dynamic isotropy. Across more than 1000 simulated morphologies, we found that higher dynamic symmetry consistently improved trajectory tracking, task success, robustness, resiliency, and energy efficiency, with the benefits becoming most pronounced as dynamic isotropy approached its theoretical limit. To study this regime systematically, we developed Argus, a family of spherical robots designed to explore the effects of increasing dynamic symmetry. Members of the Argus family vary...

论文介绍 研究探索了如何将对称性作为统一设计策略应用于机器人的动态驱动能力层面。论文提出了“动态对称性”概念,即机器人可达到的质心加速度的均匀性,并通过“动态各向同性”度量进行形式化。通过对超过1000种仿真形态的研究表明,更高的动态对称性能够持续提升轨迹跟踪、任务成功率、鲁棒性和能效。为系统研究此规律,研究团队设计了Argus系列球形机器人,并通过实验验证了高动态对称性设计的实际效益。

Learning and Adaptation in Wire Arc Additive Manufacturing Bead Geometry Control

第一作者: Chen-Lung Lu · 方向: 具身智能 · 来源: cs.RO

Abstract:Robotics Wire Arc Additive Manufacturing (WAAM) is governed by complex and nonlinear process dynamics coupling thermal field to the build geometry. The process may be regarded as a multi-input/multi-output dynamical system with welding torch speed and wire feed rate as inputs and weld bead deposition height and width as outputs. In this paper, we use the input/output data to learn a data-driven model and use it for weld planning and control. We show that a simple recurrent neural network architecture and one-step-ahead predictive control can improve the process performance in terms of height and width consistency. To account for the changing thermal conditions during the printing process, we update the learning model using prediction error from the previous layer. This adaptation step further improves the prediction accuracy and controller performance. Experiments on a robotic...

论文介绍 本文研究了机器人电弧增材制造中,复杂非线性过程动力学对焊道几何形状的影响。研究将焊接过程建模为多输入多输出动态系统,并利用输入输出数据学习一个数据驱动模型。采用简单的循环神经网络架构和单步预测控制器,以改善焊道高度和宽度的一致性。为应对打印过程中热条件的变化,研究还利用前一层的预测误差对模型进行在线更新,从而进一步提高了预测精度和控制器性能。

Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

第一作者: Petras Swissler · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Swarm and field robotics face significant barriers to real-world validation due to the high cost and development time to deploy hardware. This paper introduces the ``Bionic Swarm,'' a novel system that lowers these barriers by abstracting away many of the tasks that are difficult to implement on robots but which do not contribute to the overall algorithm evaluation, giving these tasks to human users. These human users take directions from a smartphone web-app that takes measurements from Bluetooth-connected sensors and relays them to a centralized server. This server runs the swarm algorithm and directs actions to the human users. We evaluate this system through the experimental validation of a geotechnically-focused search algorithm named Score-Biased-Search, which functions by assigning a ``score'' to each location on a reconstructed map, then biases search patterns through...

论文介绍 为降低群体和现场机器人硬件部署的高成本与开发时间,本文提出了“仿生群”系统。该系统通过将部分对算法评估非关键的机器人任务抽象化,并交由人类用户执行。人类用户通过手机应用接收指令,利用蓝牙传感器进行测量,并将数据传回中央服务器。服务器运行群体算法并指导用户操作。该系统通过一项地质技术导向的搜索算法“Score-Biased-Search”的实验验证,展示了其在实际世界土壤测绘任务中的可行性。

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

第一作者: Luzhou Ge · 方向: 多模态具身 · 来源: cs.RO

Abstract:Integrating open-vocabulary semantic information into dynamic 3D scene representations is essential for long-term embodied scene understanding. However, existing methods often suffer from fragile instance association due to incomplete cross-view cues, while their limited ability to handle object-level topological changes restricts long-term robotic task execution. Moreover, current 3D scene understanding methods either rely on simple feature matching without explicit spatial reasoning or assume offline ground-truth 3D geometry. To address these challenges, we present DGSG-Mind, a hybrid instance-aware 3D Gaussian dynamic scene graph system with an embodied reasoning agent. Our system couples a probabilistic voxel grid with explicit 3D Gaussians to enable robust cross-modal instance fusion and incremental semantic mapping. It handles dynamic changes through Gaussian-based...

论文介绍 为解决现有方法在动态3D场景表示中处理物体拓扑变化能力有限、导致长期机器人任务执行困难的问题,本文提出了DGSG-Mind系统。该系统是一种混合实例感知的3D高斯动态场景图,结合了概率体素网格与显式3D高斯表示。它能实现稳健的跨模态实例融合和增量式语义建图,并通过高斯表示与概率体素结合来处理场景中的动态变化,从而支持对开放词汇语义信息的集成和长期场景推理。

Energy-Aware NECO for Single-Pass Pixel-wise Out-of-Distribution Detection in Semantic Segmentation

第一作者: Boyuan Zhang · 方向: 具身智能 · 来源: cs.RO

Abstract:Reliable semantic segmentation for mobile robots requires both accurate dense prediction and robust uncertainty estimation under distribution shift. Strong uncertainty baselines such as Monte Carlo Dropout often require repeated stochastic forward passes and are difficult to deploy on edge platforms. We propose Energy-Aware NECO, a single-pass pixel-wise out-of-distribution (OOD) detector for semantic segmentation. The method combines a centered NECO-style geometric ratio computed from decoder features with a logit-based Energy score. Both components are standardized using statistics fitted on a pure in-distribution validation split and fused through a convex combination. We evaluate the method on the miniMUAD subset using true pixel-level OOD labels. The proposed hybrid score achieves an AUROC of 0.8539, outperforming NECO-only (0.8280), Energy-only (0.8171), and an ensemble...

论文介绍 针对移动机器人在分布偏移下需要可靠语义分割的问题,现有基于蒙特卡洛丢弃的强不确定性方法需要重复的随机前向传播,难以在边缘平台部署。本文提出了Energy-Aware NECO,一种单遍像素级分布外检测方法。该方法结合了基于解码器特征的NECO风格几何比率和基于逻辑的能量分数,两者通过纯分布内验证集拟合的统计量进行标准化,并以凸组合方式融合。实验表明,这种混合分数在像素级OOD检测上表现优异。

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

第一作者: Mohammadreza Teymoorianfard · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework...

论文介绍 本文研究了集成推理能力的视觉-语言-动作模型在自动驾驶中的鲁棒性。研究表明,这些模型对现实的文本输入扰动非常脆弱。通过使用NVIDIA的Alpamayo模型进行黑盒研究,发现在闭环仿真中,扰动可导致推理失败率高达89%,轨迹操纵成功率达72%,进而增加碰撞率并降低安全指标。研究引入了一个推理感知的评估框架来系统性地探测这些模型的漏洞,揭示了其在推理与轨迹生成耦合假设下的脆弱性。

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

第一作者: Jiyao Zhang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelligence in embodied 3D environments. To systematically evaluate these foundational perceptual capabilities, the benchmark includes 6 task categories divided into two core groups: Spatial Structural Understanding (Grounding, Spatial Relation Prediction, and Multi-view Correspondence) and Interaction-Oriented Perception (Affordance Prediction, Grasp Point Prediction, and Trajectory Prediction). The benchmark spans 12 subcategories and contains over 21k high-quality question-answer pairs. We evaluate 13 state-of-the-art models, and the results show that while current models exhibit relatively strong high-level spatial reasoning, such as understanding...

论文介绍 为评估视觉语言模型在3D具身环境中的低级空间智能,本文提出了一个以机器人为核心的基准测试Embodied3DBench。该基准包含空间结构理解和面向交互的感知两大类共6个任务,如空间关系预测、抓取点预测等,涵盖12个子类别和超过2.1万个高质量问答对。对13个先进模型的评估显示,当前模型在高级空间推理上表现相对较好,但在低级感知能力(如抓取点预测)上仍存在明显不足。

Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems

第一作者: Daniel Albiero · 方向: 具身智能 · 来源: cs.RO

Abstract:Tropical forests worldwide are under intense deforestation pressure driven by economic and political interests, and scientific evidence suggests this deforestation contributes to climate change. This paper proposes a novel logging method for tropical forests, Ultra-Reduced-Impact-Encased-Logging (URIEL). This new method is based on heli-logging techniques combined with intensive use of robotics and AI integrated with post-harvest silvicultural treatments performed by drones. The concept of appropriate equipment for this method was developed, dimensions were determined, details were completed in a digital proof of concept, and an effective digital simulation and economic feasibility analysis were carried out for various helicopter-timber-distance combinations. The results demonstrated that a URIEL method has high economic viability and makes it possible to virtually eliminate...

论文介绍 针对热带森林面临的巨大毁林压力,本文提出了一种名为URIEL的新型伐木方法。该方法结合了直升机伐木技术与机器人和人工智能的密集使用,并由无人机执行采伐后的林分抚育处理。研究开发了该方法的概念设备,确定了尺寸,并通过数字仿真和经济可行性分析进行了验证。结果表明,URIEL方法具有较高的经济可行性,并能几乎消除传统伐木作业对森林环境的破坏性影响。

VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

第一作者: Mingjian Gao · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irrelevant or weakly textual information can interfere with action prediction, while autoregressive text decoding adds too much latency for real-time closed-loop execution. We present VISUALTHINK-VLA, a visual intermediate-reasoning framework for accurate, low-latency VLA policies. Our bootstrapping philosophy is to guide action with effective visual thinking: VISUALTHINK-VLA bootstraps action prediction through a compact visual-evidence interface that preserves spatial precision while avoiding decoding overhead. Besides, to further improve performance and efficiency, VISUALTHINK-VLA adopts a tailored selective routing mechanism to learn the visual evidence tokens, enabling low-latency inference while...

论文介绍 本文针对视觉语言动作(VLA)策略中链式思考文本解码的高延迟和干扰问题,提出VisualThink-VLA框架。该框架通过紧凑的视觉证据接口引导动作预测,保留空间精度同时避免解码开销,并采用选择性路由机制优化推理效率,旨在实现准确、低延迟的实时闭环机器人控制。

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

第一作者: Shilin Ma · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across execution steps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce an...

论文介绍 为解决视觉语言动作模型中视觉令牌修剪基于浅层线索可能导致关键信息丢失的问题,本文提出SAFE-Pruner框架。该框架利用语义注意力一致性,通过前向策略预测深层令牌的显著性,从而在修剪时保留重要信息,实现更稳定的加速推理,适用于机器人操控等实时应用。

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning

第一作者: Kyujin Lee · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Vision-Language-Action (VLA) models have emerged as a promising framework that unifies perception, reasoning, and control for robot manipulation by adapting pretrained vision-language models (VLMs) to action prediction. However, VLM-derived representations are often insensitive to subtle visual distinctions required for low-level control, causing state aliasing between visually similar states that require substantially different actions. Prior VLA studies improve visual understanding by generating visual or reasoning outputs, such as future frames, 2D grounding points or traces, or intermediate spatial reasoning steps, but these objectives typically shape the vision encoder only indirectly through end-to-end prediction and do not explicitly analyze state aliasing in the learned visual feature space. To mitigate state aliasing, we introduce inverse dynamics learning as an...

论文介绍 视觉语言动作模型常因视觉表示对细微差异不敏感而产生状态别名,影响动作预测。本文引入逆动力学学习作为显式目标,直接分析视觉特征空间,以缓解状态别名问题。该方法旨在提升VLA模型在相似视觉状态下做出不同动作的能力,增强机器人操控的准确性。

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

第一作者: Xueqing Wu · 方向: 策略学习 · 来源: cs.CV

Abstract:Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a bottleneck for end-to-end visual reasoning. To investigate this gap, we introduce a controlled diagnostic framework with two synthetic tasks that disentangle perception from reasoning. Our analysis reveals a consistent perception-reasoning asymmetry: posttraining improves reasoning more substantially than perception, though the underlying mechanism differs by training paradigm. For supervised fine-tuning (SFT), this asymmetry stems from token imbalance in chain-of-thought supervision, where perception occupies fewer tokens and thus receives a weaker training signal. Dynamically reweighting the loss mitigates this imbalance and boosts end-to-end performance by up to 18.2. For reinforcement learning (RL), the asymmetry instead...

论文介绍 研究视觉语言模型后训练中推理与感知优化的不对称现象。通过控制诊断框架,发现监督微调中token不平衡和强化学习中奖励设计问题是关键。提出损失重加权等策略,以平衡感知和推理的训练信号,提升端到端视觉推理性能。

AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference

第一作者: Yilin Feng · 方向: 多模态具身 · 来源: cs.LG

Abstract:Vision-Language Models (VLMs) process thousands of visual tokens per image alongside comparatively few text tokens, yet existing compression methods treat both modalities uniformly. We observe that the two modalities have fundamentally different properties: vision tokens are spatially redundant and dominate prefill, while text tokens are causally dependent and accumulate during decoding. Based on this asymmetry, we propose and empirically evaluate AsymVLM, which applies aggressive pruning to vision tokens before prefill using a learned importance scorer with per-sample adaptive budgeting, and temporal threshold-based eviction to text tokens only when they exceed a fixed budget. Our experiments indicate that AsymVLM achieves the highest FLOPs savings (up to 54%) among state-of-the-art methods while outperforming existing approaches by 2--3% on document and chart understanding...

论文介绍 针对视觉语言模型中视觉和文本令牌处理的不对称性,本文提出AsymVLM方法。该方法对视觉令牌进行基于重要性评分的激进修剪,对文本令牌采用阈值驱逐策略,以优化推理效率。实验表明,AsymVLM在保持性能的同时实现高FLOPs节省。

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

第一作者: Haoyuan Shi · 方向: VLA 通用模型 · 来源: cs.AI

Abstract:Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic framework that analyzes VLA models through a unified evidence chain from representation dynamics to causal control attribution and behavioral manifestation. It specifically combines cross-modal and checkpoint-drift centered kernel alignment (CKA) to trace representation evolution, attention knockout interventions to identify modality-specific control pathways, and rollout-level behavioral probes to examine grounding, shortcut dependence, and semantic following. Experiments on $\pi_{0.5}$ and OpenVLA reveal three key findings. First, the two models exhibit distinct modality-specific adaptation dynamics during VLA finetuning. Second, they rely on different multimodal routing strategies and layer-wise...

论文介绍 理解视觉语言动作模型如何将多模态知识转化为控制是关键挑战。本文提出VLA-Trace诊断框架,通过表示动态追踪、注意力干预和行为探针,分析模型从表征到行为的完整链路。该框架揭示不同VLA模型的适配差异,为模型改进提供 insights。

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

第一作者: Tianzhuo Yang · 方向: 数据集与评测 · 来源: cs.AI

Abstract:Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias...

论文介绍 现有机器人世界模型评估侧重于视觉保真度,缺乏对动作条件可靠性的系统评测。本文提出MiraBench基准,定义动作条件可靠性为核心目标,分解为物理一致性、动作遵循保真度和乐观偏差三个层次,旨在标准化评估世界模型的预测可靠性。

市场总览

美股技术面呈现强势但分化的格局:SPY和QQQ的RSI分别高达74.1和77.2,处于超买状态,且接近52周高点,显示上行动能强劲但回调压力增大;QQQ出现MACD金叉信号,短期动量延续。个股如MSFT因MACD金叉和显著日涨幅5.45%显示技术突破,而AAPL虽维持多头排列但RSI超买限制上行空间。加密货币板块整体疲软,辅助背景显示恐慌贪婪指数仅23(极度恐慌),总市值2.56万亿美元,BTC主导率57.5%;BTC、ETH、SOL均呈空头排列,RSI偏弱(35.5、32.2、40),技术面偏下行。中概股普遍承压,BABA、PDD、JD触发MACD死叉和空头排列,5日跌幅显著,如PDD达13.65%,下行趋势明确。商品外汇中,黄金期货RSI47.3中性,趋势震荡;原油期货下跌;美元兑人民币空头排列且接近52周低点。整体而言,市场技术面分布不均,美股超买强势,加密与中概股弱势,商品外汇分化。

今日关注

MSFT Microsoft (MSFT)
偏上行

MACD出现金叉信号,当前价450.24高于20日均线417.59和50日均线402.83,短期均线呈多头支撑;1日涨幅5.45%显示强劲动量,RSI14为69.6处于正常区间,无超买压力,但价格低于200日均线458.46,长期阻力仍需观察。

AAPL Apple (AAPL)
中性

RSI14高达78.8进入超买状态,暗示短期回调风险;但价格312.06高于所有关键均线(SMA20 297.54、SMA50 275.28、SMA200 263.24),保持多头排列,且接近52周高点仅差0.93%,上行空间受限,技术面呈现矛盾信号。

BTC-USD Bitcoin (BTC-USD)
偏下行

趋势为空头,价格73474.33低于所有关键均线(SMA20 77580.28、SMA50 77207.27、SMA200 79809.73),空头排列明确;RSI14为35.5虽未超卖但偏弱,MACD值-924.77低于信号线-311.03,下跌动能持续。

PDD 拼多多 (PDD)
偏下行

触发MACD死叉信号,价格84.44低于所有均线(SMA20 95.79、SMA50 98.42、SMA200 113.23),空头排列稳固;RSI14为32.3接近超卖但未确认,5日跌幅达13.65%显示强烈下跌动量,技术面偏弱。

USDCNY=X 美元/人民币
偏下行

MACD死叉信号出现,价格6.77低于均线(SMA20 6.80、SMA50 6.83、SMA200 6.99),空头排列显著;接近52周低点仅差0%,RSI14为30.6偏弱但未超卖,下跌趋势延续。

全部资产

^VIX

VIX 恐慌指数

$15.32 -2.67%
5 日
-8.26%
距 52w 高
-56.6%
RSI(14)
35.4
趋势
中性
SMA 20 / 50 / 200
17.25 / 19.95 / 18.38
MACD / 信号
-0.913 / -0.846
MACD 死叉 (1 天前)

^TNX

10Y 美债收益率 (%)

$4.45 -0.04%
5 日
-2.90%
距 52w 高
-10.9%
RSI(14)
49.8
趋势
多头
SMA 20 / 50 / 200
4.48 / 4.39 / 4.20
MACD / 信号
0.040 / 0.055
MACD 死叉 (2 天前)多头排列

DX-Y.NYB

美元指数 DXY

$98.94 -0.08%
5 日
-0.25%
距 52w 高
-1.7%
RSI(14)
51.8
趋势
多头
SMA 20 / 50 / 200
98.72 / 98.90 / 98.57
MACD / 信号
0.140 / 0.096
接近 52 周高多头排列

SPY

S&P 500 ETF

$756.48 +0.25%
5 日
+1.85%
距 52w 高
-0.2%
RSI(14)
74.1
趋势
多头
SMA 20 / 50 / 200
739.34 / 703.65 / 681.17
MACD / 信号
12.698 / 12.873
RSI 超买接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$738.31 +0.37%
5 日
+3.33%
距 52w 高
-0.4%
RSI(14)
77.2
趋势
多头
SMA 20 / 50 / 200
709.04 / 652.93 / 617.85
MACD / 信号
21.470 / 21.438
MACD 金叉 (今天)RSI 超买接近 52 周高多头排列

AAPL

Apple

$312.06 -0.14%
5 日
+2.32%
距 52w 高
-0.9%
RSI(14)
78.8
趋势
多头
SMA 20 / 50 / 200
297.54 / 275.28 / 263.24
MACD / 信号
10.392 / 9.777
RSI 超买接近 52 周高多头排列

MSFT

Microsoft

$450.24 +5.45%
5 日
+7.43%
距 52w 高
-18.9%
RSI(14)
69.6
趋势
中性
SMA 20 / 50 / 200
417.59 / 402.83 / 458.46
MACD / 信号
5.802 / 4.124
MACD 金叉 (今天)

NVDA

Nvidia

$211.14 -1.45%
5 日
-3.81%
距 52w 高
-10.7%
RSI(14)
49.4
趋势
多头
SMA 20 / 50 / 200
215.46 / 199.35 / 187.65
MACD / 信号
3.809 / 5.974
多头排列

GOOGL

Alphabet

$380.34 -2.51%
5 日
-1.89%
距 52w 高
-6.9%
RSI(14)
52.9
趋势
多头
SMA 20 / 50 / 200
391.15 / 347.57 / 299.92
MACD / 信号
9.641 / 13.543
多头排列

TSLA

Tesla

$435.79 -1.43%
5 日
+4.29%
距 52w 高
-12.6%
RSI(14)
60.0
趋势
中性
SMA 20 / 50 / 200
421.39 / 391.80 / 412.13
MACD / 信号
12.067 / 11.367
MACD 金叉 (2 天前)

META

Meta

$632.51 -0.44%
5 日
+4.14%
距 52w 高
-20.6%
RSI(14)
55.4
趋势
中性
SMA 20 / 50 / 200
613.33 / 618.53 / 666.57
MACD / 信号
-1.092 / -4.297
MACD 金叉 (2 天前)
加密恐慌贪婪
23
极度恐慌
加密总市值
$2.56 T
+0.33% / 24h
BTC 主导率
57.5%
ETH 9.5%
24h 成交量
$88.9 B
活跃币 17,410

BTC-USD

Bitcoin

$73,474.33 -0.08%
5 日
-4.56%
距 52w 高
-41.8%
RSI(14)
35.5
趋势
空头
SMA 20 / 50 / 200
77,580.28 / 77,207.26 / 79,809.73
MACD / 信号
-924.765 / -311.031
空头排列

ETH-USD

Ethereum

$2,015.00 +0.37%
5 日
-3.96%
距 52w 高
-59.3%
RSI(14)
32.2
趋势
空头
SMA 20 / 50 / 200
2,152.74 / 2,251.80 / 2,511.45
MACD / 信号
-65.386 / -52.263
空头排列

SOL-USD

Solana

$82.42 +0.53%
5 日
-3.32%
距 52w 高
-67.4%
RSI(14)
40.0
趋势
空头
SMA 20 / 50 / 200
87.27 / 86.47 / 105.16
MACD / 信号
-1.335 / -0.683
空头排列

BABA

阿里巴巴 (BABA)

$124.22 -1.54%
5 日
-5.51%
距 52w 高
-35.5%
RSI(14)
37.7
趋势
空头
SMA 20 / 50 / 200
134.18 / 131.07 / 149.62
MACD / 信号
-1.893 / -0.440
空头排列

PDD

拼多多 (PDD)

$84.44 +1.70%
5 日
-13.65%
距 52w 高
-39.4%
RSI(14)
32.3
趋势
空头
SMA 20 / 50 / 200
95.79 / 98.42 / 113.23
MACD / 信号
-3.204 / -1.799
MACD 死叉 (4 天前)空头排列

JD

京东 (JD)

$28.83 -1.06%
5 日
-8.39%
距 52w 高
-21.8%
RSI(14)
38.8
趋势
空头
SMA 20 / 50 / 200
30.88 / 30.03 / 30.44
MACD / 信号
-0.087 / 0.314
MACD 死叉 (4 天前)空头排列

0700.HK

腾讯控股 (0700.HK)

HK$427.20 +0.52%
5 日
-2.69%
距 52w 高
-37.5%
RSI(14)
31.0
趋势
空头
SMA 20 / 50 / 200
454.80 / 484.61 / 574.68
MACD / 信号
-16.238 / -14.448
接近 52 周低空头排列

GC=F

黄金期货

$4,569.90 +1.57%
5 日
+0.66%
距 52w 高
-18.2%
RSI(14)
47.3
趋势
中性
SMA 20 / 50 / 200
4,590.16 / 4,630.46 / 4,370.90
MACD / 信号
-50.745 / -47.034

CL=F

WTI 原油期货

$87.76 -1.28%
5 日
-8.92%
距 52w 高
-26.5%
RSI(14)
38.3
趋势
中性
SMA 20 / 50 / 200
98.53 / 97.87 / 72.04
MACD / 信号
-1.669 / 0.244

USDCNY=X

美元 / 人民币

¥6.77 -0.20%
5 日
-0.42%
距 52w 高
-6.2%
RSI(14)
30.6
趋势
空头
SMA 20 / 50 / 200
6.80 / 6.83 / 6.99
MACD / 信号
-0.014 / -0.013
MACD 死叉 (2 天前)接近 52 周低空头排列
风险提示

本报告基于公开技术指标数据自动生成,过去走势不代表未来表现。所有内容仅供技术指标解读参考,不构成投资建议,市场存在不确定性,读者应独立评估风险。

Trump’s Boat Strikes Have Failed to Curb Cocaine Flow to U.S., Experts Say

Despite the rising body count off the South American coast, researchers say cocaine is as easy to get in many parts of the United States as it was before the strikes began.

中文摘要 专家称,针对南美海岸的军事打击未能遏制流入美国的可卡因。尽管该行动导致伤亡增加,但在美国多地获取可卡因仍如行动开始前一样容易。

Ghana parliament passes anti-LGBTQ+ bill

Same-sex acts are punishable by jail terms under Ghana's new bill targeting those identifying as gay, lesbian or transgender.

中文摘要 加纳议会通过了一项反LGBTQ+法案。根据新法案,同性行为将面临监禁处罚。该法案旨在针对自认为是男同性恋、女同性恋或跨性别人群。

Louisiana lawmakers pass congressional map favouring Republicans

Louisiana approves new congressional map eliminating a majority-Black district after an April Supreme Court ruling.

中文摘要 路易斯安那州立法者通过了一份有利于共和党的新国会选区划分图。该图取消了一个以黑人为主的选区,此举是在四月最高法院裁决后进行的。

Iran War Updates: Trump Puts Off ‘Final Determination’ on Iran Proposal

The president met with aides for two hours at the White House about a possible cease-fire extension, according to a senior administration official. Earlier, Mr. Trump had suggested in a social media post that he was ready to make a decision.

中文摘要 特朗普在与高级助手进行了两小时会面后,推迟了对伊朗提案的“最终决定”。会议讨论了可能的停火延期。此前,特朗普曾在社交媒体暗示已准备好做出决定。

Canadian Man Pleads Guilty to Aiding 14 Suicides

Kenneth Law, who ran an online business that shipped toxic salt to customers in 40 countries, also admitted to causing the deaths of 79 people in Britain, prosecutors said.

中文摘要 加拿大男子Kenneth Law承认协助14起自杀事件。检察官称,他经营的在线企业向40个国家的客户运送有毒盐,并造成英国79人死亡。

First survivor rescued from flooded cave in Laos

Divers in Laos have rescued the first of five villagers trapped in a flooded cave for more than a week.

中文摘要 老挝潜水员在一座被淹洞穴中成功救出了第一名被困村民。该村民与其他四人被困在洞穴中超过一周。

Iran war live: Trump due to make ‘final determination’ on deal with Tehran

Israel pushes deeper into Lebanon just days after Israel's prime minister ordered 70 percent of Gaza to be occupied.

中文摘要 特朗普预计将就与德黑兰的协议做出“最终决定”。与此同时,以色列在总理下令占领加沙70%地区后,正进一步深入黎巴嫩。

Mexican Senate Votes to Allow Voiding Elections Over Foreign Interference

The legislation, which comes amid growing tensions between Mexico and the White House, must be approved by a majority of state legislatures and sent to the president.

中文摘要 墨西哥参议院投票通过一项法案,允许以外国干涉为由宣布选举无效。该立法仍需多数州议会批准并提交总统。此举正值墨西哥与白宫关系紧张之际。

Iran’s Hard-Liners Try to Derail Potential Deal With the U.S.

A political fight is playing out in Iran, where the small but loud faction of hard-liners has used rallies, state media and private and public statements to try to undermine negotiations.

中文摘要 伊朗国内强硬派正试图破坏与美国达成的潜在协议。这个规模不大但声音响亮的派系通过集会、官方媒体及公开声明来干扰谈判。

ICE agent arrested over shooting of Venezuelan man in US immigration raid

The charges stem from the January 14 shooting of Julio Cesar Sosa-Celis in Minneapolis during Operation Metro Surge.

中文摘要 一名美国移民与海关执法局特工因在移民突袭行动中枪击一名委内瑞拉男子而被逮捕。该事件发生于1月14日的明尼阿波利斯“地铁洪流”行动期间。

After decades risking arrest, South Korea's tattoo artists step into the limelight

Only licensed doctors were allowed to ink tattoos in Korea - breaking the law could lead to heavy fines or jail.

中文摘要 韩国纹身师数十年来面临被捕风险,如今终于步入公众视野。此前韩国法律规定只有持牌医生才能进行纹身,违法者可能面临高额罚款或监禁。

Russian Drone Hits Romanian Apartment Building, Officials Say

Romania is a NATO country, and the security alliance condemned “Russia’s recklessness” for an episode that sharply escalated tensions with Moscow.

中文摘要 官员称,俄罗斯一架无人机击中了罗马尼亚的一栋公寓楼。罗马尼亚是北约成员国,该安全联盟谴责了“俄罗斯的鲁莽行为”,此事件严重加剧了与莫斯科的紧张关系。

Samsung’s AI Bonuses Divide Workers

Samsung has narrowly averted a strike by promising some employees hefty bonuses — dividends from the windfall it’s seen from the AI boom. But that’s stoking unhappiness among other workers who don’t stand to reap equal benefits. (Source: Bloomberg)

中文摘要 三星为避免罢工,承诺给部分员工丰厚AI奖金,源于AI热潮收益,但这引起其他无法享受同等待遇的工人的不满。

Wall Street Week | Britain’s Debt Problem, Poland’s Economic Boom

This week, the UK’s debt burden and weak growth are reviving fears that financial markets could once again destabilize British politics and policy. And, investors see data centers as long-term infrastructure, but neighbors worry about noise, water use, power demand and lasting costs. Plus, is Poland

中文摘要 本周英国债务负担和增长疲软引发担忧,金融市场可能再次动摇英国政治和政策;投资者视数据中心为长期基础设施,但邻居担忧噪音、用水等问题。

Citadel Securities Loses Court Fight Over New IEX Options Venue

Citadel Securities lost its bid to block IEX Group Inc. from launching a new type of options exchange that intentionally slows orders, after a federal appeals court on Friday rejected the market maker’s challenge.

中文摘要 Citadel Securities在法庭上未能阻止IEX Group推出一种故意减慢订单的新期权交易所,联邦上诉法院驳回其挑战。

SpaceX Wins $4 Billion Contract for US Golden Dome Satellites

SpaceX has won a contract for more than $4 billion to build satellites to track foreign aircraft and missiles as part of President Donald Trump’s Golden Dome defensive shield. Bloomberg's Sana Pashanka reports. (Source: Bloomberg)

中文摘要 SpaceX赢得超过40亿美元合同,为特朗普总统的「金穹」防御系统建造卫星,用于追踪外国飞机和导弹。

'There's A Lot More To Come' In AI Says Aliaga

Stephanie Aliaga, JPMorgan Asset Management Global Market Strategist joined Bloomberg Businessweek Daily to discuss AI trends and how it is changing how we view the markets. She stated that we are still in the beginning stages of the AI boom and we will see some more volatility as markets get it foo

中文摘要 摩根大通资产管理全球市场策略师Stephanie Aliaga表示,AI热潮仍处于初期阶段,未来还有更多发展,市场将面临更多波动。

Strait of Hormuz Ship Transits Are Rising Thanks to Help From US

Shipowners are increasingly optimistic about a pickup in traffic through the Strait of Hormuz after more vessels left the waterway this week with the US providing information to aid those making the journey.

中文摘要 由于美国提供航行信息援助,霍尔木兹海峡船只通行量增加,船东对交通回暖持乐观态度。

Quantinuum Said to Weigh Boosting IPO Size and Price Range

Honeywell International Inc.-backed quantum computing company Quantinuum Inc. is considering increasing the size of its initial public offering, according to a person familiar with the matter.

中文摘要 霍尼韦尔支持的量子计算公司Quantinuum据称考虑扩大IPO规模和价格区间。

EasyJet draws takeover interest from private credit firm Castlelake

Takeover of budget airline would lead to another UK-listed company taken off the stock market

中文摘要 廉价航空易捷航空收到私募信贷公司Castlelake的收购兴趣,若成功将导致又一家英国上市公司退市。

US Jobs Report Due Next Friday

Anna Wong, chief US economist at Bloomberg Economics, joins Katie Greifeld on "Bloomberg Real Yield." Bloomberg Economics expects the May jobs report to show payrolls rising by 95,000. (Source: Bloomberg)

中文摘要 美国5月就业报告将于下周五公布,彭博经济学预计非农就业人数增加9.5万。

NYC’s Mamdani Mimics Trump, Musk DOGE With COGE Plan

Bloomberg's Nacha Cattan joins Katie Greifeld on "Bloomberg Real Yield." New York City Mayor Zohran Mamdani said he’s forming a new Commission of Government Efficiency, or COGE, to improve the way City Hall spends public funds. (Source: Bloomberg)

中文摘要 纽约市长Zohran Mamdani宣布组建政府效率委员会(COGE),模仿特朗普和马斯克的DOGE计划,旨在改善市政公共资金支出。

真不是我看不起国模。。

各位佬,真不是我看不起国模,我好不容易下定决心用DP来接手codex的活,他给我干的第一件事,我让他启动项目。然后他说项目启动失败,有个SQLite错误,然后直接删库,因为是测试数据,没在意,然后后面就查了半天,最后发现删的是一个共享数据库的库,这还好是测试数据,这要真是正式的,我都不知道咋办了 58 个帖子 - 44 位参与者 阅读完整话题

zed 用了我的 rust 库!

事情起因是因为一直用 codex 写了不少工具,年前开始想整个大活,锻炼下自己用 AI 的能力。参考 Zed 的 UI 框架做一个新的通用的自绘 UI 框架,还有配套的一些渲染方面的库,其中包括一个重写 mermaid.js 的 rust 库(因为是底层库,就不推广了)。 库目标是不依赖浏览器的能力下渲染 mermaid 图为 svg png 等格式,是一个 headless 的渲染器。最近几个月一直让 codex 对着 mermaid.js 转换出来的 svg 来作为测试来不断补全逻辑。(也算是一种 harness?) 因为是底层库,因此一直没有宣传,好几个月陆陆续续只有 10 个 star

A\的模型又在走下坡路

从A\的模型在走下坡路继续讨论 如我所料,Opus 4.8 < Opus 4.7 <GPT 5.5。初步体验 Opus4.8 后,我觉得 claude 的 Pro 也没必要开了仅存的那点 4.6 额度也扣扣搜搜的。还是那句话,c 端的 A\ 之后只会越来越 ÷,除非它快死了。祝 A\ 早日殡天吧。 首先是在协作体验上,它的体验会略好于 4.7,不及 4.6 和 5.5。最大的问题就是「君の日本語は本当に上手ですね」。家乡的语言充斥在它的思维链和中期输出中,严重干扰我对它工作进度和思考的理解判断,真绷不住了。但如果我用英语与它进行协作,它的输出又很正常。A\ 你罪大恶极啊! 但也有可圈可点之处,

【长文】Claude自称自己是Ds是不是蒸馏?从舆论场的角度整理一下我们应该如何应对“借着蒸馏的名词忽悠人污名化国模”问题的想法 (又名为了痛痛快快的怼A畜自己找的一大堆理由)

随着 Claude Opus 4.8 发布,大家又迎来了一波经典讨论的回归: Claude 到底蒸没蒸? 技术上,当然不能这么简单判断,“蒸馏”是个比较严谨的词汇,在大模型领域有明显的定义和边界,不能简单靠ai的自称来下结论。 但舆论上,这波就是一个非常漂亮的回旋镖:过去很多人拿“国模自称 Claude、GPT”来证明国模偷懒,蒸馏洋模,现在 Claude 自己也开始自称 DeepSeek。那按他们之前的标准,A\ 是不是也蒸馏 DS 了? 蒸不蒸是个技术问题一会会说,但我先下个结论, 我们承认国模的蒸馏现象,但蒸馏本来就是正常的行为。我们反对的是那些借着”蒸馏“污名化国模带节奏的人,光靠”蒸

【CHY公益站】人是复活了,但是公益站还在打复活赛啊

这几天因为要期末考试了,没有动电脑,给Trae SOLO(搭载的Mimo2.5 Pro模型)提个优化请求就没管了,结果一看记录,它为了节省空间居然把我公益站的数据库给删了,我之前也没有备份数据库内的数据,从站点数据到佬友们的用户数据全没了 佬友们可以原谅我吗 把我们辛辛苦苦攒的额度搞没了,不可原谅! 毕竟是公益嘛,额度没了可以再攒,原谅 点击以查看投票。 48 个帖子 - 47 位参与者 阅读完整话题

【干草铺公益站】临时运维公告

各位佬,今天蹬的有点狠了,我要优化一下服务,预计今天 19:00 服务暂停运维,明天上午 10 点再开放,佬们先去蹬冰佬吧,各位周末愉快~ 13 个帖子 - 13 位参与者 阅读完整话题

关于公益站的小小回复。

本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 / 否 我的帖子已经打上 公益推广 标签: 是 / 否 我的项目属于个人项目,与公司或商业机构无关: 是 / 否 我的项目不存在QQ、TG等群组引流: 是 / 否 我的项目不存在非运营必要的网站引流: 是 / 否 我的项目不存在为他人推广、AFF: 是 / 否 我的项目无关联的商业项目: 是 / 否 我的站点存在登录,并已接入 LINUX DO Connect: 是 / 否 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 / 否 以上选择我承诺是永久有效的

爱看美女是有心理问题吗?

我有老婆,但是走大街上,如果是年轻貌美的女的,总是忍不住看两眼,尤其会盯着胸部和屁股看,一直感觉这样不太好,但是总是忍不住,跟条件反射似的。佬友们是心理问题吗?如何治疗? 114 个帖子 - 108 位参与者 阅读完整话题

给公司培训完codex后,反而工作变得更多了

在公司内部我是比较前沿使用ai的,其他人都处在使用免费的ai ,豆包,deep seek,qwen这一些东西。 我本人在过去半年整体来说就是vibecoding, 基本上就是一小时开发,划水七天。前一阵vibecoding了一个工具后,领导问我使用的工具然后让我给大家培训一下。结果噩梦开始了 自从培训过后,可能是领导对于vibecoding的提效感触很大 就开始对我的任务时间进行缩短,在他的认知这个工作已经不是一周的工作量了,而是一天的。就会有更多的任务压上来。 但是分享的初衷是提效减少非必要劳动,而不是三倍速牛马啊。 无语了 5/29 17:34 可能从一开始在制作这个skill的时候我就应