每日简报

2026-05-27

← 历史归档

Lum1104/Understand-Anything

TypeScript · ★ 36,770 · 🍴 2,948 · 📈 4,697 stars today

Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.

中文介绍 将任意代码库转换为可交互的知识图谱,帮助开发者直观探索、搜索和提问以理解代码结构。它利用图谱和 AI 技术,能与 Claude Code、Cursor 等编程助手集成,适合需要快速掌握大型或不熟悉代码库的开发者进行代码理解和学习。

affaan-m/ECC

JavaScript · ★ 194,798 · 🍴 30,013 · 📈 1,915 stars today

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

中文介绍 面向 AI 代理(如 Claude Code、Codex)的性能优化与能力增强系统。它通过技能、直觉、记忆管理和安全框架,帮助开发者系统性地构建更强大、更可靠的代理应用,适用于需要开发复杂 AI 代理系统的工程师。

rohitg00/ai-engineering-from-scratch

Python · ★ 21,108 · 🍴 3,496 · 📈 2,155 stars today

Learn it. Build it. Ship it for others.

中文介绍 一个从零开始学习人工智能工程的完整项目,遵循“学习、构建、交付”的实践路径。它旨在帮助开发者或学习者系统性地掌握 AI 工程的核心概念与技能,并最终能构建出可实际使用的产品。

anthropics/knowledge-work-plugins

Python · ★ 16,812 · 🍴 1,968 · 📈 1,718 stars today

Open source repository of plugins primarily intended for knowledge workers to use in Claude Cowork

中文介绍 一个为 Claude 等 AI 助手设计的开源插件集合,主要面向知识工作者。这些插件旨在扩展 AI 在日常知识处理任务(如文档分析、信息整合)中的能力,提升工作效率。

mukul975/Anthropic-Cybersecurity-Skills

Python · ★ 10,274 · 🍴 1,202 · 📈 880 stars today

754 structured cybersecurity skills for AI agents · Mapped to 5 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND & NIST AI RMF · agentskills.io standard · Works with Claude Code, GitHub Copilot, Codex CLI, Cursor, Gemini CLI & 20+ platforms · 26 security domains · Apache 2.0

中文介绍 一个包含 754 个结构化网络安全技能的数据库,专为 AI 代理设计。技能映射到 MITRE ATT&CK、NIST CSF 等主流安全框架,采用 agentskills.io 标准,能与 Claude Code、GitHub Copilot 等集成,辅助进行安全分析与自动化操作。

hardikpandya/stop-slop

★ 5,145 · 🍴 406 · 📈 539 stars today

A skill file for removing AI tells from prose

中文介绍 一个用于改善 AI 生成文本风格的技能文件。其主要功能是移除 AI 生成内容中常见的、公式化的“AI 腔调”,使文本更自然、更像人类写作,适用于希望产出高质量、个性化文风内容的写作者或内容生成流程。

Leonxlnx/taste-skill

Shell · ★ 22,448 · 🍴 1,775 · 📈 1,430 stars today

Taste-Skill - gives your AI good taste. stops the AI from generating boring, generic slop

中文介绍 一个旨在提升 AI 生成内容“品味”的技能文件。它通过约束和引导,阻止 AI 产出枯燥、同质化的内容,帮助生成更具创意和独特性的文本、设计或创意输出。

DigitalPlatDev/FreeDomain

HTML · ★ 167,794 · 🍴 3,110 · 📈 1,219 stars today

DigitalPlat FreeDomain: Free Domain For Everyone

中文介绍 DigitalPlat 提供的一项免费域名注册服务。该项目为所有人提供免费的域名,旨在降低个人和小型项目在互联网上建立专属地址的门槛,适用于学生、初创项目或个人博客等场景。

jellyfin/jellyfin

C# · ★ 52,442 · 🍴 4,875 · 📈 83 stars today

The Free Software Media System - Server Backend & API

中文介绍 一款完全免费的开源媒体系统服务器后端与 API。它让用户能够自由地管理、流媒体播放自己的音视频媒体库,是 Plex 或 Emby 的开源替代品,适用于注重隐私、希望完全掌控媒体中心的家庭或个人用户。

Axorax/awesome-free-apps

JavaScript · ★ 5,417 · 🍴 274 · 📈 731 stars today

Curated list of the best free apps for PC and mobile

中文介绍 一份精心整理的免费应用清单,覆盖 PC 和移动端。该项目为用户快速发现优质、可靠的免费软件提供了便捷的索引,适用于寻找各类免费工具、提升效率或娱乐的普通用户。

twentyhq/twenty

TypeScript · ★ 46,970 · 🍴 6,664 · 📈 216 stars today

The open alternative to Salesforce, designed for AI.

中文介绍 一款为 AI 时代设计的、开源的 Salesforce 替代方案。它是一个现代化的客户关系管理(CRM)平台,将 AI 能力深度融入业务流程,帮助团队更智能地管理客户互动与销售管道。

Open-Dev-Society/OpenStock

TypeScript · ★ 12,255 · 🍴 1,634 · 📈 156 stars today

OpenStock is an open-source alternative to expensive market platforms. Track real-time prices, set personalized alerts, and explore detailed company insights — built openly, for everyone, forever free.

中文介绍 一个开源、免费的金融市场追踪平台,旨在替代昂贵的专业服务。它提供实时价格监控、个性化警报设置和详细的公司洞察等功能,致力于为所有投资者提供透明、免费的金融信息工具。

thedotmack/claude-mem

TypeScript · ★ 78,819 · 🍴 6,778 · 📈 352 stars today

Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

中文介绍 为 AI 代理提供跨会话持久记忆的解决方案。它能捕获代理在单次会话中的所有操作,通过 AI 进行压缩摘要,并在未来会话中注入相关上下文,从而维持任务的连续性和个性化,尤其适用于需要长期任务跟踪的复杂代理场景。

st-tech/ppf-contact-solver

Python · ★ 3,573 · 🍴 255 · 📈 170 stars today

A contact solver for physics-based simulations involving 👚 shells, 🪵 solids and 🪢 rods.

中文介绍 一个用于基于物理模拟的接触求解器,专门处理壳体、固体和杆状结构之间的复杂接触交互。该技术可用于游戏开发、计算机动画或工程仿真领域,以模拟衣物、木材、绳索等物体的真实物理行为。

How to Build a Software Factory with Claude Code That Ships Features While You Sleep

@sairahul1 · 106.0K 粉丝 · 2.4M 阅 · 1.2K 赞 · 175 转

I thought I was using AI to code. I was actually just typing faster. Here is the difference — and the 7-agent system that changed everything. Save this. It will save you months. THE PROBLEM NOBODY

中文介绍 分享基于 Claude Code 的 7-agent 系统,通过多代理协作自动化软件开发,实现「睡着也能发功能」的高效工作流,区分 AI 编码与单纯打字加速的差异。

AI Agents: The Complete Course

@sairahul1 · 106.0K 粉丝 · 203.2K 阅 · 500 赞 · 82 转

Everyone is talking about AI agents in 2026. Most people have no idea how they actually work. This changes today. I spent weeks distilling everything: courses, books, real builds, production failures.

中文介绍 提供 AI agents 的完整课程,涵盖理论、实践构建和生产失败经验,花数周时间提炼,帮助读者全面理解 AI agents 的工作原理和应用。

How to Build a Claude Research Agent That Reads the Internet Every Morning and Briefs You in 5 Mins

@cyrilXBT · 179.9K 粉丝 · 127.8K 阅 · 533 赞 · 80 转

Most people start their day the same way. They open Twitter and spend 20 minutes scrolling through noise looking for the three things that actually matter. They open their email and get pulled into

中文介绍 构建 Claude Research Agent,每天早上自动浏览互联网并在 5 分钟内提供关键信息摘要,节省时间避免信息过载,提高日常信息获取效率。

How I Use Cursor

@poteto · 26.6K 粉丝 · 86.5K 阅 · 540 赞 · 48 转

I need to get something off my chest. Before my interview @cursor_ai, I had never actually used Cursor. At Meta, Claude Code was explosively taking off. I even paid for a personal $200 a month plan

中文介绍 分享使用 Cursor 编辑器的经验,对比 Claude Code 的流行,提及个人每月支付 $200 的计划,提供工具使用心得和实际投资考量。

Step-By-Step LLM Engineering Projects (2026 Edition)

@TheAhmadOsman · 59.9K 粉丝 · 54.5K 阅 · 512 赞 · 65 转

At some point, reading about LLMs stops being enough. You need to build the stack yourself: Tokenizer first, then embeddings, position, attention, Transformer blocks, objectives, decoding, cache, long

中文介绍 提供 2026 年版 LLM 工程分步项目指南,从 Tokenizer 开始构建整个栈,包括嵌入、注意力机制等,强调动手实践以深入理解 LLM 技术。

THE DIFF THAT CHANGED EVERYTHING

@difflawb · 20.3K 粉丝 · 21.9K 阅 · 1.1K 赞 · 389 转

How a 40-line shell script became infrastructure In August 2024, Andrej Karpathy — co-founder of OpenAI, former AI Director at Tesla — published something unexpectedly small. Not a paper. Not a model.

中文介绍 讲述一个 40 行 shell 脚本如何演变为基础设施,引用 Andrej Karpathy 在 2024 年发布的小项目,展示小工具在技术生态中的巨大影响。

how i make AI videos (a beginner’s breakdown)

@0xileri · 7.3K 粉丝 · 12.2K 阅 · 533 赞 · 63 转

I’ve been getting a lot of DMs since I started posting AI videos, so I figured I’d just write it all out. Fair warning: I’m still learning too. This is just what’s been working for me. tools

中文介绍 分享 AI 视频制作的入门指南,包括工具选择和个人实践,基于收到私信,提供初学者友好的 breakdown 和仍处于学习阶段的实用经验。

The Start of the End: AI Replacement Has Begun

@ActionModelAI · 57.1K 粉丝 · 5.8K 阅 · 505 赞 · 344 转

We are witnessing the beginning of the biggest economic shift in modern history. And most people still don’t realize it. AI replacement is no longer some distant sci-fi prediction. It has started.

中文介绍 分析 AI 替代工作的趋势,指出这已开始,是现代经济最大转变的开端,大多数人尚未意识到,强调其紧迫性和对社会的深远影响。

Rethinking organizational design in the age of agentic AI

Amid rapidly growing adoption of enterprise-level AI agents, there’s a disconnect emerging between ambition and execution. Although 85% of organizations say they want to be agentic within the next three years, 76% say their current operations and infrastructure can’t support that change. They cite a

中文介绍 随着企业级AI智能体应用加速,目标与执行之间出现脱节。调查显示,85%的组织希望在三年内实现智能化,但76%表示其当前运营和基础设施无法支持这一目标。

A reality check on the AI jobs hysteria

Haven’t you heard? White-collar jobs are going away, decimated by AI. Waves of layoffs in the tech sector (most recently at Coinbase and Meta and Cisco) are said to presage what will soon come for all of us knowledge workers. But before you quit your job as a software developer or financial analyst—

中文介绍 针对AI导致白领工作大规模消失的恐慌情绪进行现实检查。虽然Coinbase、Meta和Cisco等科技公司近期裁员引发担忧,但目前总体就业保持稳定。

It’s time to address the looming crisis in entry-level work.

Artificial intelligence has not so far produced a clean story of mass unemployment. Aggregate employment in developed countries remains broadly stable, and recent assessments have found limited evidence that AI has shifted the headline numbers. But a troubling change may be hiding beneath the surfac

中文介绍 人工智能尚未导致发达国家出现大规模失业,总体就业形势基本稳定。然而,入门级工作领域可能潜伏着一场日益临近的危机,其影响尚不明显。

not much happened today

**Harness engineering** is emerging as the key differentiator for coding agents, emphasizing the stack of **model + harness + eval loop** over just stronger base models. **DeepSeek** is building a harness team to optimize interaction and verification loops, while **Google's Gemini Managed Agents** a

中文介绍 Harness工程成为编程智能体的关键差异化因素,其重要性在于「模型+Harness+评估循环」的集成。DeepSeek正组建团队优化此循环,而Google的Gemini也在应用类似技术。

OpenAI, Grupo Folha and Grupo UOL announce strategic content partnership

OpenAI partners with Grupo Folha and Grupo UOL to bring trusted Brazilian journalism to ChatGPT, expanding access to news with attribution and transparency.

中文介绍 OpenAI宣布与巴西媒体集团Grupo Folha及Grupo UOL达成战略合作,将可信的巴西新闻内容引入ChatGPT,旨在扩大新闻获取渠道并确保透明归因。

[AINews] All Model Labs are now Agent Labs

a quiet day lets us tie together a few quotes as all model labs become agent labs

中文介绍 行业趋势显示,所有主要的AI模型实验室正纷纷转型成为AI智能体实验室。这一转变标志着研发与产品重心从基础模型向能执行复杂任务的智能体应用迁移。

Google I/O showed how the path for AI-driven science is shifting

During Tuesday’s Google I/O keynote, Demis Hassabis, the CEO of Google DeepMind, proclaimed that we are currently “standing in the foothills of the singularity.” It was a striking statement—the singularity is the theoretical future moment when AI rapidly exceeds human intelligence and dramatically t

中文介绍 谷歌I/O大会展示了AI驱动科学研究的路径正在发生转变。DeepMind CEO Demis Hassabis宣称我们正处于「奇点的山脚」,暗示AI将在科学领域加速超越人类智能。

How Virgin Atlantic ships faster with Codex

How Virgin Atlantic used Codex to ship its revamped mobile app on a fixed holiday travel deadline, reaching near-total unit test coverage and zero P1 defects.

中文介绍 维珍大西洋航空利用OpenAI的Codex工具,在严格的节假日旅行截止日期前成功交付了其新版移动应用程序,实现了近乎100%的单元测试覆盖率和零P1级缺陷。

Shortest Path Problem with Subnormal Gaussian Fuzzy Costs

第一作者: Murat Moran · 方向: 安全研究

Abstract:This paper addresses the fuzzy shortest path problem in directed graphs, where edge costs are modeled as generalized fuzzy numbers with Gaussian membership functions. We interpret height as an indicator of information reliability. Based on this view, we introduce a weighted geometric mean to aggregate heights during the addition of generalized Gaussian fuzzy numbers. We employ a reliability-aware ranking that jointly considers the core, height, and standard deviation of fuzzy edge costs to determine the shortest path, thereby capturing their central tendency, reliability, and variability while keeping Dijkstra-level complexity per relaxation. The method yields routes that are not only cost-efficient but also supported by highly reliable information. To assess robustness, we construct a crisp baseline from the ranking and conduct Monte Carlo alpha-cut sampling--drawing...

论文介绍 本文研究有向图中的模糊最短路径问题,其中边成本被建模为具有高斯隶属函数的广义模糊数,并将高度解释为信息可靠性的指标。作者引入加权几何平均来聚合高度,并采用可靠性感知排序法,综合考虑模糊成本的核心、高度和标准差来确定最短路径。该方法在保持Dijkstra级别复杂度的同时,能够捕获成本的趋势、可靠性和变异性,从而得到既成本高效又有可靠信息支持的路径。为评估鲁棒性,研究基于排序构建清晰基线并进行蒙特卡洛alpha截集抽样。

Risk Averse Alert Prioritization for IDS Using Subnormal Gaussian Fuzzy Models

第一作者: Murat Moran · 方向: 网络安全

Abstract:Modern intrusion detection systems generate thousands of alerts daily, but alert fatigue severely limits security operations effectiveness due to too many false positives or low-impact events. We address this by proposing a principled framework for alert prioritization based on subnormal Gaussian fuzzy numbers, explicitly modeling three sources of uncertainty: threat severity, detection confidence, and organizational risk attitude. Each alert is represented as a fuzzy number with the core indicating severity, spread indicating uncertainty, and height reflecting detection reliability. We apply ranking indices to prioritize alerts, allowing organizations to tune security posture through a risk-attitude parameter. Experimental validation on CIC-IDS2017 and NSL-KDD demonstrates greater robustness than baselines under detector degradation (0.9963 vs 0.8215 NDCGrel@100), with...

论文介绍 现代入侵检测系统每日产生大量告警,但因误报和低影响事件导致的告警疲劳严重影响了安全运营效能。本文提出一个基于次正态高斯模糊数的框架来对告警进行优先级排序,明确建模了三种不确定性来源:威胁严重性、检测置信度和组织风险态度。每个告警被表示为一个模糊数,其核心指示严重性,标准差表示不确定性,高度反映检测可靠性。通过应用排序指标并引入风险态度参数,允许组织调整其安全态势。实验表明,该方法在探测器性能下降时比基线方法更具鲁棒性。

Landseer: Exploring the Machine Learning Defense Landscape

第一作者: Ayushi Sharma · 方向: AI 安全

Abstract:Machine learning systems face diverse threats that undermine robustness, privacy, and fairness. Although many defenses have been proposed, each typically addresses a single risk in isolation. Real-world deployments, however, require these defenses to be composed to meet multiple guarantees simultaneously. The process of composing defenses is complex and not well understood, and its impact on performance and security remains unclear. We present Landseer, a modular framework for integrating machine learning (ML) defenses into the ML lifecycle and systematically evaluating their composition. Landseer encapsulates defenses as containerized modules, allowing existing and new techniques to be plugged in with minimal effort. Its evaluation engine automates experiments across multiple metrics, supporting the study of defenses both individually and in combination. In a preliminary...

论文介绍 机器学习系统面临多样化的威胁,现有防御通常孤立地解决单一风险。然而,实际部署需要将这些防御组合起来以同时满足多项保障。本文介绍了Landseer,一个用于将机器学习防御集成到生命周期并系统评估其组合的模块化框架。Landseer将防御封装为容器化模块,便于即插即用。其评估引擎可自动化跨多种指标的实验,支持对防御进行单独和组合研究。初步结果揭示了防御组合中的复杂权衡,例如在增加针对某一威胁的防御时可能削弱另一防御的效果。

Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?

第一作者: Enoal Gesny · 方向: AI 安全

With the rapid proliferation of generative models, such as diffusion models, digital watermarking has emerged as a crucial solution for identifying AI-generated images. Modern post-hoc watermarking schemes use neural networks to achieve an extremely low false-alarm rate while remaining robust to common image transformations. However, there is a lack of comparison between these modern methods and classic ones, particularly in real-world scenarios where robustness and security take precedence over achieving an extremely low false-alarm probability. In this paper, we propose a fair comparison of robustness and security between modern and classic post-hoc watermarking across various types of classic augmentations and recent sophisticated attacks. Our experiments show that, in a realistic scenario, classic watermarking outperforms modern techniques in terms of security while maintaining...

论文介绍 随着扩散模型等生成模型的普及,数字水印成为识别AI生成图像的关键方案。现代后处理水印方法利用神经网络实现极低的虚警率,并对常见图像变换保持鲁棒。然而,这些现代方法与经典方法,尤其是在以鲁棒性和安全性为首要考虑的真实世界场景中,缺乏充分比较。本文对现代与经典后处理水印方法在多种经典数据增强和近期复杂攻击下的鲁棒性与安全性进行了公平比较。实验表明,在现实场景中,经典水印在安全性上优于现代技术。

BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

第一作者: Xuan Luo · 方向: 安全研究

Abstract:In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the...

论文介绍 本文提出BAIT(边界感知迭代陷阱),一个通过内部渐进式披露来接近恶意目标的三步越狱框架。BAIT首先要求模型识别其防护边界,然后要求其细化该边界,最后请求一个详细示例。通过将每一步建立在模型先前回复的基础上,BAIT将模型自身的推理和一致性倾向转变为一条披露路径。在多个基准测试上的实验表明,BAIT在顶级大语言模型上持续取得较高的攻击成功率。进一步分析发现,以预防为导向的框架设定显著优于直接的知识请求,且细化步骤对披露升级至关重要。

Lessons from Penetration Tests on Large-Scale Agent Systems

第一作者: Kevin Eykholt · 方向: 安全研究

Abstract:As AI systems gain increasing autonomy and execution capability, the number of discovered security vulnerabilities continues to rise. However, many of these vulnerabilities are not fundamentally novel, but instead reflect recurring classes of weaknesses long observed in prior computing systems. Execution-capable AI agents are effectively unbounded, self-modifying programs that interact extensively with multiple layers of the computing stack. This broad interaction surface imposes a significant security burden on developers, who must reason about and secure complex cross-layer behaviors. Prior research has primarily focused on vulnerabilities in open-source agents and agent frameworks. In contrast, it remains unclear whether proprietary agent systems -- developed under stricter coding standards and formal review processes -- exhibit similar security weaknesses. In this paper...

论文介绍 随着AI系统获得更强的自主性和执行能力,其安全漏洞的发现数量持续上升。本文总结了从对大规模智能体系统的渗透测试中获得的教训。研究指出,许多漏洞并非根本性新问题,而是反映了先前计算系统中反复出现的弱点类别。具备执行能力的AI智能体实际上是与计算栈多层广泛交互的、无界且可自修改的程序。这种广泛的交互面给开发者带来了巨大的安全负担。本文的研究表明,即使是在更严格的编码标准和形式化审查流程下开发的专有智能体系统,也可能展现出相似的安全弱点。

The Fault in Our Drafts: Vulnerabilities in RPKI Specification and Software

第一作者: Oliver Jacobsen · 方向: 软件安全

Abstract:The Resource Public Key Infrastructure (RPKI) secures the Internet's routing system by defining a complex trust and validation framework for certificates, Route Origin Authorizations (ROAs), manifests, and Certificate Revocation Lists (CRLs). These mechanisms are specified across dozens of RFCs. This paper presents the first comprehensive analysis of the causal link between flaws in RPKI Requests for Comments (RFCs) and vulnerabilities in implementations and real-world deployments. We reveal how vague, conflicting, or underspecified requirements in 50 RPKI RFCs propagate into inconsistent implementation behavior and operational failures. We conduct the first large-scale, impact-driven evaluation of RPKI specifications. Our methodology combines differential fuzzing of major RPKI implementations with Internet-wide crawling and validation log analysis, enabling us to trace...

论文介绍 资源公钥基础设施(RPKI)通过定义证书、路由源授权、清单和证书吊销列表等复杂框架来保护互联网路由系统。本文首次全面分析了RPKI请求注释(RFC)中的缺陷与实现及现实部署漏洞之间的因果联系。研究揭示了50份RPKI RFC中模糊、冲突或不明确的要求如何传播到不一致的实现行为和运营故障中。通过结合对主要RPKI实现的差分模糊测试以及互联网范围的爬取和验证日志分析,作者进行了首次大规模、影响驱动的RPKI规范评估。

Practical Anonymous Two-Party Gradient Boosting Decision Tree

第一作者: Huang Chenyu · 方向: AI 安全

Abstract:Structured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features across mutually distrustful parties. High speed and interpretability make GBDTs popular in finance and healthcare, where neural networks may fall short. Enabling secure computation for GBDTs poses unique challenges, requiring secure record alignment for comparison. Relying on private set intersection (PSI) is a de facto approach. Mistaking PSI for a safety measure actually exposes which record identifiers (IDs) are shared between the datasets. Although circuit-PSI could help, it is costly for generic uses. New ideas are needed to efficiently train in a "dark forest". Aiming to hide the IDs, we initiate the study of anonymous GBDT training on split data held by two parties. Dual circuit-PSI in our design lets the parties alternate as receiver...

论文介绍 梯度提升决策树(GBDT)擅长处理结构化数据,通常在互不信任的各方之间垂直划分的特征上进行训练。为GBDT启用安全计算面临独特挑战,需要安全的记录对齐进行比较。依赖私有集交互(PSI)是事实上的方法,但标准PSI会暴露哪些记录标识符在数据集间是共享的。本文开创性地研究了在双方持有的分割数据上进行匿名GBDT训练,旨在隐藏记录标识符。该设计使用双电路PSI,让双方交替作为接收方,从而实现高效的隐私保护模型训练。

Privacy-Preserving Screening for Record Linkage

第一作者: Chenyu Huang · 方向: AI 安全

Abstract:In an era dominated by big data and machine learning, establishing valuable data collaboration has never been more critical. However, such collaborations must operate under regulatory and legal constraints. Two-party Privacy-Preserving Record Linkage (PPRL) emerges to assess the potential collaboration value and also ensure the privacy and security of the involved data. Nevertheless, the substantial computational and communication overheads associated with PPRL hinder its practical adoption in data markets with numerous potential collaborators. Therefore, we present the Screening-then-Linkage framework, which incorporates a lightweight Screening phase prior to the resource-intensive PPRL phase, i.e., PPRS, to mitigate the scalability issue of PPRL. We propose a circuit-PSI-based system, named Appraisal to realize a secure, effective, and efficient PPRS. To reconcile the...

论文介绍 在数据协作中,隐私保护记录链接(PPRL)的计算与通信开销过高,限制了其可扩展性。本文提出了「筛选-链接」框架,在资源密集的PPRL阶段前引入轻量级筛选阶段(PPRS)。为此,作者设计并实现了一个基于电路PSI的安全、高效系统「Appraisal」,旨在评估潜在协作价值的同时,缓解PPRL在包含大量潜在协作者的数据市场中的应用瓶颈。

Secure UAV Swarms in Low-Altitude Wireless Networks: Challenges and Solutions

第一作者: Yuntao Wang · 方向: 系统安全

Unmanned aerial vehicle (UAV) swarms are increasingly deployed in vast low-altitude applications, owing to their capabilities in distributed sensing, flexible communication, and autonomous coordination. Nevertheless, the open and highly dynamic operating environment of UAV swarms introduces serious security risks, including GPS spoofing, insider threats, and multi-hop intrusion. These threats are aggravated by limited on-board resources, frequently changing network topology, and the presence of intelligent adversaries. To tackle these issues, this paper proposes a cloud-edge-end collaborative defense framework for UAV swarms. Based on this framework, three complementary mechanisms are developed. First, a cooperative perception scheme is designed to resist GPS spoofing via interactive attack-defense game modeling. Second, a behavior-driven authentication method with trust evaluation is...

论文介绍 无人机集群在低空应用中面临严峻安全挑战,如GPS欺骗、内部威胁和多跳入侵。本文提出了一个针对无人机集群的云-边-端协同防御框架,并基于此开发了三种互补机制:通过交互攻防博弈建模的抗GPS欺骗协同感知方案、基于行为的信任评估认证方法,以应对开放、动态环境及智能对手带来的安全风险。

Anonymous YARA Rules Are Not Anonymous

第一作者: Usman Rabiu Isah · 方向: 网络安全

YARA rules are widely shared across threat intelligence communities to enable collective defence against malware. This practice implicitly assumes that removing metadata (e.g., author fields) sufficiently protects the identity of contributing organisations. To assess the validity of this assumption, we systematically evaluate how much can be inferred from YARA rule text alone. Specifically, using a corpus of 23,305 rules from three major public repositories, we train independent classifiers along four stylometric fingerprint dimensions: individual author, source repository, malware family, and temporal drift, using three complementary methods: lexical n-grams (Burrows' Delta), syntactic AST features (Caliskan-Islam), and fine-tuned CodeBERT. Our results demonstrate that repository origin is almost perfectly recoverable (up to 99% accuracy), individual authors can be re-identified well...

论文介绍 共享威胁情报中的YARA规则在移除元数据后,其隐私保护程度仍不明确。本研究通过风格分析方法,从规则文本本身推断作者、来源仓库等信息。利用三个公共仓库的超过2.3万条规则,研究人员使用词法、句法和CodeBERT微调等多种方法进行评估,结果表明来源仓库几乎可以完美还原,个体作者也能被有效重新识别,揭示了现有匿名化实践中的隐私风险。

Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control

第一作者: Zhe Yu · 方向: 安全研究

Abstract:Retrieval-augmented generation (RAG) increasingly underpins high-stakes applications, yet remains vulnerable to Confundo-style poisoning where adversarially optimized documents manipulate generated outputs. Existing defenses assume that detecting poisoned evidence prevents harm. We show this assumption is incorrect: models exhibit a monitoring-control gap -- they can detect contradictions in retrieved evidence yet still act on poisoned claims. We introduce the Cordon Principle -- no agent capable of final synthesis may access untrusted natural-language evidence -- and realize it through CORDON-MAS, a compartmentalized framework that enforces this principle architecturally by separating evidence extraction, cross-source audit, and answer synthesis into agents with asymmetric memory privileges. Across five BEIR datasets, CORDON-MAS reduces attack success rate by 92.4\% relative...

论文介绍 检索增强生成(RAG)易受知识投毒攻击。现有防御存在监控与控制的差距:模型能检测出证据中的矛盾,但仍可能基于投毒后的主张行动。本文提出「Cordon原则」,即最终进行综合的智能体不得访问不可信的自然语言证据。该原则通过CORDON-MAS多智能体框架在架构上实现,将证据提取、跨源审计和答案合成分离到具有非对称内存权限的智能体中,从而大幅降低攻击成功率。

Certified Causal Attribution for Real-Time Attack Forensics in 6G Network Slicing

第一作者: Minh K. Quan · 方向: 网络安全

Abstract:Cross-slice attack attribution in 6G networks requires identifying causal propagation chains through shared infrastructure in under 100 ms. Existing methods struggle to satisfy this strict SLA without sacrificing accuracy, because shared resource contention creates spurious correlations that are indistinguishable from genuine causal links under standard Granger tests. We propose DA-GC, a certified causal attribution framework that integrates resource-conditioned Granger causality with an axiomatically derived Resource Contention Model (RCM) to systematically block resource-mediated confounding. On a 15-slice production-emulation 6G testbed with 1,100 attack scenarios, DA-GC achieves 89.2% attribution accuracy at 87 ms. This represents a 7.9 percentage-point improvement over the strongest baseline at 2.7x lower latency, alongside demonstrated cross-topology generalization and...

论文介绍 6G网络中的跨切片攻击溯源需要在严格的服务等级协议(如100毫秒)内识别因果传播链,而共享资源竞争导致的伪相关干扰了标准Granger因果检验。本文提出DA-GC,一个经过认证的因果归因框架。它集成了资源条件Granger因果检验与一个基于公理推导的资源竞争模型,以系统性阻断资源介导的混杂因素,在生产仿真测试中实现了接近90%的归因准确率和低延迟。

Resolving the Correct Library: A Loader-Level Defense Solution Against Shared Object Hijacking

第一作者: Can Ozkan · 方向: 安全研究

Shared library hijacking attacks in the Linux ecosystem, including embedded Linux, are a significant concern. It fundamentally exploits the dynamic linker's library-resolution semantics rather than modifying trusted libraries directly. Prior research has extensively analyzed attack vectors exploiting environment variables, embedded search paths, and dynamic loader internals, demonstrating that hijacking is rooted in fundamental loader behavior rather than isolated misconfigurations. Existing defenses either harden or replace the loader, enforce control-flow integrity after libraries are loaded, or apply file-centric integrity mechanisms such as signatures and measurement frameworks. However, these approaches fail to address a critical gap: none verify whether the shared object actually resolved by the loader is the intended and trusted one. In this paper, we argue that shared library...

论文介绍 共享库劫持攻击利用动态链接器的解析语义,而非直接修改受信库。现有防御措施未能解决一个关键缺口:即验证链接器最终解析出的共享对象是否确实是预期且受信的。本文分析了攻击的根本原因在于加载器行为,并提出了一种加载器级的防御方案,旨在解决此根本性问题,确保解析结果的可信性。

Batch Me If You Can: Coverage-guided RPKI Fuzzing at Scale

第一作者: Haya Schulmann · 方向: 软件安全

Abstract:The Resource Public Key Infrastructure (RPKI) has become essential to secure inter-domain routing. Despite its critical role, RPKI software remains largely untested beyond shallow parsing. Existing fuzzers, like AFL++ or libFuzzer, do not work well for RPKI as they assume a single, self-contained input per execution, while RPKI repositories contain hundreds of interdependent cryptographically linked objects. Existing fuzzers fail to handle this complexity and lack the ability for precise coverage attribution in multi-object repositories, breaking feedback-based exploration and thereby missing most severe vulnerabilities in RPKI validation. In this paper, we overcome these limitations through novel fuzzing techniques, including continuous sampling and using functions as side-channels for per-object coverage attribution in large input repositories. We further show how parsing...

论文介绍 用于保障域间路由安全的资源公钥基础设施(RPKI)软件缺乏深度测试。传统模糊测试工具难以处理RPKI仓库中相互依赖的密码学对象集合。本文提出新颖的模糊测试技术,包括持续采样和将函数用作每个对象覆盖引导的侧信道,从而克服了现有模糊器在大型输入仓库中精确覆盖归因的局限,能更有效地发现RPKI验证中的严重漏洞。

Control Physiology: An Agent-Based Model of FAIR-CAM Dynamics

第一作者: Jack Jones · 方向: 安全研究

Security risk analysis typically treats control effectiveness as a static input, yet controls degrade through configuration drift, depend on monitoring systems that may themselves be degraded, and compete for finite remediation budgets. The FAIR Controls Analytics Model (FAIR-CAM) provides the theoretical framework for these dynamics but has so far remained theoretical. We present the first agent-based model to operationalize the core FAIR-CAM dynamics, making control physiology computationally observable, and release the implementation as open source. The simulation implements eight agent types, a multiplicative defense-in-depth susceptibility formula, a three-source variance model, budget-constrained remediation, and a narrative causation engine that produces a complete causal trace for every loss event. In a hospital ransomware scenario (N=1,000 iterations), three organizational...

论文介绍 安全风险分析通常将控制有效性视为静态输入,而实际上控制会因配置漂移而退化。FAIR控制分析模型(FAIR-CAM)提供了相关动态的理论框架,但此前未被操作化。本文提出了首个代理模型来实现FAIR-CAM的核心动态,使控制的“生理”变化变得可计算观测。该模型在勒索软件场景仿真中展示了组织因素对风险的影响,并开源了实现。

Cordyceps: Covert Control Attacks on LLMs via Data Poisoning

第一作者: Zedian Shao · 方向: AI 安全

Abstract:Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primarily rely on fixed trigger phrases that defenses such as outlier detection, clean-data regularization, or online monitoring can neutralize. In this paper, we propose a data poisoning method that teaches an LLM an information hiding scheme reliably and stealthily through semantic associations between shared knowledge such as facts or concepts and attacker-chosen phrases. The induced hiding scheme can encode and decode arbitrary malicious instructions, thus revealing a new and subtle poisoning-induced vulnerability: covert control attacks. We precisely characterize covert control attacks and evaluate them across $5$ LLMs, $3$ backdoor defenses, and $4$ prompt injection defenses. With a small poisoned fraction, covert control attacks outperform...

论文介绍 针对大型语言模型微调数据可能被投毒的问题,本文提出一种通过语义关联实现隐蔽控制的新攻击方法「Cordyceps」。该方法使模型学会一种信息隐藏方案,可编码和解码任意恶意指令。评估表明,该攻击在多种后门防御和提示注入防御下依然有效,揭示了数据投毒导致的一种新型、微妙的安全漏洞。

GradSentry: Gradient Spectral Entropy for Backdoor Sample Filtering in Large Language Model Fine-Tuning

第一作者: Haodong Zhao · 方向: AI 安全

Abstract:Fine-tuning Large Language Models with untrusted data exposes models to backdoor attacks, where poisoned samples cause targeted misbehavior. Existing sample-filtering defenses rely on clustering, which requires sufficient data and can fail at extreme poison ratios. We propose GradSentry ({Grad}ient {Sentry}), a backdoor sample filtering method based on the spectral entropy of per-sample gradients. Our key finding is that poisoned samples produce gradients with higher spectral entropy compared to clean samples. GradSentry captures output-altering backdoor signatures using per-sample gradient spectra, avoiding pairwise sample comparisons and clustering during feature construction. Importantly, our method is training-agnostic: it works for both parameter-efficient fine-tuning methods like LoRA and full-parameter tuning, as the gradient analysis operates independently of which...

论文介绍 针对大模型微调中利用有毒数据注入后门的安全问题,本文提出GradSentry防御方法。其核心发现是,有毒样本的梯度谱熵高于干净样本。基于此,该方法利用逐样本梯度谱进行特征构建和过滤,避免了聚类需求,且适用于LoRA等参数高效微调,为模型安全提供了一种新的、与训练无关的防御思路。

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

第一作者: Hwiwon Lee · 方向: 软件安全

Large language models (LLMs) now support automated software security tasks, including vulnerability discovery and proof-of-concept (PoC) generation. Existing benchmarks do not faithfully evaluate LLMs in real-world bug hunting scenarios because they rely on fuzzing harnesses, target-specific descriptions, or vulnerability-reproduction tasks. We present SEC-bench Pro, a benchmark for measuring agent bug hunting on critical, high-complexity software systems. This work discloses reports with concrete PoC inputs and links fixes into reproducible tasks through a three-phase pipeline for vulnerability collection, environment reconstruction, and oracle-based validation. We instantiate SEC-bench Pro with 183 validated vulnerabilities across V8 and SpiderMonkey, including a V8 subset with more than $1.5 million in cumulative Google Vulnerability Reward Program awards. These instances span...

论文介绍 针对现有基准测试无法真实评估语言模型在复杂软件系统中执行安全任务能力的问题,本文提出SEC-bench Pro基准。该基准通过三阶段流程收集并验证V8和SpiderMonkey等关键系统中的高复杂度漏洞,构建了183个可复现的任务实例。它旨在衡量语言模型智能体进行长期漏洞发现与利用的能力,推动自动化软件安全研究。

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

第一作者: Xiaochong Jiang · 方向: 软件安全

Abstract:Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterprise services at runtime. This creates a safety gap in tool composition: an agent can satisfy every per-tool permission check and still produce an unsafe end-to-end effect, such as reading a confidential document, summarizing it, and sending the summary to an external endpoint. We call this failure mode permission laundering. ChainCaps addresses it with a runtime rule: every value carries a sink-specific capability budget, and tool composition propagates budgets by intersection. A value can preserve or lose authority as it moves through a tool chain, but it cannot gain new authority through composition. We implement ChainCaps as a transparent MCP proxy that requires no changes to the agent or tool servers. On 82 tasks across five...

论文介绍 针对工具使用智能体在运行时组合多种工具可能产生不安全端到端效果的「权限洗白」问题,本文提出ChainCaps框架。该框架为每个数据值关联一个能力预算,并通过交集传播确保权限在工具链中只能保留或丧失,而不能通过组合新增。ChainCaps作为透明代理实现,无需修改智能体或工具服务器,有效增强了开放环境中的组合安全性。

Aligning Provenance with Authorization: A Dual-Graph Defense for LLM Agents

第一作者: Peiran Wang · 方向: 软件安全

Abstract:LLM-based agents are increasingly deployed in high-stakes scenarios such as email management, financial transactions, and code execution, where they interact with the external world through tool calling. During execution, these agents must read external data sources (emails, webpages, files) that attackers can control; through indirect prompt injection, attackers embed malicious instructions in this data to manipulate agents into performing unauthorized operations such as transferring funds to attacker-controlled accounts. Existing defenses either perform tool-call-level value checking without tracking where parameter values originate, or analyze execution traces from a single perspective without a clean authorization baseline for comparison. We propose AuthGraph, a dual-graph alignment defense framework that constructs two complementary graphs: an injected reasoning graph...

论文介绍 针对LLM代理在调用工具时可能因数据源被注入恶意指令而执行未授权操作的问题,本文提出AuthGraph双图对齐防御框架。该框架同时构建记录参数值来源的「推理图」和定义合法操作的「授权图」,并通过对齐来验证操作是否授权。这有助于追踪值的起源并确保符合授权策略,从而防御间接提示注入攻击,提升代理在关键场景中的安全性。

Beyond Epsilon: A Principled QIF Framework for Local Differential Privacy

第一作者: Ramon G. Gonze · 方向: 隐私保护

Local Differential Privacy (LDP) has become the de facto standard for privacy-preserving data collection in large-scale systems, in particular for the purpose of estimating frequencies. However, the current research landscape lacks a systematic and principled way to compare LDP protocols. The parameter $\varepsilon$ of LDP is considered the measure of privacy, but it only bounds worst-case distinguishability. Other comparisons rely on utility-driven analyses, where mechanisms are ranked based on their ability to preserve data utility for a given privacy budget $\varepsilon$. Both such kinds of comparisons fail to account for the strength of protocols against diverse attacker models. In this paper, we propose a framework for analyzing LDP frequency estimation protocols through the lens of Quantitative Information Flow (QIF). By modeling LDP mechanisms as probabilistic channels, we...

论文介绍 本文指出仅依赖隐私参数ε来比较本地差分隐私协议存在局限,因其仅约束最坏情况区分度。为此,作者提出一个基于定量信息流的分析框架,将LDP机制建模为概率信道。该框架能够更系统、更原则地评估不同LDP频率估计协议在对抗多种攻击者模型时的隐私强度,为隐私保护数据收集协议的设计与比较提供了新视角。

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models

第一作者: Hayden Helm · 方向: 系统安全

Abstract:Evaluating and mitigating a generative system's susceptibility to jailbreak attacks is critical to its safe deployment. Given the number of deployable systems, full per-configuration evaluation and optimization is impractical. In this paper, we formalize the behavioral geometry of a population of models that, by leveraging previously evaluated and defended models, supports both efficient susceptibility prediction and effective defense transfer across a population. We apply the framework to 79 models spanning 24 providers and to 100 system configurations of a single base model. Simple methods that use the behavioral geometry reach an AUPRC of $0.94$ for susceptibility detection with $\approx98\%$ fewer probes relative to a full evaluation. Using the behavioral geometry to select which model to transfer an optimized defense from outperforms same-provider assignment ($+2\%$, $p =...

论文介绍 针对全面评估和优化生成系统对越狱攻击易感性成本高昂的问题,本文提出利用模型群体的行为几何特性。该框架通过建模已评估模型的行为关联性,支持仅使用少量探针高效预测新模型的易感性,并能有效将经过优化的防御措施从一个模型转移到其他模型,显著降低了安全评估和防护部署的开销。

Context-Aware Metric Differential Privacy for Vehicle Trajectory Data

第一作者: Gaoyi Chen · 方向: 隐私保护

Metric Differential Privacy (mDP) generalizes differential privacy by allowing privacy guarantees to be expressed with respect to an arbitrary distance metric over secrets. While mDP has been adopted in geo-location protection, most existing mechanisms perturb each location record in isolation and do not model how contextual information (e.g., recent mobility history) affects the utility of the released data. This mismatch is particularly pronounced for vehicle mobility traces, where service quality often depends on temporally correlated locations. In this paper, we propose Context-aware mDP (C-mDP), a framework for vehicle location privacy that incorporates contextual dependencies into both the utility model and the privacy notion. C-mDP treats the protected secret as a context-augmented record and enforces metric indistinguishability over this augmented domain. We formulate optimal...

论文介绍 针对车辆轨迹数据发布中,现有度量差分隐私机制忽略上下文信息(如移动历史)导致数据效用降低的问题,本文提出C-mDP框架。该框架将上下文依赖同时纳入隐私定义和效用模型,通过处理增强的「上下文增强记录」来确保度量不可区分性。它在为车辆位置提供强隐私保证的同时,更好地保留了数据的服务质量和实用性。

Intelligent Detection and Mitigation of Carpet-Bombing DDoS Attacks in SDN Using Retrieval-Augmented Generation and Large Language Models

第一作者: Mohammed N. Swileh · 方向: 软件安全

Abstract:Software-Defined Networking (SDN) provides flexible and programmable network management; however, its centralized control architecture remains highly vulnerable to Distributed Denial-of-Service (DDoS) attacks, particularly Carpet-Bombing DDoS attacks that distribute malicious traffic across multiple targets to evade conventional detection mechanisms. In this paper, a Retrieval-Augmented Generation (RAG)-based framework is proposed for real-time detection and mitigation of Carpet-Bombing DDoS attacks in SDN environments. The proposed framework combines interface-level traffic features representation, semantic embedding generation, FAISS-based similarity retrieval, and Large Language Model (LLM)-driven contextual inference to classify traffic behavior without requiring conventional supervised model training or retraining. To evaluate the effectiveness of the proposed framework...

论文介绍 本文针对软件定义网络(SDN)集中式控制架构易受分布式拒绝服务(DDoS)攻击,特别是将恶意流量分散到多个目标以规避检测的“地毯式”攻击的问题,提出了一种基于检索增强生成(RAG)的实时检测与缓解框架。该框架结合接口级流量特征、语义嵌入、FAISS相似性检索及大语言模型(LLM)驱动的上下文推理,无需传统监督模型训练即可分类流量行为,旨在提升SDN环境下的防御能力。

Sandlock: Confining AI Agent Code with Unprivileged Linux Primitives

第一作者: Cong Wang · 方向: 软件安全

Abstract:AI agents increasingly run untrusted code on developer machines: shell commands generated by language models, third-party scripts retrieved at runtime, and tool plugins of unknown provenance. Existing isolation mechanisms impose tradeoffs that fit this workload poorly: containers and microVMs add privilege, image-management, and startup costs, while ad-hoc process controls and wrappers (e.g. chroot, ulimit) provide weak guarantees and little syscall-level control. Sandlock is a lightweight Linux process sandbox organized around a simple split: static, input-independent policy is compiled into kernel-enforced rules, while a narrow supervisor handles runtime-dependent decisions and virtualized effects. This split lets Sandlock enforce filesystem, network, IPC, and syscall policies without root, cgroups, images, or mandatory namespaces. It also supports dynamic network decisions...

论文介绍 针对AI代理在开发者机器上运行不可信代码(如LLM生成的命令、第三方脚本)带来的安全问题,本文提出了Sandlock,一个轻量级的Linux进程沙箱。其核心设计是将静态、输入无关的策略编译为内核强制执行的规则,并由一个小型监督器处理运行时决策。这种方法能在不需要root权限、容器或镜像的情况下,强制执行文件系统、网络、IPC和系统调用策略,为AI代理提供了灵活且低开销的代码执行隔离环境。

AgentSecBench: Measuring Prompt Injection, Privacy Leakage, and Tool-Use Integrity in LLM Agents

第一作者: Faruk Alpay · 方向: AI 安全

Abstract:LLM agents process trusted instructions, retrieved records, and tool observations through a common generative channel. This conflates data flow with authority: an untrusted string can affect a secret-bearing response or an action proposal even when no application policy authorizes that influence. We introduce AgentSecBench as an empirical instantiation of a formal security framework for this problem. The framework defines three games-instruction-integrity, retrieval-confidentiality, and capability-integrity-under a common notion of intent-to-execution noninterference with permitted leakage. It represents an application policy as a projection onto authorized observations and capabilities, distinguishes prompt annotations from enforcing projections, and measures both adversarial advantage and whether a defense closes the relevant model-visible channel before generation. The...

论文介绍 本文提出了AgentSecBench基准,用于实证评估大语言模型(LLM)代理面临的安全威胁。该基准基于一个形式化安全框架,定义了指令完整性、检索机密性和能力完整性三类“安全游戏”,并引入“意图到执行的非干扰”概念进行评估。它旨在量化未受信输入对代理秘密信息或操作能力的影响,为评估和开发针对提示注入、隐私泄露等攻击的防御提供了标准化测量方法。

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly

第一作者: Yihe Fan · 方向: AI 安全

Abstract:LLM-based agents are increasingly used for cybersecurity tasks, but most existing systems rely on fixed, human-designed scaffolds that struggle to adapt across diverse targets and failure modes. We introduce \textsc{CyberEvolver}, a self-evolving cybersecurity agent framework that iteratively revises its own scaffold based on experience from failed execution attempts. Self-evolution in cybersecurity is challenging because the space of possible scaffold changes is largely unstructured, execution feedback is sparse and often obscured by the environment, and low-diversity updates can cause errors to compound over repeated iterations. \textsc{CyberEvolver} addresses these challenges with a four-layer evolvable agent architecture that decomposes scaffold optimization into structured components, a trace-to-diagnosis mechanism that converts noisy execution logs into actionable...

论文介绍 现有基于大语言模型(LLM)的网络安全代理通常依赖固定的、人工设计的框架,难以适应多样化的攻击目标和失败模式。本文提出了CyberEvolver,一个能够基于执行失败经验迭代修订自身框架的自进化网络安全代理。该框架采用四层可进化架构,并通过“跟踪到诊断”机制将嘈杂的执行日志转化为可操作的反馈,以结构化地优化代理的“脚手架”,提升其自主性和任务成功率。

Enhancing Autonomous Online Intrusion Detection for IoT with Balanced Learning, Reliable Pseudo-Labels, and Lightweight Architectures

第一作者: Hanzala Afzaal · 方向: 网络安全

Abstract:The rapid proliferation of Internet of Things (IoT) devices has created an urgent demand for adaptive, resource-efficient Intrusion Detection Systems (IDS) capable of handling dynamic and evolving cyber threats. This paper investigates AOC-IDS, a state-of-the-art autonomous online IDS published at IEEE INFOCOM 2024, which employs an Autoencoder (AE) with Cluster Repelling Contrastive (CRC) loss and an autonomous Gaussian-based decision module. We first successfully replicate AOC-IDS on the UNSW-NB15 benchmark, achieving 89.39% accuracy in close agreement with the published 89.19%. We then identify four key limitations: class imbalance, unreliable pseudo-label generation, limited generalization, and computational overhead for IoT deployment, and propose targeted improvements for each. Our XGBoost-BalSamp method achieves 95.45% accuracy on UNSW-NB15, a gain of 6.26% over the...

论文介绍 本文研究了针对物联网(IoT)环境的自主在线入侵检测系统(IDS)。作者首先复现了已发表的AOC-IDS系统,确认了其性能,随后识别出其在类不平衡、伪标签生成不可靠、泛化能力有限以及计算开销方面的四个关键局限。为此,论文提出了XGBoost-BalSamp等针对性改进方法,在基准数据集上显著提升了检测准确率,旨在开发更适用于资源受限、威胁动态变化的IoT场景的高效入侵检测方案。

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

第一作者: Tongxi Wu · 方向: 密码学协议

Abstract:Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshold mechanism. We challenge this assumption by revealing that safety behavior is governed by an instability region where small perturbations induce stochastic refusal decisions rather than deterministic outcomes. We develop a multi-metric diagnostic framework combining external and internal signals to characterize this instability. Through systematic experiments, we identify a characteristic diagnostic signature: inputs in unstable regimes exhibit elevated output uncertainty yet decreased internal safety activation, a decoupling phenomenon that explains why detection-based defenses fail against sophisticated attacks. Building on this framework, we introduce Furina, a jailbreak attack that deliberately induces this signature through...

论文介绍 本文挑战了大语言模型(LLM)和多模态大语言模型(MLLMs)的安全对齐机制近乎二元阈值的假设,揭示了存在一个“不稳定性区域”,其中微小扰动会导致随机的拒绝决策。作者开发了一个多指标诊断框架,并基于此提出了名为Furina的越狱攻击。该攻击通过诱导特定的诊断特征(即高输出不确定性但低内部安全激活)来利用这种不稳定性,揭示了现有基于检测的防御可能存在的盲点。

Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges

第一作者: Xianglin Yang · 方向: AI 安全

Abstract:The known stylistic biases in LLM judges, such as a preference for verbosity or specific sentence structures, present an underexplored security vulnerability. In this work, we introduce BITE (BIas exploraTion and Exploitation), a black-box adversarial framework that learns semantics-preserving edits to mislead an LLM judge and artificially inflate the scores it assigns. We cast the selection of stylistic edits as a contextual bandit problem and use a LinUCB policy to adaptively choose edits that maximize the judge's score without access to model parameters or gradients. Empirically, we test BITE across a diverse range of LLM judges and tasks, including both pointwise and pairwise comparisons on chatbot leaderboards and AI-reviewer benchmarks. BITE achieves an attack success rate exceeding 65% and raises scores by 1-2 points on a 9-point scale, all while preserving semantic...

论文介绍 本文揭示了大语言模型(LLM)评委中存在的风格偏好(如偏爱冗长或特定句式)可被利用的安全漏洞。作者提出了BITE黑盒对抗框架,将风格化编辑的选择建模为上下文老虎机问题,使用LinUCB策略自适应地选择能在不改变语义的前提下,人为抬高LLM评委评分的编辑。实验表明,BITE能在多种评委和任务上实现超过65%的攻击成功率,平均提升评分1-2分。

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

第一作者: Xuanye Zhang · 方向: AI 安全

Abstract:LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents toward inappropriate/wrong tools and enabling malicious actions. Most existing attacks primarily manipulate the tool metadata, which is easily detectable by auditing and may lose effectiveness as modern agents increasingly adopt memory modules to refine tool selection policies through accumulated experience. This paper proposes MemMorph, the first attack that bias tool selection by poisoning the agent's long-term memory. Rather than explicitly dictating the tool invocation decision, MemMorph injects a small number of crafted records that are disguised as technical facts, incident reports, and operational policies. These poisoned records reshape the agent's contextual perception and decision-making process, leading it to...

论文介绍 针对由大语言模型(LLM)驱动的代理通过记忆模块优化工具选择策略的趋势,本文提出了首个通过投毒代理长期记忆来影响其工具选择的攻击方法MemMorph。该方法不直接操控工具元数据,而是注入少量伪装成技术事实、事件报告或操作策略的精心构造记录,以此重塑代理的上下文感知和决策过程,引导其选择不适当或错误的工具,从而执行恶意操作。

On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning

第一作者: Xiaotian Ye · 方向: AI 安全

Abstract:Counterfactual tuning (CFT) has emerged as a promising paradigm for Large Language Model (LLM) unlearning by training models to generate alternative fictitious knowledge in place of undesired content. However, in this work, we find that this paradigm still underperforms other paradigms in some aspects, and identify two previously overlooked pitfalls underlying this gap: (1) knowledge conflict, where mutual inconsistencies within counterfactual corpora induce conflicting gradients that disrupt parameter optimization, and (2) hallucination spillover, where fitting false targets instills a persistent fabrication bias, inflating hallucination rates on unrelated domains. To systematically diagnose these issues, we introduce RWKU+, an extended benchmark equipped with novel trade-off metrics and gradient-level diagnostic tools. Our work further discusses the limitations and overhead...

论文介绍 本研究探讨了用于大语言模型遗忘的反事实微调范式存在的隐藏成本。作者发现该方法存在知识冲突和幻觉溢出两个未被重视的问题:不一致的反事实语料会导致梯度冲突,而拟合虚假目标则会引入持久的编造偏差。为此,研究引入了RWKU+基准,以系统地诊断这些缺陷。

Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals

第一作者: Akindoyin Akinrele · 方向: AI 安全

Abstract:Prompt injection poses a critical threat to the safe deployment of large language models, yet existing detection approaches are typically evaluated under limited settings that do not reflect real-world operating constraints. In this work, we present a deployment-aware evaluation of prompt injection detection using a multi-model and multi-regime experimental framework. We compare lexical, semantic, structural, and transformer-based detectors across multiple out-of-distribution settings, repeated data splits, and both ranking and thresholded deployment metrics. We introduce interpretable structural signals that capture hierarchy overrides, system prompt spoofing, role redefinition, and evasion patterns, and assess their contribution both within sparse models and in combination with strong encoder baselines. Our results show that detection performance is highly regime-dependent...

论文介绍 该研究对提示注入检测方法进行了部署感知评估。通过多模型、多场景的实验框架,比较了多种检测器在分布外数据上的表现。研究引入了可解释的结构化信号,并指出检测性能高度依赖于具体的部署场景和数据分布,现有评估未能充分反映这一现实。

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

第一作者: Pengzhen Chen · 方向: 安全研究

Reliable watermarking of panoramic imagery is fundamentally challenged by arbitrary 3D rotations. As panoramas are defined on the sphere, they naturally transform under the action of $SO(3)$, rendering conventional planar representations and augmentation-based robustness strategies inadequate and devoid of theoretical guarantees. To address this, we formulate panoramas as spherical signals and leverage $SO(3)$ representation theory to derive provably rotation-invariant descriptors. While spherical harmonic coefficients transform equivariantly under rotations, the natural invariant constructions are typically limited to zeroth-order statistics which eliminate directional information and severely constrain embedding capacity. In this work, we introduce a principled third-order invariant construction by coupling higher-order $SO(3)$ irreducible representations via tensor products and...

论文介绍 本文针对全景图像在三维旋转下的可靠水印问题展开研究。传统方法在球面信号上缺乏理论保证。作者利用SO(3)表示理论,通过耦合高阶不可约表示构建了具有可证明旋转不变性的三阶描述子,在保持方向信息的同时增强了嵌入容量。

FuzzPilot: Plateau-Triggered Recipe Validation for Structured Text Fuzzing

第一作者: Zhiyi Yao · 方向: 软件安全

Abstract:FuzzPilot is a controller for AFL++ that moves expensive reasoning out of the mutation hot path. When coverage plateaus, it snapshots the corpus, prepares candidate mutation recipes, evaluates them in short isolated AFL++ micro-campaigns, and promotes only recipes with positive validation reward. Recipes are JSON data, not generated code: a native custom mutator consumes operator weights, byte ranges, corpus-selection rules, and dictionary tokens. Candidate recipes can come from local rules or from a language-model agent, with Ghidra-derived constants and decompiled context as target hints. This preprint reports a deliberately narrow cJSON evaluation. We compare vanilla AFL++ and the full FuzzPilot agent over five 14,400 s repetitions per arm. cJSON is saturated: baseline AFL++ reaches the exposed 269-edge ceiling at a median of about 2,500 s. The experiments therefore do not...

论文介绍 FuzzPilot是一个针对AFL++的控制器,旨在解决覆盖平台期问题。当覆盖增长停滞时,它会评估候选的变异配方(如JSON数据)在短时隔离测试中的效果,仅奖励有效配方。该方法将昂贵的推理移出热路径,并在cJSON库的评估中展示了其基础架构。

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

第一作者: Kevin Kuo · 方向: AI 安全

Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmful usage without fine-tuning at all? In this paper, we show that open-weight safeguards are susceptible to simpler strategies that, despite being well known, have not been systematically evaluated against these safeguards. Specifically, we evaluate two low-cost attacks--abliteration and prefilling--that do not rely on gradient-based optimization. Across three harmfulness evaluation benchmarks (BeaverTails, HarmBench, and AdvBench), these...

论文介绍 现有针对开源权重大型语言模型的安全防护旨在防止微调滥用。本文研究发现,这些防护对更简单的攻击策略仍然脆弱。作者评估了消融和预填充这两种无需基于梯度优化的低成本攻击,表明它们能有效地绕过防护,利用模型固有的有害知识。

"You do understand that people don't trust technology?": Explaining Trusted Execution Environments to Non-Experts

第一作者: McKenna McCall · 方向: 软件安全

Abstract:Trusted Execution Environments (TEEs) protect confidentiality and integrity of trusted applications by creating an isolated environment for executing code. Prior work has shown that users may feel more comfortable sharing data when they know it will be protected by a TEE, especially if they understand what a TEE is. In this study, we evaluated text-based explanations introducing TEEs to non-experts. We analyzed existing TEE explanations to develop candidate explanations and evaluated them via vignette scenarios with 966 crowdworkers. The explanations that enhanced understanding most were non-technical ones that highlighted specific threats that can be prevented by a TEE. Surprisingly, even the explanations that enhanced understanding had little effect on willingness to use the TEE-enhanced technology. These results provide insights into ways to communicate technical security...

论文介绍 本研究通过包含966名参与者的实验,评估了向非专家解释可信执行环境的不同文本方式。结果表明,能增强理解的最佳解释是非技术性的、并突出TEE能防御的具体威胁的说明。然而,令人惊讶的是,这些能提升理解的解释对用户使用TEE技术的意愿影响甚微。

Device Context Protocol: A Compact, Safety-First Architecture for LLM-Driven Control of Constrained Devices

第一作者: Dongxu Yang · 方向: 密码学协议

Abstract:Large language models are increasingly used as orchestrators of external tools via the Model Context Protocol (MCP), but MCP is built for software services with megabytes of memory and does not descend to the microcontrollers that dominate the long tail of physical devices. Recent work (IoT-MCP) ports MCP to edge gateways at 74 KB peak memory; this still excludes the smallest commodity MCUs and, critically, does not address the safety problem of giving an unreliable caller (an LLM that may hallucinate or be prompt-injected) direct control of physical hardware. We present the Device Context Protocol (DCP): a sub-50-byte typical frame (6-byte header + CBOR payload + optional 16-byte HMAC), a manifest schema in which capability scoping, range and type checks, dry-run evaluation, and units-as-types are protocol-layer primitives, and a host-side Bridge that rejects malformed or...

论文介绍 本文提出了设备上下文协议,用于安全地连接大语言模型与资源受限的物理设备。协议设计了小于50字节的帧格式,并将能力范围、类型检查、试运行评估等安全机制作为协议层原语,通过主机端的桥接器拒绝不规范或超出范围的请求。

SolarChain: Bridging Physical Law, Verifiable Trust, and Sustainable Markets for Urban Energy Resilience

第一作者: Shilin Ou · 方向: 系统安全

Abstract:Urban decarbonization requires scaling rooftop solar across millions of fragmented producers, yet cities face a fundamental tension: energy data is easily manipulated, and economic incentives often reward speculation rather than actual infrastructure deployment. We present SolarChain, a platform that resolves both problems by anchoring digital accountability to the thermodynamic limits of solar energy conversion. Using real-time meteorological data, geospatial coordinates, and first-principles calculations of solar yield, the system establishes a hard physical boundary for every panel's maximum possible output; any reported generation exceeding this limit is automatically rejected before entering the shared ledger. This trustless verification enables a peer-to-peer marketplace with programmatic reward structures that continuously reinvest value into equipment maintenance and...

论文介绍 SolarChain是一个解决城市太阳能数据可信性与市场激励问题的平台。它通过实时气象数据和物理第一性原理计算,为每块光伏板设定一个最大可能输出的硬性物理边界。任何超出此边界的报告发电量在进入共享账本前都会被自动拒绝,从而建立了基于物理定律的信任验证机制。

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

第一作者: Xintong Hu · 方向: VLA 通用模型 · 来源: cs.RO

Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair trajectories with coarse goal-level language, leaving execution-critical details such as active arm, approach direction, and contact region unspecified. This limits steerable policy learning and robotic video understanding. We introduce FineVLA, an open framework for action-aligned fine-grained VLA supervision. The framework includes: (1) a data construction tool that unifies 972,247 trajectories across 85K tasks from 10 open-source robot datasets and builds FineVLA-Data, a human-verified dataset of 47,159 fine-grained trajectories; (2) a held-out benchmark with 500 videos, 10,816 atomic facts, and 1,030 VQA questions; (3) a robotics-specialized VLM annotator for...

论文介绍 现有的视觉-语言-动作(VLA)模型常缺乏对执行细节的指令遵循能力。FineVLA 提出了一个开放框架,用于构建动作对齐的细粒度VLA监督数据。该工作整合了多个开源数据集,构建并验证了一个包含4.7万余条细粒度轨迹的基准数据集,并配套了评估基准与专用标注工具,旨在提升机器人策略的可控性与视频理解能力。

VR-DAgger: Immersive VR for Dexterous Data Collection and Uncertainty-Guided On-Policy Correction

第一作者: René Zurbrügg · 方向: 机器人操作 · 来源: cs.RO

Learning from demonstrations is effective for robotic manipulation, but collecting sufficient task-specific data remains a major bottleneck. Under distribution shift, small errors compound, performance degrades, and expert time is often spent on redundant, low-value corrections instead of the few critical failure cases.

论文介绍 从演示中学习机器人操作时,收集足够且高质量的数据仍是瓶颈。VR-DAgger 提出利用沉浸式虚拟现实进行灵巧的数据收集,并引入不确定性引导的在线策略修正方法。该方法旨在减少分布偏移下的复合误差,使专家能更专注于修正关键失败案例,从而提升示教学习的效率和策略鲁棒性。

Manipulating Tangible Virtual Object Dynamics to Promote Learning of Precision Force Generation

第一作者: Alberto Garzás-Villar · 方向: 具身智能 · 来源: cs.RO

Robotic haptic devices combined with virtual reality offer novel opportunities to train fine force generation, an essential yet overlooked component of post-stroke rehabilitation. This study proposes that manipulating the rendered dynamics of tangible virtual objects can be leveraged to train precise force control while engaging the somatosensory system. We conducted an experiment with fifty healthy participants who performed a curling-inspired task in which they had to stretch a virtual spring to generate a target release force to propel the stone to a predefined location on the ice sheet. During training, the spring's force-elongation relationship was modeled as either a linear or non-linear function, i.e., a Gaussian or antisymmetric Gaussian (AS-Gaussian) function with zero derivative at the release target force. Results indicate that the AS-Gaussian group consistently achieved...

论文介绍 该研究探索了利用机器人触觉设备与虚拟现实来训练精细力量生成,特别是针对中风后康复。通过操纵虚拟物体的力-伸长关系(如采用反高斯模型),研究人员发现可以更有效地训练参与者在释放目标力量时的精确控制能力,为利用具身认知原理进行康复训练提供了新思路。

RCSP: Risk-Sensitive Conjectural Scenario Planning for Safe Dynamic Robot Navigation

第一作者: Zhengye Han · 方向: 导航与运动 · 来源: cs.RO

Mobile robots can fail before they collide: a velocity that is safe now may commit the robot to a passage that moving obstacles will soon close. We study this predictive near-miss commitment problem and propose Risk-Sensitive Conjectural Scenario Planning (RCSP), a planning layer that evaluates candidate commands against plausible short-horizon obstacle futures. RCSP maintains a lightweight belief over local motion conjectures, samples future interactions, penalizes high-risk tails, and executes through a local safety check. In controlled MuJoCo bottleneck tasks, the RCSP planner reaches the goal without collisions and yields higher secondary safety and path-quality point estimates than a non-adaptive predictor, with additional latency. In ROS2/Gazebo, adding the local safety layer to a standard Nav2 stack reduces dynamic near-miss failures. On official DynaBARN/Jackal transfer, tuned...

论文介绍 针对移动机器人在动态环境中可能因当前安全速度而陷入未来危险通道的「预测性近失承诺」问题,本文提出了风险敏感猜想场景规划(RCSP)。该规划器通过评估候选命令在假设障碍物未来运动下的风险,执行带局部安全检查的路径,从而在模拟和实物平台中显著降低了动态导航的碰撞与近失失败。

PhyPush: One Push is All You Need for Sensorless Physical Property Estimation with Physics-Guided Transformers

第一作者: Koyo Fujii · 方向: 机器人操作 · 来源: cs.RO

Accurately estimating object mass and friction is fundamental to achieving reliable and adaptive robotic manipulation. Although interactive perception provides a powerful mechanism for inferring such properties, most existing approaches depend on specialized hardware such as force/torque sensors, tactile arrays, or multi-camera motion-capture systems, limiting scalability and deployment. This paper presents PhyPush, a physics-guided Transformer framework that estimates an object's mass and friction coefficient using only kinematically derived end-effector velocity from a single push. This typically requires data available on standard robotic arms. The model incorporates constraints from Newton's second law and the Coulomb friction model through a physics-guided loss, improving physical consistency and generalization to unseen objects and surfaces. Across diverse simulation and...

论文介绍 准确估计物体的质量和摩擦系数对可靠自适应操作至关重要,但现有方法常依赖昂贵的专用传感器。PhyPush 提出一个物理引导的Transformer框架,仅利用单次推动产生的末端执行器速度(通常可从标准机械臂获取),就能估计物体的物理属性,通过融入牛顿定律和库仑摩擦模型提升泛化能力。

Learning to Balance Motor Thermal Safety and Quadrupedal Locomotion Performance with Residual Policy

第一作者: Yuhang Wan · 方向: 导航与运动 · 来源: cs.RO

Abstract:Motor thermal management is often overlooked in the context of electrically-actuated robots, particularly legged robots, but motor overheating is a key factor that limits long-duration locomotion especially under payload conditions. This paper integrates a whole-body thermal model of a quadruped robot into the reinforcement learning pipeline to update motor temperatures, and proposes a two-stage training framework for motor thermal management. In this framework, a nominal policy is first pre-trained as a locomotion baseline capable of traversing diverse terrains. A residual policy is then trained on top of the nominal policy to provide corrective actions based on the robot's thermal state, ensuring high performance under low-temperature conditions and preventing motor overheating under high-temperature conditions. Simulation results demonstrate that the proposed policy...

论文介绍 电机过热是限制电驱足式机器人长时运动和负重能力的关键因素,但常被忽视。本研究将机器人的全身热模型集成到强化学习流程中,并提出一个两阶段训练框架:首先预训练一个可穿越多地形的基础运动策略,然后在此基础上训练一个残差策略,根据电机热状态进行修正,以平衡运动性能与热安全。

Towards Shared Embodied Intelligence in Humanoid Robots through Optimization Development and Testing of the Human Aware ergoCub Robot

第一作者: Carlotta Sartore · 方向: 具身智能 · 来源: cs.RO

Abstract:Collaboration is central to human behavior, enabling tasks beyond individual capability. This ability arises from coordinating actions through internal representations of others, a concept known as shared intelligence. Additionally, humans are characterized by physical bodies and cognitive abilities that are optimized in response to their environment, a phenomenon referred to as embodied cognition. Designing humanoid robots that collaborate safely and effectively with people requires unifying these principles. Here we propose an architecture that integrates shared intelligence and embodied cognition to enable robots to physically collaborate with humans, where robot hardware and control are optimized for human metrics, using representations of the human body and motion intelligence. The ultimate goal is to achieve a form of shared embodied intelligence. Specifically, our...

论文介绍 安全有效的人机物理协作要求机器人兼具对人类动作的理解(共享智能)以及与环境适配的身体认知能力。本文以 ergoCub 机器人为例,提出一种整合上述原则的架构,其硬件和控制均针对人类指标进行优化,并利用人体表征和运动智能,旨在实现一种共享的具身智能形式,促进人机物理协作。

A Bioinspired Underwater Robot with a Latch-Mediated Soft Bistable Mechanism

第一作者: Chongze Bi · 方向: 具身智能 · 来源: cs.RO

Abstract:Underwater robotics has advanced significantly over recent decades. however, the development of miniaturized underwater robots remains limited by low energy densities of traditional power sources. Nature offers compelling solutions-organisms like mantis shrimps and fleas utilize latch-mediated spring actuation (LaMSA) systems that achieve rapid movements through a decoupled energy storage and release mechanism. Despite extensive studies of LaMSA, replicating such rapid, asymmetric actuation within simple, compact structures remains challenging. In this work, we introduce a bioinspired, soft bistable actuator with an integrated latch mechanism that enables asymmetric energy input and release using a single motor. Coupled with fin structures, this design facilitates efficient underwater propulsion and maneuverability. Experimental results demonstrate stable periodic flapping...

论文介绍 针对微型水下机器人能源密度低的限制,该研究受虾蛄等生物利用锁存介导弹簧驱动(LaMSA)机制实现快速运动的启发,设计了一种软体双稳态执行器。该设计通过集成锁存机构,利用单电机实现不对称的能量存储与释放,配合鳍结构,实现了稳定、高效的水下拍动推进与机动。

Learning Compositional Symbolic Task Rules from Demonstrations with Inductive Logic Programming

第一作者: Oleh Borys · 方向: 模仿学习 · 来源: cs.RO

Abstract:Learning from Demonstration~(LfD) should capture not only how a task is executed, but also its high-level task structure that explains the demonstrated behavior. As robots become more autonomous, such task representations must be inspectable, reusable, and human-interpretable. To address this, we study how to represent and learn robotic tasks with inductive logic programming~(ILP) by decomposing a complex task into a series of simpler learning objectives at different abstraction (ontological) levels. The system infers symbolic rules from demonstrations and prior (domain) knowledge, and reuses learned rules when learning higher-level task structure. We evaluate the approach in a synthetic block-assembly scenario and show that the learned abstractions are interpretable and support strong generalization to harder, held-out tasks with unseen objects. These results provide...

论文介绍 本研究旨在学习机器人任务的高层次、可解释结构,以提升自主系统的可检查性、可重用性和人机交互性。作者提出一种基于归纳逻辑编程的方法,将复杂任务分解为不同抽象层次的子目标。系统能够从演示和领域知识中推断出符号化的任务规则,并在学习更高层任务结构时复用已有规则。在合成的方块装配场景中评估表明,该方法学习到的抽象概念具有可解释性,并能良好泛化至未见过的物体和更难的任务。

Can VLA Models Learn from Real-World Data Continually without Forgetting?

第一作者: Jiarun Zhu · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) models provide a promising foundation for general-purpose robotics. However, their successful deployment in real-world scenarios requires the ability to continually acquire new skills while retaining previously learned behaviors. While pioneering research has studied the continual learning of VLA models in narrowly simulated environments, this challenge remains largely unexplored under realistic conditions. To address this limitation, we construct a real-world continual learning dataset comprising four sequential manipulation tasks, spanning rigid-object pick-and-place, contact-rich pressing, and deformable-object folding. Using this dataset, we conduct comprehensive experiments and find that VLA models suffer significant catastrophic forgetting when continually learning from heterogeneous real-world demonstrations. We then systematically evaluate...

论文介绍 视觉语言动作模型是通用机器人的有前景基础,但在真实场景部署中需要持续学习新技能而不遗忘旧技能。该研究构建了一个涵盖四类顺序操作任务的真实世界持续学习数据集,包含刚性物体抓放、接触丰富的按压和柔性物体折叠。通过全面实验发现,VLA模型在从异构的真实世界演示中持续学习时,会遭受严重的灾难性遗忘。论文进一步系统评估了多种持续学习方法在应对这一挑战时的效果。

Look Further: Socially-Compliant Navigation System in Residential Buildings

第一作者: Akira Shiba · 方向: 导航与运动 · 来源: cs.RO

Abstract:The distance at which a mobile robot reacts to a person strongly impacts various qualities of the human-robot interaction. In this paper, we focus on the navigation of a mobile delivery robot platform in a residential indoor hallway environment. Social navigation methods typically focus on avoiding uncomfortable human-robot interactions, such as when a robot encroaches on someone's personal space. Since personal space has been shown to be in the range of just a few meters, social navigation methods typically focus on deconflicting and resolving these short-range interactions. In this work, however, we demonstrate that by extending the reaction distance to over eight meters, far beyond the typical interaction distance, we can improve the human's perception of the robot's motion. We introduce the Proactive Lane-Changing (PLC) motion pattern and a navigation system that leverages...

论文介绍 本研究聚焦于移动配送机器人在住宅室内走廊环境中的导航问题。传统社交导航方法主要关注避免短距离内侵犯个人空间带来的不适。本文证明,将机器人的反应距离主动延长至八米以上,远超典型交互距离,可以显著改善人类对其运动的感知。为此,论文提出了一种「主动变道」运动模式及相应的导航系统,旨在实现更早、更符合社交礼仪的路径规划与避障。

On the Generalization Capabilities, Design Choices and Limitations of Keypoint Imitation Learning

第一作者: Thomas Lips · 方向: 机器人操作 · 来源: cs.RO

Abstract:RGB-based imitation learning requires many demonstrations to generalize to unseen objects or scenes, motivating research into intermediate representations to improve generalization for robotic manipulation. Visual foundation models enable one-shot extraction of keypoints to provide such representation. However, it remains unclear how to integrate them into imitation learning optimally and when they outperform alternative representations. We combine approaches from previous works on keypoint imitation learning (KIL) and investigate several design choices to provide practical guidelines. Using over 2000 real-world rollouts, we also assess the generalization capabilities of KIL to unseen objects and scene variations. KIL achieves a 75% overall success rate across five tasks, significantly outperforming the RGB baseline (47%) and performing on par with S2-diffusion (73%). Finally...

论文介绍 基于RGB图像的模仿学习通常需要大量演示才能泛化到新物体或场景。视觉基础模型提供了从单张图像中提取关键点的能力,作为一种中间表征,有望改善泛化性。本文结合先前关键点模仿学习的研究,探讨了多种设计选择并提供实践指南。通过超过2000次真实世界实验,评估了关键点模仿学习在未见物体和场景变化下的泛化能力。结果显示,该方法在五项任务中总体成功率达75%,显著优于RGB基线,并与S2-diffusion方法表现相当。

L-Learning : A Lyapunov-Based Approach Leveraging Lagrangian Mechanics for Efficient and Stable Robot Tracking

第一作者: Quan Quan · 方向: 导航与运动 · 来源: cs.RO

Abstract:This paper presents L-Learning, a novel data-driven control framework for robotics that integrates Lyapunov stability theory with Lagrangian mechanics to enhance trajectory tracking performance. While traditional control methods often suffer from performance degradation in dynamic and uncertain environments, data-driven approaches, while more adaptable, are frequently limited by high sample complexity and a lack of rigorous stability guarantees. L-Learning mitigates these challenges by explicitly learning the system's energy function from data, thereby optimizing performance while ensuring closed-loop stability intrinsically. Characterized by superior control accuracy, theoretical stability guarantees, and high sample efficiency, L-Learning represents a promising solution for practical robotic applications.

论文介绍 本文提出L-Learning,一种新颖的机器人数据驱动控制框架。该框架将李雅普诺夫稳定性理论与拉格朗日力学相结合,旨在提升轨迹跟踪性能。传统控制方法在动态不确定环境中性能易下降,而纯数据驱动方法则常受限于样本复杂度高且缺乏严格的稳定性保证。L-Learning通过从数据中显式学习系统的能量函数来缓解这些挑战,在优化性能的同时能内在保证闭环稳定性。该框架具有高控制精度、理论稳定性保证和高样本效率的特点。

HyperSim: A Holistic Sim-To-Real Framework For Robust Robotic Manipulation

第一作者: Junyi Dong · 方向: 机器人操作 · 来源: cs.RO

Abstract:Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, transferring robotic manipulation policies from simulation to the real world (sim-to-real) remains a formidable challenge due to the domain gap. This paper presents HyperSim, a holistic framework spanning from synthetic data generation to policy training and seamless real-world deployment. To systematically bridge the sim-to-real gap, HyperSim is realized through three core pillars: high-fidelity environment synthesis, adversarial trajectory generation, and sim-and-real co-training. Collectively, these modules address domain discrepancies by enhancing visual fidelity, expanding data coverage, and enforcing domain-invariant representations. We rigorously validate HyperSim through a large-scale...

论文介绍 扩大数据规模和多样性对于具身智能的泛化至关重要。本文提出了HyperSim,一个从合成数据生成、策略训练到无缝真实世界部署的整体框架,以系统性地弥合仿真到现实的差距。该框架通过三个核心支柱实现:高保真环境合成、对抗性轨迹生成以及仿真与现实的协同训练。这些模块通过增强视觉保真度、扩大数据覆盖范围和强制学习领域不变表征来解决领域差异。论文通过大规模实验严格验证了HyperSim框架的有效性。

Enabling Extensible Embodied Capabilities with Tools

第一作者: Xueyang Zhou · 方向: 具身智能 · 来源: cs.RO

Abstract:Most existing embodied intelligence methods formulate perception, reasoning, planning, and control within a unified parameterized policy. Yet these capabilities are inherently hierarchical and heterogeneous, making them difficult to reliably learn and modularize within a single model. We propose a capability externalization approach that decouples heterogeneous capabilities into independently optimized tools, dynamically invoked at inference time. To this end, we introduce Embodied Tool Protocol (ETP), a standardized protocol for embodied tool registration, discovery, invocation, and execution, and curate 100+ validated tools spanning perception, cognition, reasoning, and execution as the tool base. Building on this, we construct EmbodiedToolBench to evaluate both whether tool augmentation improves embodied performance and how well current models use tools across...

论文介绍 现有具身智能方法通常将感知、推理、规划与控制统一于一个参数化策略中,但这些能力本质上是分层且异质的,难以在单一模型中可靠学习和模块化。本文提出一种能力外化方法,将异质能力解耦为独立优化的「工具」,并在推理时动态调用。为此,论文引入了「具身工具协议」,用于工具的标准化注册、发现、调用和执行,并构建了包含100多个经过验证工具的工具库。研究还构建了评估基准,以检验工具增强是否能提升具身任务表现。

Provably Safe Motion Planning Under Unknown Disturbances

第一作者: Ibon Gracia · 方向: 导航与运动 · 来源: cs.RO

Abstract:We present a provably safe sampling-based motion planning algorithm for robotic systems affected by random disturbances of unknown distribution. We consider systems with linear or linearizable dynamics evolving in workspace with arbitrary-shaped obstacles subject to state and control constraints. Safety requirements are formulated as chance-constraints. Our approach leverages data from trajectories of the system to learn a Wasserstein ambiguity tube, i.e., a sequence of ambiguity sets, which contains the trajectory of the system's state distribution with high confidence. This ambiguity tube is then used in a probabilistically complete algorithm to grow a sampling-based motion planning tree that respects the constraints of the problem. We show that learning several lower-dimensional ambiguity tubes instead of a single high-dimensional one effectively reduces the conservatism...

论文介绍 本文提出了一种可证明安全的采样式运动规划算法,用于处理受未知分布随机扰动影响的机器人系统。该研究考虑具有线性或可线性化动力学、在存在任意形状障碍物和约束的工作空间中运动的系统。安全要求被表述为机会约束。方法利用系统轨迹数据学习Wasserstein歧义管,即一系列歧义集,以高置信度包含系统状态分布的轨迹。随后,该歧义管被用于一个概率完备的算法中,以构建一个尊重问题约束的采样规划树。

Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient

第一作者: Haoxiang You · 方向: 机器人操作 · 来源: cs.RO

Abstract:We present the stochastic decoupled policy gradient (SDPG), a lightweight visual reinforcement learning (RL) method that trains diverse visuomotor control policies end-to-end within a few hours on a single NVIDIA RTX 4080 GPU. SDPG estimates policy gradients via random perturbations of trajectory rollouts, requiring orders of magnitude fewer batch-rendered environments and substantially reducing compute and memory overhead. On visual MuJoCo benchmarks, SDPG consistently outperforms baseline methods in training time, memory usage, and rewards. Finally, to support future research, we introduce a suite of realistic visual robotics benchmarks spanning dexterous manipulation, challenging locomotion, and demonstrate effective sim-to-real transfer on physical hardware.

论文介绍 研究高效训练视觉强化学习策略的问题。提出随机解耦策略梯度(SDPG),通过随机扰动轨迹估计梯度,减少批量渲染环境需求,降低计算和内存开销。在视觉MuJoCo基准测试中,该方法在训练时间、内存使用和奖励方面优于基线。引入视觉机器人基准套件,支持模拟到现实迁移,适用于机器人操作等应用。

Robust Koopman Control Barrier Filters for Safe Actor-Critic Reinforcement Learning

第一作者: Dhruv S. Kushwaha · 方向: 策略学习 · 来源: cs.RO

Abstract:Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints during both training and deployment. Control barrier functions (CBFs) provide a principled mechanism for enforcing forward invariance through minimally invasive safety filters, but their use in model-free RL is limited by the need for accurate dynamics and hand-designed barrier certificates. We propose Robust Koopman-CBF SAC, a safety-filtered actor--critic framework that learns a finite-dimensional Koopman predictor from data, constructs affine CBF constraints in the lifted space, and enforces them through a quadratic-program safety layer. To account for finite-dimensional Koopman approximation error, the CBF condition is tightened using a projected residual margin estimated from held-out rollout data. The critic is trained on the...

论文介绍 研究机器人安全强化学习中满足状态和输入约束的问题。提出鲁棒Koopman-CBF SAC框架,从数据中学习有限维Koopman预测器,在提升空间构造仿射控制屏障函数约束,并通过二次规划安全层施加。为处理近似误差,使用投影残差余量收紧约束条件,旨在提升任务性能的同时确保安全。

Multi-Robot Box Transport over Different Surfaces with Decentralized Role-based Proportional Control

第一作者: Aditya Bhatt · 方向: 机器人操作 · 来源: cs.RO

Abstract:Collaborative transport of objects via pushing by multiple robots has many applications, ranging from construction and warehouse environments to post disaster debris clean-up. Achieving collaborative transport over surfaces with different inclination and friction properties however poses unique challenges. To address these challenges, this paper presents an asynchronous decentralized task and motion planning approach for transporting rectangular boxes of varying mass over flat, uphill and downhill terrain. Such a decentralized approach alleviates communication, synchronization and consensus needs and mitigates single point of failure issues. Our approach, called R2P2 or Roles with Rules and Proportional-control Primitive, assigns roles (e.g., push, support and prevent) to robots based on rules cognizant of the mode of manipulation needed (box rotation vs translation); this is...

论文介绍 研究多机器人在不同地形上协作运输矩形箱的挑战。提出异步去中心化任务和运动规划方法,基于角色和规则分配推、支撑、防止等角色,考虑操作模式如箱体旋转或平移。该方法减少通信和同步需求,适用于仓库、建筑和灾后清理等应用,降低单点故障风险。

Closing the Loop in Teleoperation: Episode-Level Data Quality Assessment and Feedback for High-Quality Demonstration Collection

第一作者: Gokul Narayanan · 方向: 模仿学习 · 来源: cs.RO

Abstract:Industrial automation is at a pivotal moment, as Physical AI is driving a transition from rigid, hand-engineered automation systems toward more flexible and adaptive systems. This shift has created a growing demand for large-scale, real-world robot demonstration data, making teleoperation an increasingly important mechanism for data collection. However, high-quality teleoperated demonstrations remain difficult to obtain in practice, as novice operators often produce episodes that are task-successful but suboptimal for downstream use due to inefficient motion, repeated corrections, or operation near robot joint limits. We present a Data Quality Assessment and Feedback (DQAF) framework that closes the loop in teleoperation by providing immediate post-episode feedback grounded in semantic task progress and robot telemetry. The framework extracts quality relevant signals such as...

论文介绍 研究遥操作中获取高质量演示数据的挑战。提出数据质量评估和反馈(DQAF)框架,通过分析语义任务进度和机器人遥测数据,在 episode 级别提供即时反馈。该框架提取质量相关信号,帮助操作者改进演示,提升物理AI数据收集效率,支持工业自动化中的灵活系统。

NightSight: Passive Computation for Navigation in Dark Using Events

第一作者: Deepak Singh · 方向: 导航与运动 · 来源: cs.RO

Abstract:Small aerial robots are particularly well-suited for search and rescue in confined and hazardous environments due to their agility, low cost, and ability to traverse through cluttered spaces that are inaccessible to larger platforms. However, enabling autonomous navigation in complete darkness remains a significant challenge, because small aerial robots cannot easily accommodate perception systems that demand substantial payload, power, or computation. In this work, we present a lightweight perception approach that combines a monocular event camera, a coded aperture lens, and an infrared dot projector to enable navigation in such conditions. The projected pattern, when imaged through the coded aperture, produces depth dependent blur signatures that implicitly encode scene geometry. We train a convolutional neural network to decode these signatures into dense depth maps using...

论文介绍 研究小型飞行器在完全黑暗环境中的自主导航问题。提出轻量级感知方法,结合单目事件相机、编码孔径镜头和红外点投影仪,投影图案通过编码孔径产生深度相关模糊特征。训练卷积神经网络解码这些特征为密集深度图,适用于搜索和救援等受限环境。

Collaborative Navigation and Exploration with $β$-Sparse Gaussian Processes

第一作者: Evangelos Psomiadis · 方向: 导航与运动 · 来源: cs.RO

Abstract:Collaborative navigation of heterogeneous robots in unknown environments poses significant challenges due to sensing, communication, and computational limitations. In this work, a lead robot navigates toward a target while a mobile sensor robot (e.g., a drone) assists by transmitting information about its locally observed environment under bandwidth constraints. We propose a framework that enables the sensor to jointly select its transmitted map points and navigation actions online, while also predicting unexplored regions of the environment. To this end, we present $\beta$-Sparse Gaussian Processes, a novel and robust variational sparse Gaussian Process model for task-aware inducing point selection. Furthermore, we develop an action-selection strategy that balances task relevance with exploration. Simulations on Mars and Earth maps show that the framework can reduce path cost...

论文介绍 研究异质机器人在未知环境中协作导航,受限于带宽约束。提出框架使传感器机器人在线选择传输的地图点和导航动作,并预测未探索区域。核心是β-稀疏高斯过程,用于任务感知的诱导点选择,平衡任务相关性和探索,在火星和地球地图模拟中显示路径成本降低。

YOLO26-RipeLoc Lite: A lightweight architecture for tomato ripeness detection and picking point localization in greenhouse robotic harvesting

第一作者: Rajmeet Singh · 方向: 机器人操作 · 来源: cs.RO

Abstract:In greenhouse tomato production, automated harvesting requires accurate detection of ripe tomatoes, ripeness classification, and precise picking-point localization for robotic end-effectors. This paper proposes YOLO26-RipeLoc Lite, a lightweight deep learning architecture based on YOLO26 for simultaneous detection, ripeness classification, and center-point localization of greenhouse tomatoes. The model introduces three modifications: (1) a Lightweight Feature Pyramid Network (LFPN) with depthwise separable convolutions for efficient multi-scale fusion, (2) a Ripeness-Aware Attention Module (RAAM) with dual pooling and a learnable ripeness bias vector for enhanced color-texture discrimination, and (3) a Compact Detection Head (CDH) with shared convolutions and an integrated center-point regression branch for direct grasp planning. The model is evaluated on a custom dataset of...

论文介绍 研究温室番茄自动采摘中成熟度检测和中心点定位的问题。提出YOLO26-RipeLoc Lite轻量级架构,基于YOLO26进行修改,包括轻量化特征金字塔网络、成熟度感知注意力模块和紧凑检测头。该模型支持同时检测、成熟度分类和中心点回归,旨在提高农业机器人采摘效率和准确性。

The Sensation Modulating Network:Haltability as the architectural ground for object-directed phenomenology

第一作者: G. Nagarjuna · 方向: 具身智能 · 来源: cs.RO

Abstract:Cognitive science remains split between cognitivism - which accounts for recursion and language but cannot ground formal symbols in meaning - and 4E approaches - which ground cognition in the body but rarely specify the body's architecture in enough detail to support generativity. We argue the impasse stems from an incomplete account of the embodied agent's architecture, and propose one: the Sensation Modulating Network (SMN), the cognitive agent conceived as the whole body, organized at every anatomical scale by opponent dynamics, built from Sensation Modulators that sense and act through one substrate, paired into Coordinated Action Zones routed by a body-wide broadcast network. Three commitments give the SMN its purchase. Haltability - the recruitment of antagonistic affordance into co-activated equilibrium - provides the architectural locus that object-directed...

论文介绍 研究弥合认知主义和4E认知方法差距的问题。提出感觉调制网络(SMN),将认知体视为由感觉调制器组成的全身网络,基于对立动态和可停顿性。可停顿性作为架构基础,支持对象导向现象,为具身智能提供详细理论框架,探索认知体的组织结构。

OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes

第一作者: Regina Kurkova · 方向: 机器人操作 · 来源: cs.RO

Abstract:Semantic mapping methods are increasingly used as intermediate scene representations for downstream robotic reasoning and manipulation, yet their evaluation is still largely tied to fixed benchmark datasets with limited coverage of manipulation-relevant corner cases. In this work, we extend OSMa-Bench toward controllable benchmarking with prompt-generated synthetic indoor scenes. Our pipeline automatically generates scene descriptions, synthesizes corresponding environments with SceneSmith, and adapts the resulting assets into an OSMa-Bench-compatible simulation format. This adaptation requires a nontrivial intermediate layer, including semantic normalization, material and texture repair, shader fallback policies, floor handling, navigation setup, and controlled lighting configuration. A key advantage of the proposed setup is that the original scene-generation prompt is known...

论文介绍 语义映射方法常作为机器人操作中的场景表示,但其评估依赖固定数据集,缺乏操作相关边缘案例。本文扩展OSMa-Bench,利用提示生成合成室内场景,自动化流程包括场景描述生成、环境合成及仿真格式适应,涉及语义归一化和材质修复等。该设置允许可控基准测试,提升机器人操作中语义映射方法的评估覆盖度和真实性。

Towards Real-World Identification of Fatigued Muscle Groups via Musculoskeletal Simulation

第一作者: Jenishkumar Chauhan · 方向: 具身智能 · 来源: cs.RO

Abstract:Contactless diagnosis of musculoskeletal disorders can potentially improve population health as well as robot behaviours in collaborative settings. However, current diagnosis methods require an in-person physical examination in which a trained physician senses, through contact, the force applied by various muscles. Simulation tools exist, but their use for diagnosis with real data is under-explored. In this paper, we propose an algorithm for identifying which upper-limb muscle group is fatigued. Our algorithm compares the realworld free-space motion of the subject with that of a simulated musculoskeletal model, and is therefore contactless: preventing the need for invasive sensing or in-person assessment. Our algorithm simulates various fatigue conditions using a physics-based musculoskeletal model and extracts diagnostic motion features from both real and simulated data...

论文介绍 肌骨疾病诊断通常需要接触式物理检查,现有仿真工具用于真实数据诊断尚待探索。本文提出算法识别上肢疲劳肌肉群,通过比较真实自由空间运动与仿真肌骨模型的运动,实现非接触诊断。算法仿真多种疲劳条件,从真实和模拟数据中提取诊断运动特征,避免侵入式传感或现场评估,适用于健康改善和机器人协作场景。

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

第一作者: Xiaoyuan Cheng · 方向: 策略学习 · 来源: cs.LG

Abstract:Model-based reinforcement learning (RL) can be effectively supported at scale through the use of world models. However, in practice, scaling such approaches remains fundamentally limited. A commonly recognized challenge is model bias and error compounding, which degrade long-horizon predictions. Beyond these issues, we identify a more critical yet underexplored bottleneck: a structural misalignment between search and value learning in existing world model approaches. In particular, policy improvement often relies on value functions induced by a separate, non-search policy, resulting in training inconsistency and ultimately suboptimal learning. To address this limitation, we propose Model-Based Diffusion Policy Optimization (MBDPO) in world models, a framework that unifies search and policy optimization through diffusion policy representations, thereby unlocking the potential...

论文介绍 基于模型的强化学习在世界模型支持下可有效扩展,但面临模型偏差、错误累积以及搜索与价值学习结构性错位等挑战。本文提出基于模型的扩散策略优化(MBDPO)框架,通过扩散策略表示统一搜索和策略优化,以解决训练不一致性问题。该方法旨在提升世界模型强化学习的扩展性和性能,减少长期预测误差。

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

第一作者: Heriberto Cuayahuitl · 方向: 数据集与评测 · 来源: cs.AI

Abstract:Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, their application to textual or spoken medical consultations is still an open research problem. This paper proposes MeDial-Speech, a novel speech dataset for training and evaluating Med-AIs that can carry out consultations with patients. It was collected in realistic environments from robot-patient and doctor-patient dialogues, contains 111+ hours of speech data (without data augmentation), and covers four health conditions: Lewy body dementia, heart failure, shoulder pain, and angina. In addition, we propose a dialogue benchmark via sentence selection (with 20 options) to evaluate three state-of-the-art LLMs: GPT-5 mini, DeepSeek-V3, and Claude Sonnet 4. Experimental results reveal that Claude Sonnet 4 is the best in sentence...

论文介绍 大型语言模型在医疗语音咨询中的应用仍是开放问题。本文提出MeDial-Speech数据集,基于机器人-患者和医生-患者对话的语音数据,包含111小时以上内容,覆盖四种健康状况。通过句子选择基准评估GPT-5 mini等模型,为训练和评估医疗AI进行患者咨询提供资源,支持语音处理任务。

市场总览

美股方面,主要ETF如S&P 500(RSI 71.3超买)和Nasdaq 100(RSI 75.3超买)价格接近52周高点,呈多头排列,但部分出现MACD死叉信号,显示短期上涨动能可能放缓。加密货币市场情绪极度恐慌(恐慌贪婪指数25),总市值2.61T USD,24小时下跌1.19%,BTC主导率57.9%;主要币种BTC、ETH和SOL均处于空头排列,RSI偏低(BTC 41.7,ETH 35.8),技术面偏弱。中概股板块整体疲软,多数标的如阿里巴巴(RSI 43.5空头排列)、腾讯(RSI 32.5接近超卖)呈下行趋势,近5日跌幅显著。商品与外汇中,原油期货CL=F近5日暴跌14.68%,MACD死叉;黄金期货GC=F中性;美元指数DXY偏强,接近52周高点并多头排列。整体市场技术面分化明显,美股超买但动量减弱,加密恐慌蔓延,中概下行压力,商品波动加剧,需关注关键支撑阻力位。

今日关注

QQQ Nasdaq 100 ETF
偏上行

Nasdaq 100 ETF的RSI 14指标读数为75.3,已进入超买区域。当前价格730.28,距离52周高点仅差-0.12%,显示强劲上升动能。移动平均线呈现多头排列(价格高于SMA20 698.21、SMA50 644.84、SMA200 615.47)。然而,MACD指标已出现死叉信号(MACD值20.898低于信号线21.6008),暗示短期动量可能减弱,存在技术性回调风险。

0700.HK 腾讯控股 (0700.HK)
偏下行

腾讯控股的RSI 14指标为32.5,接近超卖阈值30。当前价格434.8,较52周低点仅高0.65%,且处于空头排列(价格低于SMA20 459.86、SMA50 489.46、SMA200 575.85)。近5个交易日跌幅达-5.48%,MACD值为-14.8447且低于信号线-13.5343,确认下跌趋势。技术面显示下行压力较大,需关注是否形成底部支撑。

CL=F WTI 原油期货
中性

WTI原油期货的RSI 14指标为42,处于中性范围。价格在91.95,近5日大幅下跌-14.68%,但仍在200日均线71.81之上。MACD出现死叉(MACD值-0.1205低于信号线1.2199),短期动量转弱。整体趋势中性,关键支撑位在200日均线附近,阻力位在50日均线98.25,需观察价格在均线附近的反应。

全部资产

^VIX

VIX 恐慌指数

$17.01 +2.53%
5 日
-5.81%
距 52w 高
-51.8%
RSI(14)
43.1
趋势
中性
SMA 20 / 50 / 200
17.52 / 20.44 / 18.37
MACD / 信号
-0.763 / -0.834
接近 52 周低

^TNX

10Y 美债收益率 (%)

$4.49 -1.43%
5 日
-2.81%
距 52w 高
-10.1%
RSI(14)
53.6
趋势
多头
SMA 20 / 50 / 200
4.47 / 4.37 / 4.20
MACD / 信号
0.065 / 0.062
多头排列

DX-Y.NYB

美元指数 DXY

$99.08 -0.07%
5 日
-0.22%
距 52w 高
-1.5%
RSI(14)
54.4
趋势
多头
SMA 20 / 50 / 200
98.66 / 98.93 / 98.56
MACD / 信号
0.153 / 0.064
接近 52 周高多头排列

SPY

S&P 500 ETF

$750.59 +0.66%
5 日
+1.62%
距 52w 高
-0.2%
RSI(14)
71.3
趋势
多头
SMA 20 / 50 / 200
733.36 / 698.44 / 679.44
MACD / 信号
12.481 / 13.146
RSI 超买接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$730.28 +1.78%
5 日
+3.46%
距 52w 高
-0.1%
RSI(14)
75.3
趋势
多头
SMA 20 / 50 / 200
698.21 / 644.84 / 615.47
MACD / 信号
20.898 / 21.601
MACD 死叉 (4 天前)RSI 超买接近 52 周高多头排列

AAPL

Apple

$308.33 -0.16%
5 日
+3.52%
距 52w 高
-1.1%
RSI(14)
77.4
趋势
多头
SMA 20 / 50 / 200
291.38 / 271.72 / 261.99
MACD / 信号
10.090 / 9.212
RSI 超买接近 52 周高多头排列

MSFT

Microsoft

$416.03 -0.61%
5 日
-1.77%
距 52w 高
-25.1%
RSI(14)
52.3
趋势
中性
SMA 20 / 50 / 200
416.17 / 400.85 / 459.88
MACD / 信号
3.414 / 3.995

NVDA

Nvidia

$214.86 -0.22%
5 日
-3.36%
距 52w 高
-9.2%
RSI(14)
53.3
趋势
多头
SMA 20 / 50 / 200
214.66 / 197.50 / 187.20
MACD / 信号
6.098 / 7.441
MACD 死叉 (2 天前)多头排列

GOOGL

Alphabet

$388.88 +1.54%
5 日
-2.03%
距 52w 高
-4.8%
RSI(14)
61.0
趋势
多头
SMA 20 / 50 / 200
387.41 / 342.87 / 297.15
MACD / 信号
12.693 / 16.258
多头排列

TSLA

Tesla

$433.59 +1.78%
5 日
+5.76%
距 52w 高
-13.1%
RSI(14)
60.9
趋势
中性
SMA 20 / 50 / 200
412.00 / 389.18 / 410.58
MACD / 信号
10.700 / 10.795
MACD 死叉 (3 天前)

META

Meta

$612.34 +0.34%
5 日
+0.18%
距 52w 高
-23.1%
RSI(14)
46.4
趋势
空头
SMA 20 / 50 / 200
615.79 / 617.79 / 668.68
MACD / 信号
-6.561 / -6.169
空头排列
加密恐慌贪婪
25
极度恐慌
加密总市值
$2.61 T
-1.19% / 24h
BTC 主导率
57.9%
ETH 9.6%
24h 成交量
$96.6 B
活跃币 17,406

BTC-USD

Bitcoin

$75,520.29 -2.28%
5 日
-2.60%
距 52w 高
-40.2%
RSI(14)
41.7
趋势
空头
SMA 20 / 50 / 200
78,540.28 / 77,070.68 / 80,266.42
MACD / 信号
-345.306 / 117.149
空头排列

ETH-USD

Ethereum

$2,071.78 -1.88%
5 日
-2.80%
距 52w 高
-58.2%
RSI(14)
35.8
趋势
空头
SMA 20 / 50 / 200
2,196.78 / 2,263.35 / 2,533.99
MACD / 信号
-53.787 / -42.044
空头排列

SOL-USD

Solana

$83.64 -1.61%
5 日
-4.03%
距 52w 高
-67.0%
RSI(14)
41.9
趋势
空头
SMA 20 / 50 / 200
88.61 / 86.57 / 106.38
MACD / 信号
-0.797 / -0.170
空头排列

BABA

阿里巴巴 (BABA)

$129.47 -0.41%
5 日
-2.84%
距 52w 高
-32.8%
RSI(14)
43.5
趋势
空头
SMA 20 / 50 / 200
134.93 / 131.66 / 149.54
MACD / 信号
-0.482 / 0.546
空头排列

PDD

拼多多 (PDD)

$96.64 +2.24%
5 日
+1.76%
距 52w 高
-30.7%
RSI(14)
46.6
趋势
空头
SMA 20 / 50 / 200
97.83 / 99.51 / 113.68
MACD / 信号
-1.056 / -0.972
MACD 死叉 (1 天前)空头排列

JD

京东 (JD)

$29.99 -1.74%
5 日
-5.09%
距 52w 高
-18.6%
RSI(14)
44.9
趋势
中性
SMA 20 / 50 / 200
30.98 / 29.98 / 30.48
MACD / 信号
0.378 / 0.565
MACD 死叉 (1 天前)

0700.HK

腾讯控股 (0700.HK)

HK$434.80 -0.96%
5 日
-5.48%
距 52w 高
-36.3%
RSI(14)
32.5
趋势
空头
SMA 20 / 50 / 200
459.86 / 489.46 / 575.85
MACD / 信号
-14.845 / -13.534
接近 52 周低空头排列

GC=F

黄金期货

$4,497.80 -0.06%
5 日
-0.19%
距 52w 高
-19.5%
RSI(14)
38.4
趋势
中性
SMA 20 / 50 / 200
4,597.21 / 4,647.90 / 4,359.31
MACD / 信号
-52.982 / -42.577

CL=F

WTI 原油期货

$91.95 -2.07%
5 日
-14.68%
距 52w 高
-23.0%
RSI(14)
42.0
趋势
中性
SMA 20 / 50 / 200
100.46 / 98.25 / 71.81
MACD / 信号
-0.121 / 1.220
MACD 死叉 (3 天前)

USDCNY=X

美元 / 人民币

¥6.78 -0.22%
5 日
-0.51%
距 52w 高
-6.0%
RSI(14)
35.0
趋势
空头
SMA 20 / 50 / 200
6.81 / 6.84 / 6.99
MACD / 信号
-0.012 / -0.013
接近 52 周低空头排列
风险提示

本报告仅基于公开行情数据计算的技术指标进行客观描述,过去走势不代表未来表现,不构成任何投资建议。所有解读仅供技术指标解读参考,投资者应自行判断并承担风险。

Australia politics live: PM plugs Tim Wilson’s ‘terrific’ book as question time debate over CGT changes resumes

Anthony Albanese has taken another shot at Tim Wilson’s book. Follow today’s news live Get our breaking news email, free app or daily news podcast Syrian children should be treated ‘sensitively and gently’, Ryan says Independent MP Monique Ryan says the women and children returning from a Syrian cam

中文摘要 澳大利亚总理安东尼·阿尔巴尼斯在质询时间辩论中再次提及蒂姆·威尔逊的书,涉及资本利得税政策讨论。政治辩论聚焦预算和税收变化。

Iceland’s foreign minister fears ‘Brexit moment’ in country’s EU accession referendum

Þorgerður Katrín Gunnarsdóttir accuses opponents of fearmongering amid warnings over misinformation and AI Iceland’s foreign minister has said she fears her country faces a “Brexit moment” in its looming EU referendum amid warnings over misinformation, foreign interference and AI. With just over thr

中文摘要 冰岛外长担心即将到来的欧盟加入公投可能出现「英国脱欧时刻」,并警告错误信息和人工智能的风险。

Paul Keating urges Labor to stick with capital gains tax overhaul and avoid exemptions that would hurt economy

Exclusive: Former PM says changes to tax rates are ‘so marginal that no entrepreneurial initiative is likely to be thwarted’ Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Paul Keating has urged Labor to stick to its guns on controv

中文摘要 澳大利亚前总理保罗·基廷敦促工党坚持资本利得税改革,避免会损害经济的豁免。他称税率变化微乎其微。

Justin Stevens resigns as ABC director of news after four years in role

Managing director Hugh Marks says Stevens made ‘incredible commitment’ to broadcaster over 19 years Justin Stevens has resigned as ABC director of news after four years in the role, citing personal and professional reasons. ‌ABC managing director Hugh Marks said Stevens had made an “incredible commi

中文摘要 贾斯汀·史蒂文斯辞去ABC新闻总监职务,任职四年。总经理休·马克斯称赞其19年的贡献。

Iceland Warms to Europe

The country has always stood apart from the continent. Then President Trump started threatening Greenland.

中文摘要 冰岛开始对欧洲暖化,此前特朗普总统威胁格陵兰岛。冰岛传统上与欧洲保持距离。

Police were outgunned during Bondi beach massacre due to lack of long-arm rifles, royal commission hears

Officers were ‘placed at significant risk, being in a gunfight armed with 9mm Glocks against long-arms’, NSW police deputy commissioner says Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Police were outgunned at the Bondi massacre

中文摘要 皇家委员会听取称,警方在邦迪海滩枪击案中因缺乏长枪而火力不足。官员表示警方用9毫米手枪对抗长枪,面临巨大风险。

Dozens killed in Lebanon as Israel intensifies strikes

Israel says it struck 100 Hezbollah infrastructure sites and fighters in Lebanon, after PM Benjamin Netanyahu vows to "crush" the Iran-backed group.

中文摘要 以色列加强打击黎巴嫩,造成数十人死亡。以色列称袭击了100个真主党基础设施和人员,内塔尼亚胡誓言「粉碎」该组织。

Can Nigel Farage’s Right-Wing Party Win It All in Britain?

Nigel Farage’s anti-immigrant, populist agenda has helped his Reform U.K. Party emerge from the fringe of British politics. But it faces an uphill climb to win power.

中文摘要 奈杰尔·法拉奇的右翼政党Reform U.K.凭借反移民议程在英国崛起。该党面临夺取权力的挑战。

Where Time Is Always 15 Minutes Apart From Everywhere Else

In Nepal, a nation wedged between India and China, a unique time zone is just one expression of a singular national identity.

中文摘要 尼泊尔采用独特时区,比其他地方快15分钟,体现其国家认同。该国位于印度和中国之间。

Inventor of the Basque Cheesecake Plans to Retire. His Secret: He Prefers Chocolate

Santiago Rivera is widely credited with creating the “burnt” cheesecake in the 1980s, though he doesn’t love the spinoffs it has spawned. Decades later, he’s preparing to hand over his kitchen to his children.

中文摘要 巴斯克芝士蛋糕发明者圣地亚哥·里维拉计划退休。他于1980年代创建该蛋糕,正准备将厨房交给子女。

EU Chamber on Business Confidence in China

Jens Eskelund, President of the European Union Chamber of Commerce in China, tells Bloomberg’s The China Snow that the business group's latest survey shows an "inflection point" in confidence. (Source: Bloomberg)

中文摘要 欧盟商会中国区主席Jens Eskelund向Bloomberg表示,该商会最新调查显示在华企业商业信心已出现“拐点”,表明市场情绪可能转变。

Everybody Wants Fed Independence, Rubenstein Says

Carlyle Co-Founder and Co-Chairman David Rubenstein discusses his outlook for interest rates and his expectations of the Federal Reserve under Chair Kevin Warsh. He also shares his views on investing with Stephen Engle from the sidelines of the UBS Asian Investment Conference in Hong Kong. (Source:

中文摘要 凯雷集团联合创始人David Rubenstein在瑞银亚洲投资会议上强调美联储独立性至关重要,并讨论了利率前景及对新任主席Kevin Warsh的政策预期,分享投资见解。

Trad bookies ♥ prediction markets

C-c-c-c-combo rakers

中文摘要 标题指出传统博彩公司对预测市场持积极态度,但具体细节未明。

Ex-Fed Official Evans on What's Next for Monetary Policy

Former Chicago Fed President Charles Evans discusses what's next for monetary policy. Evans, who ran the Chicago Fed for 16 years until 2023, emphasizes that monetary policy ahead will likely remain cautious, balancing inflation risks with growth concerns. He speaks with Stephen Engle from the sidel

中文摘要 前芝加哥联储主席Charles Evans表示,未来货币政策可能保持谨慎,在通胀风险和增长担忧之间寻求平衡,他管理该行16年至2023年。

SK Hynix, Micron Join $1 Trillion Market Cap Club | The Asia Trade 5/27/2026

"Bloomberg: The Asia Trade" brings you everything you need to know to get ahead as the trading day begins in Asia. Bloomberg TV is live from Tokyo and Sydney with Shery Ahn and Haidi Stroud-Watts, getting insight and analysis from newsmakers and industry leaders on the biggest stories shaping global

中文摘要 Bloomberg报道,SK Hynix和Micron市值均突破1万亿美元,加入该俱乐部,反映半导体行业在2026年的强劲增长。

FirstFT: Is the Board of Peace in limbo?

Also in today’s newsletter: BP boardroom shake-up and EU car subsidies

中文摘要 FirstFT新闻通讯质疑“Board of Peace”是否陷入僵局,并涵盖BP董事会变动及欧盟汽车补贴等议题。

ASX Shares Eye Decade Low After Cost Woes Fuel Profit Downgrades

ASX Ltd. shares have tumbled to their lowest level in almost a decade after rising costs prompted several brokers to trim their earnings outlooks for Australia’s beleaguered exchange operator.

中文摘要 澳大利亚交易所运营商ASX Ltd.股价跌至近十年低点,因成本上升促使多家经纪商下调其盈利前景,该公司面临运营压力。

Defense Billionaire’s Father Builds New Business on CSG Model

The founder of CSG NV, which completed the biggest ever initial public offering for a pure defense company four months ago, is seeking to raise as much as €200 million ($233 million) this year to expand his industrial group.

中文摘要 CSG NV创始人计划今年筹集最多2亿欧元(2.33亿美元)以扩展工业集团,此前该公司完成纯国防公司最大IPO。

Korean Stocks Surge 100% in 2026 to Surpass Dotcom Era Gains

The breathtaking rally in South Korean stocks took gains for 2026 to 100%, eclipsing even the historic run-ups seen before the dotcom bubble burst and during the nation’s industrial boom in the late 1980s.

中文摘要 韩国股市2026年涨幅达100%,超越互联网泡沫时期和1980年代工业繁荣时期的历史涨幅,显示市场显著反弹。

Samsung workers set for $400,000 bonus after deal to share AI profits

Agreement with union ends wrangling over how to share spoils of boom at memory-chip maker

中文摘要 三星与工会达成协议,员工将获得高达40万美元奖金,以分享公司在AI领域的利润增长,结束了双方关于利润分配的争执。

Thailand Eyes $5 Billion From Notes, Loans as Bond Yields Soar

Thailand plans to raise about $5 billion through a mix of promissory notes and term loans to fund a raft of measures to ease living costs, shunning bonds after the Iran war sent sovereign yields to multi-month highs.

中文摘要 泰国计划通过本票和定期贷款筹集约50亿美元,资助生活成本缓解措施,避免债券发行,因伊朗战争导致主权收益率升至多月高位。

Ferrari shares slump after it unveils first fully electric car

The new Luce model has divided opinion on social media, and comes despite intense pressure from Chinese EV makers.

中文摘要 法拉利股价下跌,因其推出首款全电动汽车Luce,该车型在社交媒体引发争议,并面临中国电动汽车制造商的激烈竞争。

【干草铺公益站】运营情况通告

各位佬,太忙就简单说了 前两天号池不足,所以临时关闭了 5.5,现在恢复了,模型都放开了 号池够用,但不宽裕,所以鼓励大家尽量不用 5.5,后续可能会单独调高 5.5 的倍率,但是目前没这个计划,放心,不发公告我是绝对不会偷偷调倍率的 关于破限问题,我最近会开始封一批账号,这也是为了大家更多佬友的利益,希望大多数佬友理解,如果有误封可以私聊,不是误封的话就不用找了,因为破限问题发了不止一次公告了依然是关于破限,近期我会逐渐灰度一些拦截策略 关于注册和已有用户余额见底的问题,目前依然没有开放注册的计划,但是余额购买我会尽快补充小铺的 干草铺不会停止运营,但是可能会根据号池情况关闭 5.5 渠道,

关于男人必定经历的四大神圣地方,一旦染上就戒不掉了。

四大神圣:足疗->柔式->私影->商K。 我说实话,我都20大几了,马上快奔3了,也就体验过一回装修很大,能看电影的足疗店,让男的给我按了一回。体验一般般,也就那样。 就在前几月前,有一天身体很累,选了一家评价还OK的店,准备去做一下足疗。离家差不多几百米左右,走了大概十几分钟,到了,映入眼帘的是一款长长的楼梯,带拐弯的,越往里走,灯光开始慢慢变蓝,充满着幽密的气息。 到了前台,正好路过一个身着白色旗袍,身材妖娆的女性从旁边走过,带着一丝洗发水的清香。我心跳开始加速,算了,说太详细了,太墨迹了。 反正就是按了一回,开了个盲盒,有个女的给我按,穿着丝袜旗袍,先给你按手,按腿,再按脚,避免不了一些

发现个有趣&美观的个人介绍页

昨天晚上刷帖子的时候,看到一个冷门开源项目(只有几百 star ),看起来很酷炫,拿来微改了下,部署了下,效果还不错: https://me.kaisir.cn/ 19 个帖子 - 18 位参与者 阅读完整话题

雷总这是干啥?又给了 820 亿!!!佬们进来狠狠蹬(0.01 元续费直接变 1200 亿!)

拿到的佬点个赞就行,还有 70 多亿 key:tp-cpk7yqh7pwigecbwc7dt8kqeh9f17b1v72qyyymf53oh70ic 兼容 OpenAI 接口协议:https://token-plan-cn.xiaomimimo.com/v1 兼容 Anthropic 接口协议:https://token-plan-cn.xiaomimimo.com/anthropic 我勒个飞天大草,0.01 续费直接变 1200 亿,雷大圣人的恩情还不完啊 刚刚 429了,现在好像已经恢复了 40 个帖子 - 20 位参与者 阅读完整话题

追加了三千个体验兑换码cdk(领过的就没有啦,没有LD账号的,可以加群私信管理员 “C” 领取)

领取地址 (没有LD账号的,可以加群私信管理员 “C” 领取) cdk.linux.do LINUX DO CDK Linux Do 社区 CDK 快速分享平台 - 让分享变得更简单 兑换地址 ai.centos.hk New API 统一的 AI 模型聚合与分发网关,支持将各类大语言模型跨格式转换为 OpenAI、Claude、Gemini 兼容接口,为个人与企业提供集中式模型管理与网关服务。 使用文档 doc.centos.hk ShowDoc 一个非常适合IT团队的在线API文档、技术文档工具。你可以使用Showdoc来编写在线API文档、技术文档、数据字典、在线手册 35 个帖子 -

你们至于吗?满地都是MIMO

你们至于吗?满地都是MIMO,就算它不怎么样,咱也不能把token全扔地上吧,感觉这几天遍地都是MIMO,一上来就是几个亿十几个亿的。吓死人了,不愧是MI。 不好意思,大佬们,没有其它意思,就是感慨一下,确实一翻都是一页一页的送mimo 43 个帖子 - 35 位参与者 阅读完整话题

any 大善人支持 gpt-5.5了

那个男人的 马上见底了 any 大善人 居然上了 gpt 5.5 这下好了,可以猛猛的瞪了,问我为啥偏爱 5.5 因为4.7 有点糖 85 个帖子 - 47 位参与者 阅读完整话题

小米mimo史诗级大降价,降幅最高达99%,性价比比肩deepseek

尊敬的开发者, MiMo-V2.5 全系大幅调价,最高降幅 99% MiMo-V2.5-Pro(/百万 tokens) 输入(命中缓存):¥0.025 | 输入(未命中缓存):¥3 | 输出:¥6 MiMo-V2.5(/百万 tokens) 输入(命中缓存):¥0.02 | 输入(未命中缓存):¥1 | 输出:¥2 MiMo-V2.5-TTS 系列 继续限时免费 Xiaomi MiMo API Open Platform Token Plan 加量不加价 Credits 加量不加价:V2.5 系列模型用量可提升 5-8 倍;对 cache、输入、输出整体比例均有计量优化,整体更清晰。 Cred