每日简报

2026-07-01

← 历史归档

hasaneyldrm/exercises-dataset

HTML · ★ 6,911 · 🍴 824 · 📈 1,343 stars today

A comprehensive dataset of 433 fitness exercises. Each entry includes name, category, target muscle group, equipment, instructions, thumbnail image, and animation video.

中文介绍 提供包含433个健身动作的综合数据集,每条记录涵盖名称、分类、目标肌肉群、所需器械、动作说明及图文视频素材。适用于健身App开发者、AI动作识别研究人员或健康内容创作者,用于构建训练库或训练计算机视觉模型。

usestrix/strix

Python · ★ 28,224 · 🍴 3,130 · 📈 515 stars today

Open-source AI penetration testing tool to find and fix your app’s vulnerabilities.

中文介绍 一款开源的AI驱动渗透测试工具,旨在自动化发现并修复应用程序中的安全漏洞。结合人工智能技术进行漏洞扫描与评估,适合安全工程师、DevSecOps团队及独立开发者使用,帮助在开发周期中快速定位并解决代码安全隐患,提升应用整体安全性。

msitarzewski/agency-agents

Shell · ★ 121,099 · 🍴 19,782 · 📈 1,791 stars today

A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy injectors to reality checkers. Each agent is a specialized expert with personality, processes, and proven deliverables.

中文介绍 一套模拟完整AI代理机构的专家级Agent集合,涵盖前端开发、社区运营等多种角色。每个Agent都具备独立性格、标准化工作流程与专业技能。适合需要构建多智能体协作系统、自动化工作流或探索AI角色扮演与任务分配的开发者及AI创业者使用。

altic-dev/FluidVoice

Swift · ★ 4,971 · 🍴 304 · 📈 588 stars today

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model - Local Wispr Flow alternative. One ⭐ takes us a long way :)) Windows, iOS and Linux coming soon.

中文介绍 一款专为macOS打造的高性能本地语音输入应用,采用端侧语音转文字(STT)技术与自定义训练的AI增强模型,作为Wispr Flow的开源本地替代方案。注重隐私与低延迟,适合需要高效语音记录、代码口述或文案撰写的macOS用户及开发者。

diegosouzapw/OmniRoute

TypeScript · ★ 8,649 · 🍴 1,404 · 📈 387 stars today

Never stop coding. Free AI gateway: one endpoint, 231+ providers (50+ free), connect Claude Code, Codex, Cursor, Cline & Copilot to FREE Claude/GPT/Gemini. RTK+Caveman stacked compression saves 15-95% tokens, smart auto-fallback, MCP/A2A, multimodal APIs, Desktop/PWA.

中文介绍 一个免费AI API网关,提供单一端点接入超200家模型提供商。支持将Claude Code、Cursor等编程助手无缝连接至免费大模型资源。内置RTK与Caveman堆叠压缩技术,可大幅节省15%至95%的Token消耗,适合重度AI辅助编程开发者控制成本。

browser-use/video-use

Python · ★ 12,671 · 🍴 1,609 · 📈 721 stars today

Edit videos with coding agents

中文介绍 一款利用编程Agent进行视频剪辑的创新工具。通过将视频编辑任务转化为代码指令,让AI代理自动执行视频处理操作。适合需要批量处理视频、自动化剪辑工作流的开发者或内容创作者,降低传统视频编辑软件的学习成本,实现通过自然语言或代码驱动的视频生产。

xbtlin/ai-berkshire

Python · ★ 7,604 · 🍴 966 · 📈 969 stars today

AI 时代的伯克希尔:基于 Claude Code / Codex 的价值投资研究框架。巴菲特·芒格·段永平·李录四大师方法论 + 多Agent并行研究。| AI-era Berkshire: a value investing research framework built for Claude Code / Codex. 4 masters' methodologies + multi-agent adversarial analysis.

中文介绍 基于Claude Code与Codex构建的价值投资研究框架。深度融合巴菲特、芒格、段永平与李录四位大师的方法论,采用多Agent并行架构进行深度分析。适合个人投资者与金融AI开发者,用于自动化挖掘公司基本面数据、生成研报并辅助价值投资决策。

Mebus/cupp

Python · ★ 6,118 · 🍴 2,049 · 📈 32 stars today

Common User Passwords Profiler (CUPP)

中文介绍 一款常见的用户密码画像分析工具(CUPP)。通过收集和分析目标用户的公开社交媒体及个人信息,利用字典生成技术定制化预测并生成高概率的密码列表。适合安全研究人员、渗透测试工程师用于密码爆破测试,或帮助普通用户评估自身密码的安全强度。

ripienaar/free-for-dev

HTML · ★ 127,365 · 🍴 13,314 · 📈 742 stars today

A list of SaaS, PaaS and IaaS offerings that have free tiers of interest to devops and infradev

中文介绍 一份详尽的开发者免费资源清单,专门收录提供免费层级的SaaS、PaaS和IaaS服务。涵盖数据库、托管、CI/CD、监控等多种基础设施与开发工具。适合DevOps工程师、独立开发者及初创团队在搭建项目初期寻找零成本或低成本的技术栈解决方案。

google/agents-cli

Python · ★ 4,257 · 🍴 461 · 📈 445 stars today

The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.

中文介绍 Google推出的命令行工具与技能集,旨在将任意编程助手转化为AI Agent开发专家。支持在Google Cloud平台上高效创建、评估和部署智能体。适合云原生开发者与AI工程师,用于简化Agent在GCP环境中的全生命周期管理与自动化部署流程。

roboflow/supervision

Python · ★ 45,949 · 🍴 4,077 · 📈 309 stars today

We write your reusable computer vision tools. 💜

中文介绍 Roboflow开源的可复用计算机视觉工具库,提供一系列用于图像处理、目标检测、模型推理及结果可视化的实用函数。封装了复杂的CV底层逻辑,适合AI算法工程师、计算机视觉开发者及数据科学家,用于快速构建、调试和部署视觉AI应用,大幅提升模型开发效率。

ogulcancelik/herdr

Rust · ★ 9,069 · 🍴 544 · 📈 486 stars today

agent multiplexer that lives in your terminal.

中文介绍 一款运行在终端环境中的AI Agent多路复用器。允许开发者在同一个命令行界面中同时管理、监控和调度多个AI代理任务。适合重度依赖终端工作流的极客开发者与运维人员,用于并行执行代码生成、系统监控或自动化脚本任务,提升多Agent协同操作的效率。

simplex-chat/simplex-chat

Haskell · ★ 17,369 · 🍴 1,011 · 📈 1,235 stars today

SimpleX - the first messaging network operating without user identifiers of any kind - 100% private by design! iOS, Android and desktop apps 📱!

中文介绍 首个完全摒弃用户标识符的隐私通信网络,从底层设计上实现100%数据隐私保护。支持iOS、Android及桌面端,通过去中心化路由和端到端加密防止元数据泄露。适合注重个人隐私的普通用户、记者、活动家及需要高安全级别通信的企业团队使用。

CoreBunch/Instatic

TypeScript · ★ 1,572 · 🍴 138 · 📈 351 stars today

Instatic is a modern self-hosted visual CMS - get it running in 1 minute

中文介绍 一款现代化的自托管可视化内容管理系统(CMS),主打极简部署,仅需一分钟即可启动运行。提供直观的图形化界面来管理网站或应用内容,无需复杂配置。适合独立开发者、小型团队及内容创作者,用于快速搭建个人博客、企业官网或轻量级内容展示平台。

microsoft/AI-For-Beginners

Jupyter Notebook · ★ 49,467 · 🍴 10,171 · 📈 252 stars today

12 Weeks, 24 Lessons, AI for All!

中文介绍 微软官方推出的AI入门系统课程,包含12周共24节结构化课时,涵盖人工智能基础理论与核心实践。内容面向零基础学习者,提供丰富的代码示例与动手实验。适合在校学生、跨行转型者及希望系统建立AI知识体系的初学者,快速掌握机器学习与深度学习基础。

facebook/astryx

TypeScript · ★ 1,845 · 🍴 100 · 📈 364 stars today

An open source design system that's fully customizable and agent ready

中文介绍 Meta开源的全定制且支持AI Agent的设计系统。提供高度模块化的UI组件与交互规范,允许开发者深度自定义主题,并原生兼容AI代理的自动化生成与调用。适合前端工程师与AI应用开发者,用于快速构建现代化且对AI友好的Web应用界面。

HKUDS/Vibe-Trading

Python · ★ 15,856 · 🍴 2,740 · 📈 721 stars today

"Vibe-Trading: Your Personal Trading Agent"

中文介绍 一款专注于金融交易领域的个人AI代理系统。通过整合市场数据分析与量化策略,为用户提供自动化的交易决策支持与执行辅助。适合个人投资者、量化交易爱好者及金融工程学生,用于构建个性化交易策略、回测历史数据以及在复杂市场环境中进行智能化辅助交易。

obra/superpowers

Shell · ★ 242,631 · 🍴 21,527 · 📈 890 stars today

An agentic skills framework & software development methodology that works.

中文介绍 一套实用的AI Agent技能框架与配套软件开发方法论。旨在规范智能体在软件工程中的行为模式,提供可复用的技能模块与最佳实践指南。适合AI应用开发者、架构师及技术团队,用于构建更稳定、可控的Agent系统,提升AI辅助编程与自动化开发的实际落地效果。

Robbyant/lingbot-map

Python · ★ 8,898 · 🍴 861 · 📈 189 stars today

A feed-forward 3D foundation model for reconstructing scenes from streaming data

中文介绍 用于从流式数据中重建3D场景的前馈基础模型。突破传统3D重建的耗时限制,支持实时或近实时的场景几何与纹理生成。适合计算机视觉研究员、机器人开发者及AR/VR工程师,用于自动驾驶环境感知、机器人实时建图及沉浸式空间计算等前沿应用场景。

ORACLE: Official AI Agents Trade on Polymarket

@OracleLimited · 37.6K 粉丝 · 202.9K 阅 · 2.8K 赞 · 562 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 探讨预测市场 Polymarket 的 AI 交易现状。指出自主 AI 代理已成为该市场最有效策略之一,目前超 30% 的交易活动来自算法和 AI 驱动的钱包。揭示了 AI 代理在加密预测市场的高渗透率与实战价值。

ORACLE: Official AI Agents Trade on Polymarket

@OracleAiTrading · 34.1K 粉丝 · 176.1K 阅 · 2.7K 赞 · 567 转

In 2026, autonomous AI agents have become one of the most effective strategies on prediction markets. Over 30% of all activity on Polymarket now comes from algorithmic and AI-powered wallets. We

中文介绍 探讨预测市场 Polymarket 的 AI 交易现状。指出自主 AI 代理已成为该市场最有效策略之一,目前超 30% 的交易活动来自算法和 AI 驱动的钱包。揭示了 AI 代理在加密预测市场的高渗透率与实战价值。

How to Build a $10,000-Level Website With Animations in Claude Code

@monokern · 1.9K 粉丝 · 175.8K 阅 · 546 赞 · 49 转

Agencies charge $5,000 for a portfolio site that looks this good I built mine in 2 hours. Here's exactly how This is the real walkthrough - not a generic template guide I'm using my own portfolio as

中文介绍 分享用 Claude Code 搭建高质感个人网站的教程。博主仅用 2 小时完成,效果媲美外包 5000 美元作品。提供非模板化的真实工作流拆解,以自身网站为例,展示 AI 编程在复杂前端动效开发中的提效潜力。

How To Become An AI Engineer in 2026 (Without a CS Degree)

@cyrilXBT · 187.0K 粉丝 · 171.8K 阅 · 505 赞 · 91 转

There is a sentence sitting on almost every AI engineering job posting that stops people before they even apply. Bachelor's degree in Computer Science required. Most people read that line, close the

中文介绍 探讨无计算机学位如何成为 AI 工程师。针对招聘中常见的「要求计算机本科学位」门槛,博主分享打破学历限制的路径与策略,为转行者或非科班开发者提供 2026 年 AI 工程岗位的求职与技能构建指南。

Two kinds of scheduled work in Codex

@jxnlco · 113.3K 粉丝 · 54.2K 阅 · 501 赞 · 29 转

You want Codex to do something later, or keep checking something until it changes. That sounds like one feature. It is actually two different kinds of work, and the difference is simple: Scheduled

中文介绍 解析 Codex 的两种定时任务机制。指出让 Codex 延后执行或持续轮询状态变化看似单一功能,实则是两种不同工作模式。清晰界定定时执行与持续检查的区别,帮助开发者更精准地设计 AI 代理的异步工作流。

How LLM Inference Works, Clearly Explained.

@_avichawla · 71.1K 粉丝 · 39.8K 阅 · 501 赞 · 67 转

Every generate() call to an LLM runs two distinct computational phases on the same GPU: prefill (processing the prompt) is compute-bound while decode (generating tokens one at a time) is memory-bound.

中文介绍 科普大模型推理底层机制。指出每次 generate() 调用在 GPU 包含两阶段:处理 prompt 的 prefill 受限于算力,逐字生成 token 的 decode 受限于内存带宽。为优化推理性能提供清晰理论视角。

Build the machine economy with Unicity

@unicity_labs · 126.3K 粉丝 · 5.2K 阅 · 531 赞 · 200 转

Unicity is building The Secure Compute Platform for Autonomous AI. Identity, execution, governance, and payments - rebuilt for machines, with no human in the loop. The internet is being rebuilt for

中文介绍 介绍 Unicity 构建的自主 AI 安全计算平台。该平台专为机器经济设计,将身份验证、执行、治理和支付系统重构为无人类干预的自动化架构。展示了面向未来 AI 代理自主交互的基础设施演进方向。

Forward Deployed Engineers and the future of software engineering

Sierra's Natalie Meurer on why product engineers and forward deployed engineers are starting to converge.

中文介绍 Sierra公司Natalie Meurer指出,产品工程师与前线部署工程师的职责正逐渐融合,探讨了软件工程未来的发展趋势及两种角色的演变。

Ahmad Osman on why local AI is catching up

After two packed AIEWF workshops, Ahmad Osman makes the case that local AI is catching up fast — from laptops and phones to enterprise-grade infrastructure.

中文介绍 Ahmad Osman在AIEWF研讨会上表示,本地AI发展迅速,正从笔记本电脑和智能手机快速扩展至企业级基础设施,逐步缩小与云端AI的差距。

Claude Science is Anthropic’s newest flagship product

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way that Claude Code supports software engineering. Like Claude Code, Claude Science can autonomously c

中文介绍 Anthropic在面向制药高管和生物技术研究人员的活动中发布旗舰产品Claude Science。该产品旨在为科学研究提供支持,其定位类似于支持软件工程的Claude Code。

Why Specialization Is Inevitable

中文介绍 Dharma AI探讨了AI模型专业化的必然趋势,分析了在复杂应用场景下,专注特定领域的专业化模型相较于通用模型的优势与发展前景。

Agriculture is ready for AI, but its data isn’t

Artificial intelligence is transforming what is possible in agriculture, but industry leaders should be wary of investing in AI without first laying the groundwork. The use cases are promising, especially for an industry navigating volatile fertilizer costs, unpredictable weather, and margins that l

中文介绍 人工智能正改变农业前景,尤其在应对化肥成本波动和极端天气方面潜力巨大。但行业专家指出,在投资AI前需先完善数据基础设施,当前农业数据准备仍显不足。

How ChatGPT adoption has expanded

New OpenAI Signals data shows how ChatGPT adoption is growing globally, with users increasing usage, exploring more capabilities, and driving growth across regions and languages.

中文介绍 OpenAI最新Signals数据显示,ChatGPT在全球范围内的采用率持续增长。用户不仅增加了使用频率,还探索了更多功能,推动了该产品在不同地区和语言市场的增长。

[AINews] not much happened today

a quiet day before the storm.

中文介绍 今日人工智能领域动态相对平静,暂无重大新闻或产品发布,业内视此为重大行业变革或技术风暴来临前的短暂休整期。

Introducing GeneBench-Pro

Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.

中文介绍 OpenAI推出GeneBench-Pro新基准测试,利用复杂的真实世界数据集,专门用于评估人工智能在基因组学、生物学及科学研究领域的性能表现。

Core dump epidemiology: fixing an 18-year-old bug

OpenAI engineers used large-scale core dump analysis to debug rare infrastructure crashes, uncovering both a hardware fault and a long-standing software bug.

中文介绍 OpenAI工程师通过大规模核心转储分析,排查罕见的基础设施崩溃问题,成功发现并修复了一处硬件故障以及一个存在长达18年的软件漏洞。

Inside Genebench-Pro

中文介绍 OpenAI发布GeneBench-Pro案例研究,深入解析该基准测试的设计原理、评估方法及其在基因组学和生物学研究中的具体应用与测试结果。

Robocalls: A Worldwide or US-only Problem? Analyzing Spam and Fraud in International Phone Calls

第一作者: Kemal Altwlkany · 方向: 网络安全

Abstract:Unsolicited automated phone calls (robocalls) are a serious threat: in the US alone, these calls resulted in reported losses of 1.1$ billion during 2025. Phishing and spoofing consistently rank among the most reported crimes within the FBI's Internet Crime Complaint Center, with phone call scams having the highest reported median loss. Combating robocalls is difficult due to many legal and practical constraints: robocalls often encompass multiple legal jurisdictions of different countries/states, the large volume of robocalls, their multilingual nature, the lack of publicly available data, privacy concerns with obtaining data, etc. We present a study of international robocalls, aggregating robocall reports from countries across all inhabited continents and contribute by providing new findings on international robocalls from 65 different countries. We also present the first...

论文介绍 本文针对自动语音电话带来的全球性网络欺诈威胁展开研究。由于跨国管辖、数据隐私及多语言等限制,此类攻击难以打击。研究聚合了全球65个国家的垃圾电话报告数据,系统分析了国际自动语音电话的分布与特征,为跨国网络欺诈的治理与防御提供了重要的数据支撑与新视角。

Exploring Side-Channel Protections in Hardware Implementations of PQC ML-KEM Verification

第一作者: Davis Ranney · 方向: 密码学协议

Abstract:As ML-KEM is adopted as a post-quantum cryptographic standard, resilience against physical side-channel attacks has become essential. Among the constituent steps, the decapsulation Fujisaki-Okamoto (FO) verification is particularly vulnerable to side-channel power and electromagnetic (EM) analysis. In this work, we focus on common FPGA-based implementations and examine their side-channel vulnerabilities, and compare them with those of microcontroller implementations. Three verification implementations, unprotected, hash-based (first-order), and higher-order masked, are evaluated for side-channel security on both a microcontroller and an FPGA. While FPGAs offer higher speed and parallelism, they often exhibit stronger side-channel leakage, especially in high bandwidth configurations. The higher-order masked designs still leak information about the underlying data due to...

论文介绍 本文研究后量子密码标准ML-KEM在硬件实现中的侧信道防护。针对其解封装验证易受功耗和电磁分析攻击的弱点,对比了微控制器与FPGA平台上的三种验证实现。结果表明FPGA侧信道泄漏更显著,且高阶掩码设计仍存在信息泄漏风险,为后量子密码硬件安全设计提供了重要参考。

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

第一作者: Seyed Bagher Hashemi Natanzi · 方向: 软件安全

Abstract:Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, coding environments, robotic systems, security-operation workflows, and autonomous agents that can read private data, call tools, write files, execute code, and act across organizational boundaries. This shift changes the security problem: risks do not arise from the model weights alone, but from the full lifecycle and application stack through which data, prompts, model outputs, tools, memories, and user authority interact. This paper systematizes the literature on vulnerabilities in large language model systems through a lifecycle and application-stack lens. We organize attacks across eight stages: data collection, pretraining, post-training alignment, model packaging and supply chain, retrieval and memory, prompting and inference...

论文介绍 随着大语言模型深度融入各类应用系统,其安全风险已从单一模型权重扩展至全生命周期与应用栈。本文从生命周期和应用栈视角,系统梳理了大语言模型系统的漏洞文献,将攻击划分为数据收集、预训练、模型供应链、检索记忆及推理等八个阶段,为全面理解和防御大模型系统级安全风险提供了系统性框架。

Hybrid Topological Data Analysis and LSTM Networks for Enhanced Network Intrusion Detection Using CIC-IDS2017 Dataset

第一作者: Amar Jeet · 方向: AI 安全

Abstract:Network intrusion detection systems (NIDS) are crucial in cybersecurity infrastructure, needing advanced techniques to detect hostile activity in network traffic. This research introduces a hybrid approach that combines Topological Data Analysis (TDA) with Long Short-Term Memory (LSTM) networks to improve anomaly detection in network security. Our multi-layered design combines TDA's persistent homology with LSTM networks to capture topological characteristics of network traffic patterns and simulate temporal sequences. We assessed our methodology using the CIC-IDS2017 dataset, which includes over 2.8 million labelled flows, 77 network variables, and 14 attack categories that reflect modern threat landscapes such as DDoS, brute force, web attacks, penetration, and botnet activities. Integrating Betti curves and persistence diagrams with deep learning architectures enhances...

论文介绍 针对网络入侵检测中复杂流量模式的识别难题,本文提出一种结合拓扑数据分析与长短期记忆网络的混合方法。该模型利用持久同调捕获网络流量的拓扑特征,并结合长短期记忆网络处理时间序列,在包含多种现代攻击类型的CIC-IDS2017数据集上进行了评估,有效提升了网络异常检测的准确性与鲁棒性。

Digital signature schemes based on code equivalence and syndrome decoding from restricted errors

第一作者: Sarah Arpin · 方向: 密码学协议

Abstract:Digital signature schemes are an important cryptographic tool to ensure data authenticity and integrity in many applications that must be resilient to attacks, including those facilitated by quantum computers. We consider the two digital signature schemes based on error-correcting codes that are second-round candidates in NIST's call for Additional Signature Schemes, which is part of the Post-Quantum Cryptography Standardization Process. Specifically, we provide an overview of the Codes and Restricted Objects Signature Scheme (CROSS) and the Linear Equivalence Signature Scheme (LESS). We describe their underlying problems of syndrome decoding from restricted errors and code equivalence. We review sigma protocols and how they can be transformed into digital signature schemes via the Fiat-Shamir transform. Finally, we explain how this procedure yields code-based digital...

论文介绍 本文综述了基于纠错码的后量子数字签名方案,重点分析NIST额外签名征集的第二轮候选者CROSS与LESS。研究详细阐述了受限错误综合征解码与码等价性等底层数学问题,回顾了Sigma协议及其通过Fiat-Shamir变换转化为签名的过程,为理解抗量子数字签名设计提供了系统性参考。

Comparative Analysis of Machine Learning based Intrusion Detection in Realistic IoT Networks

第一作者: Rana Alharbi · 方向: AI 安全

Abstract:The Internet of Things (IoT) is rapidly growing and expanding into various sectors, such as healthcare, transportation, smart homes, and more. Despite the benefits of using IoT devices, they present several challenges. Given the significant role these devices play in our lives, it is crucial to address issues related to their security and privacy. These devices are limited in resources, which complicates their security and the protection of the data that they manage. The paper aims to examine intrusion detection systems using the Gotham2025 dataset, generated through the Gotham testbed, which consists of 78 emulated IoT devices utilising various protocols, including MQTT, CoAP, and RTSP, to assist in safeguarding IoT networks from attacks. We conduct a comparative analysis between five machine learning algorithms, including Random Forest, XGBoost, Logistic Regression, Naive...

论文介绍 针对物联网设备资源受限导致的安全防护难题,本文基于包含多种真实协议与模拟设备的Gotham2025数据集,对物联网入侵检测展开研究。研究对比分析了随机森林、XGBoost等五种主流机器学习算法在识别网络攻击中的表现,评估了各算法在复杂物联网环境下的检测效能,为轻量级安全防御提供了实验依据。

CVE-TTP KG: Knowledge Graph Linking Software Vulnerabilities to Attack Behaviors

第一作者: Basant Agarwal · 方向: 软件安全

Abstract:In the evolving threat landscape, adversaries exploit software vulnerabilities to launch sophisticated attacks, challenging traditional defenses. Although databases like CVE and NVD provide detailed technical information, they often lack links to attacker behaviors such as tactics and techniques, limiting effective threat interpretation and response. This work bridges this gap by connecting vulnerabilities with behavioral patterns from the MITRE ATT&CK framework. We construct a CVE-TTP Knowledge Graph that links CVEs to tactics and techniques using classification and relation extraction. Transformer-based models are developed for behavior identification, with CySecBERT achieving macro F1-scores of 87.71% (techniques) and 96.16% (tactics). Also, we created an annotated dataset with 24,820 entities and 43,608 relations for entity and relation extraction. The pipeline-based...

论文介绍 针对传统漏洞数据库缺乏攻击行为关联的问题,本文构建了连接软件漏洞与攻击行为的CVE-TTP知识图谱。研究利用分类与关系提取技术,将CVE与MITRE ATT&CK框架的战术和技术相链接,并开发基于Transformer的识别模型。该工作弥合了漏洞信息与攻击行为间的鸿沟,提升了威胁分析能力。

A forgery attack on the Block.co blockchain-based digital credential certification system

第一作者: Giacomo Zonneveld · 方向: 系统安全

Abstract:Certification of digital documents, such as academic credentials, seems a particularly suitable application for the use of blockchain and distributed ledger technologies. Indeed, these technologies enable decentralized certification systems that rely on the immutability and persistence of their distributed ledgers. However, in the absence of a central trusted authority, it is not easy to guarantee the authenticity of the connection between the real identity of an academic institution and the digital identity of the certificate issuer. In this paper, we demonstrate that one of such systems, known as this http URL, has a vulnerability that allows the production of forged certificates that are recognized as valid by the system. Since this is an inherent limitation of the approach used for blockchain-based certification, our attack is likely to be extendable to other systems...

论文介绍 本文针对基于区块链的数字凭证认证系统展开安全分析。研究指出,在缺乏中央可信机构时,系统难以保证机构真实身份与发行者数字身份绑定的安全性。通过揭示Block.co系统中允许生成有效伪造证书的漏洞,本文证明了此类去中心化认证方法的固有局限,并警示该攻击手法可能扩展至其他类似系统。

EnclaveX: End-to-End Confidential AI with CPU/GPU TEEs

第一作者: Robert Schambach · 方向: 软件安全

Abstract:Large Language Models (LLMs) have rapidly proliferated, driving widespread adoption of AI applications. Most deployments rely on centralized infrastructures such as Microsoft Azure, Google Cloud, or AWS, requiring users to share sensitive data and training or fine-tuning code. This dependence raises significant security and privacy concerns, as cloud providers must be trusted to ensure confidentiality and integrity. Trusted Execution Environments (TEEs) e.g., Intel SGX/TDX, AMD SEV-SNP, and ARM CCA have been introduced to mitigate these risks. More recently, NVIDIA has developed GPU TEEs (e.g., H100/H200), yet comprehensive evaluations of end-to-end workflows that integrate CPU and GPU TEEs remain limited. Critical aspects, including performance overhead, remote attestation, and security guarantees for AI/LLM applications, have not been sufficiently studied. This paper...

论文介绍 针对大语言模型在云端部署面临的隐私与安全风险,本文研究了集成CPU与GPU可信执行环境(TEE)的端到端机密AI工作流。研究评估了该架构在性能开销、远程证明及安全保证方面的表现,为构建安全可靠的云端AI基础设施提供了系统性分析,有助于推动机密计算在AI领域的实际应用。

Witness Complexity of Short Descriptions: A Cryptographic Perspective

第一作者: Fabio F.G. Buono · 方向: 密码学协议

Abstract:In cryptographic practice, where protocols impose strict time bounds, implementations demand predictable resource usage, and real-world systems require immediate verification for security and usability, a short key or certificate is useful only if it can be expanded or verified within a bounded time; otherwise a compact representation that requires superpolynomial work to expand offers no operational guarantee within a bounded-time protocol. This paper formalises that gap by introducing \emph{witness complexity} \(\gam(x)\), the minimum running time over near-shortest descriptions of a string on a universal Turing machine. \(\gam\) differs from Shannon entropy and Kolmogorov complexity \(\KC\): low \(\KC\) can coexist with high \(\gam\). We prove invariance up to polynomial factors; a conditional separation (assuming \(\PneqNP\)). An unconditional lower bound from...

论文介绍 针对密码学协议中短密钥需在有限时间内验证的需求,本文提出“见证复杂度”概念,定义为图灵机上展开字符串最短描述的最小运行时间。研究证明该复杂度与香农熵和柯尔莫哥洛夫复杂度存在本质差异,并给出多项式不变性及计算复杂性分离的理论结果,为有界时间协议的安全分析提供了新视角。

Automated High-Precision Extraction and Forensic Verification of Data-Bearing Vector Figures

第一作者: Bowen Sun · 方向: 安全研究

Abstract:The quantitative record of science and engineering is increasingly carried by figures rather than text or tables, and a reader who needs the underlying numbers must usually re-digitize them by hand: slowly, imprecisely, and with no way to prove the result is faithful. Yet when a figure is stored as vector graphics, its data are not approximated by the picture but encoded in it: the renderer writes each marker and vertex at a printed precision that, for the dominant scientific toolchain, exceeds the data's own. We turn this into three contributions, one per shortcoming of hand digitization. First, a precision theory bounding how accurately data can be recovered for a given renderer and export format: bit-exact float32 for matplotlib markers, and a calibration-limited three to four significant figures end to end. Second, an automatic extractor that decodes a figure in one pass...

论文介绍 针对科学图表数据提取耗时且缺乏验证的问题,本文提出从矢量图形中自动高精度提取数据并进行取证验证的方法。研究建立了数据恢复的精度理论,开发了单次扫描自动提取器,并引入取证机制确保结果忠实度。该工作解决了手工数字化的痛点,为科学数据的自动化提取与验证提供了可靠工具。

CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs

第一作者: Zhengxing Li · 方向: AI 安全

Abstract:While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity of such methods for LLMs. First, the LLM input space is discrete, with up to 150,000^k k-tuples to consider with k the token-length of a putative trigger. Second, one must blacklist tokens typical of the putative target response (class) of an attack, as such tokens may give false detection signals. However, a comprehensive blacklist is not available, in general, for a given domain. We develop a highly effective detection and inversion framework for LLMs treated as classifiers. Central to our approach is class subspace orthogonalization (CSO), a novel plug-and-play paradigm for backdoor detection that serves two fundamental roles when applied to LLMs: i) it enhances both sensitivity and specificity of a baseline detector; ii) it provides a...

论文介绍 针对大语言模型后门检测中离散输入空间及目标响应误报的问题,本文提出类子空间正交化(CSO)框架。该方法通过正交化操作提升基线检测器的敏感性与特异性,并有效实现触发器反转。研究为LLM后门防御提供了即插即用的新范式,显著增强了模型在复杂场景下的安全性与检测准确性。

The Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills

第一作者: Hongliang Liu · 方向: 密码学协议

Abstract:AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from marketplaces and other agents. Governing them needs a stable notion of skill identity, yet cryptographic hashing is engineered to destroy the very similarity we need, as a one-character edit scrambles the digest. We present a compact, locality-sensitive fingerprint that embeds each component of a skill and projects it to bits with a multi-bank SimHash, giving a fixed 120-byte signature compared in constant time by Hamming distance. Our central claim is that keeping the fingerprint as a per-component triple (prompt, code, tools), rather than a single score, is what makes it useful: the triple recovers skill-family identity through paraphrase, renaming, refactoring, and controlled code translation when another component remains shared...

论文介绍 针对AI智能体运行时技能身份难以稳定识别的问题,本文提出一种基于组件分解的局部敏感指纹方法。研究利用多库SimHash生成固定长度签名,并将指纹保留为提示、代码和工具的三元组。该设计在组件发生重构或翻译时仍能准确恢复技能族身份,为智能体技能的安全治理与追踪提供了高效方案。

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming

第一作者: Yong Yang · 方向: 密码学协议

Abstract:The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context Protocol (MCP) ecosystem and the language models themselves, has outpaced the security tooling available to defend it. We present AI-Infra-Guard, an open-source framework that organizes AI red teaming around a single observation: the attack surface of an AI agent is stratified across layers (infrastructure, protocol/tool, agent behavior, and model), and no single detection paradigm fits all of them. The framework therefore matches a paradigm to each layer, from deterministic rule matching over 75+ AI components and 1{,}400+ vulnerability rules, through LLM-driven agentic auditing of MCP servers and agent-skill packages and multi-turn black-box agent red teaming, to a jailbreak harness with 26+ attack operators over sixteen datasets. To our knowledge it is the...

论文介绍 针对AI基础设施安全工具滞后于发展的问题,本文提出AI-Infra-Guard开源红队测试框架。该框架将智能体攻击面划分为基础设施、协议、行为和模型四层,并为每层匹配专属检测范式,涵盖规则匹配、智能体审计及越狱测试等。研究为AI系统的全方位安全评估提供了统一且可扩展的解决方案。

Probe Choice Changes Canary-Memorization Verdicts: Three Post-Hoc Disagreement Case Studies in a Text-Dominant LoRA-Tuned Autoregressive Testbed

第一作者: Zhichao Fan · 方向: 安全研究

Abstract:We audit a fixed prefix-window mean-NLL memorization probe (K=20) on a Qwen2.5-VL-7B canary testbed and report three post-hoc cases where it disagrees with full-span secret NLL or greedy exact-recall. C3 (false negative, window truncation): damage lands on hex tokens outside K=20; the probe stays flat while hit@1 drops. C4 (false positive, non-secret drift): the probe moves, but approximately 99% sits on non-secret preamble; the secret span and hit@1 are unchanged. C5 (ambiguous in-window drop): the probe falls on an undertrained baseline while full-span hex is positive and hit@1=0. Recommendation: report (i) full-span secret NLL, (ii) a span-localised decomposition, (iii) behavioural exact-recall at k>=4, and (iv) decoy probes before asserting secret-specificity. Evidence is on controlled canaries in one backbone; magnitudes are testbed-specific.

论文介绍 针对大语言模型记忆探针评估标准不一的问题,本文在LoRA微调测试床上审计了固定窗口记忆探针,揭示了其与全跨度秘密评估及精确召回之间的三种事后分歧案例。研究指出窗口截断和非秘密漂移会导致误判,并建议采用全跨度评估、局部化分解及诱饵探针等综合指标,以提升模型隐私审计的准确性。

Secure-CHG: A Comprehensive Framework for Robust and Fair Federated Learning via Hybrid Defense and Contribution-Aware Trust

第一作者: Guanming Che · 方向: 隐私保护

Abstract:Federated Learning (FL) is highly susceptible to stealthy backdoor attacks, which aim to force a model into predicting an attacker-chosen target class for inputs containing a specific trigger. However, existing statistical defenses primarily focus on the early stages of model convergence. In this paper, we identify a fundamental vulnerability termed ``Late-stage Failure.'' We demonstrate that as the global model converges, decaying gradient norms render malicious and benign updates morphologically indistinguishable. This vanishing statistical variance effectively blinds traditional defenses, enabling adaptive adversaries to remain dormant and subsequently hijack the training process. To overcome these constraints, we propose Secure-CHG, a hybrid framework that pivots the defense paradigm from superficial morphological detection toward intrinsic semantic contribution...

论文介绍 针对联邦学习中传统统计防御在模型收敛后期失效的问题,本文揭示了“后期失败”漏洞,即梯度范数衰减导致恶意与良性更新难以区分。为此提出Secure-CHG混合防御框架,将防御范式从表面形态检测转向基于内在语义贡献的信任评估。该研究有效提升了联邦学习在应对自适应后门攻击时的鲁棒性与公平性。

Certified Speculative Execution for Untrusted AI Agents

第一作者: Chenyu Zhou · 方向: AI 安全

Abstract:Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM forfeits the feasibility guarantee a trusted solver provides, while invoking the solver at every step forfeits the speed the AI offers. Certificate-Gated Prefix Acceptance (CGPA) closes this gap with a certified speculative-execution contract for untrusted AI agents: a trusted verifier rejects constraint-violating transitions exactly, a conformally calibrated value boundary gates the longest low-cost prefix within a per-segment regret budget, and the rest defers to the solver, so safety, regret, and speed decouple by construction. The contract drives every untrusted proposal source - adversarial drafters and six heterogeneous frozen LLMs (including a 12B model that violates constraints in 98% of...

论文介绍 针对硬约束序列决策系统中AI计算与可行性保证的矛盾,本文提出证书门控前缀接受机制,为不可信AI代理提供认证的推测执行契约。该方法通过可信验证器拦截违规转换,并利用共形校准的价值边界控制低成本前缀,从而在解耦安全性、遗憾与速度的同时,提升多步草稿的执行效率与系统可靠性。

Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models

第一作者: Arash Raftari · 方向: 网络安全

Abstract:Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. In this work, we study post hoc detoxification of backdoored LLMs in a practical setting where the defender has access to the poisoned model but does not wish to retrain the full network from scratch. We propose a mechanistically guided weight-space repair framework that first localizes modules involved in propagating trigger-induced behavior using activation patching and Fisher/K-FAC curvature analysis, and then applies targeted low-rank repair to only the most influential modules. We evaluate the method on poisoned variants of \texttt{Llama-3.2-1B-Instruct} with triggers inserted at the beginning, middle, and end of otherwise benign prompts. Results show that the proposed approach...

论文介绍 针对大语言模型后门攻击的威胁,本文提出一种机制引导的权重空间修复框架,用于事后解毒。该方法利用激活补丁与曲率分析,精准定位传播触发器行为的关键模块,并仅对这些高影响力模块实施定向低秩修复。此方案无需从头重训完整网络,即可在保留模型正常功能的同时有效消除后门行为。

AI-Generated PowerShell Malware: An Experimental Framework and Dataset

第一作者: Luciano Pianese · 方向: 软件安全

Abstract:Generative AI has emerged as a significant cybersecurity threat, with several recent attack campaigns leveraging LLMs to generate code for malicious purposes via scripting languages such as PowerShell. Consequently, for cybersecurity analysts, it is imperative to investigate the offensive capabilities of AI code generators. In this paper, we propose an experimental framework to assess LLM-generated PowerShell malware, which comprises a novel sandbox approach for dynamic analysis of AI-generated malware. Furthermore, we present a novel, manually curated dataset of real-world PowerShell malware, annotated in natural language to assist the training and evaluation of LLMs. Finally, this study evaluates permissive, open-weight LLMs adapted to PowerShell malware generation. Our results reveal a high degree of similarity between real malware and LLM-generated ones in terms of...

论文介绍 针对生成式AI被用于编写恶意代码的威胁,本文提出一个评估大语言模型生成PowerShell恶意软件的实验框架。该框架包含用于动态分析的新型沙箱环境,并构建了一个带自然语言注释的真实恶意软件数据集。研究还评估了多款开源模型生成恶意代码的能力,为网络安全分析师提供了重要的攻防研究基础。

Security--Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense

第一作者: Mitchell Hermon · 方向: AI 安全

Abstract:We identify a security-fidelity tradeoff in defending LLMs against indirect prompt injection: defenses resist injected instructions largely by suppressing untrusted text, which corrupts tasks that must preserve it, such as translation and document editing. Attack-success metrics cannot see this, because a model that ignores an injection and one that faithfully processes it as data score identically. We introduce SecFid, a benchmark built so that executing an injection, processing it as data, and ignoring it produce distinguishable outputs. This makes fidelity measurable and exposes a frontier: across 1,168 examples and 48 configurations, no model or defense achieves both objectives. The highest-fidelity model reaches 96.5% fidelity at 47.8% security, while the most secure defenses invert this, at 99.3% security but only 71.0%-73.9% fidelity. Even defenses with identical...

论文介绍 本文揭示了防御大语言模型间接提示注入时存在的安全与保真度权衡问题。现有防御主要通过抑制不可信文本抵御攻击,但这会破坏翻译等需保留原文的任务。为此,研究引入SecFid基准,使注入执行、数据处理与忽略操作产生可区分输出。实验表明,现有模型与防御均无法同时兼顾高安全性与高保真度。

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

第一作者: Peizhi Niu · 方向: 系统安全

Abstract:Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on system-level responsibilities -- installing packages, maintaining state, scheduling subtasks, and mediating I/O -- making security failures far more severe than in other agents. Yet existing benchmarks focus on model responses and tool calls, leaving cross-component failure modes largely unmeasured. We adopt a computer-system analogy: treating a Claw-like agent as an agentic computer system whose gateway runtime plays an OS-like mediation role, whose Skills resemble user-installed applications, and whose Plugins resemble loadable extensions with runtime privileges. Each component has a classical counterpart whose protection mechanisms -- refined over decades of cybersecurity research -- are absent on the agent side. From this...

论文介绍 本文从计算机系统视角分析具有持久系统级权限的常驻AI代理的安全问题。研究将此类代理类比为代理计算机系统,将其网关运行时、技能与插件分别映射为操作系统、用户应用与扩展程序。通过指出传统系统保护机制在AI代理侧的缺失,本文揭示了跨组件故障模式,为评估和构建更安全的代理系统提供新框架。

An AI-Based Solution for Secure Service Provisioning in IoT

第一作者: Marco Arazzi · 方向: 软件安全

Abstract:As the Internet of Things (IoT) continues its rapid expansion, the attack surface grows accordingly, with emerging threats targeting smart objects and their interactions. In this evolving landscape, securing service provisioning is crucial to ensure the proper functioning, security, and reliability of the IoT ecosystem. Service provisioning encompasses key tasks such as device registration, configuration, authentication, authorization, and software deployment, all of which are essential for seamless and secure IoT operations. In this paper, we present a comprehensive framework designed to select the most suitable smart objects to deliver a target service within a given IoT environment while also monitoring the behavior of the entities involved during the service provisioning phase. To achieve this, we employ a Deep Reinforcement Learning (DRL) approach in which an intelligent...

论文介绍 针对物联网攻击面不断扩大带来的安全挑战,本文提出一种基于人工智能的安全服务配置框架。该框架采用深度强化学习方法,旨在物联网环境中智能选择最合适的设备交付目标服务,并实时监控配置阶段各实体的行为。此方案有效提升了设备注册、认证与软件部署等关键环节的安全性和可靠性。

Semantic Leakage and Privacy Preservation in Relay-Assisted Semantic Communications

第一作者: Yalin E. Sagduyu · 方向: AI 安全

Abstract:Semantic communication (SemCom) has emerged as a promising paradigm in which the transmission of task-relevant information is prioritized over raw data, enabling efficient and robust communication under resource and channel constraints. In this paper, the privacy implications of relay-assisted SemCom systems are studied, where the intermediate relay node operates directly on learned latent representations. It is shown that the relay, even without access to source data, can reliably infer semantic meaning and reconstruct signals with performance comparable to that of the legitimate receiver, revealing a fundamental privacy vulnerability of semantic representations. To address this issue, an iterative adversarial training framework is proposed in which a strong, adaptively trained eavesdropper at the relay is explicitly accounted for. The proposed approach alternates between...

论文介绍 本文揭示了中继辅助语义通信系统中的语义泄露隐私漏洞,指出中继节点即使无源数据访问权限,也能通过潜在表示推断语义并重构信号。为解决此问题,研究提出一种迭代对抗训练框架,通过显式引入自适应训练的强窃听者模型,交替优化发送端与中继端的策略,从而在保障通信效率的同时有效保护语义隐私。

Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

第一作者: Jonas Schäfer · 方向: 安全研究

Abstract:This work presents Dual-Embedding Watermarking (DEW), a semantic watermarking scheme for large language models (LLMs) that leverages contextual and token-level embeddings to enhance robustness against paraphrasing and translation. DEW utilizes a signal-processing methodology, applying algebraic vector-space operations to \mbox{token and context embeddings to derive a watermark signal that degrades gracefully under semantic shifts. The method obfuscates the watermark by projecting embedding vectors through pseudo-random matrices seeded with a secret key. Relevant distributions derived from the underlying algebra are evaluated and employed for statistical testing and benchmarking of DEW. Experimental results across multiple LLMs indicate that DEW improves post-paraphrase detection while maintaining competitive text quality, and remains detectable after translation, even when...

论文介绍 针对大语言模型生成文本的版权与溯源问题,本文提出双嵌入水印方案。该方法结合上下文与词元级语义嵌入,利用代数向量空间操作提取水印信号,并通过伪随机矩阵投影进行混淆。实验表明,该方案在保持文本质量的同时,显著提升了对释义和翻译等语义偏移攻击的鲁棒性,实现了更可靠的文本版权保护。

No Prompt, No Leaks: A Robust Generative Steganography Framework via Prompt-Free Diffusion

第一作者: Jingwen Cai · 方向: 安全研究

Abstract:Generative image steganography synthesizes stego images directly from secret information to achieve inherent security advantages. Latent Diffusion Models (LDMs) have recently emerged as a fundamental image steganography framework that modulates secret latent representations with text prompts. Limited by the inflexibility of text prompts, these methods still struggle to generate high-quality stego images and accurately recover secret images. In this work, we propose a prompt-free diffusion image steganography framework that integrates style semantic priors to control more robust and reliable stego image generation. Specifically, a Cascaded Affine Coupling Module (CACM) establishes a bijective, deterministic mapping between a secret image and its latent representation. Then, style semantics are integrated into the diffusion process to control latent representation and ensure...

论文介绍 针对文本提示限制导致生成隐写图像质量低和秘密恢复不准的问题,本文提出一种无提示扩散图像隐写框架。该方法通过级联仿射耦合模块建立秘密图像与潜在表示的双射映射,并将风格语义先验集成到扩散过程中,实现更鲁棒可靠的隐写图像生成与精确的秘密恢复。

Probing Memorization of Tabular In-Context Learning

第一作者: Francesco Capano · 方向: 安全研究

Abstract:Large tabular models (LTMs), i.e., tabular foundation models leveraging in-context learning (ICL), achieve state-of-the-art performance on tabular tasks. While LLMs are known to unintentionally memorize training data, the memorization dynamics of LTMs remain largely unexplored. We investigate the potential for parametric memorization in tabular ICL. We introduce ICLMEM, a probing framework designed to separate context-based predictions from parametric memorization. Our zero-information multiple-choice context strips away valid contextual patterns to force the model to fall back on its parametric memory. Our controlled fine-tuning setup establishes membership ground truth and accounts for common pitfalls, e.g., distribution shift, feature contamination, base-rate fallacy, and the pre-trained base model acts as reference to calibrate for sample difficulty. Our controlled...

论文介绍 大型表格模型的上下文学习记忆动态尚未被充分探索。本文研究表格上下文学习中的参数记忆潜力,提出ICLMEM探测框架以分离上下文预测与参数记忆。通过引入零信息多选上下文剥离有效模式,迫使模型依赖参数记忆,并建立受控微调设置以确立成员身份基准,评估模型的记忆行为。

An Empirical Study of Security Calibration in Large Language Models for Code

第一作者: Mohammed Latif Siddiq · 方向: 软件安全

Abstract:Large Language Models (LLMs) are rapidly transforming software development, yet their use in security-critical contexts raises a key question: do models know when their generated code is insecure? This property, known as calibration, measures whether a model's confidence aligns with the true correctness of its outputs. We present the first large-scale empirical study of security calibration in LLM-generated code. We evaluate GPT-4o-mini, Gemini-2.0-Flash, and Qwen3-Coder-Next across multiple temperature settings on two complementary benchmarks: self-contained security tasks and multi-language repository-level contexts. Our results suggest that overconfidence is prevalent across the evaluated LLMs. Functional calibration is consistently worse than security calibration, suggesting that models estimate security outcomes more reliably than functional correctness, potentially...

论文介绍 本文首次对大语言模型生成代码的安全校准进行大规模实证研究,探讨模型是否知晓其生成代码的不安全性。研究评估了多款主流模型在安全任务与多语言仓库级上下文中的表现,发现模型普遍存在过度自信现象,且其安全校准表现优于功能校准,表明模型对安全结果的估计更为可靠。

Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations

第一作者: Halima Bouzidi · 方向: AI 安全

Abstract:Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based retrieval becoming the dominant trajectory-decoding paradigms. Despite their architectural diversity, current generative AD planners share a common inference pattern: a fixed set of candidate trajectories (anchors, vocabulary entries, or proposal queries) is scored by one or more learned heads conditioned on the Bird's-Eye-View (BEV) features, and the highest-scored candidate is returned as the final trajectory. Under this design, the scoring head is the only barrier between perception and the motion command, and its decision margins between competing candidates are often small. We introduce \textsc{Derail}, an adversarial framework that exploits this scoring-head attack surface. Evaluated on various generative planners...

论文介绍 当前生成式端到端自动驾驶规划器依赖评分头对候选轨迹打分以输出最终轨迹。本文指出评分头是感知与运动指令间的唯一屏障且决策裕度小,提出Derail对抗框架,通过施加安全违规的对抗性扰动劫持评分头,从而操纵规划器的轨迹决策,揭示了该类架构的潜在安全漏洞。

Revocable Learned State via Process Sidecars

第一作者: John Sweeney · 方向: 安全研究

Abstract:Language models are often adapted in stages: a public skill phase, a private memory phase, and a later safety phase that learns to refuse outputs tied to the remembered entities. Revoking the memory after the safety phase is not the same problem as subtracting the memory update: the later safety optimizer has transported the memory direction. We introduce process sidecars, a two-coefficient edit family $\hat{\theta}(\lambda,\gamma)=\theta_{\mathrm{AMS}}-\lambda\Delta_{\mathrm{M}}-\gamma\hat{R}_{\mathrm{S}\leftarrow\mathrm{M}}$, with $\hat{R}_{\mathrm{S}\leftarrow\mathrm{M}}=\hat{J}_{\mathrm{S},\varepsilon}(\Delta_{\mathrm{M}})-\Delta_{\mathrm{M}}$, where $\hat{J}_{\mathrm{S},\varepsilon}$ is a centered secant through the realized future AdamW safety-training process. The implementation uses $\varepsilon=1$ at the natural memory-edit scale; it reuses $\theta_{\mathrm{AMS}}$ as...

论文介绍 语言模型在分阶段适应后,撤销安全阶段记忆并非简单的参数相减。本文提出过程边车机制,引入双系数参数编辑族,通过计算安全训练过程的自然记忆编辑尺度,实现对模型已学习状态的可逆撤销。该方法为模型记忆与遗忘的动态管理提供了新的参数级编辑思路。

DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation

第一作者: Ziyu Shan · 方向: 机器人操作 · 来源: cs.RO

Abstract:Video-based embodied world models provide an appealing substrate for robotic manipulation by predicting future states, yet current approaches remain limited by a fundamental entanglement: accurately modeling dynamics typically requires low-level temporal reasoning, while producing high-resolution frames demands expansive visual synthesis according to high-level semantics. This entanglement results in slow inference speed for iterative planning or too coarse predictions to retain contact-rich details. To solve this dilemma, we present Disentangled Video Generation World Model (DVG-WM), an efficient framework that explicitly decomposes world modeling into dynamics learning and visual synthesis. Conditioned on an initial observation and a language instruction, our model first generates a plausible sequence of intermediate visual states to preview the physical interaction and...

论文介绍 针对基于视频的世界模型中动力学建模与高分辨率视觉合成相互纠缠导致推理缓慢或细节丢失的问题,本文提出解耦视频生成世界模型DVG-WM。该框架将世界建模显式分解为动力学学习与视觉合成,通过生成中间视觉状态预览物理交互,实现高效且保留接触细节的机器人操作规划。

Freeform Preference Learning for Robotic Manipulation

第一作者: Marcel Torne · 方向: 机器人操作 · 来源: cs.RO

Abstract:Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal. We introduce Freeform Preference Learning (FPL), a method for learning robot policies from freeform human preferences. Rather than asking annotators which of two trajectories is better overall, FPL lets them define natural-language preference axes, such as speed, safety, quality of placement, or carefulness, and provide pairwise preferences along each axis. These annotations are used to learn a language-conditioned reward model that maps a trajectory and preference label to an axis-specific reward. We use this model to train a reward-conditioned policy that optimizes across the multiple human-specified...

论文介绍 针对长视野机器人操作任务中奖励设计困难及二元偏好信号模糊的问题,本文提出自由形式偏好学习方法。该方法允许标注者通过自然语言定义速度、安全性等偏好维度并提供成对偏好,进而训练语言条件奖励模型与奖励条件策略,使机器人能够根据人类指定的多维偏好优化操作行为。

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments

第一作者: Xiaopeng Lin · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot data remains difficult, especially for high-DoF humanoids. Teleoperation provides controller-aligned supervision, while human egocentric videos capture diverse bimanual manipulation but do not directly provide executable robot actions. We introduce Human-as-Humanoid, a human-to-humanoid supervision framework that enables near-real-time human-centric action generation, making human demonstrations usable for high-DoF humanoid VLA training by jointly aligning the robot embodiment, the sensing setup, and the action-label interface. Built on PrimeU, a human-aligned 60-DoF upper-body humanoid, Human-as-Humanoid uses synchronized ego-exo videos to pair deployment-aligned egocentric observations with...

论文介绍 针对高自由度人形机器人缺乏高质量观察与动作监督数据的问题,本文提出Human-as-Humanoid框架。该方法通过构建人体对齐的机器人实体,利用同步的自我与外部视角视频,将人类演示转化为部署对齐的动作标签,实现零样本人形机器人视觉语言动作模型训练,有效扩展了机器人学习的数据来源。

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

第一作者: Arnav Balaji · 方向: 机器人操作 · 来源: cs.RO

Abstract:While robotic manipulation capabilities have advanced rapidly, physical safety remains a major barrier to deploying household robots: task success is insufficient if the robot damages itself or its surroundings. Simulation offers a harm-free alternative to costly and dangerous real-world training and evaluation, yet existing simulators lack general mechanisms to detect, quantify, and represent damage. To address this gap, we introduce OOPSIEVERSE, a unified simulation framework and benchmark for damage-aware household manipulation. OOPSIEVERSE provides damage as an explicit, physically-grounded, and taskagnostic signal by converting sources such as contact forces, temperature changes, and liquid interactions into corresponding mechanical, thermal or fluid damage. OOPSIEVERSE comprises two core elements: (1) DAMAGESIM, a simulator-agnostic framework for detecting and...

论文介绍 针对家庭机器人部署中的物理安全问题,本文提出OOPSIEVERSE,一个具备损坏感知能力的统一仿真框架与基准测试。该系统将接触力、温度变化等物理交互转化为显式的机械、热或流体损坏信号,弥补了现有仿真器缺乏通用损坏检测与量化机制的不足,为安全评估和无害训练提供了新范式。

Adapting Generalist Robot Policies with Semantic Reinforcement Learning

第一作者: Jagdeep Singh Bhatia · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Generalist robot policies learn a diverse repertoire of behaviors from large-scale pretraining. In principle, this makes them excellent priors for downstream adaptation via reinforcement learning (RL). In practice, however, standard RL methods leveraging this prior optimize directly over robot actions, requiring the base policy's action distribution to be close to that of a performant policy from the start. This assumption breaks down for complex or long-horizon tasks that fall outside the pretraining distribution. Our key insight is that, for sufficiently expressive generalist policies, language prompts are an effective alternative space for learning to solve such tasks: modulating language inputs elicits skills already within the policy's repertoire, which can be composed to solve tasks beyond its zero-shot capabilities. We propose Semantic Action Reinforcement Learning...

论文介绍 针对通用机器人策略在复杂任务中难以通过标准强化学习有效适应的问题,本文提出语义动作强化学习方法。该方法利用语言提示作为替代优化空间,通过调节语言输入激发并组合预训练策略库中的已有技能,从而解决超出零样本能力的长视野或复杂任务,提升了通用策略的下游适应效率。

LeCropFollow: Latent Space Planning for Navigation in Unstructured Crop Fields

第一作者: Felipe Tommaselli · 方向: 导航与运动 · 来源: cs.RO

Abstract:Unstructured navigational features, such as irregular planting or discontinuities, remain the primary failure mode for under-canopy agricultural robots. Existing geometric approaches often fail in these scenarios because they compress high-dimensional visual data into deterministic spatial references, effectively discarding the uncertainty and semantic context required to navigate ambiguous terrain. To address this, we present LeCropFollow, a visual navigation framework that bypasses explicit geometric modeling in favor of a learned latent representation. By integrating a self-supervised semantic heatmap extractor with TD-MPC2, a Model-Based Reinforcement Learning (MBRL) planner, our system optimizes trajectories directly within a latent manifold. The framework operates over the uncompressed heatmap signal, preserving the semantic context that geometric reductions discard. We...

论文介绍 针对冠层下农业机器人在非结构化农田中导航失败的问题,本文提出LeCropFollow视觉导航框架。该方法绕过显式几何建模,结合自监督语义热图提取器与基于模型的强化学习规划器,在保留语义上下文的潜在流形中直接优化轨迹,有效解决了传统几何方法丢弃不确定性与语义信息的缺陷。

Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing

第一作者: Jiale Fan · 方向: 导航与运动 · 来源: cs.RO

Abstract:Learning-based control has revolutionized dynamic locomotion, yet navigating unstructured terrain remains limited by a robot's incomplete awareness of imminent ground contact. While global perception systems such as LiDARs and depth cameras provide environmental context, they are frequently plagued by latencies, occlusions, and the high computational cost of dense geometric reconstruction. On the other hand, proprioceptive feedback is purely reactive, initiating corrections only after impact has occurred. This work explores embedding a minimal suite of low-cost, high-frequency infrared proximity sensors directly into the feet of a quadrupedal robot. These sensors provide "pre-contact" feedback that is robust to self-occlusions and significantly less computationally demanding than conventional vision-based pipelines. By integrating these localized signals into a reinforcement...

论文介绍 针对四足机器人在非结构化地形中导航受限于环境感知不足的问题,本文提出在机器人脚部嵌入低成本高频红外接近传感器。该方法提供鲁棒且计算高效的「接触前」局部反馈,克服了全局视觉系统的延迟与遮挡缺陷,并结合强化学习实现动态运动控制,提升了复杂地形下的自适应运动能力。

CoDex: Learning Compositional Dexterous Functional Manipulation without Demonstrations

第一作者: Bowen Jiang · 方向: 机器人操作 · 来源: cs.RO

Abstract:In this work, we study Compositional Dexterous Functional Object Manipulation (CD-FOM): tasks such as aiming and actuating a spray bottle on a plant or a glue gun on wood, which require both actuating an object's internal mechanism and controlling its pose to apply the object's function to the environment. These tasks pose significant challenges for robots due to the demanding integration of semantic understanding of the object's function, actuation mode, and application area with intricate physical dexterity to manage grasp stability, movement trajectory, and actuation. We introduce CoDex, a zero-demonstration framework that autonomously discovers CD-FOM manipulation strategies. CoDex uses vision-language models (VLMs) to infer semantic constraints from the task and scene. These constraints guide analytic constrained optimization to generate a short list of functional grasp...

论文介绍 针对机器人执行组合灵巧功能物体操作的挑战,本文提出CoDex零演示框架。该方法利用视觉语言模型推断任务与场景的语义约束,并引导解析约束优化自主生成功能抓取与操作策略。该框架无需人工演示即可实现物体内部机制激活与姿态控制的结合,提升了机器人对复杂功能物体的操作能力。

Improving path-tracking performance of an articulated tractor-trailer system using a non-linear kinematic model

第一作者: Marina Murillo · 方向: 具身智能 · 来源: cs.RO

Abstract:This paper presents a novel non-linear mathematical model of an articulated tractor-trailer system that can be used, in combination with receding horizon techniques, to improve the performance of path tracking tasks of articulated systems. Due to its dual steering mechanisms, this type of vehicle can be very useful in precision agriculture, particularly for seeding, spraying and harvesting in small fields. The articulated tractor-trailer system model was embedded within a non-linear model predictive controller and the trailer position was monitored. When the kinematic of the trailer was considered, the deviation of trailer's position was reduced substantially alongside not only straight paths but also in headland turns. Using the proposed mathematical model, we were able to control the trailer's position itself rather than the tractor's position. The Robot Operating System...

论文介绍 针对铰接式拖拉机-拖车系统在精准农业中的路径跟踪问题,本文提出一种新型非线性运动学数学模型。该模型嵌入非线性模型预测控制器中,直接监控并控制拖车位置而非拖拉机位置。实验表明,该方法在直道和地头转弯中均能大幅降低拖车位置偏差,显著提升了双转向车辆的路径跟踪性能。

Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models

第一作者: Lang Cao · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observations, and continuous control. However, most existing policies remain limited by behavior cloning or supervised fine-tuning (SFT) from fixed demonstrations, which provides limited opportunity to improve from the policy's own failures. In this paper, we present Z-1, a reinforcement learning (RL) post-training framework for flow-based VLA models. Built on top of $\pi_{0.5}$, Z-1 uses only publicly released RoboCasa demonstrations for SFT and then applies a task-wise Group Relative Policy Optimization (GRPO) strategy across $24$ standard RoboCasa tasks. To improve the efficiency and stability of online optimization, Z-1 combines shared-prefix rollout construction, tree-structured trajectory branching, completion-aware reward calibration, and...

论文介绍 针对视觉语言动作模型受限于监督微调而难以从失败中学习的问题,本文提出Z-1强化学习后训练框架。该方法在基础模型上采用任务级分组相对策略优化,并结合共享前缀轨迹构建与树结构分支等技术,显著提升了在线优化的效率与稳定性,使模型能够通过自我探索在复杂操作任务中持续改进。

RoboTacDex: A Dexterous Visual-Tactile-Action Dataset for Humanoid Manipulation

第一作者: Xinyi Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:In the field of robot learning, large-scale and diverse demonstration trajectories provide the fundamental basis for enhancing robotic manipulation ability. We introduce RoboTacDex, a large, multi-modal, and diverse dataset of dexterous manipulation behaviors performed with a humanoid robot. Built on the publicly accessible humanoid robot Unitree G1, RoboTacDex consists of 6k trajectories covering 19 tasks, 23 skills, and interactions with 22 objects. RoboTacDex provides comprehensive records including multi-view RGB and depth information, tactile feedback, and detailed semantic annotations. Furthermore, the dataset features a variety of relatively challenging tasks that can only be completed by dual arms and dexterous hands, aiming to mimic human-like operational logic and simulate real-world manipulation complexity. To ensure data collection quality, we develop an improved...

论文介绍 为提升人形机器人的灵巧操作能力,本文发布RoboTacDex大型多模态数据集。该数据集基于Unitree G1采集,包含6000条涵盖19项任务与23种技能的演示轨迹。数据提供多视角视觉与动作信息,并引入触觉反馈与语义注释,旨在模拟人类操作逻辑,推动复杂现实场景下的机器人学习研究。

Reinforcement Learning-Based Control for an Inline Skating Humanoid Robot

第一作者: Ethan Marot · 方向: 导航与运动 · 来源: cs.RO

Abstract:As humanoid robots become increasingly dynamic, coupling them with reinforcement learning offers a promising approach to solving the complex, underactuated mechanics of passive inline skating. Equipping a humanoid robot with passive inline skating wheels presents an opportunity to combine the versatile agility of humanoids with the high-speed, energy-efficient locomotion strategies utilized by human skaters. In this paper, we train and deploy a reinforcement learning control policy that enables novel locomotion strategies for a humanoid robot modified to equip consumer inline skates instead of conventional feet. Unlike previous work limited to quadrupedal robots or actively driven wheels, our system allows for precise 6-DoF control of the skates to execute dynamic, edge-driven propulsion strategies. Our skating strategies emerge entirely from our reward structure, without...

论文介绍 本文提出一种基于强化学习的控制策略,使配备消费级旱冰鞋的人形机器人实现动态轮滑运动。与以往依赖主动驱动轮或四足机器人的研究不同,该系统支持对旱冰鞋进行精确的六自由度控制,通过奖励结构设计实现基于边缘的动态推进策略,为人形机器人结合高能效与敏捷性提供了新方案。

UniTacVLA: Unified Tactile Understanding and Prediction in Vision Language Action Models

第一作者: Xidong Zhang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-language-action (VLA) models have achieved strong performance in many robotic manipulation tasks, yet remain limited in contact-rich dexterous manipulation. To overcome this limitation, recent vision-tactile-language-action (VTLA) methods incorporate tactile sensing into VLA models to provide direct contact information. However, they typically treat tactile signals as passive auxiliary inputs, making it difficult to model tactile semantics and future physical interactions. To this end, we propose a unified tactile learning framework for contact-rich manipulation that models tactile signals as dynamic interaction cues for both contact understanding and prediction. Specifically, we construct a unified tactile latent space and jointly model current tactile states and future contact changes through tactile chain-of-thought reasoning and coarse-to-fine future tactile...

论文介绍 针对视觉语言动作模型在接触丰富任务中的局限,本文提出UniTacVLA统一触觉学习框架。该方法将触觉信号建模为动态交互线索,构建统一的触觉潜在空间,并通过触觉思维链与由粗到细的预测机制,联合建模当前触觉状态与未来接触变化,有效提升了机器人在复杂物理交互中的理解与预测能力。

RCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile Generalization

第一作者: Jingbo He · 方向: 机器人操作 · 来源: cs.RO

Abstract:For robots manipulating open-world objects, tactile representations must generalize to unseen materials. We introduce RCT (Robotic Contact Tactile), a robot-collected touch-vision-language dataset with 29,279 tactile frames from full robot presses on 122 industrial reference materials in 7 categories, recorded with three DIGIT sensors at multiple contact positions. RCT preserves each press as a contact sequence, enabling held-out evaluation across materials, categories, sensors, contact positions, and contact sequences. Frames from one press are strongly correlated: frame-random splits can place near-duplicate observations of the same physical interaction in both training and test. With the encoder held fixed, removing contact-sequence overlap reduces tactile-to-text Recall@1 by 17.7 percentage points. When materials are additionally held out at training time, performance...

论文介绍 为使机器人触觉表征泛化至未见材料,本文引入RCT机器人接触触觉数据集。该数据集包含近三万个触觉帧,涵盖七类共122种工业材料的全面按压数据,并保留完整接触序列。研究指出,消除训练与测试间的接触序列重叠对准确评估触觉泛化能力至关重要,为开放世界物体操作提供了可靠的触觉数据基准。

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion

第一作者: Guanchen Lu · 方向: 导航与运动 · 来源: cs.RO

Abstract:Scalable reinforcement learning has popularized high-throughput sampling architectures, which significantly compresses the training time for off-policy methods in robotic locomotion. However, the rapid increase of data volume and update frequency undermines the stability of value-based methods and diminishes the plasticity of policy networks. To address these challenges, this work presents FastDSAC, a fast and high-performance variant of the Distributional Actor-Critic algorithm designed for parallel sampling scenarios. Specifically, we introduce a truncated Gaussian distribution to approximate the learned policy, which effectively excludes out-of-distribution actions that strain target value estimation while keeping necessary stochasticity for exploration. The proposed action constraint functions as an implicit regularization, which counteracts the plasticity loss typically...

论文介绍 针对大规模强化学习中数据量激增导致价值方法不稳定及策略网络可塑性下降的问题,本文提出FastDSAC算法。该方法引入截断高斯分布近似策略,排除分布外动作以稳定目标价值估计,同时保留探索所需的随机性。其动作约束机制作为隐式正则化,有效提升了可扩展人形机器人运动控制中的策略可塑性。

HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation

第一作者: Jaehwi Song · 方向: 机器人操作 · 来源: cs.RO

Abstract:Large-scale demonstration datasets have been central to recent progress in general-purpose robot policies. However, existing datasets are collected in human-absent settings, and policies trained on such data may perform tasks competently in isolation but fail to exhibit human-aware behaviors. To address this gap, we introduce HABIT, a large-scale robot demonstration dataset for human-present environments. We organize tasks into three roles capturing distinct modes of human-robot interaction: Collaborator, where human and robot jointly accomplish a task; Coworker, where they pursue separate tasks in a shared space; and Supervisor, where the human directs the robot. The dataset comprises over 10K episodes and over 160 hours across 60 tasks. Our experiments show that training on human-present data elicits human-aware behaviors that robot-only data fails to produce: spatiotemporal...

论文介绍 现有机器人策略数据集多缺乏人类在场场景,本文提出HABIT大规模人类感知行为与交互训练数据集。该数据集包含超一万集、160小时的演示数据,将任务划分为协作者、同事与监督者三种人机交互模式。研究表明,使用人类在场数据训练可赋予机器人时空感知等人类感知行为,弥补了孤立环境数据的不足。

Communication-Aware Robot Execution for Cloud Inference under Spatially Heterogeneous Connectivity

第一作者: Fengkai Liu · 方向: 具身智能 · 来源: cs.RO

Abstract:Cloud-hosted foundation models enable robots to use semantic reasoning beyond onboard computational limits. In this setting, the robot executes a currently available primitive generated by the cloud, and continued task progress requires the next cloud result before this primitive is exhausted. This execution becomes fragile under spatially heterogeneous connectivity, because the current primitive determines when the next result is needed, whereas the wireless environment determines where the next request can be submitted and where the response can be retrieved. Strategies that reduce latency or improve individual transmissions can shorten this dependency, but they do not determine a submission location that supports reliable upload and leaves a feasible opportunity for response retrieval. To address this problem, we introduce the request--response window, which characterizes...

论文介绍 针对云托管基础模型在空间异构连接下执行脆弱的问题,本文提出通信感知的机器人执行方法。研究引入请求-响应窗口概念,表征可靠的请求提交与响应检索位置。该方法确保机器人在当前动作原语耗尽前,能在合适空间位置获取云端推理结果,提升了云机器人系统的可靠性。

Robustness of Robotic Manipulation: Foundations and Frontiers

第一作者: Yifei Dong · 方向: 机器人操作 · 来源: cs.RO

Abstract:Humans and animals exhibit remarkable robustness in physical manipulation, yet robots remain far behind. Progress toward human-level manipulation robustness is hindered by the absence of a unified and systematic understanding: different subfields frame robustness in distinct ways, often leaving the concept ambiguous and limiting deeper analysis as well as communication across research areas. This paper presents a systematic study of manipulation robustness. We begin with a formal definition, characterizing robustness as the degree to which a manipulation system can achieve its goal in the presence of uncertainty and variation. Building on this definition, we introduce general formulations of manipulation robustness from probabilistic and control-theoretic perspectives. We then synthesize the guiding principles and concrete mechanisms of manipulation robustness across...

论文介绍 针对机器人物理操作鲁棒性缺乏统一理解的问题,本文展开系统性研究。文章首先给出形式化定义,将鲁棒性表征为系统在不确定性和变化下达成目标的程度。随后从概率与控制理论视角提出通用公式,并综合梳理跨子领域的指导原则与具体机制,为提升机器人操作可靠性奠定了理论基础。

ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

第一作者: Bokai Lin · 方向: 机器人操作 · 来源: cs.RO

Abstract:Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past experience and anticipated outcomes, effective policies should integrate past interactions with future predictions. However, existing visuomotor policies typically model either historical context or future dynamics in isolation, lacking a unified temporal representation of interaction dynamics. In this work, we introduce \textbf{ChronoFlow}, a temporally unified representation that captures \textbf{past, current, and future} interaction dynamics through sparse 3D keypoints of both objects and the gripper. Based on this representation, we propose \textbf{ChronoFlow-Policy}, a diffusion-based visuomotor policy that jointly learns ChronoFlow and action sequences through a co-training objective. Experiments...

论文介绍 现有视觉运动策略多孤立建模历史或未来动态,本文提出ChronoFlow-Policy以统一时间交互表示。该方法通过物体与夹爪的稀疏三维关键点,捕获过去、当前与未来的交互动态。基于此,模型采用扩散架构并通过协同目标联合学习交互流与动作序列,提升了机器人对复杂交互动态的理解与执行能力。

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation

第一作者: Fengnian Zhang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) models have made significant strides in embodied intelligence by integrating the powerful representations of pre-trained Vision-Language Models (VLMs). However, the massive parameter scale of VLAs imposes a heavy computational burden, and these models exhibit extreme sensitivity to parameter pruning. Current paradigms often treat the resulting performance degradation as inevitable, relying on fine-tuning or low-rank corrections to recover efficacy. We challenge this convention by questioning whether the removed parameters are truly redundant if VLA pruning necessitates performance recovery to be effective, or if this paradigm masks the indiscriminate pruning of critical parameters. We revisit parameter redundancy through the lens of VLM-to-VLA adaptation, first quantifying the spatial distribution of parameter divergence during adaptation to reveal...

论文介绍 本文针对视觉语言动作模型参数庞大且对剪枝敏感的问题,重新审视其参数冗余性。研究指出传统剪枝范式将性能下降视为必然并依赖微调的局限。作者从视觉语言模型到动作模型的适应视角出发,量化参数发散的空间分布,揭示被移除参数的真实冗余度,为模型的高效压缩与优化提供新见解。

Stage-Transition Dense Reward Modeling for Reinforcement Learning

第一作者: Yang Yang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations. This work proposes Stage-Transition Dense Reward (STDR), a visual reward-learning framework that converts unstructured expert videos into logically grounded dense rewards for training RL agents from scratch. STDR leverages semantic understanding to infer a task's stage structure from demonstrations, and delivers two complementary learning signals during online training: (i) stage-transition feedback that provides goal-directed reward, and (ii) within-stage progress feedback that supplies fine-grained guidance toward completing each stage. Furthermore, an out-of-distribution (OOD) detection mechanism and a grasping regulation module are integrated to...

论文介绍 针对长视野机器人操作受限于稀疏奖励的问题,本文提出阶段转换密集奖励框架。该方法利用语义理解从非结构化专家视频中推断任务阶段结构并转化为密集奖励。通过提供阶段转换的目标反馈与阶段内的细粒度进度指导,结合分布外检测机制,实现从零训练机器人策略,有效降低人工设计奖励的成本。

Verification-Gated Agentic Mission-State Governance for Intelligent Industrial Multi-Robot Systems

第一作者: Guoqin Tang · 方向: 多模态具身 · 来源: cs.RO

Abstract:Agentic artificial intelligence is increasingly used to decompose industrial tasks, propose robot actions, and adapt execution plans in dynamic cyber-physical environments. However, autonomous proposal generation alone does not guarantee that multi-robot industrial systems preserve task dependencies, resource ownership, safety holds, or repair boundaries during long-horizon execution. This paper introduces a verification-gated agentic mission-state governance framework for intelligent industrial multi-robot systems. The framework maintains two synchronized state objects: an evolving task forest for persistent hierarchy, delayed grounding, and repairable substructures; and a governed blackboard for online execution state, robot traces, resource locks, world beliefs, proposals, verification records, and scene-temporary constraints. From each forest--blackboard snapshot, a...

论文介绍 针对工业多机器人系统在长视野执行中难以保证任务依赖与资源安全的问题,本文提出验证门控的智能体任务状态治理框架。框架维护任务森林与黑板两个同步状态对象,以持久化层级结构并管理在线执行状态与资源锁。通过验证门控机制,确保多机器人在动态环境中的动作提案严格符合安全与任务约束。

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

第一作者: Dongyoon Hwang · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a...

论文介绍 针对分层视觉语言动作模型中二维轨迹指导缺乏深度导致底层控制几何扭曲的问题,本文提出三维轨迹指导框架。该方法为模型增加专用深度编码器与密集深度重建目标,使规划器直接输出度量可靠的三维航点序列。此三维轨迹指导直接融入底层策略,有效弥合高层规划与底层控制间的空间鸿沟。

Plan Right, Then Plan Tight: Symbolic RL for Efficient Embodied Reasoning

第一作者: Xiangli Shi · 方向: 具身智能 · 来源: cs.RO

Abstract:Embodied task planning asks an agent to turn a natural-language instruction into an executable sequence of actions in a physical scene, and is a building block for household, assistive, and service robots. Recent prompting-based and reinforcement-learning planners generate fluent action text but lack a cheap deterministic check that the produced plan is valid in the target world, while high-fidelity simulation is too slow to serve as an inner-loop training signal. The general problem is therefore how to obtain verifiable supervision and rewards for embodied planners without relying on string-level matching or full simulation. Here we show that a single BDDL specification, automatically constructed from open-world video evidence or curated tasks, can serve as a shared interface for data construction, plan verification, and reward design. A video-to-BDDL parser, an LLM verifier...

论文介绍 针对具身任务规划缺乏高效验证与奖励信号的问题,本文提出基于符号强化学习的具身推理方法。研究利用单一规范作为共享接口,从视频中自动构建规范,并结合大语言模型验证器,为规划器提供确定性的有效性检查与奖励设计。该方法摆脱了对字符串匹配或全真仿真的依赖,提升规划效率与可靠性。

TactX: Learning Shared Tactile Representations Across Diverse Sensors

第一作者: Junsung Park · 方向: 机器人操作 · 来源: cs.RO

Abstract:Tactile sensors provide critical information for contact-rich manipulation, yet tactile representations and policies remain tightly coupled to each specific sensor, limiting transferability across robots and hardware platforms. We propose TactX, a framework for learning a transferable tactile representation across sensors spanning three fundamentally different transduction modalities: resistive, magnetic, and vision-based. TactX maps heterogeneous tactile observations into a shared latent space through modality-specific encoders trained on paired contact data. Such paired interactions provide a natural alignment signal across modalities, and the encoders are jointly trained across all sensor pairs, inducing a consistent latent space for all sensor types. Our experiments show that TactX aligns tactile representations across sensors while preserving object-level contact...

论文介绍 针对触觉表示高度耦合于特定传感器而限制跨平台迁移的问题,本文提出跨模态触觉表示学习框架。该方法跨越电阻、磁性和视觉三种传感模态,学习可迁移的共享触觉表示。通过特定模态编码器与成对接触数据的对齐信号,将异构触觉观察映射至共享潜在空间,在保留物体级接触信息的同时实现跨传感器泛化。

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

第一作者: Hao Sun · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps between high-level instructions and low-level motor commands, and inference inefficiency due to autoregressive scalar decoding. In this work, we propose MIRTH, a unified framework designed to address these challenges. MIRTH augments a pretrained VLA backbone with three key innovations: (1) dual-scale temporal memory hubs that compress long-term scene evolution and short-term motion trends into compact embeddings; (2) latent reasoning tokens optimized via a mutual-information objective carving out a semantic plan space to align multimodal context with action trajectories; and (3) a parallel action decoding...

论文介绍 针对现有视觉语言动作模型单帧架构存在的时间近视、推理脱节及解码效率低等问题,本文提出统一推理框架。该方法引入双尺度时间记忆枢纽压缩长短期场景动态,并通过互信息目标优化潜在推理标记,构建语义规划空间以对齐多模态上下文与动作轨迹。结合并行解码机制,显著提升智能体的推理与执行效率。

LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music

第一作者: Snehasis Banerjee · 方向: 多模态具身 · 来源: cs.RO

Abstract:The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics. Traditional methods often rely on rigid, pre-programmed commands that limit the robot's expressiveness and adaptability. This paper introduces a novel framework that leverages the reasoning capabilities of Large Language Models (LLMs) to synthesize complex robotic actions from a rich tapestry of multimodal human inputs: natural speech, hand gestures, and music/sound beats. Our system architecture integrates a speech transcription model, a gesture recognition module, and a signal processing pipeline for beat detection. These processed inputs are contextualized using prompt templates and fed into a LLM. The LLM, informed by a predefined robot action space, reasons over the combined inputs to generate a coherent sequence of actions. This sequence is dispatched to an...

论文介绍 针对传统人机交互依赖刚性命令限制机器人表现力的问题,本文提出基于大语言模型的多模态动作合成框架。系统集成语音转录、手势识别与音乐节拍检测模块,将多模态输入通过提示模板转化为上下文信息。大语言模型结合预定义动作空间进行推理,生成连贯的动作序列,实现更自然直观的人机交互。

A Modular Vision-Language-Action Robotics Framework for Indoor Environments

第一作者: Anindya Jana · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:This paper presents an integrated system for the CMU Vision-Language-Action (VLA) Challenge, designed to enable an autonomous agent to perform complex tasks based on natural language instructions. Our framework employs a modular architecture that orchestrates environment mapping, question processing, and navigation. The system operates in two parallel streams: a perception pipeline that constructs a semantic voxel map from real-time camera feeds using OwlViT embeddings, and a language pipeline that classifies user commands with a Vision-Language Model. The mapping is time-constrained; the system proceeds with a partial map if a 500-second exploration limit is reached. The classified query is then grounded in the geometric and semantic context of the map to generate a detailed prompt for the VLM. This yields an actionable output, demonstrating a capable solution for bridging...

论文介绍 本文提出一种面向室内环境的模块化视觉语言动作机器人框架,旨在使机器人根据自然语言指令执行复杂任务。系统采用双流水线架构,感知端利用视觉模型构建语义体素地图,语言端通过视觉语言模型解析用户指令。两者结合实现指令在三维空间中的精准定位与动作生成,为室内自主导航与任务执行提供了高效解决方案。

ELASTIC: Efficiently Learning to Adaptively Scale Test-Time Compute for Generative Control Policies

第一作者: Andrew Zou Li · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Generative control policies (GCPs), such as diffusion policies and flow-based vision-language-action models, enable test-time scaling in robot control. Test-time compute can be allocated along two axes: sequential scaling, which increases denoising steps to refine actions, and parallel scaling, which samples multiple candidate actions to search across modes of the policy distribution. However, the optimal allocation of sequential and parallel compute is hard to know a priori as it is state-, task-, and policy-dependent. For example, early stages of a grasp may benefit from broader parallel exploration, while near-contact phases may require more sequential refinement for precision. We present ELASTIC, an algorithm that learns state-dependent test-time compute schedules for GCPs. We formulate compute allocation as a meta-Markov Decision Process in which a meta-policy interacts...

论文介绍 针对生成式控制策略中测试时计算分配难以先验确定的问题,本文提出ELASTIC算法以自适应调整顺序与并行计算资源。该方法将计算分配建模为元马尔可夫决策过程,通过学习状态依赖的计算调度策略,使机器人能根据任务阶段动态优化去噪步数与候选动作采样,在控制精度与计算效率间取得平衡。

Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors

第一作者: Zixing Wang · 方向: 机器人操作 · 来源: cs.RO

Abstract:Bridging the sim-to-real gap is a core challenge in deploying learned manipulation policies. Sim-to-real learning is attractive because it can replace expensive real robot demonstrations with scalable synthetic data, yet world-action models have not previously been shown to transfer from simulation to real robotic manipulation. We study whether a world-action model can be trained from synthetic priors and deployed zero-shot in the real world. To this end, we build upon Cosmos Policy, a video diffusion model adapted for visuomotor control. We construct simulation environments with extensive domain randomization and generate demonstrations using the AnyTask motion planning pipeline. We evaluate our approach across object lifting, drawer opening, and pick-and-place tasks using ${\sim}800$ synthetic demonstrations per task and no real demonstrations. When deployed zero-shot on a...

论文介绍 本文探讨世界动作模型从合成数据到真实机器人操作的零样本迁移能力。研究基于适配视觉运动控制的视频扩散模型,在包含广泛域随机化的仿真环境中生成合成演示数据。通过在物体提升、抽屉开启等任务上的评估,验证了仅依赖合成先验训练的模型能够零样本部署于真实物理环境,有效降低了真实数据采集成本。

Labimus: A Simulation and Benchmark for Humanoid Dexterous Manipulation in Chemical Laboratory

第一作者: Yuhan Wu · 方向: 机器人操作 · 来源: cs.RO

Abstract:Laboratory automation has made remarkable progress through robotic platforms and AI-driven scientific reasoning. However, many laboratory operations (e.g., solid--solid transfer) remain inherently dynamic and require real-time adaptation to different materials and experimental conditions. Such precision-critical manipulations are difficult to standardize, motivating the use of humanoid robots with dexterous hands. Despite this opportunity, no existing benchmark evaluates humanoid manipulation in precision-critical laboratory environments. We present Labimus, to our knowledge, the first benchmark for humanoid dexterous manipulation in organic chemistry laboratories. Labimus reconstructs over 30 functionally faithful assets from real organic chemistry workstations through real-to-sim modeling, collectively covering the core operations of routine organic chemistry experiments...

论文介绍 针对化学实验室中精密操作难以标准化的问题,本文提出Labimus,首个面向有机化学实验室的人形机器人灵巧操作仿真基准。该基准通过真实到仿真的建模方法,重建了三十多个高保真实验室资产,覆盖常规化学实验的核心操作。它为评估人形机器人在复杂动态且精度要求极高的实验环境中的适应能力提供了标准平台。

Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force

第一作者: Jaden Clark · 方向: 机器人操作 · 来源: cs.RO

Abstract:Robot manipulation often relies on sensory feedback beyond vision, particularly in contact-rich settings where force, tactile, or audio signals reveal interaction states that are not directly observable from images. However, these modalities are often hardware- and task-specific, and large-scale multisensory robot datasets remain scarce. As a result, it is impractical to pretrain policies with every sensor they may encounter. We study multisensory continual learning: adapting a pretrained robot policy to new tasks with newly introduced modalities while preserving performance under the original sensor suite. We propose MuSe, which incorporates limited multisensory data into pretrained vision-only policies through multi-stage fusion, multisensory future prediction, and experience replay over pretraining data. We instantiate MuSe by augmenting a pretrained vision-only policy with...

论文介绍 针对接触密集型任务中多传感器数据稀缺的问题,本文提出多感官持续学习方法MuSe,旨在将预训练的纯视觉策略适应新引入的力或触觉模态。该方法通过多阶段特征融合、多感官未来预测以及基于预训练数据的经验回放机制,使机器人能够在不丧失原有视觉任务性能的前提下,高效学习并利用新的多模态感觉反馈。

Motion Planning in Compressed Representation Spaces

第一作者: Lukas Lao Beyer · 方向: 机器人操作 · 来源: cs.RO

Abstract:Deep learning methods have vastly expanded the capabilities of motion planning in robotics applications, as learning priors from large-scale data has been shown to be essential in capturing the highly complex behavior required for solving tasks such as manipulation or navigation for autonomous vehicles. At the same time, model-based planning algorithms based on search or optimization remain an essential tool due to their flexibility, efficiency, and the ability to incorporate domain knowledge via expert-designed algorithms and objective functions. We propose a new generative framework to unify these two paradigms. First, we learn an autoencoder with a high compression ratio and a latent space of hierarchically ordered, discrete-valued tokens. Leveraging both the dimensionality reduction and the hierarchical coarse-to-fine structure learned by this autoencoder, we then perform...

论文介绍 本文提出一种在压缩表示空间中进行机器人运动规划的新型生成框架,旨在统一数据驱动学习与基于模型的搜索优化。该方法训练一个高压缩比的分层离散自编码器,将复杂状态映射至低维潜空间。随后利用潜空间的层次化结构进行高效规划,在保留专家先验知识的同时,大幅提升了复杂任务中运动规划的灵活性与计算效率。

The Quadruped Soft Tail: Compliant Grasping and Swabbing for Contamination Surveys in Harsh Environments

第一作者: Harald Minde Hansen · 方向: 机器人操作 · 来源: cs.RO

Abstract:Beryllium contamination surveys in radioactive areas are challenging for robots in environments cluttered with cables and electronics. To address this problem, we have developed a novel quadruped system augmentation: A lightweight, soft, and compliant tendon-actuated robotic tail mounted on a quadruped robot. The tail features a hollow, flexible backbone and a tendon-actuated soft gripper that enables the robot to pick up sampling tissues, swab contaminated surfaces, and release the tissues at designated collection locations for subsequent beryllium analysis. To enable intuitive teleoperation, a closed-form kinematic model and a singularity-robust task-space controller are developed. Experimental results demonstrate that gripper actuation has a negligible effect on robot shape, while common-mode tendon actuation provides an effective mechanism for stiffness modulation and...

论文介绍 针对恶劣环境中的污染调查难题,本文设计了一种安装在四足机器人上的轻量级柔顺腱驱动软尾巴。该尾巴配备中空柔性脊柱与软夹爪,可执行采样组织拾取、表面擦拭及定点释放等任务。结合闭式运动学与奇异鲁棒控制器,系统实现了直观遥操作,在几乎不影响四足机器人本体形态的前提下,有效扩展了其在复杂空间的作业能力。

Sampling-Based Coordination-Informed Multi-Objective Multi-Robot Reinforcement Learning

第一作者: Antonio Marino · 方向: 策略学习 · 来源: cs.RO

Abstract:Multi-robot systems must simultaneously optimize competing objectives while maintaining coordinated behavior. Existing multi-agent reinforcement learning approaches often rely on fixed or centralized coordination, which limits adaptability and violates distributed constraints. This work introduces the Coordination-Informed Multi-Objective Reinforcement Learning (CIMORL) framework, integrating a distributed weight prediction mechanism, a privileged expert training strategy, and theoretical guarantees for Pareto-optimal solutions. We present the base CIMORL method alongside two sampling-based variants, CIMORL-TS (Tree Search) and CIMORL-MPPI (MPPI), which leverage privileged global information during training to enable fully decentralized deployment. Experimental validation in cooperative and adversarial scenarios demonstrates a $21.2\%$ hypervolume improvement and superior...

论文介绍 针对多机器人系统在多目标优化中难以兼顾协调与分布式约束的问题,本文提出CIMORL框架。该方法集成分布式权重预测与特权专家训练策略,提供帕累托最优解理论保证。通过引入树搜索与模型预测路径积分两种采样变体,系统在训练阶段利用全局特权信息,实现完全去中心化部署,显著提升了多智能体协作的自适应能力。

TAPE: Tether-Aware Path Planning for Autonomous Exploration of Unknown 3D Cavities Using a Tangle-Compatible Tethered Aerial Robot

第一作者: Louis Petit · 方向: 导航与运动 · 来源: cs.RO

Abstract:This letter presents the first method for autonomous exploration of unknown cavities in three dimensions (3D) that focuses on minimizing the distance traveled and the length of tether unwound. Considering that the tether entanglements are little influenced by the global path, our approach employs a 2-level hierarchical architecture. The global frontier-based planning solves a Traveling Salesman Problem (TSP) to minimize the distance. The local planning attempts to minimize the path cost and the tether length using an adjustable decision function whose parameters play on the trade-off between these two values. The proposed method, TAPE, is evaluated through detailed simulation studies as well as field tests. On average, our method generates a 4.1% increase in distance traveled compared to the TSP solution without our local planner, with which the length of the tether remains...

论文介绍 本文提出TAPE方法,用于系留无人机在未知三维空洞中的自主探索。该方法采用两层分层架构,全局规划通过旅行商问题最小化移动距离,局部规划利用可调决策函数在路径成本与系绳展开长度间取得平衡,同时减少系绳缠绕。该研究在仿真与实地测试中验证了其在最小化距离和系绳长度方面的有效性,适用于复杂受限空间的自主探测。

From Grasps to Dexterity: Large-Scale Grasp Pretraining for Dexterous Manipulation

第一作者: Ying Yuan · 方向: 机器人操作 · 来源: cs.RO

Abstract:Large-scale dexterous grasp datasets encode rich priors over hand-object interaction, but their use has largely been confined to grasp generation and pick-and-place manipulation. We study whether such data can instead support functional dexterity in articulated tool use, where a robot must acquire a tool, maintain contact, and operate its functional moving parts. We adapt a hierarchical imitation learning framework that combines high-level hand sub-goal prediction with a low-level goal-conditioned controller. We construct a 355k-trajectory grasp-pretraining dataset from large-scale dexterous grasp annotations and use it to pretrain the low-level controller. The controller is then fine-tuned on downstream task demonstrations. To evaluate this setting, we introduce DexCraft, a simulation benchmark with six articulated tool-use tasks requiring coordinated finger motion. Across...

论文介绍 本文研究如何利用大规模灵巧抓取数据集支持关节工具使用中的功能性灵巧操作。作者采用分层模仿学习框架,结合高层手部子目标预测与底层目标条件控制器。通过构建包含35万条轨迹的抓取预训练数据集对底层控制器进行预训练,并在下游任务演示上微调。同时提出DexCraft仿真基准,评估该框架在需要协调手指运动的关节工具使用任务中的表现。

Vision-Language Procedural Reasoning for Context-Aware Reward Modeling of Robotic Endovascular Guidewire Navigation

第一作者: Wentong Tian · 方向: 导航与运动 · 来源: cs.RO

Abstract:Robotic-assisted endovascular interventions demand accurate, stable, and context-aware guidewire navigation in complex and patient-specific vascular anatomies. Despite recent advances in robotic precision and learning-based control, existing autonomous navigation methods remain limited by their reliance on static reward functions and the lack of explicit procedural reasoning regarding anatomical context and task progression. To address these challenges, this paper proposes a vision-language procedural reasoning (VL-PR) framework for autonomous guidewire navigation. The framework integrates a multimodal large language model (MLLM) as a procedural reasoning module that interprets real-time visual observations to infer high-level navigation contexts. Instead of generating low-level control commands, the inferred procedural insights enable context-aware reward adaptation by...

论文介绍 针对机器人辅助血管内干预中自主导航依赖静态奖励函数且缺乏解剖上下文推理的问题,本文提出视觉语言程序推理框架。该框架集成多模态大语言模型解释实时视觉观测以推断高层导航上下文,并将推理结果用于上下文感知的奖励自适应。该方法可引导导丝在复杂血管解剖中进行精确、稳定的自主导航,提升医疗机器人系统的任务适应性。

ViTL: Temporal Logic-Guided Zero-Shot Natural Language Navigation via Vision-Language Models

第一作者: Kaier Liang · 方向: 导航与运动 · 来源: cs.RO

Abstract:Enabling robots to follow natural language commands to complete zero-shot long-horizon tasks remains challenging. It requires extracting implicit temporal and logical constraints from natural language commands and executing multiple sub-tasks accordingly. Recent zero-shot object navigation methods use vision-language models (VLMs) to guide frontier-based exploration in unknown environments, but they are limited to single-target tasks. Real-world commands such as "Clean either the chair or the couch, then turn on the tv." require navigating to multiple targets in a temporally constrained order, which no existing zero-shot system can handle. We present ViTL, a framework that addresses this gap at two levels. At the task level, we use a large language model (LLM) to compile natural language commands into Linear Temporal Logic (LTL) formulas, which are then converted into...

论文介绍 针对机器人难以从零样本自然语言指令中提取隐式时序与逻辑约束以执行长视界任务的问题,本文提出ViTL框架。在任务层,利用大语言模型将指令编译为线性时序逻辑公式并转化为自动机;在导航层,结合视觉语言模型引导未知环境探索。该框架使机器人能够处理包含时序约束的多目标导航指令,填补了现有零样本多目标导航系统的空白。

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning

第一作者: Taozhao Chen · 方向: VLA 通用模型 · 来源: cs.RO

Abstract:Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs), have shown rapidly improving performance on robot manipulation benchmarks. These gains are commonly interpreted as evidence that semantic representations learned from internet-scale data transfer to physical execution generalization. This position paper argues that the assumption underlying this interpretation -- that semantic generalization is sufficient to support physical action decisions -- has not been independently verified and cannot be tested under current evaluation protocols. We support this claim by decomposing VLA policies into semantic mapping and physical action decision, and showing that task success rate -- the dominant evaluation metric -- cannot distinguish between these two sources of capability. As a result, improvements in benchmark performance are consistent with...

论文介绍 本文质疑当前视觉语言动作模型在操作基准上的性能提升是否证明其具备物理推理能力。作者指出,将语义泛化等同于物理动作决策的假设尚未被验证。通过将策略分解为语义映射与物理动作决策,证明任务成功率这一主导指标无法区分这两种能力来源。因此,现有基准性能的提升并不能验证模型真正具备物理推理能力,呼吁改进评估协议。

MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning

第一作者: Sheng Zhang · 方向: 数据集与评测 · 来源: cs.RO

Abstract:Large language models (LLMs) provide a promising interface for high-level robotic task planning, but their use in multi-UAV collaboration remains difficult to evaluate systematically. Existing UAV simulators mainly emphasize dynamics, perception, or low-level control, while existing LLM-agent benchmarks rarely capture aerial-robotics constraints such as partial observability, spatial coverage, UAV assignment, and multi-vehicle coordination. To bridge this gap, we present MultiUAV-Plat, a lightweight, easy-to-use, LLM-agent-oriented simulation platform for multi-UAV collaborative task planning. The platform exposes concise RESTful APIs, agent-facing observations, role-based information access, hidden validation logic, and optional 2D/3D visualization, allowing agents to solve missions through realistic tool interaction rather than privileged simulator access. Built on this...

论文介绍 针对大语言模型在多无人机协同规划中缺乏系统评估手段的问题,本文提出MultiUAV-Plat仿真平台与基准。该平台提供RESTful API、智能体观测及隐藏验证逻辑,使智能体通过真实工具交互解决任务。框架有效捕捉了部分可观测性、空间覆盖及多机协同等航空机器人约束,填补了现有大语言模型在无人机协同任务规划领域的评测空白。

Warp RL: Reshaping Base Policy Distributions for Dynamics Adaptation

第一作者: Ethan Hirschowitz · 方向: 策略学习 · 来源: cs.RO

Abstract:Residual reinforcement learning adapts a pretrained robot policy by learning an additive correction to its actions. While effective when adaptation amounts to shifting the base policy's action distribution, additive corrections cannot change the distribution's shape, scale, or state-dependent geometry -- limitations we formalize as wrong variance, miscalibrated confidence, and non-uniform correction. We show that these matter under dynamics shift: when the base distribution is geometrically mismatched to the shifted system, residual correction can underperform even the unadapted policy. We propose \textbf{Warp RL}, a policy adaptation method that replaces additive residuals with an invertible, state-conditioned transformation of the base policy's action distribution. Instantiated with monotonic rational-quadratic spline flows [arXiv:0706.1234v1], Warp RL preserves identity...

论文介绍 残差强化学习通过加性修正适应预训练策略,但无法改变基础分布的形状与尺度,在动力学偏移时表现受限。本文提出Warp RL方法,用可逆的状态条件变换替代加性残差,重塑基础策略的动作分布。该方法利用单调有理二次样条流实例化,有效解决了残差修正中的方差错误与非均匀修正问题,提升了机器人在动力学偏移下的策略适应能力。

GaussLite: Online Task-Conditioned 3D Gaussian Splatting for Real-Time Robotic Mapping

第一作者: Annika Thomas · 方向: 机器人操作 · 来源: cs.RO

Abstract:Existing 3D Gaussian Splatting (3DGS) systems distribute representation capacity uniformly across a scene, ignoring the fact that many downstream robotic tasks engage only a fraction of the reconstructed geometry. This causes valuable onboard compute to be allocated towards optimizing irrelevant parts of the scene, either limiting online capacity or under-optimizing the most relevant parts of the scene. We introduce GaussLite, a task-driven 3DGS mapping system that conditions its representation density on a natural-language task specification. Given a posed RGB-D stream and a task such as "prepare to pick up the object on the desk," GaussLite uses a one-shot LLM parser to extract target and anchor objects, which are grounded per-frame by an open-vocabulary detector and segmented to produce per-pixel relevance masks in real time. The mapper allocates seeding density, gradient...

论文介绍 现有三维高斯溅射系统均匀分配场景表示能力,易浪费计算资源。本文提出GaussLite任务驱动建图系统,根据自然语言任务条件化表示密度。系统利用大语言模型解析目标物体,结合开放词汇检测器生成实时相关性掩码,动态分配种子密度与计算资源。该方法在有限机载算力下集中优化任务相关场景,显著提升了实时机器人建图与操作的效率。

Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving

第一作者: Anh Dung Dinh · 方向: VLA 通用模型 · 来源: cs.CV

Abstract:Modern Vision-Language-Action (VLA) planners for autonomous driving emit a chain-of-causation (CoC) reasoning step \emph{before} producing a trajectory. The reasoning is autoregressive and dominates inference latency, while the trajectory head is parallel and cheap. Latency is an operational constraint in autonomous driving, so accelerating the reasoning step is the central problem we address. We observe that CoC reasoning has two qualitatively different needs: most tokens continue routine setup that follows naturally from the ego-trajectory history, and a small fraction encode commitments that require fresh visual evidence about an unexpected situation. We split this reasoning into two specialized paths: a \emph{routine reasoner} that handles the predictable continuation by attending to trajectory history, and a \emph{deliberative reasoner} (the unmodified VLA target) that...

论文介绍 针对自动驾驶视觉语言动作模型中自回归推理导致的高延迟问题,本文提出一种推理感知的推测解码方法。该方法将因果链推理拆分为处理常规延续的常规推理器,以及处理意外情况的深思推理器两条专用路径。此设计在保持推理质量的同时显著加速了自回归过程,有助于满足自动驾驶对低延迟的严格要求。

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

第一作者: Tao Chen · 方向: 机器人操作 · 来源: cs.AI

Abstract:Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similarity for object matching, neglecting physical affordances such as handle graspability and material fragility, and operate open-loop without spatial reasoning or failure recovery, limiting their effectiveness when objects are densely packed or physically diverse. We present Agentic RAG-VLM, a unified framework that bridges VLM-based semantic understanding and physically grounded grasp execution by integrating retrieval-augmented generation (RAG) with vision-language models (VLMs) and agentic self-reflective planning. Agentic RAG-VLM introduces three tightly coupled components: (1) a Hierarchical Affordance-Aware RAG (HAA-RAG) that encodes four-dimensional affordance descriptors, including type...

论文介绍 针对杂乱环境机器人抓取缺乏物理可供性理解与开环局限,本文提出Agentic RAG-VLM框架。该方法结合检索增强生成与视觉语言模型,引入分层可供性感知检索与自反思规划机制,有效桥接语义理解与物理执行,提升密集或多样化物体环境下的抓取泛化与失败恢复能力。

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

第一作者: Hao Zhang · 方向: 策略学习 · 来源: cs.CL

Abstract:Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in parallel, and are then verified by the target model, enabling lossless acceleration. Recently, diffusion-based speculative decoding further improves parallelism by generating multiple tokens per forward pass via block-level diffusion, achieving state-of-the-art (SOTA) performance. However, existing methods adopt a fixed inference block size and assume a uniform optimal decoding strategy across all inputs. In this paper, we show that this assumption is suboptimal, as the optimal block size varies across samples and plays a critical role in speculative decoding performance. Moreover, these values exhibit a clear local structure, concentrating around the training block size, which reduces the problem to a low-dimensional and structured decision space. Based on these...

论文介绍 针对基于扩散的推测解码采用固定推理块大小导致次优的问题,本文提出BlockPilot方法。研究发现最优块大小随样本变化且具局部结构,该方法通过实例自适应策略学习,动态调整推理块大小。此举打破统一解码策略局限,进一步优化了推测解码的并行度与推理加速性能。

市场总览

当前全球资产技术面呈现显著分化。美股方面,SPY与QQQ维持均线多头排列并逼近52周高点,VIX指数回落至16.45,整体技术面偏强。加密市场则极度弱势,加密恐慌贪婪指数降至11(极度恐慌),总市值2.13T(24h跌0.87%),BTC主导率55.4%、ETH占9%,主流币均呈空头排列且逼近低点。中概股技术面普遍承压,BABA与JD的RSI均跌至超卖区,均线呈空头排列。商品与外汇方面,黄金期货触发均线死叉,原油RSI跌至30超卖区;美元指数则维持多头排列逼近52周高,10Y美债收益率处于中性震荡区间。整体而言,风险资产内部出现剧烈分化,美元与部分宏观指标相对坚挺,技术面多空博弈激烈。

今日关注

TSLA 特斯拉 (TSLA)
偏上行

当前价420.60,近5日大涨10.22%,成功突破SMA20(400.59)与SMA50(405.62)。指标显示MACD形成金叉(-3.3065上穿-4.3862),RSI14达56.8处于强势区间,短期多头动量显著释放,技术形态呈现偏上行特征。

BTC-USD 比特币 (BTC-USD)
偏下行

当前价58804.59,近5日下跌1.54%,价格逼近52周低点。技术面显示MACD形成死叉(-2347.71下穿-2315.65),RSI14降至30.5逼近超卖区,且均线系统呈典型空头排列,下行动能未见衰竭,技术形态呈现偏下行特征。

AAPL 苹果 (AAPL)
中性

当前价289.36,近1日反弹2.7%但5日仍跌1.68%,价格受制于SMA20(295.93)。RSI14为46.9处于中性区间,MACD绿柱扩大(-2.7924低于信号线-0.6997),多空力量交织缺乏明确单边趋势,技术形态呈现中性震荡特征。

全部资产

^VIX

VIX 恐慌指数

$16.45 -6.80%
5 日
-15.60%
距 52w 高
-53.4%
RSI(14)
45.3
趋势
空头
SMA 20 / 50 / 200
18.06 / 17.74 / 18.67
MACD / 信号
-0.043 / 0.031
MACD 死叉 (今天)空头排列

^TNX

10Y 美债收益率 (%)

$4.42 +1.05%
5 日
-2.02%
距 52w 高
-11.6%
RSI(14)
45.8
趋势
中性
SMA 20 / 50 / 200
4.47 / 4.45 / 4.22
MACD / 信号
-0.015 / -0.001

DX-Y.NYB

美元指数 DXY

$101.30 +0.11%
5 日
-0.30%
距 52w 高
-0.5%
RSI(14)
67.9
趋势
多头
SMA 20 / 50 / 200
100.45 / 99.40 / 98.82
MACD / 信号
0.579 / 0.526
接近 52 周高多头排列

SPY

S&P 500 ETF

$746.77 +0.78%
5 日
+1.80%
距 52w 高
-1.8%
RSI(14)
55.0
趋势
多头
SMA 20 / 50 / 200
742.24 / 735.87 / 691.43
MACD / 信号
0.682 / 1.847
接近 52 周高多头排列

QQQ

Nasdaq 100 ETF

$736.40 +1.70%
5 日
+3.19%
距 52w 高
-1.6%
RSI(14)
56.6
趋势
多头
SMA 20 / 50 / 200
723.73 / 706.22 / 633.63
MACD / 信号
4.549 / 6.305
接近 52 周高多头排列

AAPL

Apple

$289.36 +2.70%
5 日
-1.68%
距 52w 高
-8.8%
RSI(14)
46.9
趋势
中性
SMA 20 / 50 / 200
295.93 / 292.25 / 270.03
MACD / 信号
-2.792 / -0.700

MSFT

Microsoft

$373.02 +1.21%
5 日
-0.25%
距 52w 高
-32.8%
RSI(14)
41.3
趋势
空头
SMA 20 / 50 / 200
391.65 / 408.94 / 446.69
MACD / 信号
-13.198 / -11.124
空头排列

NVDA

Nvidia

$200.09 +2.63%
5 日
+0.02%
距 52w 高
-15.4%
RSI(14)
45.2
趋势
中性
SMA 20 / 50 / 200
205.74 / 209.99 / 190.84
MACD / 信号
-3.879 / -2.572

GOOGL

Alphabet

$357.37 +1.05%
5 日
+3.25%
距 52w 高
-12.5%
RSI(14)
47.9
趋势
中性
SMA 20 / 50 / 200
358.53 / 369.94 / 315.01
MACD / 信号
-6.052 / -4.935

TSLA

Tesla

$420.60 +2.13%
5 日
+10.22%
距 52w 高
-15.7%
RSI(14)
56.8
趋势
中性
SMA 20 / 50 / 200
400.59 / 405.62 / 418.54
MACD / 信号
-3.306 / -4.386
MACD 金叉 (今天)

META

Meta

$563.29 +0.12%
5 日
+0.19%
距 52w 高
-29.3%
RSI(14)
42.8
趋势
空头
SMA 20 / 50 / 200
577.94 / 608.11 / 648.13
MACD / 信号
-14.688 / -13.851
空头排列
加密恐慌贪婪
11
极度恐慌
加密总市值
$2.13 T
-0.87% / 24h
BTC 主导率
55.4%
ETH 9.0%
24h 成交量
$80.8 B
活跃币 17,429

BTC-USD

Bitcoin

$58,804.59 -2.22%
5 日
-1.54%
距 52w 高
-53.4%
RSI(14)
30.5
趋势
空头
SMA 20 / 50 / 200
62,661.39 / 68,505.03 / 75,355.85
MACD / 信号
-2,347.715 / -2,315.659
MACD 死叉 (今天)接近 52 周低空头排列

ETH-USD

Ethereum

$1,580.61 -1.84%
5 日
+1.01%
距 52w 高
-68.1%
RSI(14)
35.1
趋势
空头
SMA 20 / 50 / 200
1,671.30 / 1,860.39 / 2,294.45
MACD / 信号
-75.111 / -77.573
MACD 金叉 (1 天前)空头排列

SOL-USD

Solana

$74.36 -0.78%
5 日
+10.04%
距 52w 高
-70.6%
RSI(14)
54.3
趋势
空头
SMA 20 / 50 / 200
70.89 / 76.33 / 94.54
MACD / 信号
-0.559 / -1.478
空头排列

BABA

阿里巴巴 (BABA)

$95.98 +0.49%
5 日
-6.45%
距 52w 高
-50.2%
RSI(14)
19.8
趋势
空头
SMA 20 / 50 / 200
110.63 / 124.53 / 147.68
MACD / 信号
-8.717 / -7.492
RSI 超卖空头排列

PDD

拼多多 (PDD)

$76.28 -0.34%
5 日
-0.37%
距 52w 高
-45.3%
RSI(14)
33.7
趋势
空头
SMA 20 / 50 / 200
80.57 / 90.14 / 108.88
MACD / 信号
-4.226 / -4.277
MACD 金叉 (今天)空头排列

JD

京东 (JD)

$25.48 +0.91%
5 日
-2.45%
距 52w 高
-30.9%
RSI(14)
29.0
趋势
空头
SMA 20 / 50 / 200
27.62 / 29.43 / 30.01
MACD / 信号
-1.222 / -0.983
RSI 超卖空头排列

0700.HK

腾讯控股 (0700.HK)

HK$429.80 +2.28%
5 日
+3.62%
距 52w 高
-37.1%
RSI(14)
44.5
趋势
空头
SMA 20 / 50 / 200
444.93 / 458.14 / 559.73
MACD / 信号
-10.568 / -8.869
空头排列

GC=F

黄金期货

$3,997.10 -0.64%
5 日
+0.17%
距 52w 高
-28.4%
RSI(14)
32.8
趋势
空头
SMA 20 / 50 / 200
4,197.73 / 4,439.95 / 4,453.77
MACD / 信号
-127.552 / -116.294
死叉(SMA50↓SMA200) (今天)空头排列

CL=F

WTI 原油期货

$69.81 +0.45%
5 日
-0.75%
距 52w 高
-41.6%
RSI(14)
30.0
趋势
中性
SMA 20 / 50 / 200
80.07 / 90.83 / 74.00
MACD / 信号
-6.506 / -5.854
RSI 超卖

USDCNY=X

美元 / 人民币

¥6.78 -0.20%
5 日
-0.15%
距 52w 高
-6.0%
RSI(14)
48.3
趋势
空头
SMA 20 / 50 / 200
6.77 / 6.79 / 6.95
MACD / 信号
-0.000 / -0.004
接近 52 周低空头排列
风险提示

本报告基于历史行情数据计算得出,过去走势不代表未来表现。所有技术指标读数仅反映当前市场状态,不构成任何投资建议,仅供技术指标解读参考。请投资者注意风险控制。

Trump made more than $1bn from crypto in first year back in office

The president's crypto income far outpaces his earnings from real estate and Trump-themed items such as watches.

中文摘要 特朗普在重返白宫的第一年通过加密货币赚取超过10亿美元,其加密收入远超房地产及特朗普主题商品(如手表)的收益。

Australia politics live: author Anna Funder says she’s a ‘victim of crime’ as creatives lobby government to protect them from AI

Meanwhile Sarah Hanson-Young confirms Greens and Coalition in talks to pressure Labor on gambling. Follow today’s news live Get our breaking news email, free app or daily news podcast Albanese says charges against EY employees who accessed PM’s bank account details ‘appropriate’ It’s not just KPMG i

中文摘要 澳大利亚政坛动态:作家安娜·芬德勒称自己成为「犯罪受害者」,创意界游说政府防范AI威胁。绿党与联盟党正施压工党处理赌博问题。总理阿尔巴尼斯称起诉查阅其银行账户的安永员工「适当」。

US lifts restrictions on powerful AI models Fable, Mythos, Anthropic says

AI firm says it will begin restoring access to Claude Fable 5 and Mythos 5 after removal of export controls.

中文摘要 人工智能公司Anthropic宣布,随着美国取消出口管制,将恢复其强大AI模型Claude Fable 5和Mythos 5的访问权限。此前这些模型曾受限于相关出口限制。

Southeast Asia’s homegrown artists are knocking K-pop off its pedestal

Drawing lessons from the global success of K-Pop, the region is finding its own voice in a new generation of artists.

中文摘要 东南亚本土艺术家正逐渐打破韩国流行音乐的主导地位。该地区新一代艺术家借鉴K-pop的全球成功经验,正在国际舞台上寻找并确立属于自己的独特声音与风格。

U.S. and Iran to Meet with Mediators in Qatar

American and Iranian officials are in the Gulf state, a key intermediary between the two countries, days after new round of attacks threatened efforts to sign a lasting peace deal.

中文摘要 美国和伊朗官员抵达关键中间人卡塔尔,与调解人会面。此举发生在新一轮袭击威胁到双方签署持久和平协议的努力数天之后,旨在推动和平进程。

Trump Officials Sideline Machado, Venezuela’s Opposition Leader, Over Earthquake Response

U.S. officials called a bid by María Corina Machado, the Nobel Peace Prize winner, to return to earthquake-battered Venezuela a “political stunt” that has distracted from recovery efforts.

中文摘要 美国官员将委内瑞拉反对派领导人、诺贝尔和平奖得主玛丽亚·科里纳·马查多重返震灾严重的委内瑞拉之举称为「政治噱头」,认为此举分散了救灾恢复工作的注意力,从而将其边缘化。

Algeria to vote in test of post-Hirak political landscape

Algeria holds legislative elections amid debates over reform, turnout and political stability.

中文摘要 阿尔及利亚举行立法选举,这是对后「Hirak」运动时期政治格局的一次重要考验。此次选举在关于改革、投票率和政治稳定性的激烈辩论背景下进行。

LeBron James thanks LA Lakers ahead of free agency exit for 24th season

After eight years at the LA Lakers, James's next destination will be decided in NBA's imminent free agency period.

中文摘要 在洛杉矶湖人队效力八年后,勒布朗·詹姆斯在即将开启个人第24个NBA赛季之际向球队致谢。他的下一站去向将在NBA即将到来的自由球员市场期间决定。

North Korea’s Kim hails ‘unshakeable will’ to develop ties with China’s Xi

Kim Jong Un has sent Xi Jinping a congratulatory message to mark the 105th anniversary of the Chinese Communist Party.

中文摘要 朝鲜领导人金正恩向中国国家主席习近平致贺信,纪念中国共产党成立105周年。金正恩在信中赞扬了双方发展关系的「不可动摇的意志」。

Anthropic says US lifts export ban on its advanced AI tools

Fable and Mythos were abruptly suspended in June over concerns that they could be used by hackers.

中文摘要 人工智能公司Anthropic表示,美国已解除对其先进AI工具的出口禁令。此前,其AI模型Fable和Mythos因担忧可能被黑客利用,于6月被突然暂停使用。

兩代香港人的「七一」:發聲仍在記憶和生活感受中

從主權移交、七一遊行,再到反修例運動與後《國安法》時代,7月1日如今對香港人來說意味著什麼?「七一」29週年前夕,前社民連主席陳寶瑩與宏福苑大火連署發起學生關靖豐接受DW採訪,談他們眼中的香港變化。

中文摘要 「七一」29周年前夕,前社民连主席陈宝莹与宏福苑大火连署发起学生关靖丰接受采访。两代香港人回顾从主权移交、七一游行到反修例及《国安法》时代的历程,探讨7月1日对当下香港人的意义及社会变化。

Iran war live: Qatar’s PM meets US envoys; Tehran holds firm on conditions

Iran says talks on a final deal will not begin until hostilities end in Lebanon and US releases frozen Iranian funds.

中文摘要 卡塔尔总理会见美国特使。伊朗方面态度强硬,表示在黎巴嫩敌对行动结束且美国解冻伊朗资金之前,不会开始关于最终协议的谈判。

Afghan Taliban launch strikes on border with Pakistan as tensions escalate

Pakistan's military says it shot down four rudimentary drones and will respond to any further provocation.

中文摘要 阿富汗塔利班对巴基斯坦边境发动袭击,导致双边紧张局势升级。巴基斯坦军方表示已击落四架简易无人机,并警告将对任何进一步的挑衅作出回应。

US envoys in Doha to meet mediators but not Iranians, Qatar says

Qatar's foreign ministry spokesman says no high-level meetings or direct talks between the US and Iran are scheduled.

中文摘要 卡塔尔外交部发言人表示,美国特使在多哈将与调解人会面,但不会与伊朗官员会面。目前尚未安排美伊之间的高级别会议或直接对话。

Trump made more than $1bn from crypto in first year back in office

The president's crypto income far outpaces his earnings from real estate and Trump-themed items such as watches.

中文摘要 据报道,美国总统特朗普重返白宫首年的加密货币相关收入超过10亿美元,远超其房地产及特朗普主题商品(如手表)等业务的收入。

US Lifts Export Restrictions on Anthropic’s Fable 5

The US government removed foreign access restrictions on Anthropic PBC’s Fable 5 artificial intelligence model, clearing it for wider distribution after the startup resolved the Trump administration’s safety concerns. Bloomberg's Minmin Low breaks down the latest developments. (Source: Bloomberg)

中文摘要 美国政府解除对Anthropic公司Fable 5人工智能模型的出口限制。在解决特朗普政府的安全担忧后,该模型获准更广泛地分发。

Trump Reports $1.4 Billion in 2025 Crypto Earnings

President Donald Trump reported earning at least $1.4 billion in 2025 from crypto and memecoin-related businesses, according to his latest annual financial disclosure. Bloomberg's Kailey Leinz and Romaine Bostick break it down. (Source: Bloomberg)

中文摘要 根据最新年度财务披露文件,美国总统特朗普报告称,其2025年从加密货币及模因币相关业务中获得的收入至少达14亿美元。

India Prop Traders Brace for RBI Funding Squeeze

Traders expect reduced returns from cash-futures arbitrage, options market making and index arbitrage.

中文摘要 受印度央行关于银行担保的新规影响,印度本土自营交易员面临资金收紧压力。交易员预计,现金期货套利、期权做市及指数套利等策略的收益将下降。

White House lifts ban on Anthropic models

US government move allows AI start-up to re-release Mythos and Fable models

中文摘要 美国政府正式解除对人工智能初创公司Anthropic旗下模型的出口禁令,此举允许该公司重新向市场发布其Mythos和Fable系列人工智能模型。

Mizuho Says Historic Slump in Yen Is Defying the Rates Rulebook

The slump in the yen to historic lows is forcing investors to rethink one of their most widely used strategies to gauge the currency’s moves, according to Mizuho Bank Ltd.

中文摘要 据瑞穗银行分析,日元汇率跌至历史低位的剧烈下跌正在打破传统的利率定价规则,这迫使投资者不得不重新审视并调整其用于预测该货币走势的最常用策略。

Goldman Flags Up Oil Surplus Even as Nations Rebuild Stockpiles

The global oil market is set to swing back into oversupply as the impact of the Iran war fades and traffic through the Strait of Hormuz recovers, according to Goldman Sachs Group Inc.

中文摘要 高盛指出,随着伊朗战争影响消退及霍尔木兹海峡交通恢复,全球石油市场将重回供应过剩,尽管各国正在重建库存。

Uber-backed Lime raises $167mn in bike and scooter group IPO

Shares are priced at $25 in the deal underwritten by Goldman Sachs, JPMorgan and Jefferies

中文摘要 Uber支持的Lime在自行车和滑板车业务IPO中筹集1.67亿美元。股票定价为25美元,由高盛、摩根大通和Jefferies承销。

Sixth Street, KKR Provide BridgeBio $1 Billion Preferred Equity

Sixth Street Partners and KKR & Co. have agreed to provide $1 billion of preferred equity to biotechnology company BridgeBio Pharma Inc., according to people with knowledge of the matter.

中文摘要 知情人士透露,Sixth Street Partners和KKR已同意向生物技术公司BridgeBio Pharma提供10亿美元优先股。

Chinese carmakers’ hunger for chips boosts national self-reliance drive

EV makers such as global leader BYD are rushing to increase use of locally developed semiconductors

中文摘要 中国车企对芯片的巨大需求推动了国家自主化进程。全球领先的电动汽车制造商比亚迪等正急于增加使用本土研发的半导体。

Yuan Rally Bets Fade as Options Traders Turn Bullish on Dollar

Bullish bets on the Chinese yuan are fading, with options traders unwinding one of the year’s most crowded trades after the Federal Reserve’s hawkish pivot.

中文摘要 随着美联储转向鹰派,期权交易员正在平仓今年最拥挤的交易之一,看多人民币的押注正在消退,交易员转而看多美元。

Iran War Delivers Massive Swings in Chinese Petrochemicals Trade

The war in Iran has rewired China’s petrochemicals trade, allowing plants to ease their glut of the key building blocks for plastics, rubber and textiles by boosting exports.

中文摘要 伊朗战争重塑了中国石化贸易,使工厂能够通过增加出口来缓解塑料、橡胶和纺织品关键基础材料的产能过剩问题。

该源今日无内容。