microsoft/markitdown
Python · ★ 138,422 · 🍴 9,428 · 📈 3,034 stars today
Python tool for converting files and office documents to Markdown.
中文介绍 微软开发的 Python 格式转换工具,能将 PDF、Word、PPT 等多种文件及办公文档批量转换为 Markdown 格式,方便开发者和技术作者处理文档,集成到知识库或内容管理系统中。
Python · ★ 138,422 · 🍴 9,428 · 📈 3,034 stars today
Python tool for converting files and office documents to Markdown.
中文介绍 微软开发的 Python 格式转换工具,能将 PDF、Word、PPT 等多种文件及办公文档批量转换为 Markdown 格式,方便开发者和技术作者处理文档,集成到知识库或内容管理系统中。
Python · ★ 11,309 · 🍴 1,452 · 📈 945 stars today
Hermes WebUI: The best way to use Hermes Agent from the web or from your phone!
中文介绍 为 Hermes Agent 打造的 Web 用户界面,支持从网页或手机端便捷访问和管理智能体,提供了直观的交互方式,适合需要远程操作和监控 AI 代理的用户。
TypeScript · ★ 23,979 · 🍴 2,143 · 📈 647 stars today
Memory engine and app that is extremely fast, scalable. The Memory API for the AI era.
中文介绍 一款面向 AI 时代的高性能记忆引擎与应用,旨在提供极快的读写速度和可扩展性,通过 Memory API 为 AI 应用实现持久化的长期记忆功能,适用于构建需要上下文记忆的复杂 AI 产品。
Python · ★ 76,834 · 🍴 10,903 · 📈 3,375 stars today
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
中文介绍 基于 AI 大模型的自动化短视频生成工具,只需提供文本或主题,即可一键合成包含配音、字幕和画面的高清视频,大幅降低内容创作门槛,适合自媒体和营销人员快速产出内容。
Python · ★ 58,040 · 🍴 5,634 · 📈 1,486 stars today
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
中文介绍 自适应 Web 爬虫框架,能优雅处理从单次请求到大规模分布式爬取的各类任务,自动应对反爬机制,为数据工程师和研究人员提供了强大、灵活的网页数据采集解决方案。
JavaScript · ★ 32,696 · 🍴 1,781 · 📈 485 stars today
The design language that makes your AI harness better at design.
中文介绍 一套设计语言与系统,旨在提升 AI 工具生成设计的质量和一致性,为设计师和 AI 应用开发者提供指导,确保 AI 产出的视觉内容更专业、更符合品牌调性。
Python · ★ 23,035 · 🍴 2,466 · 📈 249 stars today
Fully automatic censorship removal for language models
中文介绍 针对语言模型的自动化审查移除工具,通过技术手段自动识别并移除模型输出中的安全过滤和审查内容,主要用于研究和测试模型的原始能力及边界。
TypeScript · ★ 19,109 · 🍴 1,428 · 📈 417 stars today
Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more
中文介绍 适用于 Claude Code、Codex、Cursor 等流行 AI 编码工具的官方插件,通过集成额外的技能和工程方法论,增强 AI 助手的代码生成、重构和调试能力,提升开发效率。
Python · ★ 81,772 · 🍴 15,891 · 📈 299 stars today
TradingAgents: Multi-Agents LLM Financial Trading Framework
中文介绍 基于大语言模型的多智能体金融交易框架,通过多个专业化代理(如分析师、交易员)协作,实现市场分析、策略生成和自动执行,旨在探索 AI 在复杂金融决策中的应用。
HTML · ★ 5,148 · 🍴 690 · 📈 524 stars today
A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use.
中文介绍 一种“元技能”框架,能够根据特定领域需求,自动设计、定义和生成由多个专业化代理组成的团队及其所需技能,帮助开发者快速构建和部署领域专用的 AI 代理系统。
C++ · ★ 111,642 · 🍴 25,510 · 📈 77 stars today
Godot Engine – Multi-platform 2D and 3D game engine
中文介绍 功能强大且开源的跨平台游戏引擎,支持 2D 和 3D 游戏开发,提供一体化的编辑器、脚本语言和丰富工具链,是独立开发者和中小团队制作多平台游戏的热门选择。
TypeScript · ★ 9,466 · 🍴 760 · 📈 335 stars today
⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more
中文介绍 一款运行在终端的 AI 编码代理,具备哈希锚定的代码编辑、优化工具集成、语言服务器协议支持以及 Python 和浏览器交互能力,旨在为命令行爱好者提供强大的 AI 辅助编程体验。
Python · ★ 24,246 · 🍴 2,794 · 📈 888 stars today
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
中文介绍 无需分词器的端到端语音合成模型 VoxCPM2,支持多语言语音生成、创意性声音设计和高保真声音克隆,致力于提供更自然、更灵活的高质量文本转语音解决方案。
Jupyter Notebook · ★ 3,718 · 🍴 516 · 📈 861 stars today
A straightforward method for training your LLM, from downloading data to generating text.
中文介绍 一份从零开始训练大语言模型的实践指南,详细讲解了从数据下载、预处理到模型训练和文本生成的完整流程,旨在帮助开发者通过亲手复现来深入理解 LLM 的原理。
Jupyter Notebook · ★ 17,870 · 🍴 5,184 · 📈 93 stars today
Code for Machine Learning for Algorithmic Trading, 2nd edition.
中文介绍 《机器学习与算法交易》第二版的配套代码库,涵盖了从数据获取、特征工程到模型构建和回测的实战代码,为读者学习如何将机器学习应用于量化交易提供了可运行的实例。
Rust · ★ 7,176 · 🍴 298 · 📈 135 stars today
The fastest and the most accurate file search toolkit for AI agents, Neovim, Rust, C, and NodeJS
中文介绍 专为 AI 代理、Neovim 等场景优化的超快、超准文件搜索工具包,支持 Rust、C、NodeJS 等多语言调用,极大提升了在大型代码库中查找文件的速度和准确性。
Markdown · ★ 510,425 · 🍴 48,392 · 📈 1,212 stars today
Master programming by recreating your favorite technologies from scratch.
中文介绍 一个通过从头复现各类流行技术(如数据库、容器、Web 服务器)来深度掌握编程的实践项目集合,为开发者提供了丰富的动手教程,是提升底层系统理解能力的绝佳资源。
该源今日无内容。
Our approach to AI policy and political advocacy, transparency, support for thoughtful regulation and AI safety, and that no outside political group speaks on the company’s behalf.
中文介绍 OpenAI阐述其AI政策与政治倡导方法,强调透明度、支持审慎监管和AI安全,并声明无外部政治团体代表公司发声。
中文介绍 JetBrains发布Mellum2,这是一个拥有120亿参数的混合专家模型。
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and why Grok Imagine is so underrated. For the first time, we do a deep dive with the guy who led it!
中文介绍 xAI的Ethan He分享Grok Imagine在3个月内构建的经验,比较Videogen与世界模型,并探讨视频智能体模型的未来潜力。
中文介绍 IBM研究指出,企业AI的可扩展采用需超越大型语言模型,依赖智能体逻辑。
OpenAI breaks ground on a 1GW data center project in Michigan as part of Stargate, building AI infrastructure to expand access, create jobs, and support communities.
中文介绍 OpenAI在密歇根州破土动工一个1GW数据中心项目,作为Stargate计划的一部分,构建AI基础设施以扩大访问、创造就业并支持社区。
OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new path to build with OpenAI through the AWS environments, controls, and procurement workflows they already use. Customers can get started with OpenAI on AWS and move faster from evaluation to production.
中文介绍 OpenAI的前沿模型和Codex现已在AWS上普遍可用,企业可通过AWS现有环境、控制及采购流程更快速地集成OpenAI技术。
中文介绍 NVIDIA推出Cosmos 3,这是首个用于物理AI推理和行动的开放全模型。
中文介绍 报道Copilot超级应用泄露、Minimax M3发布及Nvidia N1X等AI动态。
a quiet day lets us highlight the new AIE WF focuses
中文介绍 AINews聚焦创始人与前方部署工程师,并介绍新AIE WF的重点领域。
Boston Children’s Hospital uses OpenAI technology to improve patient care, reduce operational burden, and help diagnose more than 40 rare disease cases.
中文介绍 波士顿儿童医院采用OpenAI技术提升患者护理、降低运营压力,并成功辅助诊断逾40例罕见病病例。
How Braintrust engineers use Codex with GPT-5.5 to run experiments and code faster.
中文介绍 Braintrust工程师利用Codex与GPT-5.5将客户需求转化为代码,加速实验与编码进程。
Pope Leo XIV’s new encyclical on artificial intelligence includes a statement that warrants serious attention from technologists and policymakers: “Technology is never neutral.” Magnifica Humanitas (“Magnificent Humanity”) is a clarion call to all people to act with courage and solidarity as we ente
中文介绍 教皇Leo XIV发布AI通谕《Magnifica Humanitas》,指出“技术从未中立”,并呼吁技术人员和政策制定者关注AI时代的行动。
**Anthropic** rolled out **Claude Opus 4.8**, which shows incremental improvements but mixed benchmark results, including better cooperation and coding behavior but some regressions in document parsing. Platform updates include mid-conversation system instructions enhancing long agent sessions, thou
中文介绍 Anthropic发布Claude Opus 4.8,该版本在合作与编码行为上有所改进,但文档解析出现退化,基准测试结果不一。平台更新引入对话中系统指令。
OpenAI launches Rosalind Biodefense, expanding trusted access to GPT-Rosalind for vetted developers and U.S. government partners advancing biodefense, public health, and pandemic preparedness through frontier AI.
中文介绍 OpenAI启动Rosalind Biodefense计划,扩大经审核开发者及美国政府合作伙伴对GPT-Rosalind的访问,以推进生物防御、公共卫生和疫情准备。
Total Anthropic victory!
中文介绍 Anthropic完成9650亿美元H轮融资,并发布Opus 4.8及动态工作流和ultracode。
第一作者: Davis Brown · 方向: 软件安全
分布式攻击代理监控状态在线监控实时聚类网络安全
Abstract:Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection, attackers frequently distribute their misuse, splitting a harmful task across many user accounts so each individual transcript looks benign. Because safety monitors score only one agent context at a time, they are structurally blind to misuse that is only visible in aggregate, across many accounts. We show this gap is real by building, to our knowledge, the first distributed agent attack, a multi-agent scaffold that completes hard cybersecurity tasks while hiding the harmful objective across subagents with limited contexts, evading a standard monitor that catches it only a fifth as often as prior agent attacks. Towards a defense, we develop an online stateful monitor that uses real-time clustering to collect weak suspiciousness...
论文介绍 该研究针对分布式代理攻击规避安全检测的问题。攻击者将恶意任务分散到多个账户,使单个上下文看起来良性,而现有监控器只能分析单个上下文,无法跨账户检测。作者构建了首个分布式代理攻击示例,并开发了一种在线状态监控器,通过实时聚类收集弱可疑信号来识别此类攻击,这有助于提升代理系统的安全监控能力。
第一作者: Sunday Ajayi · 方向: 隐私保护
移动货币USSD自动化生物识别安全无障碍访问隐私保护
Abstract:Financial inclusion has expanded significantly across Africa through mobile money services delivered primarily via USSD technology. However, visually impaired individuals continue to face accessibility and security barriers when conducting financial transactions. Current USSD systems are not designed for non-visual interaction, forcing users to rely on third-party assistance even for PIN entry, thereby increasing fraud exposure and reducing transaction confidence. Although alternative assistive technologies such as screen readers exist, they are not compatible with USSD operations, often causing sessions to time out before the user can complete a transaction. This paper presents an Android-based intelligent middleware that automates USSD transactions, integrates biometric-secured PIN injection, and introduces a privacy-preserving screen-dimming mechanism: Blackout Mode. The...
论文介绍 本文解决视障用户在使用USSD进行移动货币交易时面临的无障碍和安全障碍。现有USSD系统设计不支持非视觉交互,导致用户依赖第三方,增加欺诈风险。作者提出一个Android智能中间件,自动化USSD事务,集成生物识别安全PIN注入和隐私保护屏幕调暗机制,以增强交易的安全性和可访问性。
第一作者: Fabio De Gaspari · 方向: AI 安全
加密数据分类压缩数据多模态学习网络安全数据片段
Abstract:Reliable identification of encrypted data fragments is essential in cybersecurity, with applications to ransomware detection, digital forensics, and large-scale data analysis. Distinguishing encrypted from compressed fragments is particularly challenging, as short fragments lack structural data and exhibit low statistical redundancy. Traditional statistical methods based on byte-level distributions show limited effectiveness on this task. Recent machine learning approaches improve performance by learning subtle patterns from raw bytes, but predominantly rely on single-modal representations, implicitly assuming that a single view of the data is sufficient for accurate classification. This paper shows that this assumption becomes a fundamental limitation in low-information settings, when only small fragments of data are available (512--2048 Bytes). We propose Triumvir, a...
论文介绍 可靠识别加密数据片段在网络安全中至关重要,但区分加密与压缩片段尤其困难,因为短片段缺乏结构。传统统计方法效果有限,而现有机器学习方法依赖单模态表示。本文提出Triumvir,一种多模态分类方法,在低信息设置中表现更优,适用于勒索软件检测和数字取证等应用。
第一作者: Dominik Roy George · 方向: 密码学协议
MUDThread网络访问控制IoT安全网格网络
Abstract:The IETF standard Manufacturer Usage Description (MUD) enables manufacturers to equip IoT devices with certified URLs that provide traffic profiles for those devices, helping administrators enforce network access control. However, MUD assumes devices operate on full IP stacks and therefore does not account for constrained IoT devices running Thread--the dominant low-power mesh networking standard--which lacks complete TCP/IP functionality. While prior work proposes extensions to support MUD in Thread environments, these approaches are limited to simple topologies with a single border router and do not scale to realistic deployments with multiple, heterogeneous border routers. We introduce MeshGuard, a framework enabling MUD-based access control in complex Thread networks, with any number of border routers. MeshGuard extends the Mesh Link Establishment (MLE) protocol to deliver...
论文介绍 MUD标准帮助管理员执行IoT设备的网络访问控制,但Thread网络缺乏完整TCP/IP支持。现有扩展不适用于复杂拓扑。MeshGuard框架通过扩展MLE协议,使MUD能在具有多个异构边界路由器的大规模Thread网络中实现访问控制,提升IoT网络的安全性。
第一作者: Ransika Gunasekara · 方向: 密码学协议
加密流量分析协议无关元学习时间序列少样本学习
Abstract:Traditional traffic analysis is being fundamentally challenged by the rapid adoption of encryption, tunnelling, and privacy-preserving protocols, which increasingly obscure packet payloads and limit the usefulness of Deep Packet Inspection (DPI). Although machine learning has advanced encrypted traffic analysis, existing approaches often remain tied to protocol-specific header features, depend on large labelled datasets, and degrade when deployed across heterogeneous network environments. We present GETA, a protocol-agnostic framework for encrypted traffic analysis that models network flows as multivariate time series using only traffic metadata, thereby avoiding reliance on packet payloads or header semantics. GETA combines meta-learning, embedding refinement, and self-attention to support few-shot adaptation to previously unseen domains with minimal labelled data. Across...
论文介绍 加密流量分析因协议限制和数据标签不足而挑战重重。GETA是一个协议无关框架,将网络流建模为多元时间序列,使用元学习、嵌入细化和自注意力技术,支持少样本适应到新环境。这减少了对协议特定特征和大量标签数据的依赖,增强在异构网络中的部署能力。
第一作者: Ziqing Yang · 方向: AI 安全
后门攻击提示学习骨干模型AI安全双层优化
Abstract:Prompt learning is a new machine learning paradigm that has attracted ample attention due to its simplicity and proven efficacy. Despite its growing adoption, the security vulnerabilities associated with this paradigm remain underexplored. In this work, we take the first step to propose BadBone, a stealthy and adaptive backdoor attack against prompt learning using bi-level optimization. Instead of backdooring the prompt learning process, we aim to compromise a backbone model such that only target downstream tasks employing prompt learning inherit the backdoor vulnerability. Extensive experiments on three different models and three datasets from various domains show that our targeted/untargeted backdoored models achieve high attack performance while maintaining utility on both pre-training and downstream tasks. Moreover, we evaluate our approach against six state-of-the-art...
论文介绍 提示学习虽有效,但其安全漏洞未充分探索。BadBone是首个针对视觉提示学习中骨干模型的后门攻击。通过双层优化,攻击使骨干模型被破坏,从而只有使用提示学习的下游任务继承后门漏洞。实验表明该攻击高效且隐蔽,揭示了提示学习范式的安全风险。
第一作者: Zekeri Adams · 方向: 网络安全
动态恶意软件分析本体论MAECSTIX模块化本体
Abstract:Capturing dynamic malware behavior in a practical but still semantically precise manner remains a significant challenge in cyber threat intelligence. While standards such as MAEC and STIX provide widely adopted vocabularies for describing malware artifacts and observations, they represent data with considerable complexity in structures that often obscure important ontological distinctions. In particular, they tend to conflate enduring malware artifacts with the events generated during execution, thereby flattening distinctions that are central in foundational standards for ontology design. In this paper, we conduct a foundational ontological analysis of core MAEC and STIX constructs relevant to dynamic malware analysis relying on Unified Foundational Ontology (UFO) as a theoretical lens. Our analysis reveals some ontological mismatches arising from the conflation of artifacts...
论文介绍 捕获动态恶意软件行为在威胁情报中重要,但现有标准如MAEC和STIX存在本体混淆。本文基于统一基础本体进行分析,揭示了结构问题,并提出MAECO-Lite,一个模块化本体,以更精确地表示恶意软件行为事件和产物,支持更清晰的本体区分。
第一作者: Yu Li · 方向: AI 安全
黑盒防御大语言模型共同进化经验引导网络安全
Abstract:Large Language Models (LLMs) remain highly vulnerable to diverse attacks, particularly in black-box settings where the internals of target models are inaccessible. Existing black-box defenses typically rely on pre-defined filtering heuristics, which often fail to generalize to unseen attack types and target model architectures. We introduce EvoDefense, an experience-guided co-evolving black-box defense paradigm. EvoDefense employs a guard LLM to detect malicious queries and an experience memory module to accumulate defense knowledge from previous interactions. At the core of EvoDefense is a continuous attack-defense evolution loop, where an attack generator and the guard model iteratively refine their attack strategies and defense policies through experience-guided optimization. This design enables EvoDefense to generalize across unseen attacks and target models without...
论文介绍 大语言模型在黑盒设置下易受攻击,现有防御方法泛化能力差。EvoDefense引入经验引导的共同进化范式,使用守卫LLM检测恶意查询,并通过经验记忆模块积累知识。攻击生成器和守卫模型在持续循环中优化策略,使防御能适应未见攻击和目标模型。
第一作者: Tianhe Lu · 方向: 软件安全
Java安全API误用大语言模型代码安全外部知识
Abstract:The misuse of Java security APIs is a serious security problem in software development. Research in 2024 has shown that this problem is widespread in LLM-generated code. However, it remains unclear whether this phenomenon persists in current models and how external security knowledge affects it. This paper presents a scoped replication and extension of Mousavi et al.'s study on the Java Cryptography Architecture (JCA) and Java Secure Socket Extension (JSSE) APIs. We focus on two complementary settings: GPT-5.5 as a frontier proprietary coding model, and Llama-3.3-70B-Instruct as a strong open-weight model relevant to self-hosted deployment. The results show that although newer LLMs perform better in using Java security APIs, the problem of Java security API misuse has not been eliminated. External security knowledge substantially improves the measured outcome, but its effect...
论文介绍 本文对大语言模型(LLM)生成代码中滥用 Java 安全 API(JCA/JSSE)的现象进行了范围复制与扩展研究。通过评估前沿专有模型 GPT-5.5 和强开源模型 Llama-3.3,发现尽管新模型表现有所改善,但 API 滥用问题并未消除。研究表明,引入外部安全知识能显著提升模型的安全 API 使用表现,但效果受限于知识的充分性。
第一作者: Jiejun Tan · 方向: AI 安全
AI智能体提示注入木马后门安全基准多步攻击
Abstract:LLM agents are evolving from conversational chatbots to operational tools in real-world workspaces. In local agentic harnesses, an LLM can read and write files, call tools, and reuse workspace state across sessions. While such capabilities enhance utility, they also expose a new attack surface for attackers. Attackers can embed a prompt injection within a file or tool output. Agents may read this hidden instruction, store it, and execute it later. In this multi-step trojan attack paradigm, no individual step appears malicious on its own, but these steps can collectively turn untrusted text into persistent control content. However, existing defenses often inspect each step in isolation. As a result, they can block a clear harmful action, but fail to detect the earlier write operation that plants the backdoor. To reveal this threat, we introduce ClawTrojan, a benchmark designed...
论文介绍 随着LLM智能体具备读写文件、调用工具和持久化状态的能力,攻击者可通过多步操作将隐藏的提示注入植入工作区,形成长期控制的木马。本文揭示了这种将非受信文本转化为持久控制内容的攻击范式,并提出了 ClawTrojan 基准测试,旨在系统评估现有防御措施在此类多步、持续性威胁下的有效性。
第一作者: Henrique B. Brum · 方向: 网络安全
TLS安全加密协议网络安全配置分析出站连接
Abstract:Despite the widespread use of Transport Layer Security (TLS), its security guarantees are frequently compromised by outdated versions and misconfigurations. To analyze this problem, we collected more than 50 million TLS handshakes over a two-week period at our research institution, Fondazione Bruno Kessler, and analyzed three server-selected parameters against the recommendations of four TLS guidelines. Our analysis shows that while the use of insecure or outdated options is minimal, it remains persistent. More importantly, servers are adopting the latest TLS advancements much faster than official guidelines can be updated to provide directives for them. These findings, combined with the difficulty of configuring TLS clients due to their ephemeral, ubiquitous and server-dependent nature, leave users vulnerable to non-standard or outright insecure connections. To address this...
论文介绍 本文通过对研究机构内超过5000万次TLS握手的分析,发现尽管不安全或过时的配置比例较低,但其依然持续存在。更重要的是,服务器采纳最新TLS技术的速度远快于安全指南的更新速度。这种现状,结合TLS客户端配置的固有困难,使得用户容易受到非标准或不安全连接的威胁。研究提出了一个网关系统来帮助用户强制执行安全策略。
第一作者: Shengchen Ling · 方向: 密码学协议
x402协议机器支付安全分析原子性缺陷并发竞争
Abstract:The agentic economy demands programmatic financial rails, positioning the x402 protocol as the de facto standard for machine-to-machine payments. However, bridging synchronous HTTP requests with asynchronous blockchain finality introduces profound state synchronization challenges. In this work, we perform the first comprehensive security analysis of the x402 ecosystem. By formalizing five Security Invariants, we reveal that current implementations fail to enforce transactional atomicity and cryptographic context binding, leading to systemic vulnerabilities. We identify a semantic gap in signature design enabling cross-resource substitution, where payment proofs are transplanted to other unauthorized contexts. Furthermore, we expose a temporal gap where concurrency race conditions allow probabilistic service duplication. In the AI inference domain, we demonstrate how dynamic...
论文介绍 x402协议旨在为机器间支付提供标准化程序化接口。本文首次对其生态系统进行了全面的安全分析,通过形式化五个安全不变量,揭示了现有实现在交易原子性和密码学上下文绑定上的失败。研究发现了签名设计中的语义漏洞(允许支付证明跨资源复用)以及并发条件下的时序漏洞(导致概率性的服务重复),这些漏洞可能被利用进行“免费搭便车”攻击。
第一作者: Wanju Kim · 方向: 网络安全
虚拟化混淆自动化反混淆恶意软件分析语义提取VMPredator
Abstract:Virtualization obfuscation is a more powerful obfuscation technique compared to other obfuscation methods, and as it is increasingly being applied to malware, it demands significant effort and time from analysts. This study analyzes virtualization obfuscation and proposes a tool called VMPredator that automatically extracts semantic units. The proposed tool performs various analyses including memory analysis and trace analysis, while minimizing dependency on the specific internal structure of virtual machines in order to handle diverse forms of virtualization obfuscation that existing tools are unable to process. Experimental results demonstrate that the length of obfuscated programs was reduced by approximately 85%, and it was verified through validation that small-scale programs were fully restored to semantics identical to the original.
论文介绍 虚拟化混淆是一种强大的代码保护技术,也常被用于恶意软件,给分析人员带来巨大挑战。本文提出了一个名为 VMPredator 的自动化工具,通过内存分析和执行轨迹分析,提取混淆程序的语义单元。该工具旨在最小化对特定虚拟机内部结构的依赖,从而处理更多样的虚拟化混淆形式。实验显示,经处理后混淆程序的长度可减少约85%,并能将小程序完全恢复到原始语义。
第一作者: Churui Zeng · 方向: AI 安全
LLM智能体逃狱攻击任务分解安全威胁TRACE框架
Abstract:The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplored and underestimated since (i) safety alignment prevents LLMs from directly generating harmful instructions, and (ii) most existing jailbreak methods cannot consistently induce agents to execute malicious operations. In this paper, we propose TRACE, a practical agentic jailbreaking framework to further reveal the risks of this threat surface. To conceal the malicious intent, TRACE decomposes a malicious task into multiple subtask sequences under different schemes and selects the sequence with the fewest explicitly harmful subtasks. TRACE then disguises the remaining harmful subtasks as benign-looking instructions by embedding them in task-aware scenarios with related roles, environments...
论文介绍 LLM智能体的规划与执行能力带来了新的安全威胁。现有逃狱方法难以稳定诱导智能体执行恶意操作。本文提出了TRACE框架,通过将恶意任务分解为多个子任务方案,并选择显式有害子任务最少的序列。剩余有害子任务被伪装成嵌入在特定场景和角色中的良性指令,从而绕过安全对齐。该框架旨在揭示智能体环境下的安全风险。
第一作者: Ziwen Li · 方向: 软件安全
文本匿名化隐私保护大语言模型效用保持再识别风险
Abstract:Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-identification, yet those same details also carry downstream analytic value of the text. Existing defenses either remove explicit identifiers, perturb text for formal privacy, or test rewritten text against non-web inference models, leaving underexplored the operating region between resistance to agentic web-search re-identification and utility retention. We introduce AURA (\textbf{A}nonymization with \textbf{U}tility-\textbf{R}etention \textbf{A}daptation), an LLM-powered \textit{mask-reconstruct} framework that decouples privacy localization from utility-preserving reconstruction and selects candidates with adversarial privacy and utility-retention checks. We evaluate AURA on real-user interview transcripts using re-identification...
论文介绍 具备网络搜索能力的智能体LLM改变了文本匿名化的威胁模型,弱上下文线索可能成为再识别的证据。现有防御方法未能充分平衡对抗智能体再识别与保留文本效用。本文提出了AURA框架,这是一个由LLM驱动的“掩码-重建”系统,将隐私定位与保效用重建解耦,并通过对抗性检查筛选候选方案,以在隐私保护和效用保留之间取得更好的权衡。
第一作者: Shuhao Zhang · 方向: AI 安全
提示注入自适应防御检测器分配安全基准SCOUT框架
Abstract:Prompt-injection detectors are heterogeneous: each is strong on a different slice of attacks, and none is always reliable. Yet existing systems still treat detection as a fixed single-detector pipeline, committing every request to one detector's blind spots. We reframe defense as detector allocation: given a heterogeneous pool, decide per request which detectors to run and whether to escalate to an LLM judge. Our framework SCOUT (Scalable and Controllable Outcome-prediction for Uncertainty-aware Triage) makes this decision dynamic by predicting each detector's per-sample reliability and latency from how it behaved on similar past inputs, and exposes a single safety-utility threshold to the operator (where utility bundles benign-pass rate and wall-clock). To evaluate this setting, we build SCOUT-450, a benchmark that captures the structurally complex, agent-facing injections...
论文介绍 现有的提示注入检测器各有所长,但无一是全能的。然而,当前系统通常采用固定的单一检测器流水线。本文将防御问题重新定义为检测器分配问题:为每个请求动态决定运行哪些检测器,以及是否需要升级到LLM裁判。提出的SCOUT框架通过预测每个检测器在每个样本上的可靠性和延迟来动态决策,并暴露一个统一的安全-效用阈值供操作员调整。
第一作者: Fengyu Gao · 方向: AI 安全
差分隐私偏好对齐大语言模型数据合成隐私保护
Abstract:Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-training on real human preference data raises privacy concerns, as these datasets often contain sensitive user prompts and human judgments. To address this, we propose DPPrefSyn, a novel algorithm for generating differentially private (DP) synthetic preference data to enable privacy-preserving preference alignment. DPPrefSyn is a principled framework grounded in the Bradley-Terry preference model and the intrinsic geometric structure of pairwise human preference data. It first learns an underlying preference model from private data with formal differential privacy guarantees, and then leverages the learned model together with public prompts to synthesize high-quality preference data. It exploits the shared linear structure of...
论文介绍 针对大语言模型偏好对齐过程中使用真实人类偏好数据引发的隐私泄露问题,本文提出了DPPrefSyn框架。该方法基于Bradley-Terry模型,在提供严格差分隐私保证的前提下从敏感数据中学习潜在偏好模型,并利用该模型和公开提示合成高质量的隐私保护偏好数据。其核心是利用成对偏好数据的内在几何结构来生成合成数据,旨在实现无需直接访问原始敏感数据的模型对齐。
第一作者: Ian Dardik · 方向: 系统安全
系统安全分析STPA自动化形式化方法不安全控制行为
Abstract:The System-Theoretic Process Analysis (STPA) is a well-established hazard analysis technique that has been applied to a wide range of safety-critical systems. Despite its popularity, there is relatively little automation support for STPA, and most of its steps are carried out manually by a human analyst, which can be time consuming and error prone. This paper investigates the potential use of model-based engineering and formal methods to assist human analysts in efficiently and accurately carrying out STPA. The proposed tool, called FASR (Formalizing and Automating STPA with Robustness), enables automated, complete identification of unsafe control actions (UCAs), leveraging recent advances in robustness analysis to identify UCAs as undesirable deviations in the controller's actions. The use of the tool is demonstrated on a case study involving a Braking System Control Unit...
论文介绍 系统理论过程分析(STPA)是一种广泛使用的危害分析技术,但其步骤通常依赖人工分析,效率低且易出错。本文提出了FASR工具,旨在利用基于模型的工程和形式化方法来辅助分析。该工具通过鲁棒性分析技术,实现了对不安全控制行为(UCAs)的自动化、完整识别。研究在制动系统控制单元案例上展示了该工具的应用,旨在提升安全分析的效率和准确性。
第一作者: Wenjie Jacky Mo · 方向: 安全研究
安全防护栏大语言模型威胁分诊路由专家框架鲁棒性评估
Abstract:Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains challenging because safety risks span heterogeneous threat domains, while existing datasets cover only fragmented risk subsets and rely on inconsistent taxonomies. Consequently, it remains unclear whether current guardrails can generalize beyond narrow evaluation settings. To better understand the robustness of guardrail models, we first introduce GuardZoo, a unified human-annotated benchmark with 32,460 samples covering 15 distinct unsafe categories. Evaluation on GuardZoo reveals that monolithic guardrails suffer from task interference: different threat domains require distinct decision boundaries that are difficult to compress into a single model. We therefore propose RouteGuard, a router-expert framework that triages each...
论文介绍 构建鲁棒的安全防护栏对于部署大语言模型至关重要,但现有数据集覆盖有限且分类不一致。本文引入了GuardZoo统一评估基准,揭示了单一防护栏模型在多威胁域间存在任务干扰问题。为此,研究提出了RouteGuard框架,它通过一个路由器将输入分派给不同的专家模型,对威胁进行分诊处理,旨在提升防护栏在多样化现实场景中的泛化能力和鲁棒性。
第一作者: Mohammadreza Rashidi · 方向: AI 安全
提示注入ReAct智能体工具调用安全攻击面分析
Abstract:ReAct agents that interleave chain-of-thought reasoning with tool calls are increasingly deployed for real tasks such as scheduling, file retrieval, and data access. Their tool observation loop creates a direct attack surface: an adversary who controls any tool's return value can embed instructions that redirect the agent away from the user's goal, a threat known as indirect prompt injection. Existing benchmarks evaluate attack success rate (ASR) at a fixed injection position under fixed conditions, leaving three risk dimensions unexplored: where in the tool sequence the payload appears (injection depth), what rhetorical register it uses (framing), and how many turns the agent is permitted (turn cap). We conduct four controlled studies on 20 scenarios spanning five attack categories, totalling 460 trials against GPT-4o-mini and Claude Haiku at a combined API cost under 0.36...
论文介绍 ReAct智能体的工具调用机制带来了间接提示注入的攻击面。现有研究多在固定条件下评估,忽略了注入深度、提示表述和轮次限制等关键风险维度。本文通过四个对照实验,系统研究了攻击载荷在工具序列中的位置(深度)、所使用的修辞手法(表述)以及智能体可用回合数(轮次预算)对攻击成功率的影响,旨在更全面地揭示此类攻击的风险特征。
第一作者: Brian Crawford · 方向: 软件安全
软件逆向工程AI智能体提示注入防御攻击检测混淆
Abstract:Agentic software reverse engineering systems are vulnerable to prompt injection attacks placed into the source code of executable binary files. This research demonstrates defensive tactics for detecting the presences of prompt injection strings in the decompiler output of adversarial example programs. Methods for obfuscating these attacks and subsequent methods for defending against these obfuscations are also explored. This research advances the understanding of risk and security of agentic software analysis systems necessary for their deployment into production-level cyber workflows.
论文介绍 基于智能体的软件逆向工程系统容易受到嵌入在二进制文件源代码中的提示注入攻击。本文研究了检测反编译器输出中提示注入字符串的防御策略,并进一步探索了对此类攻击进行混淆以及针对混淆的后续防御方法。该研究旨在增进对智能体软件分析系统在生产级网络工作流中安全部署所需风险与安全性的理解。
第一作者: Brian Crawford · 方向: 软件安全
软件逆向工程大语言模型对抗攻击遗传算法恶意软件分析
Abstract:Software tools for reverse engineering executable binary files, such as Ghidra, enable malware analysts to safely conduct robust static analysis without having access to original source code. Coupled with the analytic power of large language models (LLM), agentic systems enabled with tools, such as GhidraMCP, can allow analysts to automate a previously human driven process. Although this automation can increase the productivity of a single malware analyst, it also introduces a new area of vulnerability for malware obfuscation. This paper presents an adversarial technique using genetic algorithm-based prompt generation, a modification of an adversarial attack known as AutoDAN, to demonstrate the ability to deceive LLM-powered disassembly and decompilation systems into misinterpreting binary executables, effectively corrupting their analytical output. This proof-of-concept...
论文介绍 大语言模型增强了软件逆向工程工具的能力,但也引入了新的安全脆弱性。本文提出了一种基于遗传算法的对抗性提示生成技术,该技术修改了AutoDAN攻击方法,旨在欺骗由大语言模型驱动的反汇编和反编译系统,使其误解二进制可执行文件,从而破坏其分析输出。这证明了此类自动化系统在对抗性输入面前的潜在脆弱性。
第一作者: Qingwen Zeng · 方向: AI 安全
可信AI金融科技AI安全对抗性机器学习生命周期
Abstract:Artificial intelligence is now embedded as a primary decision engine in continuously operated financial AI pipelines spanning training and updating, deployment and inference, and operation with monitoring and feedback. The automation and scale that make these pipelines effective also create novel attack surfaces, where small algorithmic perturbations can amplify into persistent, system-level financial harm. Existing surveys, however, either treat AI as a defensive tool or analyse adversarial machine learning in a domain-agnostic manner, abstracting away finance-specific constraints such as accounting plausibility, non-IID federated data, continuous retraining, and automation-amplified downstream effects. We address this gap with a unified, lifecycle-centric and mechanism-driven framework. We partition financial AI into three lifecycle stages: training and updating, deployment...
论文介绍 人工智能已深度嵌入金融AI流程的各个阶段,但其自动化和规模化也带来了新的攻击面。现有综述往往将AI仅视为防御工具或从领域无关角度分析对抗性机器学习。本文提出了一个统一、以生命周期为中心且机制驱动的框架,将金融AI划分为训练与更新、部署与推理、以及监控与反馈三个阶段,并聚焦于金融特定约束(如会计合理性、非独立同分布数据)下的安全风险分析。
第一作者: Lingfeng Yao · 方向: 安全研究
音频水印扩散模型黑盒攻击知识产权保护水印移除
Abstract:With the rise of AI-generated audio, watermarking has become widely used for detecting misuse and protecting intellectual property. However, adversaries may try to remove these watermarks, making it critical to evaluate how well watermarking schemes withstand removal attacks. Existing attacks are often impractical: they either noticeably degrade perceptual quality or require access to the watermarking scheme. We propose DiffErase, a black-box watermark removal attack that assumes no knowledge of the target watermarking scheme while maintaining perceptual quality. DiffErase perturbs watermarked audio to an intermediate diffusion noise level and regenerates it using a pretrained denoising model, effectively suppressing watermark signals. Theoretical analysis and extensive experiments demonstrate that inaudible audio watermarks are highly vulnerable: across multiple audio...
论文介绍 音频水印技术用于检测滥用和保护知识产权,但其抗移除攻击的能力有待评估。现有攻击方法通常需要知晓水印方案或会明显降低音频质量。本文提出DiffErase,一种黑盒音频水印移除攻击方法。它在不了解目标水印方案的情况下,通过将带水印音频扰动至中间扩散噪声水平,并使用预训练去噪模型进行再生,有效抑制水印信号,同时保持良好的感知质量。
第一作者: Ryan Fahey · 方向: AI 安全
提示缓存时序攻击API网关LLM安全数据泄露
Abstract:Over the past year, prompt caching in Large Language Models (LLMs) has become increasingly more popular across inference APIs. Prompt caching helps save precious compute resources and speeds up response times by reusing parts of the KV cache of a specific prompt for another request. However, many implementations of prompt caching are not secure against timing attacks or even basic metadata disclosure. Gu et al. (ICML 2025) develop a method to audit prompt caching in LLMs. This paper investigates whether OpenRouter's API gateway architecture introduces prompt caching vulnerabilities that bypass provider-level prompt cache isolation guarantees. Most LLM inference providers implement per-account or per-organization prompt caching to prevent data leaks, but does routing through OpenRouter with shared organizational credentials inadvertently create global cache sharing across all...
论文介绍 本文研究了大语言模型(LLM)推理API中提示缓存的安全性。它聚焦于审计OpenRouter等API网关架构,探究在使用共享组织凭证进行路由时,是否会无意中绕过提供商级别的提示缓存隔离保障,从而在所有用户间创建全局缓存共享,引发潜在的数据泄露风险。
第一作者: George Fatouros · 方向: AI 安全
LLM智能体运行时架构网络安全合规安全上下文
Abstract:Regulated cybersecurity workflows lack a runtime substrate that enforces organization-level scope across retrieval, tool calls, memory, findings, reports, and audit while remaining model-agnostic and locally deployable. Recent large language model (LLM) agent systems report strong results on isolated cybersecurity tasks, yet they do not by themselves define an auditable platform architecture for regulated security operations centre (SOC) and compliance workflows, where a single analyst may trigger actions that bind the organization, and where the runtime must integrate with existing SIEM/XDR stacks as a primary source of context and alert-driven triggers rather than operate as a standalone analytical layer. This paper proposes an organization-scoped LLM agent runtime architecture for financial cybersecurity. The contribution is a typed Security Context that is created at every...
论文介绍 针对受监管的金融网络安全运营,本文提出了一种组织范围的大语言模型(LLM)智能体运行时架构。该架构旨在强制执行涵盖检索、工具调用、内存和审计的全流程组织级别安全范围,支持模型无关和本地部署,并能与现有SIEM/XDR系统集成,以满足安全运营中心(SOC)的合规性需求。
第一作者: Xiaoyong · 方向: 软件安全
对抗补丁场景鲁棒性评估框架AI视觉安全操作包络
Abstract:Adversarial patches are physical patterns attached to real objects to mislead AI vision systems. Their real-world risk is not determined by a single successful prediction, but by whether they remain effective after deployment under changing viewpoints, distances, and scene conditions. We refer to this property as scene robustness, the effectiveness of a deployed patch across conditions in a real environment. Yet existing evaluations do not measure scene robustness well: real image benchmarks are realistic but fixed, while simulators are controllable but not grounded in a specific real scene. We present AdvScene, a scene-grounded framework for measuring the scene robustness of adversarial patches in reconstructed real environments. AdvScene reframes evaluation as operational measurement: given a fixed deployed patch, it characterizes the patch's operational envelope - where and...
论文介绍 本文指出,现有对抗补丁的评估方法无法有效衡量其在真实环境变化下的“场景鲁棒性”。为解决此问题,研究者提出了AdvScene框架,它通过在重建的真实场景中部署固定补丁,来系统性地表征补丁的有效作用范围,从而提供更贴近实际风险的评估。
第一作者: Nima Dorzhiev · 方向: AI 安全
提示注入攻击动态分隔符多态提示组装LLM安全
Abstract:Polymorphic Prompt Assembling (PPA) defends LLM agents against prompt injections by randomly selecting separator pairs from a fixed pool to isolate user input from system instructions. Although effective, static pool reuse exposes a blast-radius vulnerability: once a separator leaks, it can be exploited in future requests. We propose a dynamic per-request separator generation using domain-separated SHA-256 digests keyed on the timestamp, session identifier, and cryptographic nonce. Each assembled prompt receives a unique (BEGIN, END) canary pair, thereby limiting leakage exposure to a single request. We evaluated our extension against 16 injection payloads on Llama-3.3-70B-Instruct-Turbo, with cross-model validation on DeepSeek-V4-Flash model. Against the M1 obfuscation payload (leetspeak + urgency), the dynamic mode reduces the Attack Success Rate (ASR) from 0.88 to 0.38...
论文介绍 多态提示组装(PPA)使用固定分隔符池防御LLM的提示注入攻击,但存在静态池泄漏风险。本文提出为每次请求动态生成唯一的分隔符对,利用时间戳、会话ID和加密随机数生成域分离的SHA-256摘要。评估表明,该方法能有效降低特定混淆攻击的攻击成功率。
第一作者: Shifat E Arman · 方向: AI 安全
提示注入攻击表面工具描述LLM代理安全评估偏差
Abstract:Tool-augmented LLM agents are vulnerable to prompt injection: a third party who controls part of the agent's context can plant instructions that the agent then executes as if they came from the user. Current evaluations report a single attack success rate per model on one channel, the tool output and treat that number as the model's vulnerability. But tool descriptions, which the agent reads at every turn before any tool is called, are themselves an injection surface that the attacker can choose instead. We hold the injection payload byte-identical and deliver it through both surfaces across 13 LLMs from six families and four task suites. The same bytes invert in success rate across models: GPT-4.1 is 96 percent vulnerable on tool outputs but only 4 percent on tool descriptions, while GEMINI-3-FLASH shows the mirror pattern at 20 percent and 98 percent. A variance...
论文介绍 本文揭示了工具增强型LLM代理的一个关键评估盲点:相同的提示注入负载,通过工具输出与工具描述两个不同表面注入,其成功率在不同模型间差异巨大。这表明,仅测试单一通道会严重低估或误判模型的脆弱性,为安全评估提供了新的视角。
第一作者: Yifan Liao · 方向: AI 安全
对抗攻击深度伪造检测歌声合成自监督学习黑盒攻击
Abstract:Recent Singing Voice Synthesis (SVS) advances enable highly realistic but potentially malicious AI covers, making singing voice deepfake detection (SVDD) crucial. Self-Supervised Learning (SSL)-based detectors achieve state-of-the-art performance by fine-tuning speech SSL backbones to capture singing-specific spoof artifacts. Existing adversarial attacks often fail against SSL-SVDD, creating a false impression of inherent robustness. We reveal this stems from two challenges. First, at the objective level, attacks optimize cross-entropy on local surrogates, crossing surrogate-specific boundaries rather than suppressing shared spoof evidence. Second, at the method level, attacks follow the surrogate's dominant gradient direction. In SSL-SVDD, this aligns with fine-tuned artifact-sensitive directions, limiting transferability to unseen detectors - a geometric failure we term the...
论文介绍 针对基于自监督学习的歌声深度伪造检测器,本文研究了其为何对现有对抗攻击表现出“虚假鲁棒性”。研究揭示了攻击在目标函数和方法论上的局限性,并提出了一种通过探索模型流形绕道来生成更具迁移性扰动的新方法,旨在更全面地评估检测系统的安全性。
第一作者: Maksuda Bilkis Baby · 方向: 软件安全
凭证泄漏源代码安全三分类CodeBERT误报率
Abstract:Credential leakage in public source code repositories poses a critical security threat, with over 23.8 million secrets exposed in 2024 alone. Existing detection tools suffer from high false-positive rates because rigid pattern matching and binary classification schemes fail to distinguish genuine credentials from placeholder or weak credentials. We propose a three-class classification framework that explicitly models placeholder or weak credentials as a distinct class, leveraging CodeBERT-based semantic understanding combined with character-level pattern recognition. We evaluate our approach on a newly constructed dataset of 9,426 samples spanning 10 programming languages. Our model achieves a Matthews Correlation Coefficient of 0.86 and a macro F1-score of 0.90, achieving 93% recall and 89% precision for genuine credential leaks while reducing high severity alerts by 33.0%...
论文介绍 为降低公共代码仓库中凭证泄漏检测的高误报率,本文提出了一个三分类框架。该框架结合CodeBERT的语义理解和字符级模式识别,将占位符或弱凭证作为独立类别进行区分。实验结果显示,该方法在保持对真实凭证高召回率的同时,显著减少了高严重性警报。
第一作者: Alexandru Gheorghiu · 方向: 安全研究
伪纠缠量子电路纠缠熵量子复杂性量子学习
Abstract:We construct a family of 2D-local constant-depth quantum circuits that output states whose entanglement entropy across a specified cut cannot be estimated in quantum polynomial time. As constant-depth quantum circuits can be learned from polynomially many quantum samples, our resulting pseudoentangled states are implicitly public-key and not pseudorandom. This separates pseudoentanglement from pseudorandomness in the shallow-circuit regime: the former is possible, while the latter is not. The construction is based on the quantum intractability of the Dense-Sparse Learning Parity with Noise problem introduced in [DJ25] and uses a bounded-fan-in, bounded-fan-out classical randomized encoding for linear maps $\mathbf{x} \mapsto \mathbf{Mx},$ which could be of independent interest. As applications, we obtain quantum hardness for the problem of learning the entanglement structure...
论文介绍 本文构造了一族二维局部常数深度量子电路,其输出的态在特定切割下的纠缠熵无法在量子多项式时间内估计。这证明了在浅层电路情况下,伪纠缠与伪随机性是可分离的。该工作基于一个量子难解问题,并为学习纠缠结构等问题提供了量子难度结果。
第一作者: Massimo Bartoletti · 方向: 区块链安全
智能合约形式化验证大语言模型反例生成
Abstract:Recent large language models (LLMs) incorporate reasoning capabilities that allow them to perform well in predicting whether a smart contract respects a certain property, suggesting a complementary approach to traditional formal-methods-based techniques for smart contract verification. However, the application of LLMs in such context has two major issues: 1) properties expressed in natural language are intrinsically ambiguous, and 2) answers returned by LLMs have no guarantee of correctness. In this paper, we address both issues simultaneously by: 1) introducing a new formal specification language that extends Solidity with abstract types, and 2) designing a workflow that combines LLMs with type checking and concrete execution to generate and validate violation witnesses (i.e., counterexamples). The key idea is to represent a specification as a Solidity test with...
论文介绍 针对大语言模型在智能合约验证中存在属性描述模糊且答案无正确性保证的问题,本文提出一种混合验证流程。该流程引入一种扩展了抽象类型的形式化规范语言,并设计了结合大语言模型、类型检查与具体执行的工作流,旨在生成并验证违反规约的见证(反例)。核心是将规范表示为具有断言的Solidity测试,从而利用LLM的推理能力与形式化方法确保可靠性。
第一作者: Ei Hmue Khine · 方向: AI 安全
对抗攻击黑盒攻击决策边界语义流形几何搜索
Abstract:While decision-based black-box adversarial attacks present a severe security threat, current methodologies suffer from fundamental limitations. Pixel-wise attacks frequently introduce unnatural, high-frequency visual artifacts, while latent-space frameworks are confined by the limited search space of low-dimensional manifolds and inherent reconstruction flaws. To resolve these limitations, we propose Latent Geometric Chords (LGC) for Query-Efficient Decision-Based Adversarial Attacks alongside a variant, LGC-H. At its core, LGC navigates decision boundaries by executing a curvature-aware geometric search within a compressed semantic manifold. To guarantee high visual fidelity and circumvent dimensionality bottlenecks, we introduce a Residual-based Adversarial Generation (RAG) mechanism. RAG isolates semantic perturbations as geometric chords and superimposes them directly onto...
论文介绍 为解决基于决策的黑盒对抗攻击中像素级扰动不自然、以及潜空间方法受限于低维流形和重建缺陷的问题,本文提出了潜在几何和弦(LGC)方法。LGC的核心是在压缩的语义流形内执行曲率感知的几何搜索以遍历决策边界,并引入基于残差的对抗生成机制(RAG)来分离语义扰动。该方法旨在以较少的查询次数,生成具有高视觉保真度的对抗样本。
第一作者: Shangyi Shi · 方向: 密码学协议
全同态加密CKKS方案异构计算通信优化
Abstract:CKKS, an emerging fully homomorphic encryption (FHE) scheme, has been promising in privacy-preserving applications by enabling SIMD fixed-point computations on ciphertexts. Despite its strong security guarantees, CKKS involves both compute-intensive operators (ComOps) with high computational cost and memory-intensive operators (MemOps) with large memory footprints, making existing ASIC-based or NMP-based acceleration approaches suffer from high hardware overhead and limited efficiency. This observation motivates the integration of the architectural advantages of both paradigms into a heterogeneous xPU (ASIC)-xMU (NMP) architecture. However, in such a design, frequent and long-latency heterogeneous communication caused by the dominant keyswitch operator remains a key performance bottleneck. In this paper, we propose $HE^2$, a communication-light xPU-xMU heterogeneous FHE...
论文介绍 针对全同态加密(FHE)中CKKS方案计算密集型与内存密集型算子并存,导致现有加速方案效率受限的问题,本文提出了一种名为HE²的轻量通信异构架构。该架构整合了专用集成电路(ASIC)和近内存计算(NMP)的优势,并专门优化了主导性能瓶颈的密钥切换算子的通信效率,旨在降低硬件开销并提升FHE的整体处理效率。
第一作者: Ali Abdolrahimi Zarnagh · 方向: 软件安全
伪随机数生成器LFSRMersenne Twister可预测性分析
Abstract:Generating reliable random and pseudo-random sequences is important in many electronic and signal processing systems, such as secure communications, radar, spread-spectrum methods, and autonomous platforms. Although true and quantum random number generators provide stronger unpredictability, classical pseudo-random number generators, including Linear Feedback Shift Registers (LFSRs) and the Mersenne Twister (MT), are still widely used because they are efficient and easy to implement. This work introduces a user-friendly software platform for generating, analyzing, and evaluating the predictability of pseudo-random bit sequences. The software supports two main functions: generating sequences using classical PRNGs and hybrid combinations, and analyzing input sequences through statistical measures and data-driven methods. In particular, hybrid LFSR-MT structures are studied to...
论文介绍 本文介绍了一个用户友好的软件平台,用于生成、分析和评估伪随机比特序列的可预测性。平台支持使用经典伪随机数生成器(PRNG)及其混合组合(特别是LFSR与Mersenne Twister的混合结构)来生成序列,并可通过统计度量和数据驱动方法对输入序列进行分析,旨在为通信、雷达等系统提供实用的随机序列研究工具。
第一作者: James Bartusek · 方向: AI 安全
量子位置验证轨迹验证不可克隆状态量子密码学
Abstract:While quantum position verification aims to certify a prover's location using quantum information, existing security definitions only guarantee that part of the successful adversarial party is in the claimed location. This leaves open the possibility that a distributed team of adversaries can jointly simulate a prover in a way that defeats the intended meaning of ``being at a location'' in position-based cryptography. We introduce stronger notions of position verification that we call quantum localization, which requires that there is a specified, unclonable state at the verified spacetime point -- and that this state can be found nowhere else. We show that quantum localization leads naturally to a meaningful notion of trajectory verification, in which quantum information is verifiably tracked through space and time. We construct quantum localization and trajectory...
论文介绍 现有量子位置验证的安全定义仅能保证部分攻击者位于所声称位置,无法防止分布式对手团队联合模拟。本文提出了更强的量子定位概念,要求验证的时空点存在一个指定的、不可克隆的状态。由此自然引出了轨迹验证的概念,即实现量子信息在空间和时间中可验证的追踪。论文构建了量子定位与轨迹验证方案,解决了位置密码学中的一个根本性问题。
第一作者: Gudrun Schappacher-Tilp · 方向: AI 安全
隐私设计边缘AIGDPR合规视觉监控
Abstract:Visual monitoring systems that rely on cloud-based AI inference expose raw image data to external services, creating fundamental tensions with the data-minimisation principle of the General Data Protection Regulation (GDPR). This paper presents a proof-of-concept privacy-by-design pipeline that resolves this tension by confining all inference entirely to the edge device. A YOLOv5n-seg model compiled for a Hailo-8L AI accelerator delivers real-time object detection on a Raspberry Pi 5, from which raw pixel buffers are immediately discarded after inference. A stateful trigger engine forwards minimal JSON event payloads to a locally hosted instance of Phi-3 Mini (3.8B parameters, Q4_0 quantisation), which synthesises one-to-two sentence natural-language alerts for a human operator. No image data crosses the network boundary at any point; only the generated text alert is...
论文介绍 为解决依赖云推理的视觉监控系统与GDPR数据最小化原则的冲突,本文提出了一种全在边缘设备运行的概念验证流程。该流程使用在树莓派上运行的YOLOv5n-seg模型进行实时对象检测后立即丢弃图像,并利用本地部署的Phi-3 Mini模型将事件信息合成为自然语言警报。任何原始图像数据都不会离开设备,从而实现了隐私保护的视觉监控。
第一作者: Madhura Pathegama · 方向: 隐私保护
局部差分隐私相关噪声隐私预算效用优化
Abstract:We study privately estimating the sum of $n$ user-held values in the presence of an honest-but-curious server. This motivates requiring privacy not only at data release but also throughout server-side computation. We therefore adopt the local (pure) differential privacy model, in which each user transmits a noise-perturbed value. It is well known that independent local noise typically incurs a substantial utility loss compared to the centralized model, where noise is added only after aggregation. We show that this gap is not fundamental. By carefully designing correlations among the locally added noise variables, we construct $\varepsilon$-DP mechanisms whose estimation cost matches the optimal cost achievable in the centralized setting, up to an arbitrarily small error.
论文介绍 本文研究了在诚实但好奇的服务器存在下,对n个用户持有值求和的私有估计问题。在局部纯差分隐私模型中,已知添加独立噪声会导致相比集中模型显著的效用损失。本文证明这种差距并非固有,通过精心设计局部添加噪声变量之间的相关性,可以构造出估计成本与集中设置最优成本相匹配(直至任意小误差)的ε-差分隐私机制。
第一作者: Anany Kotawala · 方向: 安全研究
基准泄露记忆化大语言模型评估NumLeak
Abstract:Public numeric benchmarks appear in pretraining, so an evaluation that conditions on a date may be measuring memorized recall rather than out-of-sample skill. We introduce NumLeak, a measurement framework that combines API-boundary probes on production models with a white-box controlled validation on an open causal LM. Top-tier frontier LLMs recall the Fama-French market excess return at 3-seed pooled Pearson r=0.97-0.99 while staying within 0.15 within-25bps on the five sibling factors; comparable fidelity appears on U.S. unemployment, CPI inflation, and NOAA temperature. On a recent-release holdout, parse rate collapses to 21-57% but r stays at approximately 0.99 on months answered, the refuse-or-recall asymmetry a memorized channel predicts. The white-box experiment reproduces the dose-response, and logprob ranking detects memorization that open-ended generation misses...
论文介绍 公开的数值基准可能出现在预训练数据中,导致评估可能测量的是记忆化回忆而非泛化能力。本文引入了NumLeak测量框架,结合对生产模型的API探测和对开放因果语言模型的白盒可控验证。实验发现,前沿大语言模型能高度精确地回忆特定金融、经济指标数据,这表明模型可能记忆了这些基准。该框架为检测和量化大模型中的数据记忆化提供了方法。
第一作者: Ulf Kasolowsky · 方向: 策略学习 · 来源: cs.RO
触觉感知策略学习小物体操作强化学习机器人抓取
Abstract:We introduce and solve the novel task of controlled separation of small objects with two fingers of a multi-purpose robotic hand: after grasping into a box of small objects, the task is to drop as many of them until a desired number remains between the fingers. The objects are small compared to the width of the fingers but also in absolute terms. In our case little pellets with a diameter of only 6mm are handled. We show that the task can be performed purely tactile (no vision) using a spatially-resolved tactile skin on a fingertip. The separation policy is trained in simulation via reinforcement learning using a straightforward sparse reward, which basically checks if the desired number of objects is reached. In simulation experiments, we provide an exhaustive analysis of the benefits of using spatially-resolved tactile feedback: while an ideal (high-resolution) tactile...
论文介绍 本文提出并解决了一项新任务:使用多用途机器人手的两个手指对小物体进行可控分离。目标是在从容器中抓取一堆微小颗粒(直径仅6毫米)后,通过控制手指动作,使指定数量的物体留在两指之间。研究展示了仅依靠指尖上的高分辨率空间触觉皮肤(无视觉)即可完成此任务。其分离策略通过强化学习在仿真中训练,使用简单的稀疏奖励函数。实验证明,空间分辨的触觉反馈对精确分离至关重要。
第一作者: Yue Wang · 方向: 策略学习 · 来源: cs.RO
可微仿真刚体动力学机器人学习GPU加速PyTorch
Abstract:As robot control shifts toward large-scale reinforcement learning with in-loop dynamics computation, the community's reliance on CPU-bound libraries such as Pinocchio creates a throughput bottleneck in GPU-based training pipelines. We present BARD (Batched Articulated Rigid-body Dynamics), a self-contained PyTorch implementation of Featherstone's rigid-body dynamics algorithms, optimized for batched GPU evaluation and automatic differentiation. Three design choices make this efficient: a tiered lazy-evaluation cache that avoids redundant tree traversals, matmul-free joint transforms via pre-computed Rodrigues constants, and level-parallel propagation that reduces sequential operations to tree-depth batched steps. On five robot models (7-23 DOFs), BARD matches Pinocchio numerically while reaching up to 64x higher throughput for Forward Kinematics and 63x for Jacobians at batch...
论文介绍 随着基于强化学习的机器人控制转向大规模训练,现有依赖CPU的动力学库(如Pinocchio)成为吞吐量瓶颈。本文提出了BARD,一个完全在PyTorch中实现、针对批量GPU计算和自动微分优化的刚体动力学库。它通过分级惰性缓存、无矩阵乘法的关节变换和层级并行传播等设计,在多个机器人模型上实现了与Pinocchio相当的数值精度,同时正向运动学和雅可比矩阵计算的吞吐量分别提升了高达64倍和63倍,显著加速了GPU训练流程。
第一作者: Joonhee Lee · 方向: 具身智能 · 来源: cs.RO
机器人规划大语言模型具身推理推理优化实时决策
Abstract:Reasoning-based robotic policies using large language and vision-language models achieve strong semantic planning capabilities but mostly suffer from a high inference latency that limits practical real-time deployment. In this work, we observe that robotic reasoning workloads contain substantial temporal redundancy, where consecutive observations frequently produce identical actions and subgoals. Based on this insight, we present REIS, a human cognition inspired robotic decision-making framework that minimizes unnecessary reasoning while preserving semantic adaptability. REIS combines lightweight scene gating, KV-steered affordance routing, and deliberative reasoning to accelerate robotic control under embodied constraints. Experiments on ALFRED, and real-world robotic tasks demonstrate that REIS significantly suppresses reasoning overhead while maintaining competitive task...
论文介绍 基于大型语言或视觉语言模型的推理型机器人策略具有强大的语义规划能力,但其高推理延迟限制了实时部署。本文观察到机器人推理工作负载中存在大量时间冗余,即连续观察常产生相同的动作和子目标。受此启发,提出了REIS框架,它通过轻量级场景门控、KV引导的可供性路由和审慎推理,在保持语义适应性的同时最小化不必要的推理,从而显著降低推理开销。在仿真和真实机器人任务中,REIS在维持任务竞争力的同时,大幅减少了推理延迟。
第一作者: Mohammad Dastranj · 方向: 具身智能 · 来源: cs.RO
逆运动学关节限制力矩控制冗余机器人二次规划
Abstract:This paper proposes actuator-aware inverse kinematics for torque-controlled redundant robots under joint-limit constraints. In the considered architecture, the inverse-kinematic output is not merely a purely kinematic joint-velocity command; it is the required joint velocity supplied to a downstream torque-level controller. Therefore, a small commanded task residual may not necessarily improve realized motion. The proposed method formulates a convex quadratic programming problem whose decision variable is the joint-level required velocity. Control barrier function style bounds impose reference-level joint-limit admissibility, while the task equation is handled through a penalized slack variable. Redundancy is resolved using a controller-compatibility objective that accounts for previous-command consistency and actuator torque-capacity weighting. The method is independent of...
论文介绍 本文针对关节受限的力矩控制冗余机器人,提出了一种执行器感知的逆运动学方法。与传统方法输出纯运动学关节速度不同,该方法输出的指令将供给下游的力矩控制器。因此,其核心是构建一个凸二次规划问题,决策变量为关节速度。通过采用控制障碍函数形式的约束来确保关节限制的可容许性,并利用带罚项的松弛变量处理任务方程。同时,通过一个考虑指令一致性和执行器力矩容量的控制器兼容性目标来解决冗余。
第一作者: Shuyuan Yang · 方向: 具身智能 · 来源: cs.RO
触觉反馈微创手术力感知变压器网络机器人手术
Abstract:Robot-Assisted Minimally Invasive Surgery (RAMIS) enhances surgeon dexterity, with newer platforms leveraging haptic feedback to further improve performance. Such force information has broader potential to inform performance assessment, tactile localization, and surgical autonomy. This motivates the need for accessible approaches to integrating force sensing into RAMIS tools. This work presents a method for integrating a six-axis commercial force sensor into the distal end of a standard cable-driven surgical instrument, enabling end-effector force measurement while preserving the original mechanical functionality of the device. The proposed design emphasizes reproducibility and accessibility for research applications, requiring no specialized manufacturing tools. A transformer neural network integrates force sensor measurements with robot state information to aid estimation of...
论文介绍 机器人辅助微创手术中,力信息对于性能评估和手术自主性至关重要。本文提出一种将商用六轴力传感器集成到标准缆驱手术器械远端的方法,能够在保留原有机械功能的同时测量末端执行器的力。该设计注重可复现性和可及性,无需特殊制造工具。此外,一个变压器神经网络被用来融合力传感器测量值与机器人状态信息,以辅助估计交互动力学,从而提升力感知的准确性。
第一作者: Saksham Gupta · 方向: 具身智能 · 来源: cs.RO
自适应控制时滞估计障碍李雅普诺夫函数欧拉-拉格朗日系统不确定性补偿
Abstract:This paper addresses the challenge of simultaneously compensating for state-dependent uncertainties and enforcing time-varying state constraints in Euler-Lagrange systems, a common requirement in robotics that remains underserved by existing control designs. A novel adaptive control framework is developed that combines an artificial time-delay-based uncertainty estimation strategy, also known as time-delay estimation, with a barrier Lyapunov function to enforce constraint-aware control design. Specifically, a state-dependent upper bound on the time-delay estimation approximation error is analytically formulated, and an adaptive law is constructed to estimate its parameters online, enabling real-time state-dependent uncertainty compensation without relying on prior model knowledge. To ensure constraint compliance, the barrier Lyapunov function-based controller enforces...
论文介绍 本文解决欧拉-拉格朗日系统(常见于机器人)中同时补偿状态依赖不确定性和强制执行时变状态约束的挑战。提出了一种新颖的自适应控制框架,结合了人工时滞估计(用于不确定性估计)和障碍李雅普诺夫函数(用于约束满足控制设计)。具体而言,解析推导了时滞估计逼近误差的状态依赖上界,并构造了自适应律在线估计其参数,从而实现不依赖先验模型知识的实时不确定性补偿。
第一作者: Hannah Schieber · 方向: 导航与运动 · 来源: cs.RO
路径规划语义导航高斯溅射TSDF机器人导航
Abstract:Autonomous robots in unknown indoor environments require both reliable collision avoidance and object-level understanding. Classical representations such as TSDF support safe planning but lack semantics, while photorealistic methods like Gaussian Splatting (GS) provide rich appearance yet suffer from soft geometry, limiting precise obstacle avoidance. We present LiftNav, a hybrid navigation framework built on GSFusion's TSDF+GS dual map, augmented with a real-time pipeline of YOLO-based detection, TSDF-based 3D lifting, and B-spline trajectory optimization. This design enables flexible semantic navigation without dense 3D embeddings. We further introduce a hinge-loss-based collision penalty that improves trajectory smoothness and safety. We evaluate our approach in a simulation using the Replica dataset. Compared against a state-of-the-art radiance field baseline we show a...
论文介绍 在未知室内环境中,自主机器人需要可靠的避障和物体级理解。经典TSDF表示支持安全规划但缺乏语义,而高斯溅射等方法提供丰富外观但几何较软。本文提出LiftNav,一个混合导航框架,基于GSFusion的TSDF与高斯溅射双地图构建,并增强了实时流程:YOLO检测、基于TSDF的三维提升和B样条轨迹优化。该设计无需密集3D嵌入即可实现灵活的语义导航。引入的基于铰链损失的碰撞惩罚项也改善了轨迹的平滑性和安全性。
第一作者: Zhuoyi Lu · 方向: 机器人操作 · 来源: cs.RO
触觉探索形状估计操作规划贝叶斯优化位姿推理
Abstract:Robotics manipulation usually assumes that the shape and pose of the object are known to the robot prior to motion planning. However, precise geometric information is not always available in practice, and pose inference suffers from sensor uncertainties and view occlusion. In this work, we propose a unified model-based geometric framework integrating robotic haptic perception, modeling, and manipulation planning. Our novelties involve: \textit{i)} Introducing Bayesian Optimization (BO) to guide the haptic exploration for object shape inference, where superellipses are used to approximate geometric boundary; \textit{ii)} Adaptive formulation of manipulation potential encoding object geometry for quasi-static robot-object interaction; \textit{iii)} Proposing an online Ordinary Differential Equation (ODE) for real-time pose inference based on model prediction and tactile...
论文介绍 机器人操作通常假设已知物体的形状和位姿,但精确几何信息并非总是可用。本文提出一个统一的基于模型的几何框架,整合了机器人触觉感知、建模和操作规划。核心创新包括:引入贝叶斯优化指导触觉探索以推断物体形状(使用超椭圆近似边界);自适应地制定包含物体几何信息的操纵势函数以用于准静态交互;提出一个在线常微分方程,基于模型预测和触觉测量进行实时位姿推理。
第一作者: Sikai Guo · 方向: 机器人操作 · 来源: cs.RO
地形感知全身控制腿式操作器强化学习移动操作
Abstract:Legged manipulators integrate exceptional terrain adaptability along with mobile manipulation capabilities, which make them highly promising for deployment in human-centric environments. By coordinating the control of both legs and arms, a whole-body controller can significantly expand the operational workspace of legged manipulators. However, many existing whole-body controllers primarily depend on proprioception and do not incorporate the critical exteroception required for effective terrain topology perception. This limitation can hinder their ability to adapt to varying environmental conditions and navigate complex terrains effectively. In this paper, we introduce TA-WBC, a terrain-aware whole-body control framework for legged manipulators, which features a novel RL-based unified policy tailored to whole-body loco-manipulation tasks in various terrains. Specifically, we...
论文介绍 腿式操作器在人类中心环境中具有潜力,但现有全身控制器依赖本体感知,缺乏地形拓扑感知。本文提出TA-WBC,一个地形感知的全身控制框架,基于强化学习的统一策略,适用于各种地形中的全身移动操作任务。该框架通过协调腿和臂的控制,扩展操作工作空间,提升机器人对复杂地形的适应能力。
第一作者: Shuai Ke · 方向: 机器人操作 · 来源: cs.RO
表面约束扩散策略机器人操作动态可行性模仿学习
Abstract:Diffusion-based imitation learning methods have driven rapid progress in robot dexterous manipulation tasks. However, they have limitations when applied to tasks that involve complex free-form surface constraints because of their lack of explicit surface geometry constraint modeling and the dynamic feasibility issue, resulting in stochastic action generation that fails to achieve reliable surface alignment and maintain stable contact. To address these limitations, we propose a novel surface constraint policy (SCP) for generating robot actions that satisfy free-form surface constraints on the basis of human demonstrations and real-time visual observations. First, the surface geometry constraint is encoded using a two-dimensional weighted Gaussian kernel function that is derived from demonstrations. Building on the encoded surface geometry constraints, the diffusion-based policy...
论文介绍 针对扩散策略在表面约束任务中缺乏几何约束建模的问题,提出表面约束策略SCP。该策略使用二维加权高斯核函数编码表面几何约束,基于人类演示和实时视觉观察生成满足自由形式表面约束的机器人动作,确保动态可行性和稳定接触。
第一作者: Yifei Yang · 方向: 导航与运动 · 来源: cs.RO
自回归训练机器人导航世界模型长期预测扩散模型
Abstract:The diffusion based robot navigation world models are typically trained using parallel supervision, while autoregressive inference is employed during path planning. This results in a distribution shift between training and inference, which destabilizes the performance over long-horizon prediction. We propose AR Forcing, an autoregressive training strategy, which integrates the standard diffusion loss into the autoregressive training loop. At each step, the model uses its own predictions to update the context and optimize the single step noise prediction objective, thereby explicitly exposing the model to the inference state distribution during training. Our method does not require additional discriminators or distribution-matching losses, retains the original diffusion framework and sampler, and is easy to integrate. Experiments on multi-domain navigation datasets (RECON...
论文介绍 基于扩散的机器人导航世界模型在训练和推理间存在分布偏移,影响长期预测稳定性。提出AR Forcing,一种自回归训练策略,将扩散损失集成到训练循环中,显式暴露模型于推理状态分布。该方法无需额外判别器,易于集成,适用于长期导航规划。
第一作者: Taiyi Su · 方向: VLA 通用模型 · 来源: cs.RO
视觉语言动作基础模型可变形操作流匹配泛化
Abstract:Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and household environments. Deformable-object folding is a representative challenge, requiring robots to handle clothing items from random initial states across varying categories, geometries, materials, and scenes. However, existing VLA systems commonly train separate policies for different object categories, while naively mixed multi-task training often suffers from task interference and degraded performance. To move beyond category-specific folding policies, we introduce DeMaVLA, a VLA foundation model for generalizable Deformable Manipulation. DeMaVLA adopts a VLM backbone with an action expert and formulates continuous action generation using flow matching. To improve efficiency, the action expert is...
论文介绍 现有视觉语言动作系统针对不同物体类别训练单独策略,泛化能力有限。引入DeMaVLA,用于可变形操作的VLA基础模型,采用视觉语言模型骨干和动作专家,使用流匹配生成连续动作,提高在多类别可变形物体操作中的泛化性能。
第一作者: Luca Benfenati · 方向: 具身智能 · 来源: cs.RO
剪枝具身大型语言模型自动驾驶强化学习效率优化
Abstract:Embodied Large Language Models (LLMs) are increasingly used as reasoning modules in robotic control pipelines to improve human-robot interaction, but their memory and generation latency make real-time deployment difficult. Pruning can reduce these costs, but for controllers that undergo multiple pre- and post-training phases, the crucial question is not only how much to prune, but when pruning should occur. In this work, we propose Before Parc Fermé (BPF), a pruning strategy performed during RL that compresses embodied LLM controllers while they are still being optimized for closed-loop behavior. This allows pruning decisions to account for the task-specific supervision and closed-loop feedback that shape the final controller. We propose two variants: BPF-RL, which performs iterative pruning during RL by removing part of the model at predefined training intervals, and...
论文介绍 具身大型语言模型作为推理模块在自动驾驶中应用广泛,但内存和延迟问题阻碍实时部署。提出Before Parc Fermé策略,在强化学习训练期间进行剪枝,使剪枝决策考虑任务监督和闭环反馈,从而在压缩模型的同时保持控制器性能。
第一作者: Xiang Zhu · 方向: VLA 通用模型 · 来源: cs.RO
表示学习跨体现视觉语言动作预训练人机对齐
Abstract:Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions. While latent action models reduce the action execution gap by learning action abstractions, they still rely on visual features. Thus, misaligned human and robot visual representations can lead to inconsistencies in policy inputs and induce domain-dependent latent actions, hindering effective co-training with human videos. To address this, we propose HARP, a human-robot aligned representation learning framework for more effective VLA pretraining from human videos. Specifically, HARP uses limited paired human-robot demonstrations as cross-embodiment bridges and abundant unpaired human and robot videos as a scalable dynamics supervision data source. It trains a robot-adapted...
论文介绍 从人类视频学习视觉语言动作模型时,视觉表示不一致导致策略输入偏差。提出HARP框架,通过配对人类-机器人演示和未配对视频训练对齐表示,减少跨体现差异,提升VLA模型的预训练效果。
第一作者: Tianle Zeng · 方向: 导航与运动 · 来源: cs.RO
视觉语言导航户外导航可通行性语义中断记忆增强
Abstract:Outdoor vision-language navigation (VLN) in long-range, open-world environments is frequently disrupted by semantic-cue interruptions, where informative goal cues become sparse, occluded, or leave the field of view. Once such cues disappear, agents enter a cue-free phase and often degrade into backtracking, oscillatory headings, or aimless exploration. While memory-based methods attempt to bridge these gaps, they often fail under traversability-driven detours: the remembered cue direction may be infeasible, forcing detours that prolong cue-free phases and gradually render robot-centric cues stale and implicit histories blurred. This makes traversability a stability condition for maintaining goal-directed guidance, rather than merely a local safety concern. We propose a unified outdoor VLN framework that survives semantic-cue interruptions by maintaining...
论文介绍 户外视觉语言导航中语义线索中断会导致代理行为退化,如回溯或无目的探索。提出TARIC框架,结合记忆和可通行性感知,在语义线索中断时维持目标导向导航,通过可通行性约束确保路径可行性。
第一作者: Navin Sriram Ravie · 方向: 具身智能 · 来源: cs.RO
持续学习具身代理经验驱动野外适应异常归因
Abstract:In robotics, dangers and adversity modes are often embodiment-specific and relative to each agent. A frontier of autonomous mobile robotics is to enable agents to operate effectively in the wild in unseen unstructured environments. A significant challenge in unseen unstructured environments is that it may not be possible to predict all the dangers to the specific robot. Although recent work has used large foundation vision-language models (VLMs) to preemptively predict an exhaustive list of common-sense dangers, it remains difficult to capture possible interaction and embodiment-dependent adversities. We propose a continual learning framework for a mobile embodied agent to learn online from disturbances and attribute anomalous behaviours to causes through semantics, enabling better prediction and planning of the world in the future. Our framework, "Don't Fool Me Twice", first...
论文介绍 移动具身机器人在未知环境中难以预测所有体现相关困难。提出「Don't Fool Me Twice」框架,通过持续学习从干扰中在线学习,使用语义归因异常行为,改进未来预测和规划,增强机器人对野外环境的适应能力。
第一作者: Aravind Battaje · 方向: 具身智能 · 来源: cs.RO
机器人泛化自适应组合规则性行为生成AICON框架
Abstract:Generalization in robotics requires prior knowledge about how the world is structured, yet this structure changes from one situation to the next. This paper investigates the proposition that generalization arises from adaptively composing regularities -- predictable relationships within the robot-environment system -- into situation-appropriate structures for behavior generation. We examine this proposition by analyzing the mechanism in AICON (Active InterCONnect), a framework representing regularities as interacting processes in a differentiable network, where sensory feedback realizes composition and gradient descent generates behavior. To isolate adaptive composition as the key mechanism, we study a simple simulated problem in which all relevant regularities can be identified. We expose the resulting model to a wide range of novel conditions not considered during design...
论文介绍 本文研究机器人行为生成中的泛化问题,提出通过自适应组合规则性来构建泛化能力。核心方法基于AICON框架,将规则性表示为可微分网络中的交互过程,利用感官反馈实现组合,并通过梯度下降生成行为。通过仿真问题验证机制,使模型能适应未设计过的新条件,可能应用于提升机器人在变化环境中的适应性。
第一作者: Marcel Bartholomeus Prasetyo · 方向: 具身智能 · 来源: cs.RO
双模态3D场景图开放集任务快速与慢速模式场景表示
Abstract:Open-set task execution can significantly benefit from seamlessly switching between coarse and fine scene representations depending on the context and the evolving information as the robot explores the environment. For example, it is often sufficient to start with a coarse scene representation initially and only employ a finer, more granular scene representation when the robot encounters regions which are likely to contain the task relevant objects. Hence, in this work, we propose BiMoSG, a bimodal 3D scene graph generation approach for open-set tasks. BiMoSG employs a "fast" mode by default to efficiently generate a coarse 3D scene graph and can switch to a "slow" mode for generating a finer open vocabulary 3D scene graph of task relevant objects. We demonstrate that our proposed 3D scene graph generation approach is significantly faster than the open-source state-of-the-art...
论文介绍 针对开放集任务中场景表示的需求,提出BiMoSG双模态3D场景图生成方法。该方法默认使用快速模式生成粗粒度场景图,当遇到任务相关区域时切换到慢速模式生成精细的开放词汇场景图。核心优势在于提升效率和适应性,比现有开源方法更快,适用于机器人环境探索和任务执行。
第一作者: Tianle Zeng · 方向: VLA 通用模型 · 来源: cs.RO
视觉-语言-动作模型空地协同CARLA-Air环境闭环协调诊断任务
Abstract:Recent aerial vision-language-action (VLA) models show promising single-UAV capabilities, such as tracking moving objects and navigating to language-specified landmarks. However, it remains unclear whether these capabilities can transfer to air-ground cooperation, where a UAV and a UGV must act jointly in a shared, closed-loop physical world. We study this question with CARLA-Air, a single-process air-ground evaluation environment that unifies CARLA and AirSim inside one Unreal Engine runtime. By sharing the same world state, physics tick, and sensing pipeline, CARLA-Air enables physically consistent UAV--UGV interaction and precise measurement of simulation-timestamp alignment and effective coordination latency. Using CARLA-Air, we evaluate representative aerial VLA and planning baselines on two complementary diagnostic tasks: moving-platform landing and occlusion-recovery...
论文介绍 研究视觉-语言-动作模型在空地协同中的可行性。通过CARLA-Air环境,统一CARLA和AirSim,实现无人机与无人地面车在共享物理世界中的一致交互。评估代表性VLA和规划基线在移动平台着陆和遮挡恢复任务中的表现,以测试模型的协同能力和闭环控制效果。
第一作者: InGyu Choi · 方向: 机器人操作 · 来源: cs.RO
虚拟现实遥操作动态环境实时性碰撞处理GPU加速
Abstract:Robot teleoperation enables safe, non-contact task execution in hazardous environments where direct human access is difficult, and its application has expanded with recent VR technologies. Many VR teleoperation studies, however, have primarily served as data-collection tools for robot imitation learning, so they often do not explicitly address dynamic obstacles, workspace changes, or collision risks during operation. For real deployment aimed at operator safety, teleoperation must react to dynamic situations with low latency and remain robust to mistakes made by inexperienced operators. This paper presents a VR teleoperation framework that supports real-time manipulation while handling collisions with both static and moving obstacles. The framework integrates GPU-accelerated inverse kinematics and trajectory optimization within a VR interface to generate feasible joint...
论文介绍 提出一个实时VR遥操作框架,用于动态环境中的机械臂操作。该框架集成GPU加速的逆运动学和轨迹优化,在VR界面中生成可行的关节运动,以处理静态和动态障碍物的碰撞。旨在提升操作安全性、低延迟响应和对新手错误的鲁棒性,适用于危险环境下的远程任务。
第一作者: Zijian Zhu · 方向: VLA 通用模型 · 来源: cs.RO
视觉-语言-动作模型示范生成强化学习仿真到现实轨迹数据
Abstract:Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robot control. However, their performance remains fundamentally constrained by the availability of high-quality robot trajectory data. In current robot learning practice, such data are primarily collected through human teleoperation, which is labor-intensive, costly, and difficult to scale. In this paper, we propose RDGen, a sim-to-real reinforcement learning framework for generating high-quality robot demonstrations. Rather than employing reinforcement learning solely as the final control policy, RDGen leverages trained RL policies as a structured trajectory generator. The system consists of a VLM-based task parser that identifies task-relevant objects, a Grounding DINO-based object localizer, and an RL policy transferred from simulation to the real robot. Successful rollouts are then...
论文介绍 针对VLA模型对高质量轨迹数据的需求,提出RDGen框架,通过强化学习生成机器人示范。该框架将训练好的RL策略用作轨迹生成器,结合VLM任务解析和对象定位,从仿真转移到真实机器人。旨在减少对人工遥操作的依赖,提高数据生成效率和质量,支持通用机器人控制。
第一作者: Nina Majer · 方向: 导航与运动 · 来源: cs.RO
非通信移动机器人轨迹规划逆最优控制碰撞避免联合预测
Abstract:To enable an efficient interaction of non-communicating mobile robots in collision avoidance scenarios, we present a novel combined trajectory planning and prediction algorithm. Inverse optimal control is used to estimate unknown goal states of all robots based on observed past trajectories. Each robot also takes the perspective of other robots in considering self-prediction and solves a joint prediction problem using the estimated goal states. The resulting predictions are then considered for planning. Simulation results of scenarios with 2-8 robots show that the median of the durations until all vehicles reach their goals is 9.8 % faster compared to planning with constant acceleration based estimated goal states. Moreover, the proposed approach never leads to the solver being unable to find a solution to the planning or prediction problem.
论文介绍 为非通信移动机器人在碰撞避免场景中提出一种轨迹规划和预测算法。使用逆最优控制从观测轨迹估计机器人目标状态,并考虑其他机器人视角进行联合预测,然后用于规划。仿真显示,在多机器人场景中,该方法比基于恒定加速估计的方法更快达成目标,且无求解失败情况。
第一作者: Ryan Yu · 方向: VLA 通用模型 · 来源: cs.RO
视觉-语言-动作预训练机器人策略Wall-OSS-0.5模型开源多具身学习
Abstract:Large-scale Vision-Language-Action (VLA) pretraining is increasingly adopted as the foundation for robot policies, yet the evidence for pretrained VLAs is almost invariably reported after task-specific this http URL leaves a foundational question unanswered: does VLA pretraining itself yield executable robot behavior, or does it merely furnish a better initialization for downstream policy learning? We present Wall-OSS-0.5, an open-source 4B VLA built upon a 3B VLM backbone augmented with action-generation components, designed so that pretrained robotic capability is directly measurable on physical this http URL model is pretrained across more than 20 embodiments, processing over one million robot trajectories per epoch alongside a grounded multimodal corpus. We adopt a gradient-bridged co-training recipe in which three objectives play distinct and complementary roles: discrete...
论文介绍 介绍Wall-OSS-0.5,一个开源的4B参数VLA模型,基于3B VLM骨架并添加动作生成组件。该模型通过跨20多种具身的预训练,处理大量机器人轨迹和多模态数据,旨在直接测量预训练的机器人能力,而非仅作为下游策略的初始化,探索VLA预训练的原始行为生成潜力。
第一作者: An Li · 方向: 导航与运动 · 来源: cs.RO
电永磁足高负载密度可控吸附四足机器人磁路设计
Abstract:To enable reliable climbing locomotion of quadruped robots on ferromagnetic surfaces, this paper presents a high-load-density electro-permanent magnetic foot with controllable adhesion, featuring force-feedback circular Halbach-net electro-permanent magnet (CHN-EPM) adhesion units and a magnetization control system. Due to its three-dimensional magnetic circuit structure and flux-concentration effect, the CHN-EPM enables a distributed parallel magnetic flux path with enhanced flux utilization, resulting in reduced sensitivity to air-gap variations and allowing effective adhesion to be maintained even under partial contact conditions. The proposed CHN-EPM generates a maximum adhesion force exceeding 1000 N with a load-to-weight ratio over 200:1. A magnetization driver and a two-stage pulse current control strategy are developed to regulate the excitation current amplitude and...
论文介绍 为四足爬壁机器人设计一种高负载密度电永磁足,具有可控吸附特性。采用力反馈环形Halbach网电永磁单元和磁化控制系统,通过三维磁路结构和磁通集中效应,提高磁通利用率,降低对气隙变化的敏感性。最大吸附力超过1000N,负载重量比超200:1,适用于铁磁表面爬壁任务。
第一作者: Seongheon Park · 方向: VLA 通用模型 · 来源: cs.RO
VLA模型失败检测对比学习运行时监控
Abstract:Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliability in real-world deployment. Detecting such failures during execution is therefore critical for the robust deployment of embodied systems. Existing failure detection methods either rely on expensive action resampling or external models, while alternatives propagate trajectory-level labels uniformly across every timestep, obscuring localized failure signals. In this paper, we propose \textbf{Hide-and-Seek}, a framework that formulates VLA failure detection as a coarsely supervised learning problem. By combining inter-trajectory and intra-trajectory contrastive objectives, Hide-and-Seek localizes failure-indicative actions and induces temporally structured failure signals from...
论文介绍 本文针对视觉-语言-动作(VLA)模型在真实部署中可能发生的执行失败问题,提出了一种名为“Hide-and-Seek”的框架。该框架将失败检测形式化为一个粗监督学习问题,通过结合轨迹间和轨迹内对比学习目标,能够从稀疏的轨迹级标签中定位指示失败的局部动作,并产生时间结构化的失败信号,旨在提升具身系统的鲁棒性。
第一作者: Junyang Shu · 方向: VLA 通用模型 · 来源: cs.RO
强化学习价值估计视觉世界模型奖励塑形
Abstract:Reinforcement learning is a promising approach for improving the capabilities of vision-language-action (VLA) models while avoiding the heavy data requirements of imitation learning. However, its effectiveness for VLA models is often constrained by sparse supervision and the difficulty of designing informative reward signals for long-horizon manipulation. In this work, we present Feat2Go, a fine-grained value estimation framework for embodied reinforcement learning. Specifically, Feat2Go first derives a continuous progress target from a pretrained visual world model by measuring patch-level similarity to subgoal states and partitioning episodes into semantic stages with trend-based clustering. We then train an embodied value model to predict this structural progress from the current observation and task instruction, and use the predicted value to reshape terminal rewards...
论文介绍 本研究提出了Feat2Go,一个用于具身强化学习的细粒度价值估计框架,以改进VLA模型。该方法首先从一个预训练的视觉世界模型中,通过比较图像块与子目标状态的相似度,并基于趋势聚类将回合分割为语义阶段,从而推导出连续的进度目标。随后,训练一个具身价值模型来预测此结构化进度,并利用该预测值来重塑稀疏的终端奖励,从而为长时域操作提供更有效的监督信号。
第一作者: Nikola Raicevic · 方向: 机器人操作 · 来源: cs.RO
机器人操作非抓取操作模型预测控制长时域规划
Abstract:Long-horizon planning for non-prehensile robot manipulation is challenging due to underactuated and discontinuous interactions. We propose a hierarchical formulation of model predictive path integral (MPPI) control that guides robot-level planning with a separately computed object-level plan to achieve efficient long-horizon prediction. We first solve a simplified object-only problem, assuming the object can be actuated directly, and use the planned object trajectory as a reference in solving the joint robot-object planning problem. We evaluate our method in both simulation and hardware using a 6-DoF xArm6 manipulator to perform object pushing tasks in which the target object must reach a goal while avoiding static obstacles, necessitating non-myopic reasoning. Our object-informed MPPI increases task success by 40\% with a 26\% faster control frequency in simulation, and by...
论文介绍 本文针对非抓取机器人操作中的长时域规划难题,提出了一种分层模型预测路径积分(MPPI)控制方法。该方法的核心是通过一个独立计算的对象层规划来引导机器人层的规划。具体而言,先简化并求解一个假设对象可被直接驱动的对象轨迹规划问题,然后将此规划得到的对象轨迹作为参考,用于解决更复杂的机器人-对象联合规划问题,从而提升长时域预测的效率和任务成功率。
第一作者: Ruiqi Yu · 方向: 具身智能 · 来源: cs.RO
人形机器人行走视觉导航对称性
Abstract:Extending humanoid traversal to the open world is key to practical deployment in human environments, but remains challenging. The robot must use vision to ensure safe and reliable foot placement on heterogeneous terrain under highly dynamic motion, while producing coordinated, natural whole-body behaviors. We propose SSR, an efficient end-to-end framework for egocentric vision-based humanoid traversal that jointly learns these capabilities. SSR introduces imagined foothold guidance, which learns to model forthcoming swing-foot contacts and evaluates their support to guide pre-touchdown swings toward stable regions, reducing edge slips. It further employs equivariant latent-space symmetry augmentation to efficiently induce bilateral coordination under high-dimensional visual observations, and uses terrain-specific multi-discriminator motion priors to encourage human-like...
论文介绍 为将人形机器人的行走能力扩展到开放世界,本文提出了SSR,一个高效的端到端框架。该框架基于自我中心视觉进行学习,联合优化安全可靠的落足点选择与协调的全身运动。其关键创新包括“想象落足点引导”(学习摆动足即将到来的接触并评估其支撑性以引导稳定落足)以及“等变潜在空间对称性增强”(在视觉观测下高效诱导双侧协调),并辅以地形特定的多判别器运动先验,以生成类人的行走行为。
第一作者: Beichen Shao · 方向: 机器人操作 · 来源: cs.RO
机器人操作关节物体VLM安全交互
Abstract:Articulated object manipulation is a unique challenge for service robots. Existing methods employ end-to-end policy learning, visionmotion planning, and large-language/visual-language model (LLM/VLM), but often overlook the diversity of articulated objects and the complexity of interactions between end-effector and handle, leading to limited generalization and destructive collisions. To address this, we propose GSAM, a generalizable and safe robotic framework for articulated object manipulation. Specifically, a vision-based perceiver generates the kinematic parameters. Considering that pre-trained markers in perceiver yield raw estimations that may deviate from commonsense, we present a f ine-tuned VLM-based refiner, using chain-of-thought (COT) commonsense reasoning to refine perception. To prevent destructive collisions, we design an interaction constraint function...
论文介绍 本文提出了GSAM,一个用于关节物体操作的通用且安全的机器人框架,以应对其多样性和交互复杂性挑战。该框架首先使用一个基于视觉的感知器来生成物体的运动学参数。随后,利用一个经过微调的视觉-语言模型(VLM)作为优化器,借助思维链常识推理来精细化感知结果,使其更符合常识。为防止破坏性碰撞,框架还设计了交互约束函数,以确保末端执行器与物体手柄安全交互。
第一作者: Siwon Jo · 方向: 导航与运动 · 来源: cs.RO
控制屏障函数碰撞避免伯恩斯坦多项式安全导航
Abstract:Safe navigation often relies on well-defined conditions based on the shape of robots and obstacles, and can be challenging when they have irregular geometries. While Control Barrier Functions (CBFs) offer an efficient mechanism to enforce safe set forward invariance, common shape surrogates (e.g., spheres or super-ellipsoids) either are overly conservative in unstructured scenes or require many local primitives, which inflates constraint counts and degrades real-time performance. In this paper, we introduce a novel geometry-aware Control Barrier Function (CBF) based on Bernstein-Polynomial Signed Distance Fields (BP-SDFs). It provides a unified way to represent the obstacles and robots, so as to represent the barrier function with a unified minimum distance. Benefiting from the differentiability of the Bernstein polynomials, one can easily enforce the control constraints in a...
论文介绍 针对在不规则几何形状环境下实现安全导航的挑战,本文提出了一种基于伯恩斯坦多项式符号距离场的几何感知控制屏障函数。该方法提供了一种统一表示障碍物和机器人形状的方式,从而可以用统一的最小距离来定义屏障函数。得益于伯恩斯坦多项式的可微性,能够轻松地在优化问题中施加控制约束,以保证安全性,同时避免了传统形状代理(如超椭球)所带来的过度保守性或高计算成本问题。
第一作者: Anya Singh · 方向: VLA 通用模型 · 来源: cs.RO
VLA模型小样本学习技能组合迁移学习
Abstract:Deploying vision-language-action (VLA) policies in industrial environments requires the ability to teach new tasks at low cost, a property current VLAs lack, since each new task requires fine-tuning. We investigate whether primitive-aware training produces a transferable artifact: a learned library of sub-skills that can be composed at inference time, conditioned on a small number of demonstrations, to perform tasks the policy was never trained on. We train two VLA architectures with different inductive biases, OpenVLA and $\pi_{0.5}$, on the REASSEMBLE contact-rich assembly dataset under matched LoRA fine-tuning recipes and locked hyperparameters, varying training between flat trajectories and primitive-segmented episodes with primitive-specific language prompts. We hold out 6 object-task combinations from training and evaluate few-shot transfer: models receive $m \in \{0, 1...
论文介绍 本文探讨了通过“原始感知”训练,是否能让VLA模型学习到一个可迁移的子技能库。研究在两个VLA架构(OpenVLA和π0.5)上,使用REASSEMBLE接触丰富的装配数据集进行训练,比较了使用原始分割的片段(配合特定语言提示)与使用平滑轨迹的训练效果。结果旨在验证,经过此类训练的模型能否在推理时,仅凭少量新任务的演示,就能组合这些子技能来执行从未训练过的任务。
第一作者: Avinash Subramanian · 方向: 导航与运动 · 来源: cs.RO
状态估计因子图凸松弛稀疏优化
Abstract:Robust and efficient state estimation is crucial for perception, navigation, and control in robotics. State estimation problems are conveniently modeled using the factor-graph framework as enabled by modern software packages such as GTSAM or g2o. However, the standard solvers included in such frameworks are local and may converge to poor local minima, posing significant safety concerns. Conversely, techniques based on convex relaxations have been shown to provide a means of globally solving or certifying many state estimation problems. However, these relaxations 1) often require substantial effort to formulate, and 2) may incur significantly higher cost compared to efficient local solvers, as they require solving a large semidefinite program (SDP). In this work, we address both shortcomings by 1) creating a new procedure within the GTSAM framework for automatically...
论文介绍 本文旨在解决机器人状态估计中,基于因子图的标准求解器可能收敛到差的局部极小值,而基于凸松弛的全局方法又存在公式化繁琐和计算成本高的问题。研究提出了一种在GTSAM框架内的新程序,利用问题中的弦稀疏性结构来自动生成更紧凑、更有效的半定规划松弛。这有助于在保持全局最优性保证的同时,显著降低计算负担,为安全关键的感知和导航任务提供更可靠的状态估计。
第一作者: Emil Martens · 方向: 具身智能 · 来源: cs.RO
CUDA加速符号编程非线性优化GPU求解器机器人学
Abstract:We present Caspar, a library that makes the power of modern GPUs more accessible in robotics and provides a state-of-the-art nonlinear GPU solver that can be applied to a wide range of different optimization problems. Caspar bridges the gap between expressive symbolic programming in Python and high-performance GPU runtimes in C++ by automatically generating optimized CUDA kernels from symbolic expressions. Building on the SymForce library, users can easily define and combine symbolic expressions, including Lie group operations, to generate custom CUDA kernels. To use Caspar as a solver, users need only define the symbolic residual functions; Caspar then uses symbolic differentiation to generate the necessary GPU kernels and interfaces to perform nonlinear optimization. In this paper, we present the core components of Caspar and showcase its performance by performing bundle...
论文介绍 本文介绍了Caspar库,旨在使现代GPU在机器人学中更易使用。它通过自动将符号表达式(包括李群运算)转化为优化的CUDA内核,弥合了Python符号编程与高性能C++ GPU运行时之间的差距。用户仅需定义符号残差函数,系统即可利用符号微分生成GPU内核,用于执行非线性优化,如光束法平差。该工作提升了符号编程在大规模优化问题中的性能。
第一作者: Weizhe Ni · 方向: 机器人操作 · 来源: cs.RO
末端执行器更换工具使用机器人操作任务规划演示学习
Abstract:Robotic manipulation dexterity is often pursued by building increasingly complex high-DoF multifingered hands. While many robotic hands are designed to replicate human morphology, the functional role of human hands suggests a different perspective: much of their complexity may exist to enable tool use and tool making. This observation motivates Any-ttach, a tool-centric manipulation framework that treats quick end-effector swapping as a mechanism for dexterity with simplicity. Any-ttach combines a low-cost automatic swapping mechanism for an open-close robot interface, a handheld device for collecting human demonstrations, and a task planning framework that composes learned, parameterized, and planned tool-use skills. The system supports diverse tools and end-effector modules, including daily tools, articulated tools such as scissors, Fin Ray fingers, and a low-cost...
论文介绍 本研究提出了Any-ttach框架,旨在通过快速更换末端执行器,以实现简单结构下的操作灵巧性,而非依赖复杂的高自由度手部。该框架包含一个低成本的自动更换机制、用于收集人类演示的手持设备,以及一个组合学习、参数化及规划技能的任务规划系统。它支持多种日常工具和模块化末端执行器,展示了通过工具使用实现灵活操作的有效途径。
第一作者: Aaron Kim · 方向: 机器人操作 · 来源: cs.RO
机器人手主动过伸触觉传感精细操作肌腱驱动
Abstract:Manipulating thin objects requires precise contact geometry and reliable force perception, yet many anthropomorphic robotic hands lack the mechanical and sensing capabilities needed for such interactions. We present the ARISTO Hand, a tendon-driven robotic hand that integrates active distal hyperextension with a hybrid fingertip-sensing architecture that combines a rigid, nail-mounted force-torque sensor and a soft capacitive tactile array. Active hyperextension enables controlled fingertip engagement beyond the kinematic limits of standard flexion, increasing pull-out force by 2.76x for object thicknesses of 1-20 mm while preserving the nominal grasp capability. The rigid nail-mounted sensor provides reliable force measurements during edge contacts, where the sensitivity of proprioceptive force estimation degrades as the contact geometry approaches kinematic singularities. We...
论文介绍 针对薄物体操作需求,本文提出了ARISTO Hand,一种集成了主动远端过伸与混合指尖传感架构的肌腱驱动机器人手。主动过伸机制允许指尖运动超出标准屈曲极限,显著提升了对1-20毫米薄物体的拉出力。同时,结合刚性力矩传感器与柔性电容式触觉阵列的混合传感方案,为边缘接触等复杂情况提供了可靠的力感知能力,增强了精细操作的可靠性。
第一作者: Shivendra Agrawal · 方向: 多模态具身 · 来源: cs.RO
视觉语言模型蒙特卡洛定位全局定位语义感知移动机器人
Abstract:Global localization in geometrically aliased, quasi-static environments such as grocery stores, offices, schools, and hospitals poses a significant challenge for mobile robots. Grocery stores with parallel aisles and a long tailed distribution of products, as well as offices and labs with repetitive furniture such as chairs, desks, monitors, and doors, exemplify common indoor environments that present geometric and even semantic ambiguity. Traditional approaches rely either on distinct geometric features or on domain-specific vision pipelines that struggle with long-tail semantic distributions and transient visual clutter. We present VLM-GLoc, a method for hierarchical semantic Monte Carlo Localization (MCL) that leverages open-vocabulary Vision-Language Models (VLMs) as a unified semantic observation front-end. We hypothesize a three-fold benefit from VLMs: (1) extracting...
论文介绍 在几何结构相似且动态变化的杂乱环境(如杂货店、办公室)中,移动机器人的全局定位面临挑战。本文提出VLM-GLoc方法,利用开放词汇的视觉语言模型作为统一的语义观测前端,进行层次化的语义蒙特卡洛定位。该方法假设VLM能提供鲁棒的语义特征,有效缓解几何和语义歧义,提升在复杂准静态环境中的定位成功率。
第一作者: Zhihao Cao · 方向: 导航与运动 · 来源: cs.RO
协同SLAM单目稠密重建3D重建先验户外建图多机器人系统
Abstract:Collaborative dense SLAM is essential for multi-robot teams to achieve scalable and consistent 3D perception across large-scale outdoor environments. Existing systems typically depend on depth sensors, incurring significant payload, power, and calibration costs. Monocular RGB cameras are a lightweight alternative, but collaborative monocular dense SLAM remains difficult due to scale ambiguity, unreliable inter-agent data association, especially in outdoor scenes where low overlap and repetitive structures make traditional feature matching unreliable, motivating robust geometric information. We propose CoMo3R-SLAM, the first collaborative monocular dense RGB SLAM system that leverages robust learned feed-forward 3D reconstruction priors for outdoor multi-agent mapping. Each agent runs a prior-guided front-end for real-time tracking and local dense fusion, while a coordinator...
论文介绍 针对多机器人团队在大规模户外环境中实现轻量化协同稠密感知的需求,本文提出了CoMo3R-SLAM系统。这是首个利用鲁棒的学习型前馈3D重建先验来支持户外多智能体单目RGB稠密建图的系统。每个智能体运行一个先验引导的前端进行实时跟踪与局部稠密融合,协调器负责处理跨智能体的数据关联与全局优化。
第一作者: Zeyuan He · 方向: VLA 通用模型 · 来源: cs.RO
视觉语言动作模型4D监督时空预测机器人操作具身中心
Abstract:Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing actions from current observations, without explicitly modeling future dynamics. This limits their ability to generalize under out-of-distribution perturbations. To address this issue, we propose ELAN4D, an embodiment-centric, 4D-aware training framework that enhances VLA policies with future robot keypoint tracks as predictive spatio-temporal supervision. Using only forward kinematics from proprioceptive states, we derive 3D displacement tracks of robot keypoints, such as joints and the end-effector, with negligible preprocess cost. These tracks provide metric and compact supervision without requiring external trackers or reconstruction. A plug-and-play auxiliary branch with a lightweight track decoder injects this 4D signal into...
论文介绍 现有视觉语言动作模型多为反应式,直接回归当前观测的动作,缺乏对未来动态的显式建模,限制了其泛化能力。本文提出ELAN4D框架,通过提供具身中心的4D监督来增强VLA策略。它仅利用本体感受状态的运动学信息,推导出机器人关键点(如关节、末端)的三维位移轨迹作为预测监督,并通过一个即插即用的轻量解码器分支注入VLA模型。
第一作者: Tri-Tin Nguyen · 方向: 导航与运动 · 来源: cs.RO
室内导航学习型规划全局规划局部规划动态窗口法
Abstract:This paper presents a learning-based navigation framework for indoor mobile robots. The proposed method combines a supervised neural global planner, trained from cost-aware A* expert trajectories, with the proposed Learning-Based DWA local planner, which is formulated as discrete candidate selection over the Dynamic Window Approach (DWA) action lattice. For local planning, the policy is first trained by behavior cloning and then refined by Proximal Policy Optimization (PPO) under feasibility-aware masking. The framework is implemented and evaluated in both simulated and real-world indoor environments. Experimental results show that the proposed method generates feasible global routes and reliable local motion commands for safe goal-directed navigation in the presence of obstacles. These results demonstrate the effectiveness of integrating learning-based global planning with...
论文介绍 本文提出一个面向室内移动机器人的学习型导航框架。该框架结合了基于监督学习的神经全局规划器和所提出的基于学习的DWA局部规划器。全局规划器从带成本信息的A*专家轨迹中训练,局部规划器则通过行为克隆和可行性感知的PPO进行训练与优化。该系统在仿真和真实室内环境中被验证,能够生成可行的全局路径与安全的局部运动指令。
第一作者: Junping Wang · 方向: 数据集与评测 · 来源: cs.RO
多机器人协调通信结构模型扩展系统设计运输建图任务
Abstract:Scaling individual robot capabilities is common but costly. Here we investigate a system-level design question in real-world multi-robot coordination: given matched hardware budgets, does restructuring communication among robots yield larger gains than increasing onboard model size? Using a representative transport-and-mapping task with 10 physical robots (5 runs per condition, 60 runs total), we find that switching from fully connected to modular hierarchical interactions improves normalised performance by 47 points (0--100), whereas doubling neural network hidden size yields at most 9 points. Nested mixed-effects model comparisons show a substantially larger improvement in model fit for topology than for scale. The pattern is confirmed in independent SMAC replications; heterogeneous benchmark reanalyses provide secondary supporting consistency checks rather than primary...
论文介绍 本研究探讨了在硬件预算固定时,优化多机器人系统中的通信结构与扩大单个机器人模型规模,哪种策略更能提升整体性能。通过10个物理机器人执行运输建图任务的实验发现,将通信拓扑从全连接切换为模块化层次结构带来的性能提升(47分),远大于将神经网络隐藏层大小翻倍带来的提升(至多9分)。这表明结构化交互设计比单纯模型扩展更有效。
第一作者: Chalamalasetti Kranti · 方向: 多模态具身 · 来源: cs.RO
视觉语言模型空间推理多智能体对话结构搭建任务
Abstract:Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected to communicate this understanding through language. Vision-language models (VLMs) support robotic tasks involving visual interpretation, question answering, and instruction following, but their capabilities in collaborative dialogue tasks requiring spatial reasoning remain underexplored. We study this gap through a collaborative structure-building task that combines visual interpretation, grounding, language-guided interaction, and action generation. We develop a framework in which VLMs use dialogue to reconstruct a target structure from visual and textual inputs. We evaluate open-weight and closed VLMs across interaction settings, input modalities, and image representations. Results show that spatial reasoning over visual...
论文介绍 本文研究了视觉语言模型在需要空间推理的协作对话任务中的不足。研究者构建了一个框架,使VLM能够通过多轮对话,根据视觉和文本输入重建目标结构,以此来评估其空间推理能力。结果表明,通过对话进行协作重建可以轻微提升模型性能,揭示了该任务在机器人协作领域的潜在应用价值。
第一作者: Jun Wang · 方向: 导航与运动 · 来源: cs.RO
碰撞接地视觉语言模型人机协作安全评估基准测试
Abstract:Safe human--robot collaboration requires more than visual description: a monitor must determine whether the robot body is safely separated, already colliding with the scene or a person, or about to collide. We call this capability collision grounding: binding visual observations to robot body geometry, camera viewpoint, scene layout, human proximity, and temporal motion in order to infer present and imminent contact. We introduce TouchSafeBench, a physics-grounded benchmark for evaluating collision grounding in vision-language models (VLMs). Built in Habitat~3.0, TouchSafeBench contains 2,940 simulated indoor co-presence episodes across social navigation and social rearrangement, with synchronized multi-view RGB-D observations, top-down trajectory maps, calibrated camera metadata, and simulator-derived contact labels. We study two deployment-facing tasks: classifying the...
论文介绍 为实现安全的人机协作,监控系统必须具备“碰撞接地”能力,即判断机器人是否安全、已碰撞或即将碰撞。本文引入了TouchSafeBench基准,这是一个基于物理仿真的评估工具,专门用于测试VLM的碰撞接地能力。该基准包含室内社交导航和社交整理场景的仿真数据,可用于评估模型对接触风险的理解。
第一作者: Varun Nair · 方向: 机器人操作 · 来源: cs.RO
自我相机定向运动学耦合模仿学习视觉概念
Abstract:Recovering ego-camera orientation from manipulation video is a prerequisite for disentangling hand motion from camera motion, a key step in imitation learning from egocentric demonstrations. The obvious approach, inferring orientation from scene geometry, fails when hands occlude the frame: VGGT, a 1B-parameter scene reconstruction model, scores worse than a constant predictor on the TACO benchmark. We identify an alternative visual concept that is present precisely when scene geometry is absent: kinematic coupling dynamics, the structured physical relationship between wrist motion and camera orientation imposed by the arm-shoulder-head chain. We find that this concept is compact (4D inter-wrist features outperform 126D full hand keypoints), temporal (requiring a GRU over short windows rather than per-frame retrieval), and physically grounded (transferring zero-shot across...
论文介绍 从自我中心操作视频中恢复相机定向是模仿学习的关键步骤。当手部遮挡画面时,基于场景几何的推理会失效。本文提出,运动链(如手腕与相机之间的物理关系)所形成的运动学耦合动力学,可作为一个有效的可学习视觉概念。实验表明,利用该概念的简洁时序特征,能够实现跨任务的零样本迁移,有望提升从演示视频中学习的鲁棒性。
第一作者: Anya Singh · 方向: VLA 通用模型 · 来源: cs.RO
VLA策略测试时扩展安全保证共形推理弃权机制
Abstract:Test-time scaling for vision-language-action (VLA) policies, methods such as RoboMonkey, SEAL, MG-Select, and V-GPS, samples K candidate action chunks at inference and executes the verifier-best. When all K candidates are unsafe, the system executes a violating action with no warning. We propose BOKBO, the first conformal abstention layer for K-sample VLA inference, providing finite-sample distribution-free guarantees on executed-violation rate. We provide both global and per-task (Mondrian) variants, with the per-task variant closing the conditional gap on the hardest tasks. Our analysis exposes a structural failure of policy-internal nonconformity scores under perturbation-based K-sampling: the base-policy confidence proxy and K-sample disagreement correlate at 0.98 with the action-noise hyperparameter $\sigma$, while correlating at the noise floor with actual safety...
论文介绍 当前基于采样的视觉语言动作模型推理方法,在所有候选动作均不安全时会盲目执行。本文提出BOKBO,这是首个用于K样本VLA推理的共形弃权层,能够在有限样本下提供关于执行违规率的分布无关保证。该方法通过引入弃权机制,为机器人策略在安全关键场景下的部署提供了一种可靠的校准和风险控制方案。
第一作者: Josef Chen · 方向: 具身智能 · 来源: cs.RO
物理AI推理优化批量解码内存带宽GPU评估
Abstract:Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM serving: single-stream, batch-1 autoregressive decode, where one robot, camera feed or user session waits on the next token. This workload is usually described as memory-bandwidth-bound. Each decode step streams model weights and the active KV cache, so latency should scale with peak HBM bandwidth. We show that this account is true but incomplete. We measure batch-1 decode for three 7 to 8B-class GQA transformers across four NVIDIA GPUs: H100 SXM5, A100-80GB SXM4, L40S and L4. We evaluate context lengths from 2048 to 16384, producing 44 valid cells under a controlled bf16 SDPA setup. The achieved fraction of peak HBM bandwidth falls as peak bandwidth rises. On the headline Qwen-2.5-7B ctx=2048 cell, an L4 reaches roughly 81...
论文介绍 物理AI系统(如机器人、自动驾驶车辆)通常运行单序列、批量为1的大语言模型解码任务,该任务常被描述为内存带宽受限。本文通过系统测量多种GPU在不同配置下的性能,发现实际达到的内存带宽利用率会随峰值带宽的增加而下降。研究表明,这一推理工作负载的瓶颈不仅是内存带宽,对优化物理AI的部署至关重要。
第一作者: Weicheng Zheng · 方向: VLA 通用模型 · 来源: cs.CV
自动驾驶视觉语言动作模型元动作端到端规划强化学习
Abstract:Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise. We propose DriveMA, a Driving VLA framework built on verifiable meta-actions, which summarize future ego motion into compact language-domain intentions and can be constructed from expert trajectories with a trajectory-grounded annotation pipeline and can be verified against generated trajectories through rule-based projection. DriveMA exploits this verifiability with action-centric supervised training and a data-efficient turn-level credit assignment reinforcement learning framework, explicitly aligning high-level decisions with low-level trajectory planning through dense rewards and precise credit assignment. DriveMA sets a new state of the art on the Waymo Open Dataset Vision-based E2E Driving, achieving a Rater Feedback...
论文介绍 驾驶视觉语言动作模型旨在用语言改进端到端规划,但存在语义差距。本文提出DriveMA框架,其核心是“可验证的元动作”,即将未来车辆运动总结为紧凑的、可验证的意图标签。该框架通过基于轨迹的标注流程和规则投影验证,并结合动作中心的监督训练与高效的强化学习,实现了高层决策与底层轨迹规划的对齐,在Waymo数据集上取得了领先性能。
第一作者: Jingtao He · 方向: VLA 通用模型 · 来源: cs.CV
视觉语言动作模型自动驾驶视觉依赖扰动分析评估框架
Abstract:Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal architectures for jointly modeling perception and planning. However, how current VLA-based driving behavior is grounded in visual information remains poorly understood. Existing evaluation protocols mainly focus on aggregate performance metrics, lacking structured and practical diagnostics to quantify visual-behavior dependency. In this work, we introduce a structured multi-level visual perturbation framework to analyze visual-behavior dependency in VLA-based driving models systematically. The framework organizes controlled visual perturbations along three complementary dimensions: channellevel degradation, information-level disruption, and structurelevel modification. We apply it to VLA-based driving systems and evaluate behavioral...
论文介绍 当前自动驾驶的VLA模型行为如何依赖于视觉信息尚不清楚。本文提出了一个结构化的多级视觉扰动分析框架,从通道、信息和结构三个互补维度系统地量化视觉-行为依赖关系。通过对驾驶系统施加受控扰动,该框架能够诊断模型是过度依赖视觉细节还是忽略了关键视觉信号,为理解与改进VLA模型提供了实用工具。
第一作者: Zhiyuan Yang · 方向: 多模态具身 · 来源: cs.CV
病理报告生成视觉语言模型全切片图像效率优化病例级推理
Abstract:Generating clinically useful pathology reports for pathology cases from whole-slide images (WSIs) is challenging due to gigapixel resolution, long visual-token sequences, and the complexity of case-level reasoning, where a single case may contain multiple WSIs with heterogeneous tissues and ambiguous findings. We present a simple token-efficient vision--language model for case-level synoptic report generation that remains practical under constrained GPU memory. Our architecture follows a minimal three-component design: a frozen pathology patch encoder, a lightweight two-layer MLP vision-language aligner, and a large language model decoder, with an explicit WSI marker token to separate slides within a case. Training proceeds in two supervised stages: (1) aligner-only WSI captioning using heterogeneous WSI-text pairs, and (2) case-level supervised fine-tuning on case-report...
论文介绍 从高分辨率的全切片图像生成病例级病理摘要报告极具挑战性。本文提出了一种简单高效的视觉语言模型架构,包含冻结的病理图像编码器、轻量级对齐层和大型语言模型解码器。通过采用明确的标记分隔不同切片并进行两阶段训练,该模型能够在GPU内存受限的情况下完成从单切片描述到完整病例报告的生成,具有临床应用潜力。
第一作者: Adam J. Thorpe · 方向: 具身智能 · 来源: cs.AI
世界模型具身智能物理可行性查询条件
Abstract:World models for embodied AI must be physically viable: constructed to answer intervention queries by representing the physical structure governing action outcomes, rather than merely predicting future observations. Existing observation-predictive world models can produce visually plausible but physically wrong rollouts. This failure is structural; distinct physical systems can look identical yet diverge under intervention. We expose this problem with controlled benchmarks that fix the visible scene while varying latent physics. We show that such models may recommend infeasible actions, mispredict interaction outcomes, or certify unsafe behavior. We argue that embodied AI requires world models that identify the simplest physical abstraction sufficient to answer an intervention query. Such a model comprises modular components, including environment representation, latent state...
论文介绍 本文针对具身AI中的世界模型物理可行性问题展开讨论。现有基于观察预测的世界模型可能产生视觉合理但物理错误的预测,导致推荐不可行动作或预测失败。作者指出,世界模型应能通过表示物理结构来回答干预查询。为此,提出了查询条件世界模型,旨在识别足够应对干预查询的最简单物理抽象,以提升具身AI系统的可靠性和安全性。
美股市场技术面偏强但临近超买。标普500ETF(SPY)和纳斯达克100ETF(QQQ)均位于所有主要均线之上,呈现多头排列,但RSI(14)分别高达75.1和78.4,进入超买区域,短期存在技术性回调压力。科技股表现分化,英伟达(NVDA)单日涨6.26%但RSI处于正常范围,而特斯拉(TSLA)和Meta(META)则出现MACD死叉或处于空头趋势。加密货币市场整体疲软,恐慌贪婪指数为23(极度恐慌),总市值24小时下跌1.64%。比特币(BTC-USD)和以太坊(ETH-USD)价格均低于所有关键均线(SMA20/50/200),RSI分别处于30.1和30.9的超卖边缘,趋势为空头。中概股技术面普遍承压,阿里巴巴(BABA)、拼多多(PDD)等均呈空头排列,RSI在40以下,显示弱势。腾讯控股(0700.HK)虽单日上涨,但仍受长期均线压制。商品与外汇方面,黄金期货(GC=F)价格位于SMA200之上但低于SMA20和50,信号中性;美元指数(DX-Y.NYB)多头排列并接近52周高点,显示强势。
当前价格742.74,位于SMA20(712.47)、SMA50(655.93)和SMA200(618.66)之上,呈现多头排列。MACD(21.74)上穿信号线(21.50)形成金叉,动量指标偏强。RSI(14)读数78.4,处于超买区域。近5日涨幅为3.51%,短期趋势向上,但RSI超买提示潜在回调压力。
当前价格460.52,高于其SMA20(419.9)和SMA50(404.26),短期均线提供支撑。MACD(8.377)已上穿信号线(4.9745)形成金叉。RSI(14)为72.9,进入超买区域。近5个交易日涨幅高达10.02%,短期动能强劲,但技术指标显示超买状态。
当前价格456,接近SMA20(452.36)但低于SMA50(480.28)和SMA200(573.51),长期仍处空头排列。MACD(-13.4394)上穿信号线(-14.4408)形成金叉,显示短期下跌动能可能减弱。RSI(14)为48.5,处于中性区域。近1日上涨4.59%,但整体趋势指标矛盾,处于整理状态。
VIX 恐慌指数
10Y 美债收益率 (%)
美元指数 DXY
S&P 500 ETF
Nasdaq 100 ETF
Apple
Microsoft
Nvidia
Alphabet
Tesla
Meta
Bitcoin
Ethereum
Solana
阿里巴巴 (BABA)
拼多多 (PDD)
京东 (JD)
腾讯控股 (0700.HK)
黄金期货
WTI 原油期货
美元 / 人民币
本报告所有技术指标解读均基于历史公开数据计算,过去走势不代表未来表现。市场存在多重不确定性,技术指标可能失效。本报告内容仅供技术指标解读参考,不构成任何投资建议或决策依据。
The Bundibugyo virus, a little known type, previously had caused just two small outbreaks. Now it’s at the center of a rapidly widening epidemic in Africa.
中文摘要 科学家正加速研发埃博拉病毒疫苗和治疗方法。班迪布焦病毒此前仅引发两次小规模爆发,如今在非洲迅速扩散。
The US president said Benjamin Netanyahu had promised not to send troops to Beirut, while Hezbollah had agreed that ‘all shooting will stop’ The exchange of strikes between the US and Iran reflects the fragility of the current ceasefire, which has seen repeated violations even as American and Irania
中文摘要 特朗普宣布停止射击后,以色列称北部遭火箭弹袭击。特朗普表示内塔尼亚胡承诺不派兵至贝鲁特,真主党同意停火,但美伊交火显示停火协议脆弱。
Unions had demanded a higher 6% pay increase after last month’s budget projected inflation reaching 5% in the year to June. Follow today’s news live Get our breaking news email, free app or daily news podcast The foreign minister, Penny Wong, has announced new sanctions by the commonwealth over thre
中文摘要 澳大利亚公平工作委员会裁定,近300万最低工资工人将获得4.75%加薪。工会此前要求6%加薪,因预算预测通胀达5%。财政部长称经济问题正驱使选民转向一国党。
President Trump wrote on social media that Israel and Hezbollah had “agreed” not to attack each other. Prime Minister Benjamin Netanyahu later distanced himself from talks of a cease-fire in Lebanon.
中文摘要 特朗普在社交媒体表示,以色列和真主党已同意不互相攻击。内塔尼亚胡随后淡化黎巴嫩停火谈判,以色列和伊朗在紧张对峙后有所缓和。
Unions had demanded 6% pay increase for lowest paid after war in Middle East pushed inflation higher Follow our Australia news live blog for latest updates Get our breaking news email, free app or daily news podcast Nearly 3 million workers will receive a 4.75% pay rise, while about 100,000 of the c
中文摘要 澳大利亚公平工作委员会裁定,约300万最低工资工人将获得4.75%加薪。工会此前要求6%加薪,因中东战争推高通胀。
Death toll from Israel's attacks on Lebanon since March has reached 3,433, with 10,395 injured, says Health Ministry.
中文摘要 特朗普与真主党和以色列会谈,黎巴嫩战斗持续。据黎巴嫩卫生部数据,自3月以来以色列袭击已造成3,433人死亡,10,395人受伤。
中文摘要 暂无详细信息。
Hours after the pause was announced, the Israeli military said it had intercepted projectiles fired from Lebanon.
中文摘要 黎巴嫩宣布真主党同意停止对以色列的袭击。但停火后数小时,以色列军方称拦截了来自黎巴嫩的发射物。
Abelardo De La Espriella, a hard-line candidate on the right facing a left-wing rival in a presidential runoff, has pledged to crush drug traffickers.
中文摘要 特朗普有望在哥伦比亚选举中获得盟友。右翼候选人Abelardo De La Espriella将在决选中面对左翼对手,承诺严厉打击毒贩。
Police seen dragging ultra-Orthodox protesters from under a bus after they blocked roads in Jerusalem over conscription.
中文摘要 极端正统派抗议者因征兵问题在耶路撒冷封锁道路,与以色列警方发生冲突。警方将抗议者从公交车下拖出。
The candidate, Abelardo de la Espriella, will face a senator from the left-wing party of the departing president, Gustavo Petro, in a June runoff.
中文摘要 哥伦比亚总统选举进入决选。候选人Abelardo de la Espriella将在6月决选中面对左翼政党参议员。
Mette Frederiksen may not be nearly as popular as she once was, but she remains the Danes’ most dominant leader in decades.
中文摘要 梅特·弗雷泽里克森在丹麦组建新政府。尽管人气下降,她仍是丹麦数十年来最具主导地位的领导人。
Labour unions and student groups protested in Chile as President Kast delivered his first State of the Nation.
中文摘要 智利因政府削减社会项目爆发暴力抗议。工会和学生团体在总统卡斯特发表国情咨文期间举行抗议活动。
The decision was split, allowing the Trump administration to continue barring transgender people from enlisting.
中文摘要 美国法院维持对特朗普禁止跨性别者参军政策的禁令。裁决存在分歧,但允许特朗普政府继续实施该禁令。
Once widespread in Japan, the colorful birds went from being fairly commonplace in the country to being on the verge of extinction.
中文摘要 日本朱鹮再次飞翔,享受特殊保护。这些鸟类曾广泛分布,一度濒临灭绝,现获重新引入。
Congo reopened the main airport in the eastern province hardest hit by Ebola, after health officials reported tentative signs the outbreak may be slowing despite a continuing struggle to trace exposed contacts and investigate suspected cases.
中文摘要 刚果重新开放了受埃博拉疫情最严重打击的东部省份主要机场,卫生官员报告疫情可能放缓的初步迹象,但追踪接触者和调查疑似病例仍面临挑战。
One of the most bitter feuds in distressed-debt investing, between billionaire Patrick Drahi and some of the world’s biggest investors, has suddenly gotten even more contentious.
中文摘要 亿万富翁帕特里克·德拉希与全球最大投资者之间在不良债务投资领域的激烈争端因资产转移而进一步加剧,接近爆发边缘。
South Korea’s equity market has overtaken India’s as the world’s sixth largest, driven by a relentless surge in chip heavyweights powering the global artificial intelligence buildout.
中文摘要 韩国股市超越印度,成为全球第六大股票市场,主要受芯片巨头推动全球人工智能建设的强劲增长驱动。
To make real estate attractive for investors, prices must go lower
中文摘要 中国房地产市场可能还有进一步下跌空间,为吸引投资者,价格需进一步下降。
Chinese who spent more than ever on Hong Kong real estate in the first quarter of this year now face a hurdle to overcome as rules get stricter on wealthy people in the mainland taking cash overseas.
中文摘要 今年第一季度中国大陆投资者在香港房地产创下购买纪录,但随着内地对富裕人士现金出境的监管收紧,这一热潮面临障碍。
Andrew Left, one of the world’s most prominent short sellers, was found guilty of securities fraud by a federal jury after a landmark trial that scrutinized his use of social media to move the price of stocks.
中文摘要 知名做空者安德鲁·莱夫特被联邦陪审团裁定犯有证券欺诈罪,审判聚焦于他利用社交媒体影响股价的行为。
Bloom Energy Corp., a supplier to Oracle Corp. that’s seen its stock price double in the past two months, doesn’t see a need to sell shares to meet surging demand from data centers, its chief executive officer says.
中文摘要 Bloom Energy公司股价在过去两个月翻倍,首席执行官表示无需出售股份以满足数据中心需求的激增。
Case has thrown spotlight on dealings with Wall Street and Hollywood elites including billionaire Steve Ballmer
中文摘要 Aspiration Partners联合创始人乔·桑伯格因欺诈罪被判14年监禁,案件涉及与华尔街和好莱坞精英的往来。
vVardis Holding AG, a Swiss dental startup, is working with banks for a US initial public offering that could come this year, according to people familiar with the matter.
中文摘要 瑞士牙科初创公司vVardis正与银行合作,计划今年在美国进行首次公开募股。
Landmark fundraising plans include $10bn private placement to Berkshire Hathaway
中文摘要 Alphabet计划出售价值800亿美元的股票以资助人工智能投资热潮,包括向伯克希尔·哈撒韦进行100亿美元的私募配售。
Gold held a decline as conflicting signals from the US and Iran cast doubt over a diplomatic resolution to the war, fanning concerns over inflation and prolonged trade disruptions.
中文摘要 黄金价格维持下跌,因美国和伊朗之间的矛盾信号令外交解决战争的可能性存疑,加剧通胀和贸易中断的担忧。
Some firms are putting pressure on staff to use AI, but have not thought through their AI rollout.
中文摘要 一些公司施压员工使用人工智能,但未充分考虑AI的部署策略,导致混乱并损害企业效益。
18 回复 · 程序员 节点
7 回复 · 程序员 节点
4 回复 · 程序员 节点
9 回复 · 程序员 节点
5 回复 · 程序员 节点
4 回复 · Apple 节点
7 回复 · Linux 节点
13 回复 · Apple 节点
9 回复 · Apple 节点
4 回复 · Apple 节点
一 一切开始的时候,我只是自己用 AI,自己折腾。先是 cliproxyplus 反代 Copilot,后来反代 GPT,再后来是那些比较麻烦但是便宜的、经济的渠道。 直到有朋友说,搞不好可以做点中转。 我想了想,觉得好像也不是不行。 后来我才知道,一件事说"好像也不是不行"的时候,多半是要为它付出代价的。 二 真正商业化是 5 月 13 号。 那天之前,我手里有一些 OpenAI Plus 号,还有当时 KIRO 那边的一个口子,可以低成本做出 Pro 号。我甚至写了一个浏览器自动化注册机,除了验证码那一步,其它都能跑。一个 GitHub 账号两块钱,一张卡两三块钱,一个号成本四五块,我卖十
很久没发话题了,上一次分享好东西还是两年前,当时我在l站分享了白嫖40显存专业显卡搭建cogvideo和flux等开源大模型之类的教程。 由于部署模型和vps成本较高,暂时只开放注册24小时,以后随缘开放,还有就是以后会上架闭源视频生成大模型,grok和happyhorse等等,也是1ldc生成一次,其他类型的开源模型也会考虑上架,比如腾讯开源的3d模型生成,还有ai变音模型。 l0veyou.com l0veyou ChatGPT account pool management dashboard 声明一下,支付系统有bug,大家先别充值,先别充值!!!我会给每个新用户赠送100-1000积
6.1儿童节快乐,没什么福利可发,把在网站上节省来的钱做成口令红包给大家了,金额不大,拼手气吧 送10个openai接码(不能二验),一次性 - 福利羊毛 / 福利羊毛, Lv1 - LINUX DO 送10个openai接码(不能二验),一次性 - 福利羊毛 / 福利羊毛, Lv1 - LINUX DO 总共都没到20层,后面的我就挪到这里来了,21楼、34楼、55楼、89楼、144楼私我领码子。 35 个帖子 - 34 位参与者 阅读完整话题
呜呜呜qwq ff想要毛茸茸的大耳朵和软软的小爪爪w 还有就是要找一个温暖的猫窝w( 还要抱抱w 14 个帖子 - 13 位参与者 阅读完整话题
信源来自 minimax 官网:(截图时间 20260601 晚上 21:44) MiniMax API Platform MiniMax API 开放平台 单价:(月度订阅和年度订阅都是海外官网更便宜) Token Plan API 退款方法: https://linux.do/t/topic/2284553 minimax 想要道歉的话,可以说是海外定价策略失误,因为你们海外max的订阅性价比反而比ultra的高 不过好像也道歉不了喔,国内 ultra 是8.5RMB / 亿 tokens,国外 ultra是 8.27RMB/亿 tokens,而且 api 价格也是国外更便宜,还是专坑国人
11 个帖子 - 11 位参与者 阅读完整话题
需要申请退款,直接在小店申请即可,无需说明原因,5月10日之后购买的我都会同意退款。 14 个帖子 - 11 位参与者 阅读完整话题
国模里面,也就剩deepseek值得支持了,感谢梁圣 环顾一周,想说支持国模,真的很难 47 个帖子 - 45 位参与者 阅读完整话题
发现学校里有台校长不要的电脑,直接就把它搬到家里云里好吧 27 个帖子 - 25 位参与者 阅读完整话题
-------------------------------二编--------------------------------- CPA Usage Keeper。跑了一个20x pro 179刀掉了39% 5h(有fast有standard,有5.5也有5.4) 188 个帖子 - 118 位参与者 阅读完整话题