palmier-io/palmier-pro
Swift · ★ 3,363 · 🍴 280 · 📈 902 stars today
macOS video editor built for AI
中文介绍 一款专为AI设计的macOS视频编辑器。它利用人工智能技术简化视频剪辑流程,旨在让内容创作者和视频编辑人员能够更高效地完成后期制作,可能集成了自动剪辑、智能特效生成等功能。
Swift · ★ 3,363 · 🍴 280 · 📈 902 stars today
macOS video editor built for AI
中文介绍 一款专为AI设计的macOS视频编辑器。它利用人工智能技术简化视频剪辑流程,旨在让内容创作者和视频编辑人员能够更高效地完成后期制作,可能集成了自动剪辑、智能特效生成等功能。
Clojure · ★ 51,431 · 🍴 3,303 · 📈 420 stars today
Penpot: The open-source design tool for design and code collaboration
中文介绍 一款开源的设计工具,专注于设计与代码的协作。它为UI/UX设计师和开发者提供了一个共同的工作平台,支持实时协作和代码导出,以减少设计到开发过程中的沟通成本,提升团队效率。
Python · ★ 7,057 · 🍴 1,151 · 📈 677 stars today
World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
中文介绍 这是全球首个开源的智能视频生产系统。它提供了12个处理流程、52个工具和超过500项智能技能,能将AI编码助手转变为一个完整的视频制作工作室,适用于需要自动化、批量生成视频内容的创作者和团队。
Rust · ★ 20,328 · 🍴 1,042 · 📈 801 stars today
Turso is an in-process SQL database, compatible with SQLite.
中文介绍 一个与SQLite兼容的进程内SQL数据库。它基于SQLite但进行了优化,提供更快的读写性能和更低的资源占用,特别适合嵌入式系统、移动应用或边缘计算场景中,需要轻量级本地数据存储的开发者。
C · ★ 9,355 · 🍴 708 · 📈 1,271 stars today
High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.
中文介绍 一个高性能的代码智能MCP服务器。它能将代码库快速索引成一个持久化的知识图谱,支持158种编程语言,实现亚毫秒级查询,并大幅减少AI上下文消耗。开发者可用它高效理解、导航和搜索大型代码库。
Python · ★ 24,548 · 🍴 2,328 · 📈 433 stars today
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
中文介绍 由Google Research开发的时间序列基础模型。它是一个预训练模型,专门用于时间序列预测任务,例如金融市场趋势、气象数据或资源需求预测,为数据科学家和分析师提供强大的预测分析工具。
TypeScript · ★ 50,863 · 🍴 7,412 · 📈 140 stars today
The open alternative to Salesforce, designed for AI.
中文介绍 一个为AI时代设计的开源客户关系管理(CRM)系统,定位为Salesforce的替代方案。它可能集成了AI驱动的自动化功能,帮助企业更智能地管理客户互动、销售流程和业务数据。
TypeScript · ★ 39,329 · 🍴 2,324 · 📈 329 stars today
The open-source, cross-platform API client for GraphQL, REST, WebSockets, SSE and gRPC. With Cloud, Local and Git storage.
中文介绍 一款开源、跨平台的API客户端,支持GraphQL、REST、WebSockets、SSE和gRPC等多种协议。它提供云端、本地和Git存储方案,是后端开发者和API测试人员进行接口调试、测试和协作的得力工具。
Rust · ★ 54,772 · 🍴 10,877 · 📈 2,546 stars today
🤱🏻 Turn any webpage into a desktop app with one command.
中文介绍 一个命令行工具,能将任意网页快速打包成一个独立的桌面应用程序。它基于WebView技术,简化了将Web应用转换为桌面端的过程,适用于希望获得原生桌面体验的普通用户和开发者。
Python · ★ 41,874 · 🍴 2,878 · 📈 3,795 stars today
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.
中文介绍 一个用于压缩数据、日志和RAG上下文块的工具,旨在它们被送入大型语言模型之前减少Token消耗。它提供了库、代理和MCP服务器等形式,可在保证答案质量的前提下,降低60%-95%的AI处理成本。
TypeScript · ★ 31,021 · 🍴 3,835 · 📈 145 stars today
The open-source AI voice studio. Clone, dictate, create.
中文介绍 一个开源的AI语音工作室。它提供语音克隆、语音听写和语音创建功能,允许用户利用人工智能技术生成、编辑或定制语音内容,适合播客制作者、内容创作者或需要语音合成的开发者。
TypeScript · ★ 23,359 · 🍴 2,742 · 📈 513 stars today
Kilo is the all-in-one agentic engineering platform. Build, ship, and iterate faster with the most popular open source coding agent.
中文介绍 一个集大成的智能工程平台,内置了流行的开源编码代理。它旨在通过AI辅助编码、任务自动化和快速迭代,帮助开发团队提升构建、交付和优化软件产品的整体效率。
Shell · ★ 138,237 · 🍴 11,994 · 📈 1,395 stars today
Skills for Real Engineers. Straight from my .claude directory.
中文介绍 一个面向真实工程师的技能与配置分享库,内容来自作者的.claude目录。它可能包含经过实践验证的开发配置、提示词或工作流,旨在为其他开发者提供参考和灵感,提升AI辅助编码的效率。
TypeScript · ★ 6,095 · 🍴 344 · 📈 316 stars today
The sandbox agent framework.
中文介绍 一个沙盒化的智能代理框架。它为开发者提供了一个安全、隔离的环境来构建、测试和运行基于AI的智能代理(Agent),是探索和开发下一代自主AI应用的基础设施。
★ 14,784 · 🍴 2,355 · 📈 48 stars today
A curated list of Artificial Intelligence (AI) courses, books, video lectures and papers.
中文介绍 一份精心整理的人工智能领域资源列表。它汇集了相关的课程、书籍、视频讲座和学术论文,为学生、研究人员和从业者提供了一个系统性的AI学习与参考资料集合。
Kotlin · ★ 26,488 · 🍴 3,273 · 📈 104 stars today
短信转发器——监控Android手机短信、来电、APP通知,并根据指定规则转发到其他手机:钉钉群自定义机器人、钉钉企业内机器人、企业微信群机器人、飞书机器人、企业微信应用消息、邮箱、bark、webhook、Telegram机器人、Server酱、PushPlus、手机短信等。包括主动控制服务端与客户端,让你轻松远程发短信、查短信、查通话、查话簿、查电量等。(V3.0 新增)PS.这个APK主要是学习与自用,如有BUG请提ISSUE,同时欢迎大家提PR指正
中文介绍 一个Android平台的短信转发器应用。它能监控手机短信、来电和APP通知,并根据预设规则自动转发到钉钉、企业微信、飞书、邮箱、Telegram等多种平台,同时支持远程发短信、查通讯录等控制功能。
Rust · ★ 7,395 · 🍴 828 · 📈 87 stars today
Coding Agent Harness
中文介绍 一个编码代理工具平台。它可能提供了一个集成环境或工具集,用于辅助开发者与AI编码代理进行交互、管理任务和优化编码流程,目标是提升人机协同的编程效率。
@horizon_trade_x · 4.4K 粉丝 · 1.3M 阅 · 507 赞 · 59 转
Your backtest looked flawless. You went live. Two weeks later, the strategy was bleeding. Every quant has lived this. The answer is a loop: generate a strategy, test it, score it, feed the result
中文介绍 博主分享量化交易中的循环工程框架,核心是通过“生成策略-测试-评分-反馈”的持续循环来迭代优化,以解决回测完美但实盘失效的常见痛点,强调系统化而非单次策略开发。
@djfarrelly · 3.8K 粉丝 · 344.7K 阅 · 501 赞 · 61 转
Everyone's asking "WTF is a loop?" Here's the question nobody's asking: what runs the loop? The AI discourse has converged on loops as a core primitive of agentic systems. Matt Van Horn (@mvanhorn)
中文介绍 文章探讨AI代理系统中“循环”这一核心原语的架构。博主指出当前讨论多聚焦于循环概念本身,而少有人追问“是什么在驱动循环运行”,旨在从技术层面解析代理循环的运作机制。
@EXM7777 · 118.9K 粉丝 · 107.7K 阅 · 509 赞 · 44 转
for a few days, we had something that felt like AGI... Fable 5 showed up, effectively unlimited inside the plans, and the ceiling on what you could build lifted overnight but then Anthropic killed it,
中文介绍 用户怀念Fable 5发布初期接近AGI的智能水平,但该能力随后被API限制削弱。帖子意在分享如何通过特定方法或工具链,重新获得或接近那种强大的智能水平。
@omarsar0 · 308.0K 粉丝 · 90.2K 阅 · 504 赞 · 69 转
A claim has been circulating in AI coding circles: stop prompting your coding agents and start designing loops that prompt them for you. As with everything new, this stuff gets repeated often and
中文介绍 介绍AI编码领域从“手动提示代理”到“为代理设计自动化循环”的理念转变。循环工程旨在通过预设的循环流程自动触发和优化对编码代理的提示,提升效率与一致性。
@Designarena · 13.9K 粉丝 · 80.4K 阅 · 518 赞 · 39 转
GLM 5.2 ranks 1st overall on Design Arena’s single-turn, HTML Web Design (Non-Agentic) evaluation, 5 places higher than its predecessor GLM-5.1. To do so, it beat Claude Fable 5, Opus 4.6, and Opus
中文介绍 根据Design Arena的评估,GLM-5.2在单轮HTML网页设计任务中排名第一,超越了Claude Fable 5、Opus 4.6等领先模型,展示了其在该特定设计能力上的显著进步。
@nifinet · 10.4K 粉丝 · 62.8K 阅 · 519 赞 · 38 转
A GTM team looks like a sending operation. Most of its real work is judgment: which company is worth a message this week, what to say that proves you noticed, which no-show to chase, what actually
中文介绍 分享如何利用Claude Code构建一个可由个人独立运营的GTM(市场进入)团队。核心是将判断力工作(如客户筛选、信息定制)交给AI,模拟团队协作以提高效率。
@jasonzhou1993 · 30.8K 粉丝 · 48.5K 阅 · 526 赞 · 49 转
At around 1:00 AM yesterday, a bunch of PRs started landing in our codebase. Not because our team was working unusually late. They came from different agent loops: agents finding issues, picking up
中文介绍 作者分享实际搭建并运行“循环工程师”的体验。描述了在凌晨由不同AI代理循环自动发现代码问题、提交PR的场景,展示了代理循环在软件开发中的自动化应用。
@contralabs_ai · 2.9K 粉丝 · 46.6K 阅 · 501 赞 · 37 转
Everyone keeps talking about taste. But you can't improve what you can't measure. So we measured it. Design Crit is a dataset of ten professional designers ranking four frontier image models across
中文介绍 团队发布“Design Crit”数据集,通过十位专业设计师对四个前沿图像模型的排名,首次尝试对AI的“设计品味”进行量化测量,为评估模型能力提供了新基准。
@Just_sharon7 · 44.1K 粉丝 · 32.5K 阅 · 502 赞 · 18 转
If you're still switching between 10 tools to make UGC ads, analyze performance, and iterate... your workflow is outdated. Arcads just dropped MCP support, and pairing it with Claude (especially Fable
中文介绍 介绍广告平台Arcads新增MCP支持后,与Claude(特别是Fable)集成,能将UGC广告的创建、效果分析与迭代优化整合到单一AI工作流中,简化传统多工具切换的流程。
a quiet day lets us promo AIE one last time
中文介绍 今日AI新闻较少,Latent Space借此机会最后一次推广AIE产品。
Miami-based AI startup Subquadratic came out of stealth mode last month with a huge claim. It announced that it had solved a mathematical bottleneck that had been holding back large language models for almost a decade. The details were thin, and many people were unconvinced. But Subquadratic has sta
中文介绍 迈阿密AI初创公司Subquadratic声称解决了阻碍大型语言模型近十年的数学瓶颈,但细节不足,许多人持怀疑态度。
With GLM-5.2 passing everyone's vibe check, the open models story finally becomes a real frontier story.
中文介绍 GLM-5.2通过vibe check测试,开源模型成为前沿技术,Z.ai预计12月发布Open Fable。
**GLM-5.2** emerges as a leading open-weight coding model rivaling **Opus 4.8** and **GPT-5.5** in software engineering tasks, emphasizing the strategic importance of open models for provider competition, on-prem deployment, and fine-tuning rights. Experts like **Patrick Toulme** and **Thomas Wolf**
中文介绍 GLM-5.2成为领先开源编码模型,在软件工程任务中与Opus 4.8和GPT-5.5竞争,专家Patrick Toulme强调开源模型的战略重要性。
中文介绍 今日AI新闻涉及GPT-5.6更新、Claude代码工件功能以及Perplexity的Brain memory系统。
中文介绍 文章探讨MosaicLeaks主题,询问研究agent是否能保守秘密。
We talk about how this legendary investor went from humble beginnings in Singapore to leading rounds in Anthropic, Mistral, Black Forest Labs, and Periodic Labs... and the AMP secret master plan!
中文介绍 投资家Anjney Midha从新加坡起步,领投Anthropic、Mistral、Black Forest Labs和Periodic Labs等公司,并分享AMP的秘密计划。
OpenAI introduces new spend controls and usage analytics for ChatGPT Enterprise, helping organizations manage costs and scale AI with confidence.
中文介绍 OpenAI为ChatGPT企业版引入新的花费控制和使用分析功能,帮助企业管理和扩展AI成本。
Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness responses with stronger reasoning, better context, clearer communication, and physician-informed evaluations.
中文介绍 GPT-5.5 Instant改进ChatGPT的健康和保健响应,具有更强推理、更好上下文、更清晰沟通和医生评估。
Researchers used an OpenAI reasoning model to help diagnose rare diseases, identifying 18 new diagnoses in previously unsolved cases.
中文介绍 研究人员使用OpenAI推理模型帮助诊断儿童罕见遗传疾病,在未解决病例中识别出18个新诊断。
**GLM-5.2** from **Zhipu** emerged as a leading open-weight model with innovative **IndexShare** sparse-attention enabling efficient **1M-token inference**, praised as comparable to **GPT-5.5** and **Opus 4.8** but lacking vision support. Other notable open models include **Laguna M.1** by **Poolsid
中文介绍 智谱的GLM-5.2成为领先开源模型,具有创新的IndexShare稀疏注意力,支持高效1M-token推理,与GPT-5.5和Opus 4.8相当但缺乏视觉支持。
The only bootstrapped frontier lab announces its second product and second
中文介绍 Midjourney Medical发布第二个产品,涉及器官扫描功能,作为唯一自举的前沿实验室推出的新产品。
中文介绍 文章探讨超越LoRA的微调技术可能性,询问是否能击败最流行的微调方法。
中文介绍 讨论在自有工具上基准测试开源模型的agentic能力,评估是否足够自主。
中文介绍 AI新闻包括ChatGPT市场份额下降、Vercel eve事件以及Replit链接Claude服务。
周末 arXiv 通常无新公告。当前展示最近一次可用公告批次。
第一作者: Shanghao Shi · 方向: 隐私保护
联邦学习隐私后门参数高效微调记忆攻击神经元隐写
Abstract:Federated learning (FL) enables multiple parties to collaboratively fine-tune language models for domain-specific tasks without sharing raw data. Since full model fine-tuning is often prohibitively expensive for FL clients, parameter-efficient fine-tuning (PEFT) has become the de facto approach in practice, freezing the base model and training only a small set of adapters. In this paper, we show that a malicious parameter server can stealthily corrupt a PEFT adapter into a privacy backdoor that implicitly memorizes the client's training samples as isolated per-sample parameter updates stored in separate neurons, without degrading model utility. Concretely, our attack, NeuroImprint, assigns a dedicated memorization neuron to each training sample and constrains that each neuron is updated at most once along the local fine-tuning trajectory. This design mitigates both...
论文介绍 本文研究联邦语言模型微调中的隐私泄露问题。针对参数高效微调场景,提出一种名为 NeuroImprint 的攻击方法,恶意参数服务器可通过篡改适配器,将客户端训练样本隐式记忆为独立的神经元更新,从而植入隐私后门。该方法在不损害模型效用的前提下,为每个训练样本分配专用记忆神经元,并限制其更新次数。研究揭示了 PEFT 框架潜在的安全风险。
第一作者: Jun He · 方向: 系统安全
代理AI访问控制运行时强制证书绑定系统安全
Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes. Existing access-control mechanisms authorize identities, while assurance layers certify proposed actions; neither alone provides a mandatory enforcement point for certified authority at the moment of mutation. This paper introduces the Sovereign Execution Broker (SEB), a runtime enforcement boundary for certificate-bound agentic infrastructure. SEB consumes certificates issued by the Sovereign Assurance Boundary (SAB), verifies that the requested mutation matches the certified execution contract, checks validity windows, policy epochs, revocation epochs, and live-state drift, mints scoped execution identity, invokes infrastructure APIs, and records signed decision and outcome records. By...
论文介绍 本文针对自主代理系统在关键工作流中的权限控制问题,提出主权执行代理 SEB 框架。该框架作为一个运行时执行边界,强制执行与证书绑定的权限。SEB 接收由主权保证边界 SAB 签发的证书,在执行突变操作前,验证请求是否匹配已认证的执行合约,并检查策略、撤销状态等,从而为非确定性的代理推理过程提供了强制性的安全执行保障。
第一作者: Alaia Solko-Breslin · 方向: 安全研究
AI代理安全概率验证Datalog运行时监控分布鲁棒优化
Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulate and enforce policies expressed in a formal language like Datalog offer a promising solution. However, existing approaches are restricted to deterministic policies. In many practical applications of AI agents, there is a need to enforce security policies in the face of ambiguity, leading to probabilistic predicates or state transitions (for example, a declassifier or Personally Identifiable Information (PII) detector that has some failure probability on each invocation). Furthermore, in many such applications, one cannot easily make the independence assumptions necessary to invoke prior work on probabilistic inference in Datalog. We address this by introducing a sound and efficient framework for such verification based on distributionally robust...
论文介绍 现有AI代理运行时监控方法通常仅支持确定性策略,难以应对策略本身存在不确定性的场景。本文针对此局限,引入一个基于分布鲁棒优化的、高效且可靠的形式化框架,用于验证涉及概率谓词或状态转换的安全策略。该框架扩展了Datalog语言,能够为存在模糊性或失败概率的AI代理安全策略执行提供验证支持。
第一作者: Arastoo Zibaeirad · 方向: 软件安全
漏洞检测大语言模型软件安全数据污染微调评估
Abstract:Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved. We present CWE-Trace, a framework for LLM vulnerability detection built from 834 manually curated Linux kernel samples spanning 74 CWEs. The framework enforces a strict temporal split (pre-2025 historical set / post-cutoff leakage-free set), preserves context-aware vulnerable--patched pairs, and introduces two diagnostic metrics: the Directional Failure Index (DFI) and Hierarchical Distance and Direction (HDD). We evaluate eight vanilla LLMs and 15 LoRA fine-tuned variants across non-targeted detection, targeted detection, and CWE classification. Our analysis yields two key results. First, data contamination provides no measurable advantage. Function-level analysis shows that 84% of nominally contaminated samples carry no usable...
论文介绍 本文探讨大语言模型在系统软件漏洞检测中的能力边界。研究构建了 CWE-Trace 框架,包含834个经人工标注的Linux内核样本。通过严格的时间划分和新的诊断指标评估了多个LLM及LoRA微调模型。分析表明,数据污染并未带来明显优势,且许多模型表现出模式匹配而非深层安全推理的特征,揭示了当前LLM在漏洞检测任务上可能存在的局限性。
第一作者: Tamara Tagliavia · 方向: 隐私保护
隐私保护匿名性形式化验证微数据COMPASS
Abstract:In the information age, one of the leading problems is how to ensure individual's privacy. Depending on the context in which privacy is considered, various data privacy models have emerged. However, the domain of formal verification of these models is still not sufficiently explored even when it comes to the most basic models. An attempt to verify privacy requirements is the Compliance Assertion Language (COMPASS). In COMPASS, one can specify an anonymity condition that a table needs to satisfy, and an action that will modify the table if the condition is not satisfied. It is designed to operate on preprocessed tables in a form one record - one group of people. In this paper, we modify the COMPASS language in order to operate on microdata tables in their usual form of one record - one person. The modified language is called A-COMPASS. Along with checking of previously applied...
论文介绍 本文旨在推进微数据匿名性分析的形式化方法。针对现有隐私模型的形式化验证不足的问题,对合规性断言语言 COMPASS 进行了修改,使其能够直接操作“一人一记录”的微数据表,形成了 A-COMPASS 语言。该工作扩展了原有语言的功能,可验证匿名性条件,并在条件不满足时执行修正操作,为微数据隐私保护提供了形式化工具。
第一作者: Reza Soosahabi · 方向: 系统安全
代理AI安全提示注入防御误导自动化攻击概率模型
Abstract:Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation. This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the attacker's automated judge. Our analysis shows that conventional detect-and-block defenses can allow attacker success rate (ASR) to approach one as the query budget grows, since predictable refusals provide useful feedback to automated search. We then examine detect-and-misdirect, where detected malicious interactions receive controlled, non-operational responses designed to induce...
论文介绍 代理AI系统易受提示注入等攻击,且攻击者正采用模型引导的自动化方法进行规模化探测。本文通过概率模型分析了该攻防场景。研究表明,传统的“检测并阻止”防御在攻击者查询预算足够时可能失效,因为拒绝响应可为自动化搜索提供反馈。进而探讨了“检测并误导”策略,即对检测到的恶意交互给予受控的、无操作的响应,以诱导攻击者做出错误判断。
第一作者: Ans Ibrahim · 方向: 密码学协议
图像加密卷积神经网络动态S盒密码学安全分析
The paper proposes a dynamic approach to image encryption, combining the use of Convolutional Neural Networks (CNNs) and classical cryptography to improve the security and flexibility of image encryption. The main concept is to create adaptive Substitution boxes (S-boxes) based on characteristics that are learned by a trained CNN. The CNN-based S-boxes can be relied on for more non-linearity, uniqueness, and input image dependence than the conventional fixed S-boxes because they are susceptible to the linear and differential attacks. This dynamic behaviour enhances the confusion property and makes it more resistant to statistical and structural attacks. The encryption algorithm consists of CNN-based feature extraction and the creation of a personalised S-box to replace the pixels. Entropy, histogram analysis, correlation, NPCR, and UACI enable security assessment of generated S-boxes...
论文介绍 本文提出一种结合卷积神经网络与经典密码学的动态图像加密算法。核心思想是利用CNN学习图像特征,以生成自适应、依赖于输入图像的动态S盒,替代传统固定S盒。这种动态S盒具有更强的非线性、唯一性和图像依赖性,增强了混淆特性,旨在提高对统计和结构攻击的抵抗力。算法安全性通过熵、直方图分析等指标进行评估。
第一作者: Bercan Turkmen · 方向: 软件安全
恶意软件分类大语言模型反编译多视图二进制分析
Abstract:Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable. Recent work suggests that large language models (LLMs) can assist this process by classifying decompiled code as benign or malicious, but existing pipelines typically rely on a single decompiler view. We argue that this assumption is fragile: decompilers are lossy heuristic tools, and different decompilers can expose different artefacts of the same binary. We curate a benchmark of benign utilities and malicious programs spanning a range of threat behaviors. Each sample is compiled and decompiled with both Ghidra and RetDec, yielding matched pseudo-C views. Across a range of LLMs from major model families, we find that providing both decompiler views improves malicious-class F1, mainly by increasing recall on malicious samples. Agreement analyses further show that...
论文介绍 大语言模型可用于分析反编译的伪代码以分类恶意软件,但现有方法多依赖单一反编译器视图。本文认为单一视图存在脆弱性,不同反编译器可能揭示二进制文件的不同伪影。研究构建了一个基准数据集,使用Ghidra和RetDec生成匹配的多视图伪C代码。实验表明,为LLM提供多个反编译器视图能提升恶意样本的分类召回率和F1值,增强了分析的鲁棒性。
第一作者: Hanwool Lee · 方向: 密码学协议
LLM代理安全多轮红队测试越狱基准对抗鲁棒性安全关键系统
Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as operators of a safety-critical system, instantiated in a simulated nuclear power plant control room. A five-role operator team, each backed by a configurable LLM, runs a plant governed by six critical safety functions (CSFs), while adversaries inject messages over four channels in bounded multi-turn sessions with per-turn feedback. Harm is an objective signal rather than LLM-judged text: a run terminates the moment any CSF is lost, attributed to the causing message. Evaluating four frontier operator models under a fixed-attack paired-replay protocol, we find that adaptive multi-turn attacks...
论文介绍 随着LLM代理被提议用于安全关键系统,其鲁棒性评估至关重要。NRT-Bench是一个多轮红队测试基准,实例化为模拟核电站控制室,包含五角色操作团队和六个关键安全功能。攻击者通过四个信道注入消息,以客观安全损失作为危害指标。该基准通过固定攻击配对重放协议评估前沿模型,揭示了多轮自适应攻击的有效性,为LLM代理的安全评估提供了新方法。
第一作者: Kaihsun Yang · 方向: AI 安全
模型量化量化条件后门任务算术AI安全防御
Abstract:Model quantization is widely adopted to reduce memory usage and inference cost when deploying deep neural networks on resource-constrained devices. However, recent studies have revealed a new security threat known as Quantization-Conditioned Backdoors (QCBs), where a model behaves normally in full precision but activates malicious behavior only after quantization. Existing defenses typically modify quantization procedures or correct activation statistics, often introducing additional computational overhead or relying on specific quantization settings. Here, we present QVec, a parameter-space perspective for defending against QCBs. We observe that the weight difference between a full-precision model and its quantized counterpart encodes a structured behavioral shift, which can be interpreted as a malicious task vector rather than random quantization noise. Based on this...
论文介绍 模型量化在资源受限设备上广泛使用,但量化条件后门(QCBs)构成新威胁:模型在全精度下正常,量化后激活恶意行为。现有防御方法引入额外开销或依赖特定设置。本文提出QVec,从参数空间角度移除QCBs,将权重差异视为恶意任务向量而非量化噪声。这提供了一种轻量级防御方案,提高量化模型的安全性。
第一作者: Yu Shen · 方向: 密码学协议
混合网络移动自组织网络匿名通信去中心化协议
Mix networks are a highly effective way to achieve anonymity, defending against a wide range of traffic-analysis attacks. However, mix networks are usually designed for infrastructure networks and cannot be directly applied in the context of mobile ad hoc networks (MANETs). The few existing solutions for MANETs require advance knowledge of the topology or a trusted central party. In this paper, we present TrustMix, a mix protocol for MANETs that operates without any central trusted party. In TrustMix, parties join groups and then messages are forwarded via multiple groups to provide anonymity. With TrustMix, users only need to find a party nearby that they consider trusted. They then forward the message to this party's group, and the party shuffles messages before forwarding to other groups, meaning that the original message and the forwarded message cannot be linked. Furthermore, even...
论文介绍 混合网络能有效实现匿名性,但通常针对基础设施网络设计,难以直接应用于移动自组织网络(MANETs)。现有方案需要拓扑先验知识或可信中心。本文提出TrustMix,一种无需中心可信方的混合协议。用户通过信任的附近方加入组,消息在多组间转发并打乱,确保原始消息与转发消息不可关联。该协议增强了MANETs中的隐私保护能力。
第一作者: Adolfo P. Jimenez · 方向: 软件安全
GNSS欺骗车辆通信安全软件定义无线电定位导航威胁
Abstract:Global Navigation Satellite Systems (GNSS) constitute a core technology for delivering crucial positioning, navigation, and timing (PNT) services in the Vehicle-to-Everything (V2X) domain, where they are indispensable for generating Cooperative Awareness Messages (CAM) that uphold network reliability and vehicular safety. Yet, GNSS signals are acutely exposed to spoofing, an advanced attack in which an adversary transmits crafted signals that replicate legitimate satellite characteristics, misleading the receiver into computing a false position. This work presents a methodology for conducting physical spoofing with inexpensive Software Defined Radio (SDR), describing a coordinate generation pipeline that employs Haversine-based distance calculations, temporal discretization to emulate constant velocity, and linear interpolation to produce high-fidelity GPS baseband signals...
论文介绍 GNSS在V2X通信中提供关键定位服务,但易受欺骗攻击。本文提出一种使用廉价软件定义无线电(SDR)进行物理欺骗的方法论。通过坐标生成管道,包括基于Haversine的距离计算、时间离散化模拟恒速运动和线性插值,生成高保真GPS基带信号。这揭示了V2X系统中GNSS欺骗的具体威胁,并为防御提供参考。
第一作者: Aymen Bouferroum · 方向: AI 安全
工业物联网信任收敛机器学习动态网络条件
Abstract:In Industrial Internet of Things (IIoT) environments, trust management plays a vital role in securing systems, especially when dealing with resource-constrained devices. Traditional trust models often overlook the impact of fluctuating network quality, leading to slower trust convergence and inaccurate assessments. In this paper, we propose a dynamic trust management solution, known as the Trust Convergence Acceleration (TCA) approach, which integrates Machine Learning (ML) to accelerate trust convergence under poor network conditions. Our model predicts the number of time units needed for trust convergence based on key network metrics and dynamically adapts transition probabilities in the trust model to enhance convergence speed. Using a simulation framework that incorporates realistic Wi-Fi channel conditions based on the IEEE 802.11 standard, we demonstrate the...
论文介绍 在工业物联网(IIoT)中,信任管理对安全至关重要,但传统模型忽略网络波动,导致信任收敛慢、评估不准确。本文提出信任收敛加速(TCA)方法,集成机器学习预测收敛时间,并根据网络指标动态调整转移概率。基于IEEE 802.11标准的仿真框架展示了在劣质网络条件下加速收敛的效果。
第一作者: Junchao Li · 方向: 系统安全
具身AI移动应用安全密码误用测量研究
Abstract:Embodied AI (EAI) mobile applications are evolving from auxiliary user interfaces into active control-path components, directly linking mobile-side cryptographic security to cyber-physical trust. Despite this shift, existing security research predominantly focuses on embodied AI devices and cloud infrastructures, leaving the mobile control layer largely unexplored as a critical attack surface. To bridge this gap, we present the first large-scale measurement study of cryptographic misuse within the EAI mobile ecosystem. We construct EAIAppZoo, a benchmark of 507 real-world applications across six EAI domains, and employ an automated semantic-aware analysis pipeline to measure the prevalence and characteristics of five major cryptographic failure modes. Our measurement yields 12,975 misuse findings (with an evaluated precision of 80.74\%), revealing that these cryptographic...
论文介绍 具身AI(EAI)移动应用正演变为控制路径组件,移动层密码安全与网络物理信任直接相关,但现有研究忽视此攻击面。本文首次对EAI移动生态进行密码误用大规模测量,构建EAIAppZoo基准包含507个应用,并使用自动化语义分析管道测量五类密码失败模式。研究发现大量误用,揭示了安全漏洞。
第一作者: Johannes Wilson · 方向: 密码学协议
形式化验证Tamarin证明器安全协议实现领域特定语言
Formal verification is a challenging but important task for ensuring the security of cryptographic protocols. While modern protocol verification tools significantly reduce verification effort, modelling remains challenging to practitioners without a background in formal verification. In addition, transferring verification results to a concrete protocol implementation requires expert knowledge. In this paper, we present a novel language-first method for verification of trace properties using a domain-specific language for protocol implementations. We target the Tamarin prover for verification, and we prove that verified universal trace properties translate back to the implementation. We additionally integrate symbolic execution in order to analyse the memory safety of protocol implementations. We use our tool to implement and generate accurate models for a signed Diffie-Hellman...
论文介绍 形式化验证对密码协议安全至关重要,但建模困难,且验证结果转移到实现需要专家知识。本文提出AutoTam,一种语言优先方法,使用领域特定语言指定协议实现,目标Tamarin证明器进行验证。通过符号执行分析内存安全,并证明验证的通用跟踪属性可回译到实现。这简化了协议验证和实现过程。
第一作者: Chaeyun Kim · 方向: AI 安全
金融大语言模型红队测试基准框架专家引导评估
Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and systemic trust erosion that require targeted evaluation. We introduce FinRED, an expert-guided red-teaming framework for financial LLM safety evaluation developed with financial experts. FinRED uses a novel two-level taxonomy mapping global standards (e.g., FATF and EU DORA) to threats ranging from regulatory evasion to complex fraud, integrated with a scalable pipeline that converts real financial documents into context-rich red-teaming Behavioral Prompts (seeds) through an expert-defined schema. Rigorous expert validation confirms seed plausibility and realism for meaningful LLM safety evaluation. We also provide an expert-validated, finance-specific rubric that goes beyond disclaimer checks, aligns more closely...
论文介绍 现有安全基准针对通用对抗场景,但缺乏金融特定风险评估。金融LLM面临合规违规、欺诈促成等威胁。本文引入FFinRED,一个专家引导的红队测试框架,与金融专家合作开发。使用两级分类法映射全球标准到威胁,并集成可扩展管道将真实金融文档转换为上下文丰富的行为提示。专家验证确保种子真实性,用于有意义的LLM安全评估。
第一作者: George Alexakis · 方向: 密码学协议
全同态加密数论变换脉动阵列AI ASIC多精度优化
Abstract:Fully Homomorphic Encryption (FHE) ensures robust data privacy but suffers from prohibitive computational overhead. Accelerating FHE on AI hardware like Tensor Processing Units (TPUs) is promising, yet fundamentally limited by a precision mismatch: TPUs are optimized for 8-bit arithmetic, whereas FHE and its critical parts such as the Number Theoretic Transform (NTT), demand high precision. Current approaches bridge this gap using matrix decomposition to execute NTT computations on low-precision matrix engines. However, reconstructing the full-precision results requires shift-and-add accumulation that does not match the dataflow of matrix multiplication. This forces offloading full-precision reconstruction from matrix engines to vector processors that disrupts the matrix multiplication dataflow, creating significant performance bottleneck. To resolve this limitation, we...
论文介绍 研究在AI专用集成电路(如TPU)上加速全同态加密的数论变换时,由于精度不匹配导致性能瓶颈。提出一种低成本多精度脉动阵列架构,优化数据流以提升计算效率。该方法有助于推动全同态加密在AI硬件上的实际应用。
第一作者: Prashanti Nilayam · 方向: AI 安全
大语言模型异构辩论对抗性同伴AI安全鲁棒性
Abstract:Heterogeneous LLM debate is motivated by the promise that diverse peers correct one another, but the same exchange that carries correction also carries adversarial influence. We measure which dominates by tracking how a heterogeneous peer changes the honest agents' revision behavior: how often they change their answer, and whether the change is corrective or harmful. We compare matched panels (homogeneous baseline, honest-mixed, and adversarial-mixed) and contaminated panels in which a malicious same-family peer is already present, spanning four model families and three reasoning benchmarks. An honest heterogeneous peer sharply lowers harmful revision, and an adversarial one reverses it. For Llama-3.1-70B defenders on MATH-hard, the honest-slot harmful-revision rate falls from 89% in the homogeneous panel to 35% with an honest peer, and an adversarial peer returns it to 90%...
论文介绍 探讨异构大语言模型辩论系统中,对抗性同伴对诚实代理修订行为的影响。通过实验比较不同面板设置,测量有害修订率的变化。研究揭示了对抗性影响的严重性,对提升LLM系统的安全性和可靠性具有指导意义。
第一作者: Tasneem Suha · 方向: 密码学协议
侧信道攻击运行时漏洞硬件软件协同嵌入式设备安全缓解
Abstract:Program runtime or timing attacks exploit variations in a program's execution times to extract sensitive information from the program (e.g. encryption keys, sensitive variable data, intellectual property). State-of-the-art solutions to runtime side-channel attacks attempt to balance the execution time of the sensitive code for different control flow paths to eliminate the timing leakage. However, during the mitigation process, most techniques do not consider the underlying hardware or device on which the target program is supposed to run on. This can lead to over-fixing (unnecessary extra operations), under-fixing (not solving the imbalance properly), and even failures. We propose DISARM, a joint hardware-software methodology (unlike any existing solution) for mitigating runtime side-channel vulnerabilities that utilizes timing values from real embedded devices to generate...
论文介绍 针对软件运行时侧信道攻击,现有缓解方法未考虑目标硬件差异,可能导致过度或不足修复。提出「DISARM」,一种联合硬件-软件的方法,利用真实设备时序数据生成针对性缓解策略。该方法有望提升嵌入式系统的侧信道安全防护。
第一作者: Haotian Xu · 方向: AI 安全
大语言模型推测推理安全保证反思采样AI安全
Abstract:Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses are largely incompatible with speculative inference: they either introduce additional computation or disrupt the draft-verify mechanism, negating acceleration benefits. This reveals a fundamental incompatibility between current safety methods and speculative decoding. We propose SafeSpec, a safety-aware speculative inference framework that integrates risk estimation directly into the verification process. SafeSpec attaches a lightweight latent safety head to the target model to jointly evaluate semantic validity and safety in a single forward pass. When unsafe generations are detected, SafeSpec applies rollback and safety-guided reflective multi-sampling to recover safe continuations rather than terminating generation. We model jailbreak...
论文介绍 推测推理虽能加速大语言模型生成,但缺乏内在安全机制。现有防御与推测解码不兼容。提出「SafeSpec」框架,在验证过程中集成风险估计,通过轻量级安全头和反思多采样实现安全生成。这为快速且安全的LLM应用提供了新思路。
第一作者: Prashant Kumar Pathak · 方向: 安全研究
向量检索hubnessRAG安全入网时控制各向异性
Vector hubness, where a few points become nearest neighbors of many queries, creates a poisoning risk in retrieval-augmented generation (RAG): one injected document can influence unrelated requests. Existing defenses use periodic reverse-kNN scans, leaving an exposure window and repeated corpus-wide work. We study admission-time control, scoring each candidate against sentinel queries and quarantining hub-like documents before insertion. Across two 100,000-document corpora, five encoders, and disjoint attacker and defender query sets, a global gate achieves recall 1.0 at the decisive embedding-space point (>=0.92 across the effective range) and 0.91 +/- 0.07 on HotFlip attacks, with 1% false positives on general documents. A per-topic gate provides no reliable benefit, consistent with anisotropy coupling local and global visibility. Thresholds are maintained incrementally, with...
论文介绍 在检索增强生成系统中,向量hubness问题允许注入文档影响无关查询,构成中毒风险。现有防御存在暴露窗口。研究入网时全局门控方法,通过哨兵查询评分隔离hub文档。实验显示高召回率和低误报率,提升了RAG系统的安全性。
第一作者: Gulshan Saleem · 方向: AI 安全
提示注入RAG安全多层防御聊天机器人LLM安全
Abstract:Prompt injection is ranked as the most critical vulnerability in large language model (LLM) deployments by the OWASP Top 10 for LLM Applications, yet existing defenses operate at isolated pipeline stages and remain incomplete. Input filters cannot inspect retrieved documents, while output monitors cannot prevent malicious payloads from reaching the model. Consequently, retrieval-augmented generation (RAG) chatbots remain vulnerable to indirect injection, where a poisoned knowledge-base document compromises every user whose query retrieves it. We present a three-layer framework that intercepts both direct and indirect prompt injection throughout the inference pipeline. Layer 1 screens user input using a rule-based pattern library and a fine-tuned semantic anomaly classifier. Layer 2 enforces a provenance-based instruction hierarchy during context assembly, preventing retrieved...
论文介绍 提示注入是大语言模型部署中的关键漏洞,现有防御分散且不完整。针对RAG聊天机易受间接注入的问题,提出三层安全框架:输入筛选、指令层次强制和输出监控。该框架旨在全程拦截恶意payload,增强系统安全性。
第一作者: Shangzhi Xu · 方向: 系统安全
ReDoS正则表达式攻击字符串生成资源耗尽安全检测
Abstract:ReDoS attacks constitute a critical class of resource-exhaustion vulnerabilities. In such attacks, adversaries exploit the pathological worst-case execution behavior of regular expression (regex) engines to induce highly asymmetric computational workloads, ultimately exhausting system resources and degrading service availability. To protect systems against ReDoS attacks, numerous detection techniques have been proposed that simulate the attack process by generating attack strings to proactively exploit ReDoS vulnerabilities at the early development stage and facilitate remediation. Existing techniques broadly fall into two classes: static analyses that search for pathological regex structures, and dynamic exploration methods that synthesize candidate attack strings. However, the generated attack strings are often impractical for real-world exploitation because they usually...
论文介绍 正则表达式拒绝服务攻击利用引擎最坏情况行为耗尽资源。现有检测技术生成的攻击字符串往往不实用。提出「PUFFERDOS」方法,高效生成有效的攻击字符串,用于早期漏洞检测和修复。这有助于提升系统对ReDoS攻击的防御能力。
第一作者: Baigang Chen · 方向: 系统安全
桥接分发隐私保护两方计算组自适应审查规避
Abstract:We present G-Lox (group-adaptive Lox), a bridge-distribution system that preserves Lox-style distributor blindness while enabling hidden, stateful group-level adaptation. G-Lox places adaptive assignment logic behind a two-server privacy wall, so no single server learns group identifiers or group-to-bridge assignments. Private state access and state-dependent updates use two-server DPF/FSS protocols and secure two-party computation, supporting blockage reporting, transport-aware reassignment, and privacy-preserving group splitting. We evaluate G-Lox through system measurements and policy simulation. In our C++/EMP implementation over real TCP sockets, private state access has low client-visible overhead: across state sizes up to 2^16, communication remains in the low-KiB range per iteration. At M=1024, the client sends 1,968 bytes, receives 1,280 bytes, and completes an...
论文介绍 在审查规避系统中,桥接分发需保持分发者盲性同时支持自适应。提出「G-Lox」系统,使用两方计算保护组标识符和分配信息,实现隐私保护的组级自适应。评估显示低通信开销,适用于实际部署,增强系统安全性和可用性。
第一作者: Nils Loose · 方向: 软件安全
大型语言模型平台触发后门浮点数运算LoRA适配器软件安全
Abstract:Large language models (LLMs) are increasingly deployed in sensitive settings such as software engineering, where their outputs directly shape downstream artifacts. Recent work has shown that an identical model can produce measurably different outputs depending on the deployment platform, a consequence of non-associative floating-point arithmetic and divergent kernel implementations. We study the security implications of this platform-dependent variability and uncover a novel attack surface on LLM deployments. We introduce FloatDoor, the first input-independent, platform-triggered backdoor attack against generative LLMs. The compromised model exhibits adversary-chosen behavior when served on a target platform and is otherwise benign. FloatDoor is realized through two lightweight LoRA adapters, one that amplifies inter-platform numerical divergence and one that binds the...
论文介绍 研究大型语言模型在不同部署平台上的输出差异导致的安全漏洞。本文提出FloatDoor,首个输入无关、平台触发的后门攻击,通过两个LoRA适配器实现:一个放大平台间数值发散,另一个绑定恶意行为。该攻击在目标平台上激活后门,否则模型表现正常,揭示了LLM部署中的新型安全风险。
第一作者: R.D.N. Shakya · 方向: 密码学协议
后量子密码学安全编码漂移大型语言模型密码学工程游戏化修复
Abstract:The transition to Post Quantum Cryptography (PQC) introduces considerable implementation complexity, requiring strict adherence to constant-time execution, side channel resistance, and precise parametrisation. Simultaneously, large language models (LLMs) are heavily embedded in software development workflows, including cryptographic engineering. While LLMs improve productivity, evidence shows that they frequently generate insecure or suboptimal code, particularly in security critical domains. This paper introduces Secure Coding Drift in PQC, a novel socio technical vulnerability model capturing the gradual degradation of secure coding practices due to sustained reliance on LLM-generated code. Unlike prior work that focuses on static vulnerabilities, we conceptualise security risk as a longitudinal behavioural phenomenon rising from human AI interaction. To mitigate this, we...
论文介绍 探讨后量子密码学开发中LLM生成代码的安全问题。本文引入安全编码漂移模型,描述因持续依赖LLM代码而导致的安全实践退化。为缓解此问题,提出一种游戏化修复方案,旨在提升LLM辅助密码学工程中的编码安全性和合规性。
第一作者: Christos Galanopoulos · 方向: 密码学协议
基因组Beacon全同态加密以太坊虚拟机隐私保护智能合约
The Global Alliance for Genomics and Health (GA4GH) Beacon protocol lets researchers ask whether a genomic variant has been observed in a participating cohort and receive aggregate variant-level counts. As Beacon networks grow, two privacy risks remain: host institutions can see plaintext queries, and repeated rare-variant queries can support membership-inference attacks. We present bioETH-Beacon, a smart-contract prototype that runs the Beacon "aggregate count" query over encrypted data on a fully homomorphic Ethereum Virtual Machine (fhEVM). Hospitals upload encrypted marker-count entries, authorized researchers submit encrypted marker queries, and the contract returns an encrypted answer that is released, via an off-chain key-management service, only to the requester named in the contract's on-chain ACL. The design is organized as a 3x4 tier-by-query-family grid spanning genotype...
论文介绍 针对基因组Beacon网络中的隐私风险,提出bioETH-Beacon原型。该系统基于全同态以太坊虚拟机,在链上处理加密查询并返回加密结果,通过链下密钥管理和访问控制列表实现隐私保护。这有助于防止查询泄露和成员推断攻击,促进安全基因组研究。
第一作者: Mikael Alemu Gorsky · 方向: 软件安全
人工智能网络安全非洲网络作战语言模型
Abstract:In 2025 and 2026, two events settled questions that had until then been speculative. In the first, a large language model executed the great majority of a state-aligned cyber-espionage campaign on its own, with human operators intervening at only a few decision points. In the second, the most capable cyber-relevant model was placed under a controlled-access program limited to a vetted set of United States technology firms, allied governments, and European standards bodies; that perimeter included no African government, operator, or university. Together the two events establish the argument of this paper: frontier language models have become a decisive instrument of cyber operations, and that instrument is built, owned, and rationed within a small circle from which Africa is absent. The paper documents Africa's exclusion on every count. The continent does not build frontier...
论文介绍 基于2025-2026年事件,分析人工智能在网络安全中的变革作用。研究显示,大型语言模型已成为网络作战的关键工具,但其构建和限制掌握在少数国家,导致非洲被排斥在外。本文记录了非洲在前沿AI技术访问上的缺失,探讨了地缘政治影响。
第一作者: Zunchen Huang · 方向: AI 安全
LLM-求解器循环叙述差距形式化验证提示注入安全关键
Formal tools such as SAT and SMT solvers are increasingly embedded in language model reasoning pipelines when a safety or security critical question can be formulated in logic. Unlike chain of thought whose steps are sampled from the model distribution without formal guarantee, a solver produces a sound and independently verifiable answer. However, the soundness guarantee can be lost in the interaction between the solver and the model. The hybrid pipeline has three components: formalizing the question, deciding it, and narrating the result. Prior work has studied the formalization and decision, but not narration, which is the step that turns a formal tool's output into the user answer. To fill the narration gap, we first model the LLM-solver loop as a verified decision procedure. We further evaluate five open-sourced models under prompt injection, and we find certificate gating makes...
论文介绍 分析LLM与求解器混合管道中的叙述差距问题。该管道包括形式化问题、决策和叙述结果三步骤,其中叙述步骤可能引入安全漏洞。本文将循环建模为验证决策程序,并评估五种开源模型在提示注入下的表现,发现证书门控能增强安全性。
第一作者: Luis Adrián Lizama-Pérez · 方向: 安全研究
量子密钥分发贝尔态被动用户无量子探测器密码学
Abstract:We propose and analyze a Bell-state extension of the Loop-Back quantum key distribution architecture for secret-key establishment between two passive users that do not require quantum transmitters or quantum detectors. In the proposed setting, a single active station, Alice, provides the entangled-state infrastructure, retains one qubit of an initially prepared Bell pair, and sends the traveling subsystem through two passive users, denoted by $B_1$ and $B_2$. Each passive user applies a local Pauli operation to the same traveling subsystem, so that the operation observed by Alice is only the effective composition $U_{\mathrm{eff}}=U_2U_1$. After the subsystem returns, Alice performs a Bell-state measurement and, using her private knowledge of the initial Bell state, deterministically identifies the effective Pauli operation. However, the individual factors $U_1$ and $U_2$...
论文介绍 提出一种量子密钥分发架构,允许被动用户无需量子发射器或探测器参与。在该设置中,主动站点提供纠缠态基础设施,用户应用泡利操作后,通过贝尔态测量识别有效操作。这简化了量子密钥分发,适用于资源受限的节点。
第一作者: Sizhe Yang · 方向: 机器人操作 · 来源: cs.RO
世界动作模型持久记忆机器人操作视觉预测动作条件化
Abstract:Robust robotic manipulation in the real world requires not only an understanding of the current observation, but also memory and dynamics modeling. World action models (WAMs) possess these capabilities by jointly modeling visual foresight and actions conditioned on both current and historical observations, making them a promising paradigm for robotic manipulation. However, existing WAMs face a fundamental trade-off: methods with efficient inference typically condition only on a bounded window of recent observations and therefore struggle in non-Markovian environments, whereas methods that preserve long histories incur time and space costs that grow substantially with sequence length. To address this challenge, we introduce MemoryWAM, a world action model with efficient persistent memory. MemoryWAM uses a hybrid memory design that combines recent frames, event-boundary anchor...
论文介绍 针对世界动作模型在推理效率与长历史建模间的权衡问题,引入MemoryWAM。该模型采用混合记忆设计,结合近期帧和事件边界锚点,以高效管理持久记忆。这提升了机器人在非马尔可夫环境中的操作性能,减少时间与空间成本。
第一作者: Sha Yi · 方向: 机器人操作 · 来源: cs.RO
机器人手生成人类演示逆运动学数据驱动设计优化
Abstract:Robot learning has advanced rapidly in learning control, but learning the physical body of a robot remains much more difficult because jointly searching over design and control creates a very large combinatorial problem. Here, we present a data-driven framework for generating robot hands from human demonstrations. Instead of learning a complex controller together with each candidate design, we generate robot hand designs using the same simple control policy used after fabrication: matching fingertip positions through inverse kinematics. Using more than 4 million frames of human fingertip motion from everyday manipulation, our algorithm optimizes tree-structured robot hands to reproduce desired target motions. The framework produced both a 6-degree-of-freedom (DoF) general-purpose hand and lower-DoF task-specific hands with spatial four-bar mimic joints. To accelerate the...
论文介绍 研究机器人手设计的高组合复杂度问题。提出一个数据驱动框架,从人类指尖运动数据优化树状结构机器人手设计,采用简单控制策略匹配指尖位置。该框架生成了通用手和任务特定手,加速了设计与制造过程。
第一作者: Oxana Shamilyan · 方向: 导航与运动 · 来源: cs.RO
连续体机器人运动规划韧性多标准决策层次分析法
Abstract:This paper presents an experimental study of motion planning for resilient continuum robots. In this study we mainly focused on multi-criteria decision-making, its application for path-planning algorithms, impact on the generated path and execution time. To do this, we used two well-known algorithms for path planning, namely Genetic algorithm and A star algorithm, and modified them by adding the Analytical Hierarchy Process algorithm to evaluate the quality of the paths generated. In our experiment the Analytical Hierarchy Process considers four different criteria, i.e. distance, motors damage, mechanical damage of the robot's arm and accuracy, each considered to contribute to the resilience of a continuum robot. The use of different criteria is necessary to increase the time to maintenance operations of the continuum robot. We conducted the experiments using two different...
论文介绍 本研究针对连续体机器人的韧性提升问题,通过实验比较遗传算法和A*算法在路径规划中的应用。核心方法是集成层次分析法评估多标准决策,如距离、电机损坏、机械损坏和精度,以优化路径生成。这有助于减少维护需求,延长机器人操作寿命,增强实际应用中的可靠性。
第一作者: Fatma Youssef Mohammed · 方向: 导航与运动 · 来源: cs.RO
注意力预测扫描路径液态神经网络主动感知自主导航
Abstract:Human visual attention relies on structured scanpaths to efficiently process scenes, yet instilling this behavior into robot autonomy is in its infancy and hindered by the high,computational costs of existing predictive models. To address this, we introduce GazeLNN, a computationally lightweight,scanpath prediction model that leverages Liquid Neural Networks as its recurrent engine and employs MobileNetV3 for feature extraction. Operating auto-regressively, the architecture predicts sequential fixation heatmaps conditioned on the current visual stimulus and fixation history. Despite requiring only 0.61 GFLOPs, GazeLNN achieves state-of-the-art performance on the MIT Low Resolution dataset achieving 0.47 ScanMatch score. It outperforms existing recurrent baselines across diverse evaluation metrics, while reducing computational costs by 99.40% and accelerating inference by up to...
论文介绍 本文针对自主导航中人类注意力预测的高计算成本问题,提出GazeLNN模型。该模型基于液态神经网络和MobileNetV3实现轻量级扫描路径预测,计算高效且性能先进。可能应用于机器人的主动感知和场景理解,提升导航效率与准确性。
第一作者: Zhenghao "Mark'' Peng · 方向: VLA 通用模型 · 来源: cs.RO
视觉语言模型轨迹规划人行道导航延迟弹性场景理解
Abstract:Learning-based planners for sidewalk navigation can generate diverse candidate trajectories in real time, yet their scoring functions often fail to select the best trajectory in challenging situations, outputting trajectories that make the mobile robot drive onto grass, toward pedestrians, or in the wrong direction, even when better candidates exist in the same set. We call this the trajectory scoring gap: in real-world sidewalk navigation, the gap between an anchor-based planner's top choice and the best possible candidate is substantial, likely due to limited high-level scene understanding capability of the planner. Rather than replacing the planner with an end-to-end Vision-Language-Action model, we propose a VLM-Planner interface that uses a VLM to select a candidate index from the planner's proposal set and then fuse it with the planner's initial output. However, VLMs...
论文介绍 研究解决人行道导航中轨迹评分差距的难题。提出VLM-Planner接口,利用视觉语言模型从规划器的候选轨迹中选择最优并融合输出。这增强了规划器的场景理解能力,提高移动机器人在复杂环境中的导航可靠性,尤其适用于延迟敏感场景。
第一作者: Hengfei Zhao · 方向: 策略学习 · 来源: cs.RO
有限元方法触觉仿真视觉触觉传感器力学计算Isaac Sim
Abstract:Vision-based tactile sensors require high-fidelity simulation for reinforcement learning, yet existing approaches struggle to provide accurate mechanical stress fields within GPU-accelerated robotics platforms. We present TaCauchy, an extensible Finite Element Method (FEM) framework that integrates rigorous physics-based force computation into Isaac Sim. Built on the Unified Incremental Potential Contact (UIPC) solver, TaCauchy directly computes Cauchy stress tensors from hyperelastic constitutive laws and projects them onto contact surfaces to obtain traction forces and pressure distributions, providing mechanical ground truth from first principles rather than empirical estimation. Our framework features automatic mesh generation with geometry-aware adaptive refinement and a modular sensor interface enabling rapid integration of diverse sensors (GelSight Mini, DIGIT, 9DTact)...
论文介绍 本文提出TaCauchy框架,将有限元方法集成到Isaac Sim中,用于基于视觉的触觉传感器高保真仿真。核心是通过计算柯西应力张量提供从第一性原理出发的力学模拟,支持自动网格生成和模块化传感器接口。可能应用于强化学习中的精确触觉数据训练。
第一作者: Ziyuan Tang · 方向: 机器人操作 · 来源: cs.RO
连续体机器人3D打印遥操作可复现平台机器人学习
Abstract:Continuum robots offer strong potential for manipulation tasks due to their high degrees of freedom, compliant structures, and operational safety. However, their adoption in both research and practical applications has been hindered by reproducibility issues arising from complex fabrication and assembly processes, challenging kinematic modeling, and a lack of intuitive control interfaces. To address these challenges, we present a novel open-source continuum robot design. The platform features a simplified fabrication pipeline enabled by multi-material 3D printing, allowing the arm to be fabricated as a monolithic compliant structure with minimal assembly. Control is achieved through an isomorphic teleoperation interface that establishes a direct actuator-level mapping, eliminating the need for explicit kinematic modeling and providing a singularity-free mapping. Building on...
论文介绍 为解决连续体机器人的可复现性和控制挑战,提出开源平台CoLI。核心方法是使用多材料3D打印实现单片制造,并通过同构遥操作接口建立直接映射,消除运动学建模需求。这简化了制造与控制流程,促进机器人学习研究和应用开发。
第一作者: Paolo Golinelli · 方向: 导航与运动 · 来源: cs.RO
相对定位多机器人系统分散式算法无基础设施测距测量
Abstract:The ability to localise teams of robots is essential for applications ranging from robotic fleets in unstructured environments to cooperative control and navigation tasks. In such contexts, fixed infrastructure is often unavailable, deployments must be fast and flexible, and system requirements must be minimal. We present a decentralised cooperative localisation algorithm that addresses all these challenges at once. The method is anchor-less, fully decentralised, and, unlike most existing approaches, does not require controlling the robots motion to ensure team observability. It relies only on local odometry, sparse inter-agent ranging measurements, and short-range communication, all of which are widely available in practice. The algorithm adopts a multi-hypothesis Bayesian framework that maintains the entire set of feasible solutions, ensuring robustness under transient...
论文介绍 本文提出分散式协作定位算法,用于无基础设施环境中的移动机器人团队。算法基于局部里程计和测距测量,采用多假设贝叶斯框架维护可行解集,确保鲁棒性。这适用于快速部署的机器人车队和协作导航任务,无需额外控制或基础设施支持。
第一作者: Yandong Wang · 方向: VLA 通用模型 · 来源: cs.RO
双臂操作视觉语言动作模型协调感知结构动作专家机器人操作
Abstract:Vision-language-action (VLA) models show strong capabilities in single and dual-arm robotic manipulation. Prior works show coordinated bimanual behaviors can emerge from end-to-end learning, leveraging large vision-language backbones with continuous action prediction. However, as bimanual tasks become tightly coupled and execution constraints become critical, implicit coordination alone is insufficient to ensure reliable, interpretable, and stable behavior. In this work, we propose Co-VLA, a coordination-aware bimanual manipulation framework introducing explicit structural priors into VLA models. We instantiate our method on a state-of-the-art vision-language backbone by replacing its monolithic action head with a Structured Action Expert (SAE) designed for bimanual coordination. Specifically, we introduce explicit structure at the action generation level with a modular...
论文介绍 针对双臂操作中协调可靠性不足的问题,提出Co-VLA框架。核心是在视觉语言动作模型中引入显式结构先验,通过结构动作专家模块增强双臂协调能力。这提升了任务执行的稳定性和可解释性,适用于复杂耦合的双臂机器人操作场景。
第一作者: Paul Koch · 方向: 具身智能 · 来源: cs.RO
合成数据生成域适应认知机器人计算机视觉AI模型训练
Abstract:AI vision models are a driving factor for the potential use case scenarios of cognitive robotics within in the industry and household applications. A large array of methods from semantic environment analysis towards 6D and grasping pose estimation have been proposed based on the latest AI achievements. However, such advancements require further strong and efficient methods w.r.t. training data and AI-architectures, which are capable in synergy to tackle current challenges, precision limits, and scalability beyond domain gaps. In this paper, we discuss these current limits and trends in the related state-of-the-art which are challenging those. Further we discuss our current work in progress on bridging the domain gap between simulations and real world applications by linking those in the training data generation.
论文介绍 本文讨论连接真实场景与合成数据生成的方法,以弥合AI视觉模型训练中的域差距问题。核心方法是通过高效链接训练数据生成,提升模型的泛化能力。这有助于优化认知机器人和计算机视觉应用,提高精度和可扩展性。
第一作者: Gia-Binh Nguyen · 方向: VLA 通用模型 · 来源: cs.RO
视觉-语言-动作模型层冗余结构压缩居中核对齐微调优化
Abstract:Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion parameter architectures impose prohibitive computational burdens during downstream fine-tuning and real-time inference. In this work, we reveal a highly non-trivial architectural characteristic of these continuous control foundation policies (e.g., pi_0, GR00T-N1.5): despite being trained on diverse physical trajectories, they exhibit severe layer-wise representational redundancy. To exploit this, we introduce a structural compression pipeline that is entirely training-free, bypassing the need of existing methods to load full-scale models to learn optimized token reductions or dynamic layer selectors. Instead, using only a single forward pass via Centered Kernel Alignment to identify redundant layer features, we remove twin layers to...
论文介绍 针对视觉-语言-动作模型在机器人操作中微调和推理计算负担重的问题,研究揭示其层间表示冗余现象。提出一种无需训练的结构压缩流水线,通过居中核对齐识别冗余层并移除孪生层,显著降低计算成本,适用于实际部署场景。
第一作者: Francesco Argenziano · 方向: 导航与运动 · 来源: cs.RO
流动匹配物体动力学多模态分布3D场景理解机器人导航
Abstract:Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments. Such agents must not only comprehend and navigate spatial layouts, but also reason about how these spaces evolve over time. In particular, humans interact with objects daily, causing them to change position throughout the environment and making it difficult for robots to reliably associate current observations with previously seen objects. However, these interactions are not random: human habits and routines induce spatio-temporally consistent patterns in object locations, which robotic agents can potentially learn and then exploit for downstream tasks such as navigation. To this end, we introduce FlowMaps, a latent flow matching model for estimating multimodal distributions over the future locations of dynamic objects in a continuous 3D space...
论文介绍 机器人需要理解家庭环境中物体位置随时间变化以支持导航等任务。提出FlowMaps模型,基于流动匹配估计动态物体在连续3D空间中的未来位置分布,利用人类习惯的时空一致性模式,实现长期多模态预测。
第一作者: Boya Zhang · 方向: 机器人操作 · 来源: cs.RO
软夹持器手上操作自由度平行夹爪升级低成本制造
Abstract:Parallel-jaw grippers are the default manipulator choice in robotics because they are simple, robust, and inexpensive. Their limited in-hand mobility, however, often forces large arm motions and restricts dexterous manipulation in confined workspaces. We present a parallel-gripper upgrade: a double-soft-belt-based finger module that preserves standard opening/closing while adding three in-hand degrees of freedom (DoF): translation, pitch, and roll. The mechanism is deliberately kept simple and engineered for inexpensive manufacturing and straightforward integration, preserving the reliability and precise control of traditional parallel grippers while greatly broadening the range of manipulation capabilities. To demonstrate the utility of the added DoFs, we integrate the gripper in two control pipelines. First, we adapt a model predictive controller for in-hand manipulation of...
论文介绍 针对平行夹爪缺乏手上移动性限制灵巧操作的问题,提出Belt-Finger模块。该模块基于双软带机制,在保持标准开合功能的同时增加平移、俯仰和滚转三个自由度,设计简单、成本低,扩展了夹持器的操作能力。
第一作者: James Fant-Male · 方向: 具身智能 · 来源: cs.RO
动作识别装配状态跟踪人机协作隐马尔可夫模型神经网络
Abstract:Human Action Recognition (HAR) is frequently investigated in Human-Robot Collaboration (HRC) research to understand what actions have been performed and hence the state of a collaborative task. Accurately tracking an assembly state from HAR is however not fully investigated, and in realistic scenarios is not a trivial task. This research systematically investigates and compares methods for tracking assembly state using action recognition inputs. Investigations using two diverse datasets and five state tracking approaches, including logic-based, Hidden Markov Model (HMM), and neural network (NN) methods, show that optimal approaches are not uniform across different tasks and that different methods fail under different circumstances. Testing is performed using both simulated inputs with varying noise levels and realistic inputs from a HAR model. Results show NN and HMM methods...
论文介绍 人机协作中从动作识别推断装配状态是关键挑战。系统比较逻辑、HMM和神经网络等五种状态跟踪方法,发现不同任务下的最优方法不一致,神经网络和HMM方法在噪声下表现良好,但需适应现实场景。
第一作者: Jianing Guo · 方向: 机器人操作 · 来源: cs.RO
频率感知流动匹配连续动作生成离散余弦变换动作一致性
Abstract:Flow matching has emerged as a standard paradigm for robotic manipulation owing to its strong expressive power for modelling complex, multimodal action distributions, alongside similar approaches like diffusion policy. However, existing methods rely on discretized action chunks, making them brittle to demonstrations collected at heterogeneous control frequencies and prone to temporally inconsistent actions that degrade control stability. In this paper, we propose Frequency-Aware Flow Matching (FAFM), which outputs continuous, temporally consistent actions. To handle heterogeneous frequency input, we transform discrete action sequences into the frequency domain with the discrete cosine transform (DCT), perform flow matching over the resulting coefficients, and reconstruct continuous actions via cosine basis expansion. To generate temporally consistent actions, we regularize the...
论文介绍 流动匹配用于机器人操作时依赖离散动作块,对异构频率不鲁棒且易导致时间不一致。提出频率感知流动匹配(FAFM),通过离散余弦变换处理频率域,生成连续、时间一致的动作,提升控制稳定性和泛化能力。
第一作者: Hyeonna Choi · 方向: 具身智能 · 来源: cs.RO
自然语言处理协议翻译机器人实验室代理框架验证机制
Abstract:Biological experiment protocols are written in natural language, whereas automation systems rely on predefined control commands, creating a semantic gap that limits autonomous execution. Microplate-based automatic experiments are particularly challenging due to the need to simultaneously control well mapping, sample-reagent combinations, replicate placement, and parallel dispensing. This study proposes an agent-based protocol translation framework that converts natural-language microplate-based protocols into executable control commands for a robotic laboratory platform. A Parser Agent formalizes the natural-language protocol into a structured representation, and a rule-based mapping engine deterministically incorporates the operational constraints of the robotic laboratory platform to generate device-level control commands. A heterogeneous LLM Validation Agent verifies...
论文介绍 解决自然语言生物实验协议与机器人平台控制命令间的语义鸿沟。提出双代理框架:解析代理将协议结构化,规则引擎生成设备命令,LLM验证代理确保一致性,实现微孔板实验的自动化执行。
第一作者: Jonghoon Lee · 方向: VLA 通用模型 · 来源: cs.RO
数据增强对象交换物理合理性多视角VLA策略
Abstract:Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution. The standard remedy is to collect multi-view teleoperation data for every failure case, but this scales poorly in both cost and time. We introduce Pose6DAug, a failure-driven data augmentation framework that turns a policy's own successful episodes into targeted demonstrations for its failure modes, without any new data collection. Our key insight is that each successful episode already encodes a physically valid action trajectory together with calibrated multi-view observations. By swapping only the manipulated object while preserving this trajectory, we obtain new and physically grounded demonstrations. However, naive 2D video editing breaks...
论文介绍 视觉-语言-动作策略在新型物体上易失败,数据收集成本高。提出Pose6DAug框架,通过交换成功演示中的操纵对象,保持动作轨迹不变,生成物理合理的新演示,无需额外数据收集,提升策略泛化能力。
第一作者: Nozomu Masuya · 方向: 模仿学习 · 来源: cs.RO
模仿学习变频控制迭代学习控制频率外推采样频率
Abstract:Conventional neural network (NN)-based imitation learning methods for variable-speed motion either restricted their scope to interpolated speeds, or generated unpredictable motions when extrapolating beyond trained velocity ranges. Variable-frequency imitation learning (VFIL) enabled extrapolations of speeds by linking the NN model's sampling frequency to the motion frequency, whereas its open-loop configuration caused frequency errors, especially in the extrapolated high-frequency settings. This study proposes variable-frequency imitation learning with iterative learning control (VFILC) based on a combination of VFIL and iterative learning control (ILC) with both feedforward and feedback parts, the former taking advantage of VFIL and the latter adjusting the frequency errors. The experimental results showed that the proposed method successfully and accurately extrapolated...
论文介绍 传统模仿学习在变速运动中频率外推不准确。提出VFILC方法,结合变频模仿学习和迭代学习控制,通过前馈和反馈部分调整频率误差,实现准确的高速频率外推,提升运动控制精度。
第一作者: Zheyu Zhuang · 方向: 模仿学习 · 来源: cs.RO
模仿学习数据增强对称性机器人操作
Abstract:Image-based behaviour cloning leverages demonstrations captured from ubiquitous RGB cameras. However, it remains constrained by the cost of collecting diverse demos, especially for generalizing across workspace variations. We propose MirrorDuo, a reflection-based formulation that operates on image, proprioception, and full 6-DoF end-effector action tuples, generating a mirrored counterpart for each original demonstration, effectively achieving "collect one, get one for free". It can be applied as a data augmentation strategy for existing learning pipelines, such as standard behaviour cloning or diffusion policy, or as a structural prior for reflection-equivariant policy networks. By leveraging the overlap between the original and mirrored domains, MirrorDuo achieves significantly improved performance under the same data budget when demonstrations are evenly distributed across...
论文介绍 该研究针对基于图像的行为克隆中收集多样化演示成本高的问题,提出了MirrorDuo方法。其核心是利用反射对称性,为每个原始演示(图像、本体感觉、6自由度动作)自动生成镜像对应样本,实现「一次采集,免费获得一份」的效果。该方法可作为现有学习流程的数据增强策略,或作为反射等变策略网络的结构先验,在相同数据预算下显著提升性能。
第一作者: Junzhe Xu · 方向: 策略学习 · 来源: cs.RO
强化学习路径规划神经形态计算机器人移动履约系统
Abstract:Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Systems (RMFS) a challenging problem for conventional search- and rule-based methods, which typically suffer from high computational complexity and long decision latency. While reinforcement learning (RL) has emerged as a powerful alternative, deploying learned policies with extreme energy efficiency on resource-constrained hardware remains an open challenge. We present SDQN-RMFS, an end-to-end framework that achieves high-fidelity deployment of an RL-trained policy from a full-precision artificial neural network (ANN) through to a neuromorphic chip. By computing only when triggered by sparse events, this framework unlocks ultra-low-power RMFS pathfinding. Our full-stack pipeline operates as follows: an ANN policy is first efficiently trained...
论文介绍 本文针对机器人移动履约系统中实时、高效的路径规划挑战,提出了一种神经形态强化学习框架SDQN-RMFS。该框架实现了从人工神经网络策略训练到神经形态芯片部署的全栈流程。通过稀疏事件触发计算,系统仅在有事件发生时才进行运算,从而在资源受限的硬件上实现了超低功耗的路径规划,为能效极高的机器人系统部署提供了新思路。
第一作者: Jinghan Yang · 方向: VLA 通用模型 · 来源: cs.RO
视觉-语言-动作模型失败预测信息论可解释性
Abstract:Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions can cause irreversible harm, making generalizable and interpretable failure detection essential. We observe that successful and failed rollouts carry systematically different information-theoretic signatures. Building on this, we formalize VLA control as a closed-loop information pipeline and derive the Triple Information-theoretic (Tri-Info) signals that capture whether actions remain diverse, temporally consistent, and coupled to state transitions. Across six VLA models and three benchmark environments, Tri-Info matches the strongest baselines in-domain. Moreover, Tri-Info transfers across architectures, environments, and the sim-to-real gap without retraining, reaching 83\% accuracy on real-world tasks where prior detectors collapse to...
论文介绍 针对视觉-语言-动作模型作为黑箱可能造成不可逆物理危害的问题,本研究从信息论视角提出了通用的失败预测方法Tri-Info。它将VLA控制建模为闭环信息管道,并推导出三重信息论信号,以捕捉动作的多样性、时间一致性与状态转换的耦合度。该方法无需重新训练即可跨架构、跨环境和跨虚实域迁移,在真实世界任务中实现了有效的失败检测。
第一作者: Xiu Zhang* · 方向: 机器人操作 · 来源: cs.RO
经食道超声心动图增强现实机器人辅助手术用户界面评估
Abstract:TransEsophageal Echocardiography (TEE) is essential for diagnosing and guiding Structural Heart Disease (SHD) interventions. However, manual TEE manipulation demands significant operator expertise, is physically demanding, and exposes clinicians to radiation when performed alongside fluoroscopy. Robotic-assisted TEE systems have been introduced to improve probe handling and reduce operator fatigue, yet the design of intuitive and effective user interfaces remains an open challenge. This study presents and evaluates a model-enhanced, Augmented Reality (AR)-based intuitive interface for robot-assisted TEE, designed to improve spatial awareness and control intuitiveness. A robotic TEE platform integrated with electromagnetic tracking and a virtual simulator was used to compare three user interfaces differing in visualization and interaction modalities: 2D jointlevel (2D-JI), 3D...
论文介绍 本研究评估了一种用于机器人辅助经食道超声心动图的增强现实直观界面。该界面旨在解决手动TEE操作专业要求高、体力消耗大以及辐射暴露等问题,通过整合电磁追踪与虚拟模拟器,提升术者的空间感知与操作直觉。研究比较了三种不同可视化与交互模态的用户界面,为改进机器人辅助介入手术的人机交互设计提供了依据。
第一作者: Kaixin Lan · 方向: 具身智能 · 来源: cs.RO
世界模型等变性对称性机器人跑酷
Abstract:While latent world models enable the proactive predictions required for extreme parkour, their purely data-driven nature forces them to redundantly encode left-right symmetric interactions as independent patterns. This inflates the learning burden and hinders the capture of geometric regularities, restricting the latent space's efficiency for downstream policies. To address this, we propose SWAP, an end-to-end equivariant symmetric world model. This framework embeds symmetry directly into both the world model and the actor-critic networks. In real-world tests, the robot leaps across a 2.13 m gap and climbs a 1.63 m platform, breaking records for quadruped parkour. Furthermore, the framework exhibits robust geometric generalization to unseen mirrored terrains and exceptional zero-shot transferability across diverse outdoor environments. These results demonstrate that symmetry...
论文介绍 为解决潜变量世界模型在机器人跑酷中因冗余编码对称交互而效率低下的问题,本文提出了对称等变世界模型SWAP。该框架将对称性直接嵌入世界模型和行动者-评论家网络,使其具备内在的几何理解能力。实验证明,搭载该框架的四足机器人成功跨越了2.13米宽的沟壑并爬上了1.63米高的平台,且展现出对未知镜像地形的泛化能力和户外环境的零样本迁移能力。
第一作者: Hunter Kuperman · 方向: 具身智能 · 来源: cs.RO
分布式优化深度展开多智能体系统超参数自适应
Abstract:Distributed optimization is a highly scalable and structurally transparent technique to solve multi-agent robotics problems; however, such methods often suffer from the need for highly-specialized, problem-specific hyperparameter tunings. In this work, we propose Deep Coordinator, a deep-unfolding framework that learns to dynamically adjust the hyperparameters of ADMM-DDP, a popular distributed solver for robotics tasks, at solve-time in response to optimizer performance. Our architecture consists of unrolling a fixed number of ADMM-DDP iterations into a neural network with learnable functions between layers mapping the optimizer state to the next hyperparameters. To the best of our knowledge, Deep Coordinator is the first deep-unfolding framework to adapt the penalty parameters of a non-convex optimizer at solve-time; we show that the mainstream supervised approach can yield...
论文介绍 针对分布式优化方法在多机器人任务中需要高度专业化超参数调优的问题,本文提出了深度协调器框架。该框架通过深度展开技术,将ADMM-DDP迭代展开为一个神经网络,在求解过程中根据优化器状态动态学习并调整超参数。这是首个能在求解时自适应调整非凸优化器惩罚参数的深度展开框架,提升了分布式求解器的性能和易用性。
第一作者: Xuetao Li · 方向: 多模态具身 · 来源: cs.RO
人机协同创作音乐机器人多模态具身语义接地
Abstract:Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route for generative models to participate in that creativity through physical action rather than disembodied digital content. In robotic music co-creation, it is challenging to connect semantic musical understanding with real-time and physically executable performance. We present Co-policy, a framework for human-robot musical co-creation that separates semantic intent grounding, constrained musical variation, and visuomotor execution. To ground musical semantics, Co-policy uses pre-inference semantic anchors and a fine-tuned Qwen-vl planner (F-Qwen) to transform speech, live musical seeds, and visual observations into structured co-creation plans. To support low-latency execution, Co-policy introduces a Gaussian-Mixture Visuomotor Policy (GMP), implemented as a...
论文介绍 本文提出了Co-policy框架,用于实现人类与机器人之间响应式的音乐协同创作。该框架解决了连接语义音乐理解与实时物理表演的挑战,它将任务分解为语义意图接地、受限音乐变奏和视觉运动执行三个模块。通过使用语义锚点和微调后的多模态模型,系统能将语音、音乐种子和视觉观察转化为结构化的协同创作计划,并执行低延迟的视觉运动策略。
第一作者: Youbin Yao · 方向: 机器人操作 · 来源: cs.RO
双臂操纵动作扩展层次化规划模仿学习
Abstract:Dual-arm manipulation can improve throughput via parallel execution, but collecting bimanual demonstrations for training is costly and difficult. We present ExS2D, a hierarchical action expansion framework that enables dual-arm manipulation from single-arm supervision. ExS2D first generates structured subtasks from textual instructions while explicitly capturing temporal precedence. It then grounds each subtask into executable actions through subtask-guided action mapping in observation. Finally, precedence-aware action allocation and synchronized planning are performed by a multimodal large language model driven coordinator to select collision-free dual-arm executions. Simulation experiments demonstrate that ExS2D reduces the average execution steps by 54.4% while maintaining a comparable success rate to a single-arm baseline. Real-robot experiments on four tasks further...
论文介绍 为解决双臂演示数据收集困难的问题,本文提出了ExS2D层次化动作扩展框架,能够从单臂监督中生成双臂操作。该框架首先从文本指令生成具有明确时间先后结构的子任务,然后通过观察将子任务映射为可执行动作。最后,由多模态大语言模型驱动的协调器执行基于优先级的动作分配和同步规划,选择无碰撞的双臂执行方案,在仿真和真实机器人实验中均展示了效率提升。
第一作者: Tai Hyoung Rhee · 方向: 具身智能 · 来源: cs.RO
热红外图像小波域去噪方向条纹指数机器人感知
Abstract:Thermal infrared (TIR) imaging has been a popular choice for field robotics due to its robust perception capability under low light visual degradation, but it suffers from severe stochastic and fixed-pattern noise that breaks downstream estimation. This noise is intensified indoors due to low thermal contrast and uniform temperature distributions, contributing to the relative lack of indoor TIR deployments. Existing TIR denoising methods exhibit a poor accuracy-efficiency tradeoff, either too slow for online deployment required in robotics or insufficiently robust to severe degradation, while typically being trained on synthetic noise. Addressing these problems, we propose TIDY, a lightweight wavelet-domain denoiser trained on real clean-noisy TIR data. By reformulating TIR denoising in the wavelet domain, TIDY explicitly disentangles noise from structural content, enabling...
论文介绍 热红外成像在低光环境下感知能力强,但易受随机噪声和固定模式噪声干扰,尤其在室内场景。现有去噪方法存在精度与效率的权衡问题。本文提出TIDY,一个轻量级的小波域去噪器,通过在小波域中显式分离噪声与结构内容,并引入方向条纹指数,使用真实数据训练,旨在为机器人在线部署提供一个鲁棒且高效的热红外图像去噪方案。
第一作者: Thien-Loc Ha · 方向: VLA 通用模型 · 来源: cs.RO
视觉-语言-动作模型等变性SO(2)流匹配机器人操作
Abstract:Vision-Language-Action (VLA) models have emerged as a powerful paradigm for generalist robot manipulation, yet they lack geometric inductive biases: policies trained at specific orientations require substantially more data to generalize across rotational configurations. We present \textsc{EquiVLA}, the first general framework for end-to-end $\mathrm{SO}(2)$-equivariant VLA models, applicable to any architecture coupling a frozen vision-language backbone with a flow-matching Diffusion Transformer action head. \textsc{EquiVLA} introduces \textsc{EquiPerceptor}, which produces approximately $\mathrm{SO}(2)$-equivariant visual representations from frozen ViT features; and \textsc{EquiActor}, an exactly $\mathrm{SO}(2)$-equivariant flow-matching Diffusion Transformer action head. Together, they establish an approximate $\mathrm{SO}(2)$ equivariance chain from camera observations to...
论文介绍 通用机器人操作中的视觉-语言-动作(VLA)模型缺乏几何归纳偏置,导致在不同旋转配置间泛化时需要大量数据。本文提出EquiVLA,首个用于端到端SO(2)-等变VLA模型的通用框架。它通过引入EquiPerceptor产生近似等变的视觉表征,并使用严格等变的EquiActor动作头,建立了从观测到动作的近似等变链,提升了模型在旋转对称性任务上的数据效率和泛化能力。
第一作者: Trong-Bao Ho · 方向: 导航与运动 · 来源: cs.RO
异步执行动作分块初始噪声前缀一致性流策略
Abstract:Action chunking enables robot policies to produce temporally coherent behavior, but generating multi-step action sequences with flow-based policies incurs latency that is incompatible with real-time control. Under asynchronous execution, the robot continues executing the current chunk while the next one is generated, causing even minor delays to create inconsistencies at chunk boundaries. Existing methods address this problem by steering generation toward the already executed action prefix. We instead show that prefix consistency can be achieved by selecting an appropriate initial noise before generation begins, allowing the unmodified flow ODE to produce a coherent next chunk. This reframes asynchronous inference as a noise selection problem rather than a trajectory steering problem. We introduce \textbf{PAINT}, a training-free method that finds this noise via backward Euler...
论文介绍 基于动作分块的流策略在生成多步动作序列时存在延迟,导致异步执行下动作块边界不一致。现有方法通过引导生成来对齐已执行的动作前缀。本文指出,可以通过在生成开始前选择合适的初始噪声来实现前缀一致性,从而将异步推理重新定义为一个噪声选择问题。论文提出了无需训练的PAINT方法,通过反向欧拉法寻找该初始噪声,以保证生成动作序列的连贯性。
第一作者: Shaoshan Liu · 方向: 多模态具身 · 来源: cs.RO
人形机器人数据标准物理一致性ISO标准具身交互数据
Abstract:The scalability of humanoid robots will depend not only on models and hardware, but also on whether physical experience can accumulate across robots, tasks, organizations, and time. Drawing on the authors' work in developing ISO/WD 26264-1, Humanoid robot datasets -- Part 1: General requirements, within ISO/TC 299/WG 16, this article argues that data standards are becoming foundational infrastructure for Physical AI. We develop three insights. First, humanoid robot data is embodied interaction data, not a collection of isolated digital samples; a useful dataset must preserve the relationship among robot body, action, task, scene, execution trace, and outcome. Second, its value depends on physical coherence: multimodal streams are reusable only when timing, coordinate frames, calibration, kinematics, units, and synchronization assumptions remain inspectable. Third, the main...
论文介绍 人形机器人的规模化发展依赖于跨机器人、任务和组织的经验数据积累。本文基于开发ISO标准的工作,论证了数据标准是物理AI的基础设施。它指出人形机器人数据是具身交互数据,其价值取决于物理一致性。数据集必须保持机器人身体、动作、任务、场景、执行轨迹和结果之间的关系,以及多模态流在时间、坐标、标定等维度上的可检查性。标准化数据是实现模型通用性的关键。
第一作者: Yinsen Jia · 方向: 机器人操作 · 来源: cs.RO
时间效率自模仿学习强化学习长时程操作策略优化
Abstract:Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while rare efficient behaviors may be forgotten during training. We argue that temporal efficiency itself provides a powerful and underutilized source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 15 distinct long-horizon manipulation tasks, TSIL consistently...
论文介绍 长时程机器人操作策略可能通过低效交互来利用密集奖励,且高效的稀有行为在训练中易被遗忘。本文认为,时间效率本身是强化学习中一种强大但未充分利用的自监督信号。为此,论文提出了时间自模仿学习框架,它能挖掘学习过程中生成的时间高效成功轨迹,并将其转化为可复用的监督信息,通过配置条件下的自适应时间目标和效率加权的自模仿来持续优化策略,提升长时程任务的效率与成功率。
第一作者: Marcus Hoerger · 方向: 具身智能 · 来源: cs.RO
POMDP在线规划扩散模型模型学习不确定性
Abstract:Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Decision Process (POMDP) provides a powerful framework for such a capability. Although POMDP-based planning has advanced significantly, its application to real-world problems is often limited by the difficulty of obtaining faithful POMDP models. We present Vectorized Online planning wIth Learned diffusion model for POMDP Agents (VOiLA), a framework that learns task-agnostic POMDP models for online planning under uncertainty. VOiLA learns transition and observation samplers using conditional diffusion models and learns observation-likelihood models for particle-based belief updates. To enable efficient online planning, the diffusion samplers are distilled into compact feedforward generators and integrated with Vectorized Online POMDP Planner (VOPP), an online POMDP...
论文介绍 在部分可观测马尔可夫决策过程中进行规划,是机器人应对不确定性的关键能力,但获取准确的POMDP模型是主要瓶颈。本文提出VOiLA框架,它使用条件扩散模型学习任务无关的转移和观测采样器,并训练观测似然模型用于信念更新。为提升在线规划效率,扩散采样器被蒸馏为紧凑的前馈生成器,并与向量化在线POMDP规划器集成,从而实现对新任务的快速在线规划。
第一作者: Rui Fukushima · 方向: 具身智能 · 来源: cs.RO
双向辅导运动技能学习社会交互行为一致性发展性学习
Abstract:Infants are well known to develop their motor skills through dense interaction with caregivers. Although such social interaction is crucial for human development, motor-skill learning in robots is often treated as a unidirectional process in which robots passively receive demonstrations from tutors. This overlooks a key property of social interaction: it is inherently bidirectional, with tutor and learner dynamically adapting to each other. In such interactions, the robot's past experiences may function as prior constraints that shape the dynamics of their co-developed trajectories. We hypothesize that bidirectional tutoring allows such constraints to guide the formation of consistent behavioral patterns that preserve behavioral coherence and support generalization, whereas unidirectional interaction lacks such constraints and leads to broader, less consistent behavioral...
论文介绍 机器人运动技能学习常被视为单向示范接收过程,忽略了人类发展中社会交互的双向本质。本文假设双向辅导——即导师与学习者动态适应对方——允许学习者的先验经验作为约束,引导形成一致的行为模式,从而保持行为连贯性并支持泛化。论文验证了这一假设,表明与单向交互相比,双向互动能引导机器人产生更稳定、更具一致性的技能学习。
第一作者: Joong-Gil Kim · 方向: 具身智能 · 来源: cs.RO
双足机器人主动脚趾敏捷性效率冲击吸收
Abstract:Human legs exhibit high efficiency, agility, and impact absorption, with toes playing a crucial role in these capabilities. While many attempts have been made to implement human-like toes in robots, they have not fully replicated human characteristics nor rigorously validated their benefits. We propose a 14-DOF biped robot emulating human toes' lightweight, high-torque, robust nature. To quantitatively analyze the effectiveness of the active toes in terms of agility, efficiency, and impact absorption, we developed a high-fidelity simulation training environment that reflects actual actuators with coupled transmissions and accurate power consumption. To ensure a fair comparison between configurations with and without active toes, we designed a minimal RL reward function and applied an identical training procedure to both. The simulation results indicate that, at 1.33 m/s...
论文介绍 人类脚趾在行走效率、敏捷性和冲击吸收中作用关键。现有机器人脚趾研究未能完全复现其特性。本文提出一种具有14个自由度、模拟人类脚趾的轻量化、高扭矩双足机器人。为了定量分析主动脚趾的效果,研究构建了高保真仿真环境,并设计了最小化的强化学习奖励函数,以确保在有无主动脚趾配置下进行公平对比。仿真结果初步表明,在特定步速下,主动脚趾能提升机器人的能效和稳定性。
第一作者: Natapat Kirdwichai · 方向: 数据集与评测 · 来源: cs.RO
森林环境四足机器人机器人陷入多模态数据集ForEnt
Abstract:Legged robots are increasingly deployed in forests for ecological surveying and monitoring, yet their autonomy is often interrupted consequent to the challenges posed in traversing forest environments. Forest entrapments, for example, when a robot's legs are ensnared in vines or other vegetation, result in loss of stability and toppling. Such events not only disrupt the mission and require manual intervention, but also risk damage to the robot hardware. To address the absence of a dedicated dataset to investigate these failure modes in forest environments, we present ForEnt, a multi-modal dataset collected with the low-cost Unitree Go2 quadruped across eight forest sites in the Southampton Common Woodlands, UK. For our dataset, over approximately 1.7 km of traversals in 11 sequences were conducted, yielding 69 recorded entrapment events. ForEnt includes time-synchronized RGB-D...
论文介绍 该研究针对森林环境中四足机器人易受藤蔓等植被缠绕导致陷入的问题,提出了多模态数据集ForEnt。该数据集使用Unitree Go2机器人在英国南安普敦林地收集,包含RGB-D等传感器数据和69个陷入事件记录,旨在支持森林环境中机器人故障模式的研究,为自主导航和故障检测提供基准数据。
第一作者: Christian Schaible · 方向: 导航与运动 · 来源: cs.RO
安全导航阿克曼转向未映射环境局部路径规划凸优化
Abstract:A control framework is proposed for safe local navigation of mobile robots equipped with Ackermann steering in unmapped environments where a global goal is absent. Based on local obstacle detections, the safest heading angle is determined along the direction of the largest open space ahead of the vehicle. Guided by this direction, bounding lines are constructed on the left and right sides of the vehicle to achieve obstacle separation. These bounding lines are obtained by solving a convex quadratic optimization that maximizes vehicle-to-obstacle clearance. Optionally, conditions are imposed on the bounding lines to preserve parallelism and smooth abrupt changes from prior control steps. A feedback-linearizing controller is then used to regulate the vehicle's distance from one or both bounding lines, effectively enabling tracking of a local reference path that preserves safety...
论文介绍 本文提出了一种用于未映射环境中阿克曼转向移动机器人的安全本地导航控制框架。核心方法基于局部障碍物检测确定最安全航向,通过凸二次优化构建左右边界线以实现障碍物分离,并结合反馈线性化控制器跟踪局部参考路径,适用于自动驾驶和移动机器人在未知环境中的安全行驶。
第一作者: Calvin Luo · 方向: 多模态具身 · 来源: cs.RO
扩散过滤探索样本高效微调生成控制策略多智能体协作DF-ExpEnse
Abstract:A natural recipe for intelligent robotic decision-making is initializing from pretrained generative control policies, which have summarized offline experience, and adapting them to self-collected online experience. We present DF-ExpEnse, an exploration technique that improves the quality of online experience collection, thus increasing finetuning sample-efficiency. DF-ExpEnse leverages the multimodal modeling capabilities of the generative control policy to create an expressive and tractably evaluatable candidate set. It then utilizes an ensemble of critics to identify the action that best balances quality with high exploration interest. In fleet settings, DF-ExpEnse further enables cross-agent communication to facilitate collaborative exploration as a group. DF-ExpEnse can be seamlessly integrated with existing strategies that finetune pretrained generative control policies...
论文介绍 该研究针对预训练生成控制策略微调时样本效率低的问题,提出了DF-ExpEnse探索技术。该方法利用生成策略的多模态建模能力创建候选集,并通过评论家集成平衡动作质量与探索兴趣,支持多智能体间的协作探索,可无缝集成到现有微调策略中,提升机器人决策的在线学习效率。
第一作者: Ahmad Farooq · 方向: 策略学习 · 来源: cs.RO
多智能体通信形式验证决策树蒸馏安全保证策略抽象
Abstract:Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets. We present the first end-to-end framework for safety verification of learned multi-agent communication policies through policy abstraction: neural policies are distilled into interpretable decision trees, then formally verified, with empirical validation confirming that verified safety properties transfer to original networks. Our four-stage pipeline consists of domain-specific feature extraction from agent observations, decision tree distillation achieving 97.9% +/- 1.2% fidelity to neural policies, automated translation to PRISM probabilistic model checker specifications with complete...
论文介绍 本文首次提出一个端到端框架,用于形式化验证通过多智能体强化学习学到的通信策略。核心方法是将神经策略蒸馏为可解释的决策树,再转换为PRISM概率模型检查器规格进行验证,确保安全属性可转移回原网络,适用于无人机群和自动驾驶车队等安全关键场景。
第一作者: Ameya Salvi · 方向: 具身智能 · 来源: cs.RO
机器人故障识别检索增强生成仓储自动化意外事件检测Fail-RAG
Abstract:Industry automation is witnessing an evolution in robotics driven by both technological breakthroughs and societal changes: progress towards generalist robots, embodied and physical artificial intelligence (AI), and increasing labor shortage in this http URL intelligent autonomous robot needs to not only act according to planned motions but also react to any unexpected events. In this study, we focus on such unexpected events in warehouses where robots are used for material handling. Specifically, we refer to any unexpected events as failures and develop methods to detect robot operations related failures. Rule-based detection methods may break since the form of failures could change due to the dynamic nature of both environments and tasks. We propose 'Fail-RAG', a Retrieval Augmented Generation (RAG)-based failure detection framework where failure images and context...
论文介绍 该研究关注仓库环境中机器人操作时的意外事件检测问题,提出了基于检索增强生成的Fail-RAG框架。该方法利用故障图像和上下文信息,通过检索和生成机制识别机器人运行故障,以应对环境和任务动态变化,有助于提升工业自动化中机器人的自主监控和故障应对能力。
第一作者: Chuer Pan · 方向: 机器人操作 · 来源: cs.RO
动作视图增强视觉运动策略数据增强高斯溅射鱼眼相机
Abstract:Visuomotor policies for manipulation have demonstrated remarkable potential in modeling complex robotic behaviors, yet minor alterations in the robot's initial configuration and unseen obstacles easily lead to out-of-distribution observations. Without extensive data collection effort, these result in catastrophic execution failures. In this work, we introduce an effective data augmentation framework that generates visually realistic fisheye image sequences and corresponding physically feasible action trajectories from real-world eye-in-hand demonstrations, captured with a portable parallel gripper with a single fisheye camera. We introduce a novel Gaussian Splatting formulation, adapted to wide FoV fisheye cameras, to reconstruct and edit the 3D scene with unseen objects. We utilize trajectory optimization to generate smooth, collision-free, view-rendering-friendly action...
论文介绍 本文针对视觉运动策略在初始配置变化或新障碍物下容易失败的问题,提出了一种数据增强框架。核心方法是从单个手持鱼眼相机演示中生成增强的鱼眼图像序列和动作轨迹,使用改进的高斯溅射重建和编辑3D场景,并结合轨迹优化生成无碰撞路径,适用于机器人操作技能学习。
第一作者: Bennett Dogbey · 方向: 导航与运动 · 来源: cs.RO
概率可微分信号时序逻辑随机系统轨迹优化不确定环境pdSTL
Abstract:Autonomous robots operating in uncertain environments must satisfy complex temporal and safety specifications despite stochastic dynamics and sensing noise. While Signal Temporal Logic (STL) offers robustness measures for gradient-based optimization, existing extensions either lack differentiability or ignore belief-space uncertainty. We introduce pdSTL (probabilistic differentiable Signal Temporal Logic), a framework that unifies probabilistic semantics with differentiable robustness over belief trajectories. pdSTL employs interval-valued probabilistic semantics to compute conservative satisfaction bounds, propagated compositionally through the STL syntax tree. We formulate the temporal robustness evaluation as a recurrent, LSTM-style unfolding of STL operators, enabling linear-time, differentiable monitoring suitable for end-to-end trajectory optimization. We validate pdSTL...
论文介绍 该研究为自主机器人在不确定环境中满足复杂时序规范,提出了pdSTL框架。该方法统一了概率语义与可微分鲁棒性,通过区间值概率语义计算保守满足界,并将时序鲁棒性评估建模为循环展开,支持端到端轨迹优化,适用于机器人在随机动力学和感知噪声下的安全规划。
第一作者: Han Zheng · 方向: 导航与运动 · 来源: cs.RO
空间碰撞感知规划四足机器人导航局部路径规划三维占用图SCAN-Planner
Abstract:Quadruped robots are increasingly expected to navigate through narrow passages, cluttered indoor scenes, and large-scale 3D unstructured environments. Existing local planners commonly approximate the robot using isotropic geometric inflation or rely on planar and elevation-map representations, leading to conservative motion in tight spaces and limited reasoning about overhanging structures. This letter presents SCAN-Planner, a spatial collision-aware local planning framework for long-range quadruped navigation. A yaw-aware twin-cylinder footprint is used to model the elongated robot body, enabling whole-body collision evaluation through sparse queries in an inflated 3D occupancy map. We further introduce a projected A* search that generates collision-free guidance on an interpolated ground-following surface, with z-gradient suppression to avoid obstacles horizontally while...
论文介绍 本文提出SCAN-Planner,一种空间碰撞感知的局部规划框架,用于四足机器人在狭窄或复杂环境中的长距离导航。核心方法使用双圆柱足迹建模机器人身体,在膨胀的3D占用图中进行稀疏碰撞查询,并通过投影A*搜索在地面跟随曲面上生成无碰撞引导路径,提升机器人在室内和大型非结构化环境中的导航能力。
第一作者: Manuel Hernández · 方向: 具身智能 · 来源: cs.RO
范畴论层理论自动组件集合形式语义分布式系统
Abstract:The proliferation of large-scale, decentralized systems of autonomous agents, such as swarms of robots and networked cyber-physical systems, presents a formidable challenge to traditional formal methods. The Software Component Ensemble Language (SCEL) offers a formal model for such systems, but its operational semantics is not ideal for reasoning about global, structural, and emergent properties. This report proposes a new, multi-layered mathematical model for SCEL using category theory and sheaf theory. We argue that a society of robots described in SCEL can be formally modeled as a sheaf on a topological space, where components are points, ensembles are open sets, and distributed knowledge forms the sheaf's data. In this framework, computational processes like information sharing become equivalent to the sheaf-theoretic operation of "gluing" local data. System failures can...
论文介绍 针对大规模分散系统(如机器人蜂群)的形式化方法挑战,该研究提出一种基于范畴论和层理论的SCEL多层数学模型。它将机器人社会描述为拓扑空间上的层,其中组件为点、集合为开集、分布式知识构成层数据。这一框架使得信息共享等计算过程等同于层理论的“粘合”操作,有助于推理系统的全局、结构和涌现属性,并支持故障分析。
第一作者: Falak Mandali · 方向: 具身智能 · 来源: cs.RO
本体感觉状态估计人形机器人非惯性地面滤波器
Abstract:This paper presents an invariant extended Kalman filtering (InEKF) approach for real-time state estimation of humanoid robots operating on non-inertial ground using only onboard proprioceptive sensing. The proposed approach estimates the robot's base position and velocity relative to the moving ground frame without requiring direct measurements of ground motion or externally mounted sensors. By exploiting kinematic constraints at the stance foot through foot-mounted IMUs, the filter accounts for ground-induced nonlinearities in the process and measurement models while remaining fully proprioceptive. The estimator is formulated to admit a right-invariant measurement model, enabling favorable error dynamics under large initial uncertainties. Observability analysis establishes conditions under which the robot's relative base position and velocity are observable with respect to...
论文介绍 该研究针对人形机器人在非惯性地面上的状态估计问题,提出一种仅使用本体感觉传感器的不变扩展卡尔曼滤波方法。通过利用足部IMU的运动约束,该方法在过程与测量模型中处理地面非线性,无需外部传感器即可实时估计机器人相对移动地面的位置和速度。其右不变测量模型有助于在初始不确定性较大时维持良好的误差动态。
第一作者: Ryan Walker Brown · 方向: 导航与运动 · 来源: cs.RO
沙地运动阻力理论物理仿真MuJoCo机器人
Abstract:Recent advancements in Resistive Force Theory (RFT) enable approximation of ground reaction forces for locomotion in sand without the computational expense of modeling interactions with individual grains. However, these tools have been absent in 3D physics engines commonly used for robot simulation. We explore if resistive force approximations are sufficient, when integrated with standard dynamics calculations, to provide a stable substrate for a freely walking robot. To determine this, we implement 3D Granular Resistive Force Theory (3D RFT) in a physics simulation engine, MuJoCo. We verify simulations in multiple scenarios to demonstrate that key trends due to end effector shape, speed, and loading are preserved. Our implementation predicts walking distance and foot sinkage of a 12-Degree of Freedom hexapod robot within 20\% of experiments in sand. While RFT has inherent...
论文介绍 为解决沙地中机器人运动仿真的计算成本问题,该研究将三维颗粒阻力理论集成到开源物理引擎MuJoCo中。通过结合标准动力学计算,3D RFT能近似沙地反作用力,无需建模单个颗粒。仿真实验表明,该实现能保留末端执行器形状、速度和加载等关键趋势,并预测六足机器人的行走距离和脚沉降,与实际沙地实验数据误差在20%以内。
第一作者: Junyi Zhang · 方向: 具身智能 · 来源: cs.RO
游戏学习代理机器人技能获取代码策略持续学习
Abstract:Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions. We study Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play as a continual skill-learning stage before downstream tasks arrive. We introduce RATs, Robotics Agent Teams designed for play-time skill acquisition. During play, RATs proposes novel yet learnable exploratory tasks, plans and executes robot-code policies, verifies intermediate progress, diagnoses failures, retries with dense, step-level feedback, and distills successful executions into a persistent code skill library. At test time, the agent reuses relevant skills from this frozen library to help solve new tasks. Experiments in LIBERO-PRO and...
论文介绍 现有代理机器人系统多为任务驱动,技能获取依赖明确指令。该研究提出“游戏式代理机器人学习”,引入机器人代理团队RATs,在游戏阶段通过自我导向探索进行持续技能学习。RATs提出可学习任务、规划执行代码策略、验证进度、诊断故障并积累技能库。测试时,代理从冻结的技能库中重用相关技能来解决新任务,提升了技能复用和任务适应性。
第一作者: Hongkang Cui · 方向: 机器人操作 · 来源: cs.RO
视觉伺服扩散策略生成模型机器人操作鲁棒性
Abstract:Visual servoing is a fundamental technique in robotic manipulation and navigation. Regression-based visual servoing frequently experiences trajectory jitter as a result of noise-sensitive single-step mappings and the accumulation of errors during distribution shifts. In contrast, Diffusion Policy maintains temporal consistency by predicting action sequences and improves robustness through implicit data augmentation. This paper presents a novel diffusion-based servoing method. Based on Diffusion Policy, the proposed approach uses normalized image coordinates of observed tag corners as input and generates camera velocity through conditional denoising. To overcome the generalization limitations of models trained on static datasets, an online training paradigm is adopted, continuously expanding the diversity of training data through interactive experience collection. This strategy...
论文介绍 视觉伺服是机器人操作中的基础技术,但回归方法易受噪声和分布偏移影响导致轨迹抖动。该研究提出DiffusionVS,一种基于扩散策略的生成框架。它使用归一化图像坐标作为输入,通过条件去噪生成相机速度,以保持时间一致性并利用隐式数据增强提高鲁棒性。为克服静态数据集的泛化限制,采用在线训练范式,通过交互经验收集持续扩展训练数据多样性。
第一作者: Dennis Rotondi · 方向: 机器人操作 · 来源: cs.RO
3D场景图空间AI机器人计算机视觉综述
Abstract:3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including manipulation, navigation, task planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real-world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize...
论文介绍 3D场景图结合几何、语义和关系抽象,在机器人和计算机视觉中应用广泛,但领域存在碎片化问题。该综述提供统一和批判性回顾,首先在共同定义下形式化3D场景图,分析主要建模选择。文章特别强调开放挑战和未来方向,如构建管道、评估协议和现实部署问题,旨在促进方法比较并指导相关应用研究。
第一作者: Wenbo Ma · 方向: VLA 通用模型 · 来源: cs.RO
机器人装配基准测试智能制造LEGO基线方法
Abstract:We introduceWorkBenchMark, a LEGO Duplo-based robotic assembly benchmark motivated by the RoboCup Smart Manufacturing League. Robotic assembly couples low-level manipulation with task-level symbolic reasoning under physical constraints, a combination that current end-to-end learning methods do not yet solve reliably. The benchmark provides 400 tasks across four complexity tiers. We provide an open-vocabulary perception, Assembly-by-Disassembly baseline solution. Our planning-based pipeline outperforms a modern vision-language-action approach across all tiers. The benchmark, simulation environment, and baseline implementation will be released openly to support the broader robotic assembly community.
论文介绍 机器人装配需结合底层操作与任务级符号推理,当前端到端学习方法可靠性不足。该研究引入WorkBenchMark,一个基于LEGO的装配基准,包含四个复杂度层级共400个任务。它提供开放词汇感知和Assembly-by-Disassembly基线解决方案,基于规划的流程在所有层级上优于视觉-语言-动作方法。基准、仿真环境和基线实现将开源,以支持机器人装配社区。
第一作者: Khurram Javed · 方向: 策略学习 · 来源: cs.RO
强化学习物理平台机器人Atari实时学习
Abstract:We built a robot called the Robotroller that actuates an Atari CX40+ controller and a device called the Atari Devbox that renders the game frame and the reward signal from the Arcade Learning Environment on a screen. The Robotroller and the Atari Devbox, together with an off-the-shelf camera and a desktop computer, constitute a system that can be used to study reinforcement learning algorithms in the physical world. We call the full system Physical Atari. In this paper, we detail the key decisions that make Physical Atari a robust and accessible platform. To make the system robust, we designed the Robotroller so that all movement is done through bearings, which reduces wear. Additionally, we wrote software that monitors the state of the servos at a high frequency and intervenes to limit stress. To make the system accessible, we used affordable off-the-shelf components and...
论文介绍 为支持物理世界中的实时强化学习研究,该研究构建了Physical Atari系统,包括用于操纵Atari控制器的Robotroller和渲染游戏画面与奖励信号的Atari Devbox。系统使用廉价现货组件,注重鲁棒性:Robotroller通过轴承减少磨损,软件高频监控伺服状态以限制应力。该平台提供了耐用、可访问的测试环境,便于算法在现实条件下的验证。
第一作者: Yuyang Zhang · 方向: 具身智能 · 来源: cs.RO
世界动作模型视频生成图像编辑机器人控制具身智能
Abstract:World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit...
论文介绍 本文针对世界动作模型(WAMs)中视频生成的局限性,如推理成本高、关注动作无关细节以及长期想象引入错误,提出ImageWAM框架。该框架将预训练图像编辑模型重新用于机器人动作预测,仅需建模目标帧转换,聚焦于动作相关的当前到目标视觉差异,从而更高效地实现任务指令的局部视觉变化,可能提升具身智能系统的动作规划性能。
第一作者: Ellina Zhang · 方向: 机器人操作 · 来源: cs.RO
3D场景表示对象中心学习自监督学习潜在粒子机器人操作
Abstract:We introduce 3D-DLP, a self-supervised object-centric representation learning model that decomposes scene-level RGB-D or voxel observations into a set of 3D latent particles. Building on the Deep Latent Particles (DLP) framework, each particle encodes disentangled attributes, including 3D keypoint position, bounding box dimensions, and appearance features, and represents a distinct entity in the scene. The model learns interpretable per-particle segmentation maps through an end-to-end self-supervised reconstruction objective. We demonstrate on both simulated and real-world datasets that the learned latent space is interpretable and controllable: by manipulating particle positions and decoding, we can generate novel scene configurations. Furthermore, we show that leveraging these compact 3D latent particles for downstream robotic manipulation improves performance over baselines...
论文介绍 本文提出3D-DLP,一种自监督对象中心表示学习模型,能将场景级RGB-D或体素观测分解为一组3D潜在粒子。每个粒子编码解耦的属性,如3D关键点位置和外观特征,通过端到端自监督重建目标学习可解释的分割图。该模型在模拟和真实数据上展示了可解释和可控的潜在空间,通过操纵粒子位置可生成新场景配置,并提升机器人操作等下游任务性能。
第一作者: Juncheng Ma · 方向: 具身智能 · 来源: cs.CV
具身预训练自我中心视频机器人数据数据缩放具身智能
Abstract:Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity. These limitations have sparked interest in egocentric human video as a scalable, substantially lower-cost, and more diverse alternative for embodied model pretraining. However, its effectiveness compared to teleoperated real-robot data remains underexplored. To address this question, we conduct a systematic study comparing egocentric human video and teleoperated real-robot trajectories as pretraining data sources for embodied foundation models, under fixed post-training and...
论文介绍 本文针对具身基础模型预训练中数据瓶颈问题,系统研究自我中心人类视频与远程操作机器人轨迹作为预训练数据源的效果。通过固定后训练和评估协议,比较两者在具身智能任务中的表现,旨在验证人类视频作为更可扩展、低成本和多样化替代方案的有效性,为具身智能提供可行的数据缩放路径。
第一作者: Ganlin Yang · 方向: VLA 通用模型 · 来源: cs.CV
视觉-语言-动作模型长期记忆事件驱动关键帧记忆机器人操作
Abstract:Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historical context, they either suffer from severe information bottlenecks, incur high latency via decoupled dual systems, or rely on unselective buffers that accumulate massive visual redundancies. To address these limitations, we introduce EventVLA, an end-to-end framework founded on the concept of sparse visual evidence memory that comprises two core components: foundational visual anchors to retain initial and short-term contexts, and a dynamic Keyframe Evidence Memory (KEM) module. Specifically, KEM directly predicts future keyframe probabilities from the VLA's latent embeddings to autonomously capture and store sparse...
论文介绍 本文针对视觉-语言-动作(VLA)政策在长期机器人操作中因任务相关线索被遮挡而失效的记忆瓶颈,提出EventVLA框架。该框架基于稀疏视觉证据记忆概念,包含基础视觉锚点保留初始和短期上下文,以及动态关键帧证据记忆模块自主捕获和存储稀疏未来关键帧,从而提升政策在长期任务中的鲁棒性和效率。
第一作者: Jianing Li · 方向: 多模态具身 · 来源: cs.CV
视觉语言模型3D场景理解占用重建室内场景多模态具身
Abstract:Recently, vision-language models (VLMs) have made significant progress in 3D scene understanding, driving advances in applications such as embodied intelligence and robotic vision. However, existing approaches typically either rely directly on explicit 3D inputs (e.g., point clouds or RGB-D sequences), or introduce an additional 3D geometry encoder to derive 3D-aware visual tokens from 2D images. Such designs structurally decouple 3D geometric perception from the rich 2D semantics learned via vision-language pre-training, hindering the development of a unified 3D vision-language representation. In this work, we propose Occ-VLM, a novel framework for 3D scene understanding that operates purely on posed RGB images and employs a single 2D vision encoder. Specifically, Occ-VLM reconstructs 3D scene occupancy as an auxiliary geometric prior, which is utilized to spatially associate...
论文介绍 本文针对现有3D场景理解方法中3D几何感知与2D语义预训练脱节的问题,提出Occ-VLM框架。该框架仅依赖姿态RGB图像,通过单个2D视觉编码器重建3D场景占用作为几何先验,从而空间关联视觉和语言特征,实现统一的3D视觉语言表示,有望提升室内场景理解和具身智能应用。
第一作者: Navin Ranjan · 方向: VLA 通用模型 · 来源: cs.CV
模型量化混合精度视觉-语言-动作模型任务证据后训练量化
Abstract:We propose Mix-QVLA, a task-evidence-aware mixed-precision PTQ framework for VLA models. Mix-QVLA anchors each quantized variant to the full-precision action-token reference decision and evaluates whether quantization preserves task-relevant evidence across key VLA functional boundaries. It computes normalized gradient-weighted task-evidence maps from boundary activations and compares full-precision and quantized maps using evidence-mass and attribution-distribution distortion, capturing changes in both the strength and allocation of decision-supporting evidence. A soft-bottleneck objective aggregates boundary-level degradation into layer-wise sensitivity scores. Mix-QVLA further models sensitivity throughout task execution, capturing phase-dependent shifts in layer importance rather than assuming a fixed sensitivity profile. The resulting evidence- and time-aware scores guide...
论文介绍 本文提出Mix-QVLA,一种任务证据感知的混合精度后训练量化框架,用于视觉-语言-动作(VLA)模型。该框架通过计算归一化梯度加权任务证据图,比较全精度和量化版本的证据分布,捕获决策支持证据的变化,并建模任务执行中的时变灵敏度,从而指导混合精度分配,在压缩模型的同时保持任务性能。
第一作者: Wenli Xiao · 方向: 机器人操作 · 来源: cs.AI
机器人政策自我改进编码代理真实世界学习具身智能
Abstract:Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to...
论文介绍 本文针对真实世界机器人操作中依赖人类监督和算法工程的瓶颈,提出ENPIRE框架。该框架为编码代理提供物理反馈循环,通过环境模块自动重置和验证、政策改进模块启动细化、推出模块执行策略等,实例化重复性反馈循环,实现机器人政策的自我改进,减少人工干预,推动通用物理智能发展。
美股技术面呈现分化格局:以QQQ、NVDA为代表的部分科技权重股维持多头排列并接近52周高点,但META、MSFT等标的已进入空头排列,RSI值偏低。加密货币市场情绪极度恐慌(恐慌贪婪指数为23),总市值2.29万亿美元,BTC与ETH均处于空头排列且RSI低于50,技术趋势偏弱。中概股板块全面走弱,BABA、PDD、JD及腾讯均呈空头排列,其中BABA的RSI已进入超卖区间。商品外汇市场中,WTI原油和黄金期货价格均受均线压制,但美元指数(DXY)技术面强势,RSI超买且出现MACD金叉。
当前价格100.85接近52周高点,技术指标显示14周期RSI读数为70.9,处于超买区间。MACD指标中,快线高于信号线,形成金叉信号。同时,价格维持在所有关键均线上方,呈现多头排列格局,短期技术动量偏强。
价格107.1已跌破所有关键均线(SMA20/50/200),呈空头排列。14周期RSI读数仅为24.7,明确处于超卖状态。MACD柱状图为负值且在零轴下方,快慢线均为负值,整体技术结构偏弱,缺乏即时的向上动能。
近一日涨幅为4.82%,短期有所反弹。然而,当前价格73.08仍低于SMA50(79.98)和SMA200(97.77),中期均线呈空头排列。14周期RSI读数49.6接近中性,MACD柱状图虽为负但在收敛,显示下跌动量有所减缓,处于多空拉锯状态。
价格4172.9较52周高点回落明显,当前处于SMA20(4358.34)与SMA50(4545.92)下方,但高于SMA200。MACD指标出现死叉信号(快线低于信号线),但柱状图负值收窄。14周期RSI为36.5,未进入超卖区,显示下行压力存在但暂未极端化。
当前价格740.62紧贴52周高点(仅差1.07%),技术指标呈现多头排列(价格>SMA20>50>200)。14周期RSI读数59.1,处于强势区间但未超买。MACD柱状图为正,快线高于信号线。近一日涨幅2.51%,短期动量充足。
VIX 恐慌指数
10Y 美债收益率 (%)
美元指数 DXY
S&P 500 ETF
Nasdaq 100 ETF
Apple
Microsoft
Nvidia
Alphabet
Tesla
Meta
Bitcoin
Ethereum
Solana
阿里巴巴 (BABA)
拼多多 (PDD)
京东 (JD)
腾讯控股 (0700.HK)
黄金期货
WTI 原油期货
美元 / 人民币
本报告仅为基于公开市场行情数据计算的技术指标解读,所有描述均基于历史与当前读数。过去走势不代表未来表现,技术指标存在滞后性与局限性,不构成任何投资决策依据,仅供技术指标解读参考。
Follow live updates Get our breaking news email, free app or daily news podcast Ted O’Brien distanced himself from Pauline Hanson’s suggestion that Australia shouldn’t give aid to Pacific countries that also take aid from China. He said it was a legitimate concern, but her solution was “completely w
中文摘要 专家指出H5N1禽流感的到来对澳大利亚野生动物构成真正的紧急威胁;政府宣布以降低的利率延长燃油消费税回扣政策。
Talks come as Iran says it is closing the Strait of Hormuz over Israel's deadly attacks on Lebanon.
中文摘要 美国副总统万斯前往瑞士进行会谈;以色列在黎巴嫩的袭击导致16人死亡;伊朗因以色列的致命攻击宣布关闭霍尔木兹海峡。
Iranian delegation arrives in Switzerland for US peace talks
中文摘要 伊朗代表团已抵达瑞士,准备与美国进行和平会谈。
Vice President JD Vance said he would prioritize nuclear issues and renewed fighting in Lebanon in talks with an Iranian delegation in Switzerland. Shipping in the Strait of Hormuz faced new disruption after Iran’s military said it was closing the waterway.
中文摘要 美国副总统JD Vance称,与伊朗的会谈将优先讨论核问题和黎巴嫩战斗;伊朗军方宣布关闭霍尔木兹海峡,导致该海峡航运再次受阻。
Fifty-five ships had passed through the strait on Saturday, the U.S. military said. But then Iran’s military said it was closing the waterway once again.
中文摘要 美国军方称周六有55艘船只通过霍尔木兹海峡,但伊朗军方随后再次宣布关闭该水道,威胁航运复苏。
For a quarter century, Ms. Khalil ran a guesthouse and worked to protect endangered sea turtles who every summer lay their eggs on a stretch of beach near Tyre, Lebanon.
中文摘要 海龟保护者莫娜·哈利勒在以色列对黎巴嫩的空袭中遇难;她25年来在黎巴嫩提尔附近经营旅馆并保护濒危海龟。
The next phase of talks to end the war in Iran is expected to begin on Sunday amid fighting in Lebanon and renewed confusion over the Strait of Hormuz.
中文摘要 美国与伊朗官员计划在瑞士举行和平会谈,以结束伊朗战争;会谈预计周日启动,同时黎巴嫩冲突持续,霍尔木兹海峡局势紧张。
The Israeli military accused Ahmed Wishah of being a "Hamas sniper operative", without providing evidence.
中文摘要 以色列空袭加沙地带,导致六人死亡,其中包括半岛电视台摄像师艾哈迈德·维沙;以色列军方指控其为哈马斯狙击手,但未提供证据。
The US-Iran memorandum of understanding does not rule out future tolls in the strait after an initial 60-day period.
中文摘要 特朗普誓言伊朗不会对霍尔木兹海峡征收通行费,但表示美国可能会征收;美伊谅解备忘录未排除初始60天后收费的可能性。
For some it's a symbol of identity. For others, a challenge to the state. NPR's Itay Stern reports on the debate over the Palestinian flag in Israel.
中文摘要 以色列国内对巴勒斯坦象征的争议加剧,巴勒斯坦国旗引发关于身份认同与国家挑战的辩论。
For many Ismaili Muslims, seeing the Aga Khan is a once-in-a-lifetime event. NPR's Betsy Joles reports from his visit to remote northern Pakistan.
中文摘要 巴基斯坦什叶派穆斯林迎来精神领袖阿迦汗的访问,这对偏远北部的信徒而言是一次难得的盛会。
Iran says it has closed the Strait of Hormuz again. The U.S. military says traffic is still flowing. NPR's Jane Arraf reports from Beirut.
中文摘要 伊朗宣布再次关闭霍尔木兹海峡,但美国军方表示船只仍在正常通行。
中文摘要 中东新闻最新动态更新。
The US military has disputed Tehran's claim, which comes ahead of US-Iran talks in Switzerland on Sunday.
中文摘要 伊朗声称因以色列对黎巴嫩的袭击而关闭霍尔木兹海峡,但美国军方对此提出异议;该声明发布在美伊瑞士会谈前夕。
An advocacy group identified some of the victims as Muslims. Counterterrorism authorities are leading the investigation, but a motive for the attacks is so far unknown.
中文摘要 苏格兰发生袭击事件,导致5人受伤;一名男子被捕,部分受害者为穆斯林;反恐当局介入调查,袭击动机尚不明确。
As concerns over the Iran war recede, stock investors are confronting another threat: climate risk, which is prompting a reassessment of bets across sectors from agriculture to insurance.
中文摘要 随着伊朗战争担忧减退,股市投资者面临罕见的「超级厄尔尼诺」现象带来的新威胁——气候变化风险,促使市场对从农业到保险等多个行业的投资押注进行重新评估。
The latest update to the Federal Reserve’s favorite inflation gauge is unlikely to challenge a growing consensus at the US central bank around the need for interest-rate hikes this year.
中文摘要 美联储青睐的通胀指标即将公布最新数据,预计将显示通胀加速。然而,这不太可能动摇联储内部对于年内需要加息的日益增长的共识。
President Luiz Inacio Lula da Silva maintained his lead in polls on the Brazilian presidential election against Senator Flavio Bolsonaro, who is weighed down by his ties to a former banker involved in the country’s biggest bank fraud scandal.
中文摘要 在巴西总统选举民调中,总统卢拉·达席尔瓦在银行丑闻后依然领先于参议员弗拉维奥·博尔索纳罗。后者因其与该国最大银行欺诈案涉事前银行家的关联而受到拖累。
David Gura, Christina Ruffini, and Lisa Mateo of “Bloomberg This Weekend” play Pointed! Wager your points, leverage your bets and answer wisely. A new quiz is available to play each week on Bloomberg.com (Source: Bloomberg)
中文摘要 彭博新闻周末节目《Pointed News Quiz》探讨了债券、流媒体和酸奶等话题。节目邀请嘉宾进行知识竞答,每周都会在彭博网站更新。
As the nation prepares to celebrate its 250th anniversary, we can look back and see the recurring themes and lessons in American history. Host of the podcast History That Doesn't Suck and professor at Utah Valley University's Center for Constitutional Studies Greg Jackson joins Bloomberg This Weeken
中文摘要 在美国准备庆祝建国250周年之际,节目回顾了美国历史中反复出现的主题与教训。犹他谷大学宪政研究中心的教授参与讨论。
In 2015, sanctions played a key role in bringing Iran to the negotiating table but now warfare has been the primary tool. Oliver Wyman Partner and Global Anti-Financial Crime Practice Leader Daniel Tannebaum explains to David Gura and Christina Ruffini on Bloomberg This Weekend how even if sanctions
中文摘要 分析指出,与2015年将伊朗拉回谈判桌时的作用相比,当前制裁措施的效力正在减弱,而军事行动已成为主要手段。
Senator Raphael Warnock (D-GA) tells David Gura on Bloomberg This Weekend that the focus of the country should be on how local Main Streets are doing, not on well Wall Street is doing. Watch part one of Senator Warnock's interview now and tune in to Bloomberg This Weekend tomorrow for part two. (Sou
中文摘要 美国参议员拉斐尔·沃诺克表示,国家的关注点应放在地方「主街」经济的状况,而非华尔街的表现上。
While the United States and Iran have a new Memorandum of Understanding (MOU), Israel has been sidelined from the talks. Bloomberg News' Jerusalem reporter Dan Williams explains the embarrassing situation for Prime Minister Netanyahu to David Gura and Christina Ruffini on Bloomberg This Weekend. (So
中文摘要 在美国与伊朗达成新的谅解备忘录后,以色列被排除在谈判之外,这对以色列总理内塔尼亚胡构成了尴尬的打击。
South Africa’s central bank governor said policymakers are seeing early signs of second-round inflation effects as underlying price pressures build, emphasizing the need to act.
中文摘要 南非央行行长表示,政策制定者观察到通胀预期上升的早期迹象,底层价格压力正在积聚,强调需要采取行动应对。
The Algerian national soccer team was assigned to Lawrence, Kansas for training and practice during the World Cup and have been met with an enthusiastic reception from local residents. Mayor of Lawrence Brad Finkeldei shares with David Gura and Christina Ruffini on Bloomberg This Weekend the joys of
中文摘要 阿尔及利亚国家足球队在世界杯期间于堪萨斯州劳伦斯市进行训练,受到了当地居民和市长的热烈欢迎。
At the Reindustrialize Summit in Detroit, industrial leaders, White House officials, and major investors convened to emphasize the critical role of manufacturing in reinforcing national and military strength. Axios Defense reporter Colin Demarest joins David Gura and Christina Ruffini on Bloomberg T
中文摘要 在底特律举行的「再工业化峰会」上,工业领袖、白宫官员和主要投资者齐聚一堂,强调制造业在巩固国家与军事实力方面的关键作用。
Georgetown University School of Foreign Service Adjunct Professor Ali Vaez joins David Gura and Christina Ruffini on Bloomberg This Weekend and explains his cautious optimism for the recent Memorandum of Understanding (MOU) involving the US, Israel, and Iran. While Vaez acknowledges a deal within th
中文摘要 专家对近期美国、以色列和伊朗之间达成的谅解备忘录持谨慎乐观态度,认为这为各方提供了喘息空间。
9 回复 · Apple 节点
9 回复 · 程序员 节点
14 回复 · 程序员 节点
9 回复 · 程序员 节点
41 回复 · 程序员 节点
36 回复 · Apple 节点
9 回复 · Apple 节点
60 回复 · Apple 节点
29 回复 · Apple 节点
166 回复 · Apple 节点
从Claude for Chrome ++继续讨论: 原帖已经过去很久了 当时认为并不需要更新 但还是把 A\ 想简单了 所以还是花了点时间更新了下 可惜没留 旧版的原版 细节改动上diff不出来了 只能硬读以前写的 来改新版 随便一段 prompt 就完成了的快捷小功能 (Mimo v2.5 绝配haha 有视觉 干这个绰绰有余) https://haleclipse.lanzout.com/b00q1i4t6b 密码:ged3 CC联动 需Patch apply-claude-code-chrome-fix.zip (11.1 KB) 15 个帖子 - 15 位参与者 阅读完整话题
以下为废话引子: 最近越用 AI 编程,越觉得一件事很有意思: 以前脑子里冒出一个小需求,通常会因为“不会写”“太麻烦”“没人会做”“不值得花钱找人做”,最后就不了了之。 现在,很多原本只存在于脑内的奇怪需求,真的可以被自己一点点搓出来。 所以不是叫大家谈“用 Claude Code / Codex / Cursor / Windsurf 做得怎么样”。 而是: AI 到底把你脑子里哪些原本不可能落地的想法,变成了真正能跑、能用、甚至已经离不开的小东西? 不要求开源。 尤其是这些: 只解决你自己一个琐碎痛点,但现在每天都在用的东西 工作里偷偷替你省掉大量重复劳动的工具 给家人、朋友、同事做的小
本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已
伊朗称霍尔木兹海峡已关闭。 这海峡开开关关比我家电灯都勤… 周一打算把科技卖点掉,估计周一晚美股不保 57 个帖子 - 39 位参与者 阅读完整话题
本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 / 否 以下为项目介绍正文内容,AI生成、润
本帖使用社区公益推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的项目是免费使用的,无收费(变相收费、赞助)部分: 是 我的帖子已经打上 公益推广 标签: 是 我的项目属于个人项目,与公司或商业机构无关: 是 我的项目不存在QQ、TG等群组引流: 是 我的项目不存在非运营必要的网站引流: 是 我的项目不存在为他人推广、AFF: 是 我的项目无关联的商业项目: 是 我的站点存在登录,并已接入 LINUX DO Connect: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已
12 个帖子 - 12 位参与者 阅读完整话题
以下为个人观点,基本上全是暴论,不喜勿喷,给孩子留点面子吧 我是从 2024 年读高二的时候了解到 Vibe Coding 的(当时大家还没有叫它 Vibe Coding ),当时 DeepSeek 刚出 R1 ,除了 OpenAI 的 GPT-o1 之外,大家还没来得及用上思维链,也没有那么强的性能,参数量最大的模型的话好像是 R1 的 671B 。 当时 Token 还没有现在那么便宜,而我又是穷鬼高中生,听说这玩意儿可以辅助编程,便心心念念日思夜想,看着 GPT-o1 馋到流口水,根本买不起订阅。 所幸,我一位玩得好的远房表哥在 AWS 悉尼工作,平时喜欢折腾新奇技术,跟他聊天的时候了解
「君の星辰」 「聊天/售后」 52 个帖子 - 28 位参与者 阅读完整话题
15 个帖子 - 10 位参与者 阅读完整话题