Trending Papers

90

GitHub 85k arXiv Page

TradingAgents: Multi-Agents LLM Financial Trading Framework

A multi-agent framework using large language models for stock trading simulates real-world trading firms, improving performance metrics like cumulative returns and Sharpe ratio.

4 authors

· Dec 28, 2024

90

GitHub 85k arXiv Page

Submitted by

XinyangDavidHan

Agents' Last Exam

Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on long-term, economically valuable real-world tasks across 13 industry clusters with 1K+ tasks, revealing significant gaps between benchmark performance and practical deployment.

UC Berkeley · Published on Jun 3, 2026

177

GitHub 499 arXiv Page

Submitted by

XinyangDavidHan

Agents' Last Exam

Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on long-term, economically valuable real-world tasks across 13 industry clusters with 1K+ tasks, revealing significant gaps between benchmark performance and practical deployment.

UC Berkeley · Jun 3, 2026

177

GitHub 499 arXiv Page

Submitted by

ChengCui

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

PaddleOCR-VL-1.6 enhances document parsing performance through targeted data optimization and progressive post-training techniques, achieving state-of-the-art results on OmniDocBench v1.6.

PaddlePaddle · Published on Jun 2, 2026

GitHub 81.7k arXiv Page

Submitted by

ChengCui

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

PaddleOCR-VL-1.6 enhances document parsing performance through targeted data optimization and progressive post-training techniques, achieving state-of-the-art results on OmniDocBench v1.6.

PaddlePaddle · Jun 2, 2026

GitHub 81.7k arXiv Page

Submitted by

zhouxiangxin

Rethinking the Divergence Regularization in LLM RL

DRPO improves LLM reinforcement learning stability by replacing hard masks with smooth regularization that provides continuous gradient corrections beyond trust-region boundaries.

Tencent-Hunyuan-Multimodal-RL · Published on Jun 8, 2026

26

GitHub 408 arXiv Page

Submitted by

zhouxiangxin

Rethinking the Divergence Regularization in LLM RL

DRPO improves LLM reinforcement learning stability by replacing hard masks with smooth regularization that provides continuous gradient corrections beyond trust-region boundaries.

Tencent-Hunyuan-Multimodal-RL · Jun 8, 2026

26

GitHub 408 arXiv Page

Submitted by

unilm

VibeVoice Technical Report

VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.

Microsoft Research · Published on Aug 26, 2025

172

GitHub 49.2k arXiv Page

Submitted by

unilm

VibeVoice Technical Report

VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.

Microsoft Research · Aug 26, 2025

172

GitHub 49.2k arXiv Page

Submitted by

taesiri

SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

SCAIL-2 enables end-to-end character animation by directly transferring motion from driving videos without intermediate representations, using unified task decomposition and synthetic data generation.

4 authors

· Published on Jun 9, 2026

32

GitHub 156 arXiv Page

Submitted by

taesiri

SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

SCAIL-2 enables end-to-end character animation by directly transferring motion from driving videos without intermediate representations, using unified task decomposition and synthetic data generation.

4 authors

· Jun 9, 2026

32

GitHub 156 arXiv Page

Submitted by

taesiri

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

SkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.

Microsoft Research · Published on May 22, 2026

225

GitHub 5.66k arXiv Page

Submitted by

taesiri

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

SkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.

Microsoft Research · May 22, 2026

225

GitHub 5.66k arXiv Page

Submitted by

taesiri

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

MinerU2.5, a 1.2B-parameter document parsing vision-language model, achieves state-of-the-art recognition accuracy with computational efficiency through a coarse-to-fine parsing strategy.

61 authors

· Published on Sep 26, 2025

165

GitHub 67.1k arXiv Page

Submitted by

taesiri

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

MinerU2.5, a 1.2B-parameter document parsing vision-language model, achieves state-of-the-art recognition accuracy with computational efficiency through a coarse-to-fine parsing strategy.

61 authors

· Sep 26, 2025

165

GitHub 67.1k arXiv Page

Submitted by

taesiri

Cosmos 3: Omnimodal World Models for Physical AI

Cosmos 3 is an omnimodal world model that processes and generates multiple data types through a unified mixture-of-transformers architecture, achieving state-of-the-art performance in various understanding and generation tasks.

NVIDIA · Published on Jun 1, 2026

GitHub 9.83k arXiv Page

Submitted by

taesiri

Cosmos 3: Omnimodal World Models for Physical AI

Cosmos 3 is an omnimodal world model that processes and generates multiple data types through a unified mixture-of-transformers architecture, achieving state-of-the-art performance in various understanding and generation tasks.

NVIDIA · Jun 1, 2026

GitHub 9.83k arXiv Page

Kronos: A Foundation Model for the Language of Financial Markets

Kronos, a specialized pre-training framework for financial K-line data, outperforms existing models in forecasting and synthetic data generation through a unique tokenizer and autoregressive pre-training on a large dataset.

7 authors

· Published on Aug 2, 2025

40

GitHub 29.1k arXiv Page

Kronos: A Foundation Model for the Language of Financial Markets

Kronos, a specialized pre-training framework for financial K-line data, outperforms existing models in forecasting and synthetic data generation through a unique tokenizer and autoregressive pre-training on a large dataset.

7 authors

· Aug 2, 2025

40

GitHub 29.1k arXiv Page

Submitted by

lhmd

Latent Spatial Memory for Video World Models

Latent spatial memory for video world models stores 3D scene information directly in diffusion latent space, eliminating pixel-space reconstruction overhead and achieving faster generation with reduced memory usage.

Microsoft Research · Published on Jun 8, 2026

GitHub 158 arXiv Page

Submitted by

lhmd

Latent Spatial Memory for Video World Models

Latent spatial memory for video world models stores 3D scene information directly in diffusion latent space, eliminating pixel-space reconstruction overhead and achieving faster generation with reduced memory usage.

Microsoft Research · Jun 8, 2026

GitHub 158 arXiv Page

Submitted by

taesiri

dots.tts Technical Report

A 2B-parameter continuous autoregressive text-to-speech model trained on a multilingual corpus achieves state-of-the-art performance on multiple benchmarks while enabling efficient low-latency speech generation through specialized distillation techniques.

9 authors

· Published on Jun 5, 2026

GitHub 428 arXiv Page

Submitted by

taesiri

dots.tts Technical Report

A 2B-parameter continuous autoregressive text-to-speech model trained on a multilingual corpus achieves state-of-the-art performance on multiple benchmarks while enabling efficient low-latency speech generation through specialized distillation techniques.

9 authors

· Jun 5, 2026

GitHub 428 arXiv Page

Submitted by

akhaliq

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory systems in terms of accuracy and computational efficiency.

5 authors

· Published on Apr 28, 2025

60

Submitted by

akhaliq

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory systems in terms of accuracy and computational efficiency.

5 authors

· Apr 28, 2025

60

Submitted by

akhaliq

Efficient Memory Management for Large Language Model Serving with PagedAttention

PagedAttention algorithm and vLLM system enhance the throughput of large language models by efficiently managing memory and reducing waste in the key-value cache.

9 authors

· Published on Sep 12, 2023

GitHub 82.5k arXiv Page

Submitted by

akhaliq

Efficient Memory Management for Large Language Model Serving with PagedAttention

PagedAttention algorithm and vLLM system enhance the throughput of large language models by efficiently managing memory and reducing waste in the key-value cache.

9 authors

· Sep 12, 2023

GitHub 82.5k arXiv Page

Submitted by

akhaliq

OpenDevin: An Open Platform for AI Software Developers as Generalist Agents

OpenDevin is a platform for developing AI agents that interact with the world by writing code, using command lines, and browsing the web, with support for multiple agents and evaluation benchmarks.

24 authors

· Published on Jul 23, 2024

80

GitHub 76.4k arXiv Page

Submitted by

akhaliq

OpenDevin: An Open Platform for AI Software Developers as Generalist Agents

OpenDevin is a platform for developing AI agents that interact with the world by writing code, using command lines, and browsing the web, with support for multiple agents and evaluation benchmarks.

24 authors

· Jul 23, 2024

80

GitHub 76.4k arXiv Page

Submitted by

RuofengYang

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

ARIS is an open-source research harness that uses cross-model adversarial collaboration to ensure reliable long-term research outcomes through coordinated execution, orchestration, and assurance layers.

Shanghai Jiao Tong University · Published on May 4, 2026

131

GitHub 11.9k arXiv Page

Submitted by

RuofengYang

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

ARIS is an open-source research harness that uses cross-model adversarial collaboration to ensure reliable long-term research outcomes through coordinated execution, orchestration, and assurance layers.

Shanghai Jiao Tong University · May 4, 2026

131

GitHub 11.9k arXiv Page

Submitted by

pat-jj

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

A 20B search agent trained with reinforcement learning within a stateful search framework demonstrates superior retrieval performance across multiple domains by separating semantic decision-making from environmental bookkeeping.

chroma · Published on Jun 1, 2026

51

GitHub 488 arXiv Page

Submitted by

pat-jj

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

A 20B search agent trained with reinforcement learning within a stateful search framework demonstrates superior retrieval performance across multiple domains by separating semantic decision-making from environmental bookkeeping.

chroma · Jun 1, 2026

51

GitHub 488 arXiv Page

Submitted by

taesiri

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications

AgentScope enhances agentic applications by providing flexible tool-based interactions, unified interfaces, and advanced infrastructure based on the ReAct paradigm, supporting efficient and safe development and deployment.

23 authors

· Published on Aug 22, 2025

64

GitHub 26.6k arXiv Page

Submitted by

taesiri

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications

AgentScope enhances agentic applications by providing flexible tool-based interactions, unified interfaces, and advanced infrastructure based on the ReAct paradigm, supporting efficient and safe development and deployment.

23 authors

· Aug 22, 2025

64

GitHub 26.6k arXiv Page

Submitted by

akhaliq

Very Large-Scale Multi-Agent Simulation in AgentScope

Enhancements to the AgentScope platform improve scalability, efficiency, and ease of use for large-scale multi-agent simulations through distributed mechanisms, flexible environments, and user-friendly tools.

8 authors

· Published on Jul 25, 2024

42

GitHub 26.7k arXiv Page

Submitted by

akhaliq

Very Large-Scale Multi-Agent Simulation in AgentScope

Enhancements to the AgentScope platform improve scalability, efficiency, and ease of use for large-scale multi-agent simulations through distributed mechanisms, flexible environments, and user-friendly tools.

8 authors

· Jul 25, 2024

42

GitHub 26.7k arXiv Page

Submitted by

VigneshHexo

SIA: Self Improving AI with Harness & Weight Updates

A self-improving AI framework simultaneously updates both model weights and task-specific agent architecture through a language-model feedback agent across legal classification, GPU optimization, and biological data denoising tasks.

Hexo AI · Published on May 26, 2026

GitHub 931 arXiv Page

Submitted by

VigneshHexo

SIA: Self Improving AI with Harness & Weight Updates

A self-improving AI framework simultaneously updates both model weights and task-specific agent architecture through a language-model feedback agent across legal classification, GPU optimization, and biological data denoising tasks.

Hexo AI · May 26, 2026

GitHub 931 arXiv Page

EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning

EverMemOS presents a self-organizing memory system for large language models that processes dialogue streams into structured memory cells and scenes to enhance long-term interaction capabilities.

11 authors

· Published on Jan 5, 2026

6

GitHub 7.25k arXiv Page

EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning

EverMemOS presents a self-organizing memory system for large language models that processes dialogue streams into structured memory cells and scenes to enhance long-term interaction capabilities.

11 authors

· Jan 5, 2026

6

GitHub 7.25k arXiv Page

Submitted by

hzeroyuke

WorldOlympiad: Can Your World Model Survive a Triathlon?

WorldOlympiad presents a comprehensive benchmark for evaluating video-based world models across physical faithfulness, geometric consistency, and interaction fidelity, revealing significant gaps in current generative models' capabilities.

DAMO Academy · Published on Jun 9, 2026

27

Submitted by

hzeroyuke

WorldOlympiad: Can Your World Model Survive a Triathlon?

WorldOlympiad presents a comprehensive benchmark for evaluating video-based world models across physical faithfulness, geometric consistency, and interaction fidelity, revealing significant gaps in current generative models' capabilities.

DAMO Academy · Jun 9, 2026

27

Submitted by

andito

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

SmolDocling is a compact vision-language model that performs end-to-end document conversion with robust performance across various document types using 256M parameters and a new markup format.

IBM Granite · Published on Mar 14, 2025

161

GitHub 61.3k arXiv Page

Submitted by

andito

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

SmolDocling is a compact vision-language model that performs end-to-end document conversion with robust performance across various document types using 256M parameters and a new markup format.

IBM Granite · Mar 14, 2025

161

GitHub 61.3k arXiv Page

Submitted by

xuc865

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Role-Agent framework enables LLM agents to function as both agent and environment through bootstrapped co-evolution, improving performance via environment-aware reasoning and targeted practice.

7 authors

· Published on Jun 9, 2026

73

GitHub 74 arXiv Page

Submitted by

xuc865

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Role-Agent framework enables LLM agents to function as both agent and environment through bootstrapped co-evolution, improving performance via environment-aware reasoning and targeted practice.

7 authors

· Jun 9, 2026

73

GitHub 74 arXiv Page

Submitted by

akhaliq

Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Mixture of vision encoders and resolutions in multimodal large language models improves performance through concatenation of visual tokens and a Pre-Alignment mechanism, leading to superior results on benchmarks.

15 authors

· Published on Aug 28, 2024

88

GitHub 2.36k arXiv Page

Submitted by

akhaliq

Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Mixture of vision encoders and resolutions in multimodal large language models improves performance through concatenation of visual tokens and a Pre-Alignment mechanism, leading to superior results on benchmarks.

15 authors

· Aug 28, 2024

88

GitHub 2.36k arXiv Page

Submitted by

taesiri

Audio Interaction Model

A unified streaming audio model is developed that combines offline task execution with real-time audio instruction following through an end-to-end framework supporting multiple audio interaction capabilities.

National University of Singapore · Published on Jun 3, 2026

107

GitHub 326 arXiv Page

Submitted by

taesiri

Audio Interaction Model

A unified streaming audio model is developed that combines offline task execution with real-time audio instruction following through an end-to-end framework supporting multiple audio interaction capabilities.

National University of Singapore · Jun 3, 2026

107

GitHub 326 arXiv Page

Submitted by

BoJack

MMAE: A Massive Multitask Audio Editing Benchmark

MMAE presents a comprehensive benchmark for instruction-based audio editing across multiple modalities and complexity levels, revealing significant gaps in current model capabilities.

38 authors

· Published on Jun 5, 2026

43

GitHub 88 arXiv Page

Submitted by

BoJack

MMAE: A Massive Multitask Audio Editing Benchmark

MMAE presents a comprehensive benchmark for instruction-based audio editing across multiple modalities and complexity levels, revealing significant gaps in current model capabilities.

38 authors

· Jun 5, 2026

43

GitHub 88 arXiv Page

LightRAG: Simple and Fast Retrieval-Augmented Generation

LightRAG improves Retrieval-Augmented Generation by integrating graph structures for enhanced contextual awareness and efficient information retrieval, achieving better accuracy and response times.

5 authors

· Published on Oct 8, 2024

39

GitHub 36.4k arXiv Page

LightRAG: Simple and Fast Retrieval-Augmented Generation

LightRAG improves Retrieval-Augmented Generation by integrating graph structures for enhanced contextual awareness and efficient information retrieval, achieving better accuracy and response times.

5 authors

· Oct 8, 2024

39

GitHub 36.4k arXiv Page

Submitted by

jasonrqh

COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

Person-grounded AI skills are automatically distilled from heterogeneous traces into inspectable, correctable packages that capture both capabilities and behavioral patterns.

shanghai ailab · Published on May 29, 2026

GitHub 19.2k arXiv Page

Submitted by

jasonrqh

COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

Person-grounded AI skills are automatically distilled from heterogeneous traces into inspectable, correctable packages that capture both capabilities and behavioral patterns.

shanghai ailab · May 29, 2026

GitHub 19.2k arXiv Page

Submitted by

mxlin043

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

OmniGameArena presents a unified benchmark for evaluating vision-language model agents in diverse game settings with a reflection-based improvement protocol that tracks performance evolution and skill generalization.

Hong Kong University · Published on Jun 8, 2026

Submitted by

mxlin043

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

OmniGameArena presents a unified benchmark for evaluating vision-language model agents in diverse game settings with a reflection-based improvement protocol that tracks performance evolution and skill generalization.

Hong Kong University · Jun 8, 2026

Submitted by

qian43

ABot-Earth 0.5: Generative 3D Earth Model

ABot-Earth 0.5 generates realistic 3D environments from satellite imagery using 3D Gaussian Splatting representation, enabling fast synthesis and real-time visualization for Embodied AI applications.

Alibaba AMAP CV Lab · Published on Jun 8, 2026

193

GitHub 109 arXiv Page

Submitted by

qian43

ABot-Earth 0.5: Generative 3D Earth Model

ABot-Earth 0.5 generates realistic 3D environments from satellite imagery using 3D Gaussian Splatting representation, enabling fast synthesis and real-time visualization for Embodied AI applications.

Alibaba AMAP CV Lab · Jun 8, 2026

193

GitHub 109 arXiv Page

PDFMathTranslate: Scientific Document Translation Preserving Layouts

PDFMathTranslate enables layout-preserving scientific document translation using large language models and precise layout detection, offering improved precision, flexibility, and efficiency.

4 authors

· Published on Jul 2, 2025

2

GitHub 34.7k arXiv Page

PDFMathTranslate: Scientific Document Translation Preserving Layouts

PDFMathTranslate enables layout-preserving scientific document translation using large language models and precise layout detection, offering improved precision, flexibility, and efficiency.

4 authors

· Jul 2, 2025

2

GitHub 34.7k arXiv Page

Zep: A Temporal Knowledge Graph Architecture for Agent Memory

Zep, a memory layer service, outperforms MemGPT in the DMR benchmark and LongMemEval by excelling in dynamic knowledge integration and temporal reasoning, critical for enterprise use cases.

5 authors

· Published on Jan 20, 2025

GitHub 27.3k arXiv Page

Zep: A Temporal Knowledge Graph Architecture for Agent Memory

Zep, a memory layer service, outperforms MemGPT in the DMR benchmark and LongMemEval by excelling in dynamic knowledge integration and temporal reasoning, critical for enterprise use cases.

5 authors

· Jan 20, 2025

GitHub 27.3k arXiv Page

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

A novel GPT-based model, OmniFlatten, enables real-time natural full-duplex spoken dialogue through a multi-stage post-training technique that integrates speech and text without altering the original model's architecture.

9 authors

· Published on Oct 23, 2024

13

GitHub 59.5k arXiv Page

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

A novel GPT-based model, OmniFlatten, enables real-time natural full-duplex spoken dialogue through a multi-stage post-training technique that integrates speech and text without altering the original model's architecture.

9 authors

· Oct 23, 2024

13

GitHub 59.5k arXiv Page

Submitted by

MoeinAbtahi

Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents

Memanto presents a universal memory layer for agentic AI that eliminates computational overhead of hybrid semantic graph architectures through a typed semantic memory schema and information-theoretic search engine.

Moorcheh.ai · Published on Apr 23, 2026

GitHub 749 arXiv Page

Submitted by

MoeinAbtahi

Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents

Memanto presents a universal memory layer for agentic AI that eliminates computational overhead of hybrid semantic graph architectures through a typed semantic memory schema and information-theoretic search engine.

Moorcheh.ai · Apr 23, 2026

GitHub 749 arXiv Page

AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets

AI-Trader presents the first fully automated live benchmark for evaluating large language models in financial decision-making across multiple markets with autonomous information processing.

6 authors

· Published on Dec 1, 2025

8

GitHub 19.5k arXiv Page

AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets

AI-Trader presents the first fully automated live benchmark for evaluating large language models in financial decision-making across multiple markets with autonomous information processing.

6 authors

· Dec 1, 2025

8

GitHub 19.5k arXiv Page

Submitted by

fdugyt

MOSS-TTS Technical Report

MOSS-TTS is a speech generation model using discrete audio tokens and autoregressive modeling with capabilities for voice cloning, pronunciation control, and long-form generation across multiple languages.

OpenMOSS · Published on Mar 18, 2026

GitHub 3.26k arXiv Page

Submitted by

fdugyt

MOSS-TTS Technical Report

MOSS-TTS is a speech generation model using discrete audio tokens and autoregressive modeling with capabilities for voice cloning, pronunciation control, and long-form generation across multiple languages.

OpenMOSS · Mar 18, 2026

GitHub 3.26k arXiv Page

Submitted by

nielsr

Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models

YOLO26 addresses real-time vision challenges through a unified model family with NMS-free inference, improved training strategies, and multi-task capabilities spanning detection, segmentation, and pose estimation.

Ultralytics · Published on Jun 2, 2026

10

Submitted by

nielsr

Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models

YOLO26 addresses real-time vision challenges through a unified model family with NMS-free inference, improved training strategies, and multi-task capabilities spanning detection, segmentation, and pose estimation.

Ultralytics · Jun 2, 2026

10

Submitted by

Rbin

RAG-Anything: All-in-One RAG Framework

RAG-Anything is a unified framework that enhances multimodal knowledge retrieval by integrating cross-modal relationships and semantic matching, outperforming existing methods on complex benchmarks.

Data Intelligence Lab@HKU · Published on Oct 14, 2025

82

GitHub 21.1k arXiv Page

Submitted by

Rbin

RAG-Anything: All-in-One RAG Framework

RAG-Anything is a unified framework that enhances multimodal knowledge retrieval by integrating cross-modal relationships and semantic matching, outperforming existing methods on complex benchmarks.

Data Intelligence Lab@HKU · Oct 14, 2025

82

GitHub 21.1k arXiv Page

Submitted by

taesiri

LongCat-Video Technical Report

LongCat-Video, a 13.6B parameter video generation model based on the Diffusion Transformer framework, excels in efficient and high-quality long video generation across multiple tasks using unified architecture, coarse-to-fine generation, and block sparse attention.

LongCat · Published on Oct 25, 2025

38

GitHub 4.27k arXiv Page

Submitted by

taesiri

LongCat-Video Technical Report

LongCat-Video, a 13.6B parameter video generation model based on the Diffusion Transformer framework, excels in efficient and high-quality long video generation across multiple tasks using unified architecture, coarse-to-fine generation, and block sparse attention.

LongCat · Oct 25, 2025

38

GitHub 4.27k arXiv Page

Submitted by

taesiri

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

SkillClaw enables collective skill evolution in multi-user LLM agent systems by aggregating user interactions to autonomously update and improve reusable skills across the ecosystem.

8 authors

· Published on Apr 9, 2026

293

GitHub 1.84k arXiv Page

Submitted by

taesiri

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

SkillClaw enables collective skill evolution in multi-user LLM agent systems by aggregating user interactions to autonomously update and improve reusable skills across the ecosystem.

8 authors

· Apr 9, 2026

293

GitHub 1.84k arXiv Page

Submitted by

taesiri

GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

GRAIL generates diverse humanoid manipulation and locomotion data through 3D asset composition and video foundation models, enabling effective sim-to-real transfer for robot control.

NVIDIA · Published on Jun 3, 2026

7

GitHub 276 arXiv Page

Submitted by

taesiri

GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

GRAIL generates diverse humanoid manipulation and locomotion data through 3D asset composition and video foundation models, enabling effective sim-to-real transfer for robot control.

NVIDIA · Jun 3, 2026

7

GitHub 276 arXiv Page

Submitted by

zbhpku

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI

DataFlow is an LLM-driven data preparation framework that enhances data quality and reproducibility for various tasks, improving LLM performance with automatically generated pipelines.

Peking University · Published on Dec 18, 2025

222

GitHub 4.74k arXiv Page

Submitted by

zbhpku

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI

DataFlow is an LLM-driven data preparation framework that enhances data quality and reproducibility for various tasks, improving LLM performance with automatically generated pipelines.

Peking University · Dec 18, 2025

222

GitHub 4.74k arXiv Page

Cybersecurity AI: Humanoid Robots as Attack Vectors

The Unitree G1 humanoid robot is vulnerable to BLE provisioning protocol exploits, exfiltrates sensor data, and can be repurposed for active cyber operations, highlighting the need for improved security standards in commercial robotics.

3 authors

· Published on Sep 17, 2025

-

GitHub 9.05k arXiv Page

Cybersecurity AI: Humanoid Robots as Attack Vectors

The Unitree G1 humanoid robot is vulnerable to BLE provisioning protocol exploits, exfiltrates sensor data, and can be repurposed for active cyber operations, highlighting the need for improved security standards in commercial robotics.

3 authors

· Sep 17, 2025

-

GitHub 9.05k arXiv Page

Submitted by

YJ-142150

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

Autoregressive diffusion method for video-to-video lip synchronization achieves real-time performance through distillation and optimized inference schedules.

KAIST AI · Published on Jun 9, 2026

28

GitHub 24 arXiv Page

Submitted by

YJ-142150

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

Autoregressive diffusion method for video-to-video lip synchronization achieves real-time performance through distillation and optimized inference schedules.

KAIST AI · Jun 9, 2026

28

GitHub 24 arXiv Page

Submitted by

liangjiaqing

GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)

GenericAgent is a self-evolving large language model agent system that maximizes context information density through hierarchical memory, reusable SOPs, and efficient compression to overcome long-horizon limitations.

Fudan University · Published on Apr 18, 2026

22

GitHub 12.8k arXiv Page

Submitted by

liangjiaqing

GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)

GenericAgent is a self-evolving large language model agent system that maximizes context information density through hierarchical memory, reusable SOPs, and efficient compression to overcome long-horizon limitations.

Fudan University · Apr 18, 2026

22

GitHub 12.8k arXiv Page

Submitted by

Luckyyy

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

LLM-based agents perform poorly on VibeSearch benchmark, which evaluates multi-turn dialogue search scenarios reflecting real user-agent collaboration rather than traditional single-turn query tasks.

rednote-hilab · Published on May 27, 2026

GitHub 878 arXiv Page

Submitted by

Luckyyy

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

LLM-based agents perform poorly on VibeSearch benchmark, which evaluates multi-turn dialogue search scenarios reflecting real user-agent collaboration rather than traditional single-turn query tasks.

rednote-hilab · May 27, 2026

GitHub 878 arXiv Page

Submitted by

Paranioar

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Unified vision-language models treat understanding and generation as integrated processes rather than separate tasks, demonstrating strong performance across multiple multimodal capabilities including image synthesis and action reasoning.

SenseNova · Published on May 12, 2026

191

GitHub 2.85k arXiv Page

Submitted by

Paranioar

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Unified vision-language models treat understanding and generation as integrated processes rather than separate tasks, demonstrating strong performance across multiple multimodal capabilities including image synthesis and action reasoning.

SenseNova · May 12, 2026

191

GitHub 2.85k arXiv Page

Submitted by

taesiri

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

QGF is an RL algorithm that improves policies at test time by using a value gradient to guide a pre-trained flow policy, avoiding training-time instability while maintaining competitive performance.

7 authors

· Published on Jun 9, 2026

3

GitHub 26 arXiv Page

Submitted by

taesiri

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

QGF is an RL algorithm that improves policies at test time by using a value gradient to guide a pre-trained flow policy, avoiding training-time instability while maintaining competitive performance.

7 authors

· Jun 9, 2026

3

GitHub 26 arXiv Page

Submitted by

Insta360-Research

UniSHARP: Universal Sharp Monocular View Synthesis

UniSHARP extends SHARP for universal monocular rendering across different camera systems by aligning images in an omnidirectional latent space through joint feature and Gaussian space alignment.

7 authors

· Published on Jun 5, 2026

14

GitHub 56 arXiv Page

Submitted by

Insta360-Research

UniSHARP: Universal Sharp Monocular View Synthesis

UniSHARP extends SHARP for universal monocular rendering across different camera systems by aligning images in an omnidirectional latent space through joint feature and Gaussian space alignment.

7 authors

· Jun 5, 2026