-
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Paper • 2605.20025 • Published • 89 -
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Paper • 2605.19769 • Published • 67 -
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 37 -
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
Paper • 2605.13941 • Published • 15
Collections
Discover the best community collections!
Collections including paper arxiv:2607.08716
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 39 -
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
Paper • 2603.09229 • Published • 85 -
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
Paper • 2603.11076 • Published • 5 -
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
Paper • 2603.21065 • Published • 80
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 85 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20
-
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Paper • 2605.20025 • Published • 89 -
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Paper • 2605.19769 • Published • 67 -
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 37 -
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
Paper • 2605.13941 • Published • 15
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 85 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 39 -
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
Paper • 2603.09229 • Published • 85 -
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
Paper • 2603.11076 • Published • 5 -
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
Paper • 2603.21065 • Published • 80