🍊 Latent Atlas 🍉
Search
搜索
暗色模式
亮色模式
探索
标签: reinforcement-learning
阅读模式
Hide sidebars
Show sidebars
此标签下有9条笔记。
2026年7月15日
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
source
paper
rlvr
entropy
grpo
ppo
reinforcement-learning
reasoning
coding
2026年7月14日
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
source
paper
agent
reinforcement-learning
post-training
grpo
asynchronous-rl
2026年6月30日
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
source
paper
code-agent
swe-agent
agentless
reinforcement-learning
mid-training
Params
2026年6月30日
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
source
paper
agent
world-model
mid-training
reinforcement-learning
2026年6月30日
Qwen3-Coder-Next Technical Report
source
paper
qwen
code-agent
mid-training
reinforcement-learning
tool-use
moe
2026年6月01日
KLong: Training LLM Agent for Extremely Long-horizon Tasks
source
paper
agents
long-horizon
reinforcement-learning
sft
evaluation
2026年5月29日
Proximal Policy Optimization Algorithms
source
paper
reinforcement-learning
ppo
2026年5月28日
RLP: Reinforcement as a Pretraining Objective
source
paper
pretraining
reinforcement-learning
reasoning
chain-of-thought
2026年5月28日
Reinforcement Pretraining
pretraining
reinforcement-learning
reasoning