Posts tagged “reinforcement-learning”

10 posts found

Self-Play Explained: Opponent Pools, Verifiers, and Honest Progress
ai-tutorialstutorialself-play

Self-Play Explained: Opponent Pools, Verifiers, and Honest Progress

A practical framework for building self-play curricula without mistaking reward hacking, forgetting, or correlated self-grading for genuine capability gains.

Sep 4, 2026•23 min read
The Agents Built an Institution
ai-tutorialstutorialai-news

The Agents Built an Institution

Highlights of AI News for August 24 - 30 2026

Aug 31, 2026•27 min read
Planning Under Uncertainty: From MPC to Active Inference
ai-tutorialstutorialmodel-predictive-control

Planning Under Uncertainty: From MPC to Active Inference

A practitioner's tour of how an agent turns a doubtful model into a decision: model-predictive control and why replanning every step is the robustness, CEM and random shooting as the planners world-model papers actually run, planning through a posterior instead of a point estimate, and expected free energy with the sign convention stated the right way round.

Aug 16, 2026•10 min read
The Full Loop: World Models That Act on What They Don't Know
ai-tutorialstutorialworld-models

The Full Loop: World Models That Act on What They Don't Know

Calibrated confidence is a permission slip — this week we spend it. From PILCO's 17.5 seconds of robot experience to V-JEPA 2 planning zero-shot on a Franka arm, we trace how uncertainty becomes action. Then we look at the 2026 result that breaks the arc's own thesis: a world model can be locally well-calibrated and globally, confidently wrong.

Aug 15, 2026•13 min read
When AI Automates AI Research: Benchmarks, Risks, and Early Results
ai-tutorialstutorialmachine-learning

When AI Automates AI Research: Benchmarks, Risks, and Early Results

Trends in ICLR 2026 RSI workshop - Self-Evolving Agents

Mar 21, 2026•8 min read
Self-Evolving AI Agents: How Models Learn to Improve Without Human Data
ai-tutorialstutorialmachine-learning

Self-Evolving AI Agents: How Models Learn to Improve Without Human Data

Trends in ICLR 2026 RSI workshop - Self-Evolving Agents

Mar 19, 2026•6 min read
From Self-Play to Self-Research: The ICLR 2026 RSI Workshop and the State of Self-Improving AI
ai-tutorialstutorialmachine-learning

From Self-Play to Self-Research: The ICLR 2026 RSI Workshop and the State of Self-Improving AI

Highlights of Trends in ICLR 2026 RSI workshop

Mar 18, 2026•14 min read
Self-Training Loops for LLMs: STaR and the Self-Instruct Family
ai-tutorialstutorialmachine-learning

Self-Training Loops for LLMs: STaR and the Self-Instruct Family

How to filter for correctness, not just fluency

Mar 13, 2026•11 min read
Self-Play in AI: From Board Games to Language Models
ai-tutorialstutorialmachine-learning

Self-Play in AI: From Board Games to Language Models

With LLMs, when self-play works and when it doesn't.

Mar 11, 2026•7 min read
Self-Play for LLM Self-Evolution: From Brittle Dynamics to Sustained Improvement
ai-tutorialstutorialmachine-learning

Self-Play for LLM Self-Evolution: From Brittle Dynamics to Sustained Improvement

AI Tutorials - The trend of Self-Play

Mar 10, 2026•9 min read