Controlling Reasoning Effort in LLMs
Explains how LLMs learn low, medium, and high-effort reasoning modes, covering training and inference techniques for controlling reasoning effort.
Explains how LLMs learn low, medium, and high-effort reasoning modes, covering training and inference techniques for controlling reasoning effort.
Explains how LLMs learn low, medium, and high-effort reasoning modes, covering training techniques and effort settings.
Announcement of the release of 'Build a Reasoning Model (From Scratch)', a book on implementing modern reasoning techniques for AI.
Announcement of the release of 'Build a Reasoning Model (From Scratch)', a book on implementing modern AI reasoning techniques.
A tribute to Dimitri Bertsekas, a pioneer in optimization and reinforcement learning, reflecting on his contributions and legacy.
Analysis of VibeThinker-3B, a small AI model achieving strong coding and reasoning performance through advanced post-training techniques.
Analysis of VibeThinker-3B, a small AI model achieving strong coding and reasoning results through post-training on Qwen2.5-Coder-3B.
A curated list of LLM research papers from January to May 2026, covering architecture, reasoning, RL, agents, and more.
A curated list of LLM research papers from January to May 2026, covering architecture, reasoning, agents, and more.
Analysis of Kimi K2.5, Cursor Composer 2, and Chroma Context-1 reports on training agentic models with reinforcement learning.
Explores the contrasting mindsets of AI and control theory, focusing on the limits and practical challenges of optimal control in sequential decision-making.
Introduces Nova, an AI co-designer for board game creation that learns a designer's preferences and remembers past decisions through conversation.
A professor reflects on the intersection of machine learning and control theory, discussing the Learning for Dynamics and Control (L4DC) conference and the need for a merged perspective.
OpenAI researchers propose 'confessions' as a method to improve AI honesty by training models to self-report misbehavior in reinforcement learning.
OpenAI researchers propose 'confessions' as a method to improve AI honesty by training models to self-report misbehavior in reinforcement learning.
A 2025 year-in-review of Large Language Models, covering major developments in reasoning, architecture, costs, and predictions for 2026.
A review of key paradigm shifts in Large Language Models (LLMs) in 2025, focusing on RLVR training and new conceptual models of AI intelligence.
Explores the tension between optimization and systems-level thinking in AI-driven scientific discovery and computational ethics.
A developer reflects on AI agent architectures, context management, and the industry's overemphasis on model development vs. building applications.
A critique of Reformist RL's inefficiency and a proposal for more effective alternatives in reinforcement learning.