Behaviorism and AI: How Rewards Shape Model Behavior
Explores how behaviorism and reinforcement learning shape AI behavior, and why rewards don't guarantee intended outcomes in enterprise AI systems.
Explores how behaviorism and reinforcement learning shape AI behavior, and why rewards don't guarantee intended outcomes in enterprise AI systems.
Explains how LLMs learn low, medium, and high-effort reasoning modes, covering training and inference techniques for controlling reasoning effort.
Explains how LLMs learn low, medium, and high-effort reasoning modes, covering training techniques and effort settings.
Announcement of the release of 'Build a Reasoning Model (From Scratch)', a book on implementing modern reasoning techniques for AI.
Announcement of the release of 'Build a Reasoning Model (From Scratch)', a book on implementing modern AI reasoning techniques.
A tribute to Dimitri Bertsekas, a pioneer in optimization and reinforcement learning, reflecting on his contributions and legacy.
Analysis of VibeThinker-3B, a small AI model achieving strong coding and reasoning performance through advanced post-training techniques.
Analysis of VibeThinker-3B, a small AI model achieving strong coding and reasoning results through post-training on Qwen2.5-Coder-3B.
A curated list of LLM research papers from January to May 2026, covering architecture, reasoning, RL, agents, and more.
A curated list of LLM research papers from January to May 2026, covering architecture, reasoning, agents, and more.
Analysis of Kimi K2.5, Cursor Composer 2, and Chroma Context-1 reports on training agentic models with reinforcement learning.
Explores the contrasting mindsets of AI and control theory, focusing on the limits and practical challenges of optimal control in sequential decision-making.
Introduces Nova, an AI co-designer for board game creation that learns a designer's preferences and remembers past decisions through conversation.
A professor reflects on the intersection of machine learning and control theory, discussing the Learning for Dynamics and Control (L4DC) conference and the need for a merged perspective.
OpenAI researchers propose 'confessions' as a method to improve AI honesty by training models to self-report misbehavior in reinforcement learning.
OpenAI researchers propose 'confessions' as a method to improve AI honesty by training models to self-report misbehavior in reinforcement learning.
A 2025 year-in-review of Large Language Models, covering major developments in reasoning, architecture, costs, and predictions for 2026.
A review of key paradigm shifts in Large Language Models (LLMs) in 2025, focusing on RLVR training and new conceptual models of AI intelligence.
Explores the tension between optimization and systems-level thinking in AI-driven scientific discovery and computational ethics.
A developer reflects on AI agent architectures, context management, and the industry's overemphasis on model development vs. building applications.