Submit Blog

Sign up Sign in

Philipp Schmid • 1/17/2025

Bite: How Deepseek R1 was trained

Read Original

This technical article details how DeepSeek AI trained its DeepSeek-R1 model, an open model rivaling OpenAI's o1 in reasoning. It explains the Group Relative Policy Optimization (GRPO) algorithm, a reinforcement learning method that eliminates the need for a value function and uses group-based scoring. The summary covers the multi-stage training process, rule-based rewards, and the resulting performance gains in mathematical and coding tasks.

0 comments

#Reinforcement Learning #Deepseek #LLM Training

#Reinforcement Learning #Deepseek #LLM Training

Bite: How Deepseek R1 was trained

Comments

No comments yet

Be the first to share your thoughts!

Browser Extension

Get instant access to AllDevBlogs from your browser

Top of the Week

1

Limit token usage in Microsoft Agent Framework

Jesse Liberty • 1 votes

2

How to Roll Back AI Agents: Incident Response, Circuit Breakers, and Recovery Patterns

Paul Bryant • 1 votes

3

Avoiding Reasoning Model Failures with Microsoft Foundry

Luke Murray • 1 votes

4

When Your AI Agent Lies: Silent LLM Fallbacks

Luke Murray • 1 votes

5

Adding a custom MCP server to Claude and ChatGPT

Simon Willison • 1 votes

6

Testing AI prompts and comparing models with promptfoo

Tim Deschryver • 1 votes

7

Mitchell Hashimoto • 1 votes