OpenAI Astra and Looped Transformers
Read OriginalA lot of hype around OpenAI’s Astra model being a “recurrent depth or looped transformer”. Let’s debunk this a bit. About 2 months ago, I shared the architecture details of Nanbeige, for example, where “Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters.” Yes, that’s it. The looped transformer
Comments
No comments yet
Be the first to share your thoughts!
Browser Extension
Get instant access to AllDevBlogs from your browser
Top of the Week
No top articles yet