Muse Glimmer 30B Architecture Notes
Analysis of Meta's new open-weight Muse Glimmer 30B LLM, focusing on its architecture, KV-cache efficiency, and comparison to similar models.
Analysis of Meta's new open-weight Muse Glimmer 30B LLM, focusing on its architecture, KV-cache efficiency, and comparison to similar models.
A visual guide to attention variants in modern LLMs, covering MHA, GQA, MLA, sparse attention, and hybrid architectures.
Analysis of OpenAI's new gpt-oss models, comparing architectural improvements from GPT-2 and examining optimizations like MXFP4 and Mixture-of-Experts.