GLM-5.3-Flash Architecture Notes
Analysis of GLM-5.3-Flash architecture, a new LLM with hybrid attention and sparse MoE, compared to previous versions.
Analysis of GLM-5.3-Flash architecture, a new LLM with hybrid attention and sparse MoE, compared to previous versions.
Analysis of Meta's new open-weight Muse Glimmer 30B LLM, focusing on its architecture, KV-cache efficiency, and comparison to similar models.
Analysis of Nemotron 3 Ultra LLM with Latent MoE scaling, hybrid Mamba-Transformer architecture, and performance benchmarks.
New LLM Architecture Gallery diff tool lets you compare model architecture stacks side by side, highlighting key differences.
New LLM Architecture Gallery diff tool lets you compare two models side-by-side, highlighting key architectural differences.
A gallery showcasing and comparing architecture diagrams and technical details of recent open-weight Large Language Models (LLMs).
A gallery showcasing architecture diagrams and technical details for recent open-weight Large Language Models (LLMs).
Explains Recursive Language Models (RLMs), which are LLMs that call themselves to break complex tasks into structured, reusable steps.
A technical analysis of DeepSeek V3.2's architecture, sparse attention, and reinforcement learning updates, comparing it to other flagship AI models.
A detailed comparison of architectural developments in major large language models (LLMs) released in 2024-2025, focusing on structural changes beyond benchmarks.
A technical comparison of architectural changes in major Large Language Models (LLMs) from 2024-2025, focusing on structural innovations beyond benchmarks.