Skip to architectures
Live architecture index
Lab Notebook · Vol. 03

Inside the model stack.

Compare decoder types, attention patterns, scale and the defining choices behind today’s open-weight models.

Figures and base specifications: Sebastian Raschka’s Architecture Gallery and comparison notes.

Browse the architectures Search by model, organization, scale or design choice.
Search
—
No models match your search.
Release Timeline Every major LLM release — open-weight vs closed, dense vs MoE.
— models — open-weight — closed — MoE
Closed Models — what we (think we) know Reverse-engineered notes on the architectures of frontier proprietary models.
GPT-4
CLOSED
OpenAI · 2023-03
Rumored MoE ~1.8T total params (8×220B experts). First multimodal GPT-4 variant, launched with Vision. Architecture never officially disclosed.
MoE (rumored) Multimodal ~1.8T params
GPT-4o
CLOSED
OpenAI · 2024-05
Dense or small MoE. Natively multimodal — text, vision, audio in a single end-to-end model. Faster inference than GPT-4. Exact architecture undisclosed.
MoE (possible) Natively Multimodal 128K context
Claude 3 Opus
CLOSED
Anthropic · 2024-03
Constitutional AI training. Exact architecture undisclosed, likely dense decoder. 200K token context window. Top benchmarks on release, surpassed GPT-4.
Likely Dense 200K context Constitutional AI
Gemini Ultra
CLOSED
Google · 2024-02
Confirmed MoE. Multimodal from ground up — handles text, image, audio, video natively. 1M token context window. Backbone of the Gemini 1.5 family.
MoE (confirmed) Multimodal 1M context
Gemini 2.5 Pro
CLOSED
Google · 2025-06
Confirmed MoE. Thinking mode (extended reasoning). #1 on most benchmarks mid-2025. Deep Research and agentic task support built-in.
MoE (confirmed) Thinking mode #1 benchmarks
Grok-3
CLOSED
xAI · 2025-02
Confirmed MoE. Trained on X (Twitter) data at massive scale. 128K context. Think mode for extended reasoning. Competes directly with GPT-4o and Claude Opus 4.
MoE (confirmed) Think mode 128K context
GPT-5
CLOSED
OpenAI · 2025-08
Official August 2025 launch of a unified GPT-5 system: a router sends each request to a faster model or a deeper thinking model. Published details stop at that routing/thinking split.
Unified system Thinking Router
Use this dataset with an AI agent Markdown endpoints and ready-to-use prompts
Machine-readable endpoints
https://llmgram.app/llm-architectures.md
https://llmgram.app/ai-signal.md
Prompt examples
Fetch https://llmgram.app/llm-architectures.md and summarize the latest MoE architectures
Fetch https://llmgram.app/ai-signal.md and give me today's most important AI news