# Model Estimator research snapshot

Original collection checked **15 September 2026**; four local-model additions checked **16 September 2026**. This is a curated collection for estimating memory fit and generation speed, not a leaderboard, a promise of model quality, or a list of every downloadable model. The supplied rough estimator was used as a research lead; its blended model names, older generations and generic quantization multipliers were not adopted as facts.

## Selection

Fourteen choices trace a path from a modest local card to system-scale inference. Three current Gemma sizes serve different local needs. Qwen3.8 uses its actual 27B, Flash-Next and 2.4T-A95B releases. GLM uses 5.3 and the distinct 5.3-Flash architecture. DeepSeek uses V4-Flash-0731, which supersedes its preview. Kimi K3 supplies a current frontier comparison. The 16 September additions are Qwen3.6 35B A3B, Muse Glimmer 30B, Nemotron 3 Nano 4B and Nemotron 3.5 Lightning 30B A3B. The user explicitly selected Qwen3.6’s distinct 35B / 3B-active architecture as a practical, fast local coding option alongside Qwen3.8’s dense 27B and larger expert choices. It is retained for that different size and workload, not to fill the collection with older generations.

The collection provides **39 published format choices** within those fourteen models. Each model has three practical steps except Muse Glimmer, which uses Meta’s two official tuned GGUF builds, and Kimi, whose native mixed-precision checkpoint remains its only curated format. Google’s QAT releases are the Gemma defaults. Qwen3.8 27B starts with the publisher-recommended Q4_K_M build. Qwen3.6 35B A3B is the initial page selection and starts with UD-Q4_K_M; its smaller Q3 option makes the 24 GB class more approachable. Muse Glimmer starts with Meta’s K-Quant 17GB format, recommended by its publisher for 24 GB hardware, with K-Quant Dynamic for the 32 GB class. Nemotron Nano starts with UD-Q4_K_XL and Lightning with UD-Q4_K_M. Larger GGUF models start with compact IQ4_XS, with a smaller option and a larger or less compressed option nearby. These are editorial starting points, not measurements of popularity or a guarantee of identical quality across quantization methods. More compression can change answer quality; larger files do not establish superiority on every task.

### Published weight allocations

Sizes below are **decimal GB**, calculated by summing the selected files’ byte sizes from the Hugging Face file metadata. The original 15 September entries round to six decimal places in `models.js`; the four 16 September additions retain the exact integer-byte totals divided by 1,000,000,000. Links are pinned to the publisher revision checked. The UI converts these allocations to **GiB** for comparison with hardware memory. Each allocation includes its selected vision/input projector where present, even for text-only planning. Separately downloaded optional draft/MTP modules are excluded. Lightning’s complete GGUF allocation includes all bundled tensors, including any MTP tensors; nothing is subtracted from its published file size. No model receives a speculative-decoding speed boost. Context, runtime and extra image/audio-processing buffers are additional. **Bold** identifies the initial default.

| Model | Smaller / first choice | Middle choice | Larger choice |
| --- | --- | --- | --- |
| Gemma 4 E4B | [Q3_K_M](https://huggingface.co/unsloth/gemma-4-E4B-it-GGUF/blob/bfc15c382204943c3a8fff0c750b94ae2364d7a3/gemma-4-E4B-it-Q3_K_M.gguf) · 5.049 GB | **[QAT Q4_0](https://huggingface.co/google/gemma-4-E4B-it-qat-q4_0-gguf/blob/4b4a2c1d584be7264f87aac328a1bc739ce81b6c/gemma-4-E4B_q4_0-it.gguf) · 6.146 GB** | [Q8_0](https://huggingface.co/unsloth/gemma-4-E4B-it-GGUF/blob/bfc15c382204943c3a8fff0c750b94ae2364d7a3/gemma-4-E4B-it-Q8_0.gguf) · 9.183 GB |
| Gemma 4 12B | [Q3_K_M](https://huggingface.co/unsloth/gemma-4-12b-it-GGUF/blob/fc034cfff751157913579611efad8462ac1be606/gemma-4-12b-it-Q3_K_M.gguf) · 5.869 GB | **[QAT Q4_0](https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/blob/29d097773436b69ff9feafd636ab4cf873786537/gemma-4-12b-it-qat-q4_0.gguf) · 7.151 GB** | [Q8_0](https://huggingface.co/unsloth/gemma-4-12b-it-GGUF/blob/fc034cfff751157913579611efad8462ac1be606/gemma-4-12b-it-Q8_0.gguf) · 12.845 GB |
| Gemma 4 26B A4B | [UD-Q3_K_XL](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF/blob/c099eb48e663fd284577b04978a94ffccb261841/gemma-4-26B-A4B-it-UD-Q3_K_XL.gguf) · 14.100 GB | **[QAT Q4_0](https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf/blob/d1c082be9cf3c8a514acf63b8761f4b41935842e/gemma-4-26B_q4_0-it.gguf) · 15.634 GB** | [Q8_0](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF/blob/c099eb48e663fd284577b04978a94ffccb261841/gemma-4-26B-A4B-it-Q8_0.gguf) · 28.053 GB |
| Qwen3.6 35B A3B | [UD-Q3_K_XL](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/blob/a483e9e6cbd595906af30beda3187c2663a1118c/Qwen3.6-35B-A3B-UD-Q3_K_XL.gguf) · 17.744795328 GB | **[UD-Q4_K_M](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/blob/a483e9e6cbd595906af30beda3187c2663a1118c/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf) · 23.033812672 GB** | [Q8_0](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/blob/a483e9e6cbd595906af30beda3187c2663a1118c/Qwen3.6-35B-A3B-Q8_0.gguf) · 37.802424000 GB |
| Muse Glimmer 30B | **[K-Quant 17GB](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF/blob/70bf1b61ac09f91b24d39038091b41c582bc5d7a/Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf) · 18.157012832 GB** | [K-Quant Dynamic](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF/blob/70bf1b61ac09f91b24d39038091b41c582bc5d7a/Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf) · 21.054289760 GB | — |
| Nemotron 3 Nano 4B | [UD-Q3_K_XL](https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Nano-4B-GGUF/blob/8e81be55c5aa3d63bb82b6ceec62d50805d9e1bb/NVIDIA-Nemotron-3-Nano-4B-UD-Q3_K_XL.gguf) · 2.682625952 GB | **[UD-Q4_K_XL](https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Nano-4B-GGUF/blob/8e81be55c5aa3d63bb82b6ceec62d50805d9e1bb/NVIDIA-Nemotron-3-Nano-4B-UD-Q4_K_XL.gguf) · 3.133118624 GB** | [Q8_0](https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Nano-4B-GGUF/blob/8e81be55c5aa3d63bb82b6ceec62d50805d9e1bb/NVIDIA-Nemotron-3-Nano-4B-Q8_0.gguf) · 4.233679008 GB |
| Nemotron 3.5 Lightning 30B A3B | [UD-Q3_K_XL](https://huggingface.co/unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF/blob/f2d3fe3694501008786e81e5f20360cbf715496a/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-UD-Q3_K_XL.gguf) · 21.235202112 GB | **[UD-Q4_K_M](https://huggingface.co/unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF/blob/f2d3fe3694501008786e81e5f20360cbf715496a/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-UD-Q4_K_M.gguf) · 25.266255936 GB** | [Q8_0](https://huggingface.co/unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF/blob/f2d3fe3694501008786e81e5f20360cbf715496a/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q8_0.gguf) · 35.004643392 GB |
| Qwen3.8 27B | **[UD-Q4_K_M](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/blob/4ca720788d1e01f1bff70c033e0d0028fd02e502/Qwen3.8-27B-UD-Q4_K_M.gguf) · 17.392 GB** | [UD-Q5_K_M](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/blob/4ca720788d1e01f1bff70c033e0d0028fd02e502/Qwen3.8-27B-UD-Q5_K_M.gguf) · 20.699 GB | [Q8_0](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/blob/4ca720788d1e01f1bff70c033e0d0028fd02e502/Qwen3.8-27B-Q8_0.gguf) · 29.975 GB |
| Qwen3.8 Flash-Next | [UD-Q3_K_XL](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/38bb39ee97821de2c9009abb7e93950eec396e66/UD-Q3_K_XL) · 90.890 GB | **[UD-IQ4_XS](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/38bb39ee97821de2c9009abb7e93950eec396e66/UD-IQ4_XS) · 94.587 GB** | [UD-Q4_K_XL](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/38bb39ee97821de2c9009abb7e93950eec396e66/UD-Q4_K_XL) · 112.239 GB |
| DeepSeek V4 Flash 0731 | [UD-IQ2_M](https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF/tree/fbbb5b93fb787c21338159b0af3318bb3f4d9768/UD-IQ2_M) · 90.927 GB | **[UD-IQ4_XS](https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF/tree/fbbb5b93fb787c21338159b0af3318bb3f4d9768/UD-IQ4_XS) · 136.662 GB** | [UD-Q8_K_XL](https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF/tree/fbbb5b93fb787c21338159b0af3318bb3f4d9768/UD-Q8_K_XL) · 161.870 GB |
| GLM-5.3 Flash | [UD-Q3_K_XL](https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF/tree/621d456e93e926e4b52f85cff5f634358c1828f9/UD-Q3_K_XL) · 148.664 GB | **[UD-IQ4_XS](https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF/tree/621d456e93e926e4b52f85cff5f634358c1828f9/UD-IQ4_XS) · 157.950 GB** | [UD-Q4_K_XL](https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF/tree/621d456e93e926e4b52f85cff5f634358c1828f9/UD-Q4_K_XL) · 200.835 GB |
| GLM-5.3 | [UD-IQ3_XXS](https://huggingface.co/unsloth/GLM-5.3-GGUF/tree/346b3591c7f28d1a23716f97a065ecf12ec14771/UD-IQ3_XXS) · 281.688 GB | **[UD-IQ4_XS](https://huggingface.co/unsloth/GLM-5.3-GGUF/tree/346b3591c7f28d1a23716f97a065ecf12ec14771/UD-IQ4_XS) · 365.313 GB** | [UD-Q4_K_XL](https://huggingface.co/unsloth/GLM-5.3-GGUF/tree/346b3591c7f28d1a23716f97a065ecf12ec14771/UD-Q4_K_XL) · 467.289 GB |
| Qwen3.8 2.4T A95B | [UD-IQ2_XS](https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF/tree/567d3e6ac26c5474b18311e619c04350fb9a5556/UD-IQ2_XS) · 730.653 GB | [UD-IQ3_XXS](https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF/tree/567d3e6ac26c5474b18311e619c04350fb9a5556/UD-IQ3_XXS) · 955.535 GB | **[UD-IQ4_XS](https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF/tree/567d3e6ac26c5474b18311e619c04350fb9a5556/UD-IQ4_XS) · 1,310.876 GB** |
| Kimi K3 | **[Native MXFP4](https://huggingface.co/moonshotai/Kimi-K3/tree/f831ab66814297da540d832a5235f8e904f29d06) · 1,560.936 GB** | — | — |

Gemma QAT projectors come from the same pinned Google repositories as their weight files. Other Gemma, Qwen27B, Flash-Next and GLM Flash variants use the `mmproj-F16.gguf` file at the same pinned Unsloth revision as their selected quant. Their projector allocations are respectively 0.990373, 0.175116, 1.193059, 0.927607, 0.904004 and 1.128047 decimal GB. Google’s corresponding QAT projectors are 0.991552, 0.175116 and 1.194828 GB. These are artifact-storage sizes, not measured runtime allocations.

The selected DeepSeek GGUFs contain the base model; their separately published optional DSpark file is excluded. The official native checkpoint includes an attached speculative module, so its old 167 GB repository total is not reused for these GGUF formats. Unsloth describes `UD-Q8_K_XL` as preserving the native mixed-precision release, rather than quantizing every parameter uniformly to eight bits. [Publisher card and DSpark files](https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF/tree/fbbb5b93fb787c21338159b0af3318bb3f4d9768)

### Exact bytes for the 16 September additions

The linked formats above use these complete file allocations. No shards or required projector files are omitted.

| Model / format | Main GGUF bytes | Projector bytes | Total selected bytes |
| --- | ---: | ---: | ---: |
| Qwen3.6 / UD-Q3_K_XL | 16,845,511,648 | 899,283,680 | 17,744,795,328 |
| Qwen3.6 / UD-Q4_K_M | 22,134,528,992 | 899,283,680 | 23,033,812,672 |
| Qwen3.6 / Q8_0 | 36,903,140,320 | 899,283,680 | 37,802,424,000 |
| Muse / K-Quant 17GB | 16,756,683,904 | 1,400,328,928 | 18,157,012,832 |
| Muse / K-Quant Dynamic | 19,653,960,832 | 1,400,328,928 | 21,054,289,760 |
| Nemotron Nano / UD-Q3_K_XL | 2,682,625,952 | — | 2,682,625,952 |
| Nemotron Nano / UD-Q4_K_XL | 3,133,118,624 | — | 3,133,118,624 |
| Nemotron Nano / Q8_0 | 4,233,679,008 | — | 4,233,679,008 |
| Nemotron Lightning / UD-Q3_K_XL | 21,235,202,112 | — | 21,235,202,112 |
| Nemotron Lightning / UD-Q4_K_M | 25,266,255,936 | — | 25,266,255,936 |
| Nemotron Lightning / Q8_0 | 35,004,643,392 | — | 35,004,643,392 |

Qwen3.6 includes [`mmproj-F16.gguf`](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/blob/a483e9e6cbd595906af30beda3187c2663a1118c/mmproj-F16.gguf). Both Muse formats include Meta’s [`mmproj-Muse-Glimmer-30B-Q4_K_M.gguf`](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF/blob/70bf1b61ac09f91b24d39038091b41c582bc5d7a/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf). Meta’s [GGUF guidance](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF) recommends K-Quant 17GB as the starting point and documents the optional 1,631,208,128-byte DFlash drafter separately. That drafter is excluded from this baseline, along with its speedup. The two Nemotron models require no vision projector. Lightning uses the whole published GGUF, including any bundled MTP tensors; the speed estimate uses NVIDIA’s nominal 30B total / 3B active ratio despite publisher metadata reporting 32.91B stored parameters. This ratio is a traffic proxy, not a measured count of bytes read per token.

### Qwen Flash-Next and the 96 GB class

The former fixed 113 GB entry represented the larger `UD-Q4_K_XL` recipe. Compact four-bit `UD-IQ4_XS` contains **93,682,584,224 bytes** of GGUF shards, plus its **904,004,000-byte** projector: **94.586588 decimal GB, or 88.090625 GiB**, before context and working memory. `UD-Q3_K_XL` uses 90.890358 GB including the projector, and `UD-Q4_K_XL` uses 112.238659 GB. All three retain the 51B n-gram tables. Switching formats is therefore a real artifact change; “4-bit” alone does not specify a unique allocation. [Pinned publisher files](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/38bb39ee97821de2c9009abb7e93950eec396e66)

Unsloth’s [Flash-Next guide](https://unsloth.ai/docs/models/qwen3.8-next) lists 96–114 GB of total RAM plus VRAM, or unified memory, for four-bit choices. That is guidance across formats and placements, not a measured all-GPU allocation for every 96 GB card. The estimator retains context and runtime headroom, identifies a close fit, and does not silently assume n-gram offload. Host-table offload and MTP need explicit runtime configuration and have different allocation requirements.

### Memory units and usable capacity

A decimal **GB is 1,000,000,000 bytes**; a binary **GiB is 1,073,741,824 bytes**. The Ladder’s product names retain their advertised 32 GB, 96 GB and other capacity classes. The estimator compares allocations in GiB, treating these hardware memory classes as nominal binary capacities where no measured total is stored. It does not reinterpret a 94.6 decimal GB download as 94.6 GiB of VRAM.

NVIDIA’s [official installation guide](https://docs.nvidia.com/vgpu/19.0/grid-vgpu-user-guide/installing-configuring-grid-vgpu.html) includes an RTX PRO 6000 `nvidia-smi` example with **97,887 MiB** total framebuffer memory, equal to **95.592773 GiB** or **102.641959 decimal GB**. The estimator uses that documented total for its RTX PRO 6000 planning example. This is an example configuration, not a promise of byte-identical available memory on every card. NVIDIA’s [memory reporting documentation](https://docs.nvidia.com/deploy/nvidia-smi/index.html#fb-memory-usage) explains that ECC settings and driver reservations can reduce available framebuffer memory. Runtime allowance, system reserve and close-fit status remain necessary.

AMD’s [GPU specification reference](https://rocm.docs.amd.com/en/docs-7.14.1/reference/gpu-specs.html) explicitly labels VRAM in GiB, including MI210, MI100, MI50 and the supported Radeon memory classes. Other inherited Ladder rows use the stated nominal-binary planning convention rather than claiming a measured available allocation.

The 32 GiB and 96 GiB shortcuts choose from the curated published formats that fit the corresponding usable-memory budget at the current context. They are memory recommendations only; the user can adjust the format and hardware independently. Unified-memory devices retain a separate system reserve, so 96 GB installed unified memory and a dedicated 96 GB accelerator are not interchangeable budgets.

## Release identity and dates

- **Gemma 4 E4B and 26B A4B: 2 April 2026.** Google's [launch](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/) identifies these as distinct original Gemma4 sizes. **12B: 3 June 2026**, from its [separate launch](https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/). The current [official model card](https://huggingface.co/google/gemma-4-12B-it) distinguishes effective, total and active parameters and their architectures.
- **Qwen3.8 27B: 14 August; 2.4T-A95B: 12 August.** Dates come from the [official Qwen3.8 repository](https://github.com/QwenLM/Qwen3.8). Their model cards are [27B](https://huggingface.co/Qwen/Qwen3.8-27B) and [2.4T](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B). The user's 32GB example corresponds to a hardware-memory class, not a verified model named Qwen3.8-32B.
- **Qwen3.8 Flash-Next: 26 August.** The [official card](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) and [initial weights history](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/commits/main) establish the release. The history displayed 20 days before this snapshot; the stored day is inferred from that relative timestamp. Its [technical report](https://arxiv.org/abs/2608.30320), published 31 August, explains the separate 51B host-offload-friendly n-gram tables. The hosted Qwen3.8-Flash service and the open-weight Flash-Next artifact have different names.
- **GLM-5.3 Flash: 26 August.** The [official card](https://huggingface.co/zai-org/GLM-5.3-Flash), [official repository](https://github.com/zai-org/GLM-5), [weight upload history](https://huggingface.co/zai-org/GLM-5.3-Flash/commits/main) and Cloudflare's [dated serving release](https://developers.cloudflare.com/changelog/post/2026-08-26-glm-5.3-flash-workers-ai/) establish its availability. It is a different 320B base architecture.
- **GLM-5.3 open weights: 28 August.** Its [initial checkpoint commit](https://huggingface.co/zai-org/GLM-5.3/commits/main) is named `Initial commit 0828`; the [official repository](https://github.com/zai-org/GLM-5) identifies 744B total and 40B active. This date describes downloadable weights, not the earlier API announcement.
- **DeepSeek V4 Flash 0731: 31 July.** The [official card](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731) explicitly identifies this as the production release replacing the preview; [checkpoint history](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/commits/main) records its publication. The [V4 model card](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash) supplies base architecture counts.
- **Kimi K3 open weights: 27 July.** The [initial checkpoint history](https://huggingface.co/moonshotai/Kimi-K3/commits/main) gives the date; its [model card](https://huggingface.co/moonshotai/Kimi-K3) states 2.8T total, 104B active, 69 KDA plus 24 gated MLA layers, and native mixed MXFP4 quantization.

- **Qwen3.6 35B A3B: 15 April 2026.** The [initial public-release commit](https://huggingface.co/Qwen/Qwen3.6-35B-A3B/commit/7da1103448ba36029c34ce1a9a741dfe93ee0c50) establishes the date; the [official card](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) identifies 35B total and 3B active parameters, 40 hybrid DeltaNet/MoE layers and 262,144 native context. The later Qwen3.8 entries do not provide this same small-active-parameter size class.
- **Muse Glimmer 30B GGUF weights: 10 August 2026.** Meta’s [GGUF upload](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF/commit/b1f3e6ec2209678b3f29525bb9646286866f1675) establishes the selected formats’ availability; its [official model card](https://huggingface.co/meta-models/Muse-Glimmer-30B) identifies approximately 29.6B parameters including vision, dense local/global attention and 131,072 native context. It targets local agents, coding and multimodal tasks. The generic estimator includes the vision weights but does not reproduce Meta’s DFlash benchmark configuration.
- **Nemotron 3 Nano 4B: 16 March 2026.** NVIDIA’s [model card](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) identifies a 3.97B-parameter dense hybrid of Mamba-2, MLP and attention, with 262,144-token context. This is the small 4B release, distinct from the older 30B-A3B Nano architecture.
- **Nemotron 3.5 Lightning 30B A3B: 11 August 2026.** NVIDIA’s [model card](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16) identifies the larger requested Lightning release: a hybrid of Mamba-2, MoE and attention with a nominal 30B / 3B-active footprint. NVIDIA supports up to 1,048,576 tokens in a compatible runtime, while its checkpoint config defaults to 262,144; higher contexts need explicit runtime configuration. The current estimator’s 8K/32K/128K controls stay below both limits.

Availability was checked against producer-controlled repositories. Community conversions are used as primary evidence of their own artifact sizes, not as independent evidence of base-model quality. No model was run locally to verify output quality or hardware speed.

## Cache and working-memory method

For one conversation, text-cache planning is `kvFixedGB + kvGBPer8k × contextTokens / 8192`, with a separate runtime/headroom allocation in the estimator engine. The stored cache fields use decimal GB; the engine converts them to GiB before comparing with hardware memory. A native maximum context is an architectural limit, not a promise that a selected card can allocate it. Extended-context overrides are excluded from Qwen limits.

For conventional attention, the base calculation is full layers × KV heads × head dimension × two K/V tensors × two bytes × tokens. Sliding-window caches are fixed at their window size. Hybrid recurrent states receive a fixed allowance based on FP32 matrix state plus a small convolution allowance. These are planning estimates for a compatible memory-efficient engine. Some runtimes preallocate larger caches, duplicate latent K/V, or allocate additional buffers. The numeric estimate cannot establish runtime support.

| Model | GB per 8,192 tokens | Fixed GB | Basis and limits |
| --- | ---: | ---: | --- |
| Gemma E4B | 0.235 | 0.04 | [Config](https://huggingface.co/google/gemma-4-E4B-it/blob/main/config.json): 7 full layers, 2 KV heads, 512 global head dimension; 35 sliding layers with 256 head dimension and 512-token windows. Shared-layer savings are ignored. |
| Gemma 12B | 0.135 | 0.34 | [Config](https://huggingface.co/google/gemma-4-12B-it/blob/main/config.json): 8 full layers, 1 global KV head, 512 dimension; 40 sliding layers, 8 heads, 256 dimension, 1,024-token windows. Conservative separate K/V allocation despite equal values. |
| Gemma 26B A4B | 0.168 | 0.21 | [Config](https://huggingface.co/google/gemma-4-26B-A4B-it/blob/main/config.json): 5 full layers, 2 global heads, 512 dimension; 25 sliding layers, 8 heads, 256 dimension, 1,024-token windows. Conservative separate K/V allocation. |
| Qwen3.6 35B A3B | 0.168 | 0.07 | [Config](https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/config.json): 10 full layers × 2 KV heads × 256 dimension gives 0.16777216 GB per 8K. The 30 recurrent layers × 32 value heads × 128 × 128 FP32 state plus convolution use about 0.06685 GB, rounded up. |
| Muse Glimmer 30B | 0.1091 | 0.082 | [Config](https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/config.json): 13 global layers × 2 KV heads × 128 dimension gives 0.109051904 GB per 8K; 39 sliding layers with 2,048-token windows use 0.081788928 GB fixed. |
| Nemotron 3 Nano 4B | 0.135 | 0.085 | [Config](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16/blob/main/config.json): 4 attention layers × 8 KV heads × 128 dimension gives 0.134217728 GB per 8K. Its 21 Mamba-2 layers × 96 heads × 80 dimension × 128 state × FP32 plus convolution use 0.084209664 GB fixed, rounded up. |
| Nemotron 3.5 Lightning 30B A3B | 0.051 | 0.05 | [Config](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/blob/main/config.json): 6 attention layers × 2 KV heads × 128 dimension gives 0.050331648 GB per 8K. Its 23 Mamba-2 layers × 64 heads × 64 dimension × 128 state × FP32 plus convolution use 0.049364992 GB fixed, rounded up. No MTP speculation is assumed. |
| Qwen27B | 0.537 | 0.16 | [Config](https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/config.json): 16 full layers × 4 KV heads × 256 dimension; 48 recurrent layers, 48 value heads × 128 × 128 state. |
| Qwen Flash-Next | 0.227 | 0.13 | [Config](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/config.json): 12 QSA layers × 2 KV heads × 256 dimension plus an uncompressed index-key allowance; 36 recurrent layers. Does not count sparse selection as removing stored history. |
| DeepSeek Flash 0731 | 0.85 | 0.10 | [Config](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/config.json): conservative uncompressed bound from 43 layers, one 512-dimensional KV head and an index allowance. Its 4×/128× compressed history and 128-token windows can substantially reduce actual cache. |
| GLM Flash | 0.10 | 0.17 | [Config](https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/config.json): 11 sparse layers × 512 latent channels, with pooled index allowance; 34 KDA layers, 64 heads × 128 × 128 FP32 state. Requires compressed MLA cache. |
| GLM full | 0.80 | 0.10 | [Config](https://huggingface.co/zai-org/GLM-5.3/blob/main/config.json): 78 × (512 latent + 64 RoPE) × two-byte cache, with shared indexers and an allowance. Requires compressed MLA cache. |
| Qwen2.4T | 0.772 | 0.70 | [Config](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/config.json): 23 full layers × 4 KV heads × 256; 69 recurrent layers with 128 value heads × 128 × 128 state. |
| Kimi K3 | 0.227 | 0.50 | [Config](https://huggingface.co/moonshotai/Kimi-K3/blob/main/config.json): 24 × (512 latent + 64 RoPE) × two-byte cache; 69 KDA layers × 96 heads × 128 × 128 FP32 state. Requires latent caching. |

All experts and embeddings stay resident in this scenario. A GPU exceeding the weight allocation is not automatically a supported deployment. Flash-Next's host-table offload, CPU expert offload and speculative decoding remain separate configurations; the estimator does not assume them.

## Common generation-speed calculation

Every model, including the four 16 September additions, uses the same single-conversation decode calculation. No architecture-specific benchmark multiplier, DFlash boost or MTP boost has been added. Model names, architecture labels and the presence of a benchmark do not enable or disable it. The calculation uses the selected published weight allocation, active and total parameter counts, the full planned context cache and memory bandwidth:

```text
activeWeightGB = selectedWeightsGB × activeB / totalB
bytesReadGB = activeWeightGB + cacheGB
perDeviceBandwidthGBs = aggregateBandwidthGBs / physicalDevices
estimatedTkPerSecond = perDeviceBandwidthGBs × 15–45% / bytesReadGB
```

Weight and cache GB in this formula are decimal bytes divided by 1,000,000,000, matching bandwidth in GB/s. The **15–45% efficiency interval is an engineering assumption**, not a measured benchmark, calibrated model profile or statistical confidence interval. It provides a broad allowance for implementation overhead, including kernel launches, quant unpacking, expert routing and shared-memory contention. Actual speed can fall outside the interval. Loading, prompt processing, batching and speculative decoding are excluded; generated reasoning tokens also take time before a final answer appears.

The mechanism follows NVIDIA's explanation of [decode memory traffic and quantization](https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/) and the Transformers team's [bandwidth divided by active-weight bytes example](https://huggingface.co/blog/moe-transformers). These sources support a first-order calculation; they do not establish a common numerical efficiency for all hardware and runtimes. Active-weight traffic is a proportional proxy: mixed precision, shared tensors, embeddings and input projectors mean artifact bytes are not distributed uniformly across active parameters.

For hardware with several physical devices, a model that fits on one device uses that device alone. A larger fitting model uses the same per-device bandwidth under a **serial layer-sharding assumption**, with the full model's per-token traffic passing through its layer sequence. Aggregate memory enables the layout but does not multiply one user's generation speed. The broad efficiency interval also includes generic sharding overhead; it does not model a specific interconnect. This follows the distinction between sequential layer pipelines and concurrent tensor splitting in [llama.cpp's multi-GPU documentation](https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md). Supported placement, kernel availability and exact allocatable memory still need validation on the intended system.

The context selector is treated as a filled context for the speed scenario. The calculation charges the full planned cache as read traffic even for sparse attention, providing a conservative traffic allowance rather than pretending to know a runtime's selected history. Flash-Next illustrates the limitation: its [official model card](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) distinguishes a 125B backbone with 6B active parameters from 51B n-gram tables, and its [technical report](https://arxiv.org/html/2608.30320v1) describes sparse history reads and inexpensive table lookups. The common formula still uses the full selected allocation and the stored 176B total. It is a rough traffic proxy, not a measurement of that architecture's tensors.

For **Qwen3.8 Flash-Next UD-Q3_K_XL on Strix Halo at 128K**, the inputs are 90.890358 GB of weights, 6B active out of 176B stored parameters, 3.762 GB of planned cache and 273 GB/s of listed bandwidth. The result is **5.97–17.91 tk/s**, displayed as approximately **6–18 tk/s**. The memory estimate is 93.23 GiB within the 96 GiB usable allowance after reserving a quarter of the 128 GB unified-memory system. Neither result is a measured run on Strix Halo.

All 35 current Ladder entries have an input for this calculation, so every fitting model/context combination receives a speed range, including conditional hardware. A missing bandwidth input in a future row must be supplied before a numeric speed can be calculated. CMP 170HX uses an **estimator-only 1,300–1,600 GB/s community envelope**, rounded outward from the published measurements in the community project's [memory-bandwidth record](https://github.com/Consensus-Protocol/cmp170hx/blob/main/docs/operations/performance.md#memory-bandwidth). Its lower endpoint drives the low speed and its upper endpoint the high speed. These reports cover different cards, memory modifications and access patterns; they do not verify bandwidth on every board. The Ladder's original unknown-bandwidth label remains unchanged, and the separate estimator input is identified as community evidence. No-fit and over-limit context combinations retain N/A because the fully resident scenario cannot run as configured.


## One or two physical cards

Discrete Ladder entries can be configured as one or two copies of the same card. Aggregate capacity and input bandwidth scale with physical card count; capacity and bandwidth per device stay unchanged. Native dual-device boards therefore become four devices when two boards are selected. Unified-memory machines retain a single-machine configuration.

The calculator uses a single resident model allocation and the existing runtime allowance. When the allocation exceeds memory on one device, it assumes serial layer sharding across the selected devices, retaining the common per-device bandwidth calculation. Card count therefore creates more memory room without automatically multiplying single-user speed. Exact split balance, duplicated runtime buffers, device interconnect and host placement remain implementation-dependent. The diagram shows two vertically aligned cards in a close physical stack, with a small air gap suggesting roughly 1 cm and the upper card obscuring most lower-card tiles. Allocations retain balanced shares when sharding is assumed and a second empty card when one device is enough; the accessible description preserves both cards’ memory totals.
