02 / LOCAL AI, MADE TANGIBLE
A model. A machine.
Room to think.
Find out what fits, how much context you can carry, and what the pace might feel like.
The GPU Ladder’s hardware. A considered collection of current open models. One practical starting point.
Find your fit ↘01 / THE MODEL
Choose your model.
Quants adjust to your hardware and context.
02 / THE MACHINE
See how it fits.
THE WEIGHT FORMAT
YOUR COMBINATION
Memory fit is a planning estimate. Exact runtime and model support still matter.
03 / ROOM FOR THE CONVERSATION
Context holds the prompt, conversation and generated answer.
04 / A FEEL FOR THE PACE
From waiting
to flowing.
Single-user generation, once the prompt is processed. Loading, prompt reading and thinking time are additional.
How this is calculated
A little more memory can make room for a longer conversation. A faster memory interface can help the answer arrive sooner. The right balance depends on what you want to do.
KEEP EXPLORING
A different amount of room.
FIELDNOTES
Know what’s
behind the estimate.
The selected model & its sources
The selected hardware & its limits
How memory and pace are estimated
Download sizes use decimal GB; memory allocations on this page use GiB (1 GiB = 1,073,741,824 bytes). The hardware names retain the Ladder’s rated GB labels. Memory planning accounts for the capacity basis documented in the hardware notes. Each curated quant has its own published footprint. We add the complete selected weights, a model-specific context allowance and runtime workspace. The 32 GiB and 96 GiB starting points choose a listed quant for that usable memory budget, including context and runtime. Unified-memory systems also leave 25% for the operating system and other applications.
A model’s active parameters affect work per token; all its weights still need storage. Multi-chip boards require sharding and do not provide a single shared memory pool. Modified hardware, legacy runtimes and specialist accelerators need separate compatibility checks.
Generation ranges are broad bandwidth-based planning estimates, not benchmarks or a guarantee of accuracy. Kernel efficiency, attention, compute, quantization and drivers can move results outside the range. Every model and card uses the same calculation: memory bandwidth divided by estimated active-weight and context traffic, at 15–45% assumed efficiency. For boards with multiple devices, we use bandwidth per device and assume serial layer sharding when needed. Combinations that exceed memory or context limits retain N/A. CPU offload and tensor-parallel acceleration are outside this calculation.
Context means the total token budget for input, history and output. Multimodal inputs, batching, speculative decoding and concurrent users are outside this estimate.