Choose a dedicated GPU, unified memory, or CPU-only and select a matching preset.
Local AI Hardware Checker
Choose GPU VRAM, unified memory, or CPU and instantly see which model size makes sense.
Open tool
Choose memory type and model, then see a clear fit estimate.
Your setup
Three choices are enough for a first estimate.
Exact hardware valuesOnly needed when no preset fits.
The common local starting point for chat, coding, and agents.
Your result
Planning estimate — not a device benchmark.
Chat and coding should feel responsive.
How this estimate is calculated7B/8B Chat & Coding · Q4
Hardware guidance
Only relevant when choosing or upgrading hardware.
Data date & methodology
Good to know
What the calculator covers, where it stops, and which sources it uses.
VRAM, unified memory, RAM, download, and KV cache estimate for local LLMs
No real benchmark and no tokens-per-second guarantee
Can my PC run local AI?
Reddit questions about local LLMs almost always start with GPU, VRAM, quantization, and “will this run on my machine?” The checker turns those technical levers into a quick status.
Model memory first
For smooth local models, the first question is whether weights, KV cache, and runtime fit in dedicated VRAM or usable unified memory. Once offload is needed, speed often drops sharply.
Quantization explained
Q4, Q5, Q8, and FP16 change memory demand and quality. The checker shows which level looks realistic for your setup.
Context costs memory
Long context windows increase KV cache. This is exactly what many hardware discussions underestimate.