Choose a dedicated GPU, unified memory, or CPU-only and select a matching preset.
Local AI Hardware Checker
Choose GPU VRAM, unified memory, or CPU and instantly see which model size makes sense.
For users who want to know whether 3B, 8B, 14B, 32B, or 70B models are practical locally before downloading or buying hardware.
Open tool
Choose memory type and model, then see a clear fit estimate.
Ollama running slowly?
Narrow down the issue, then compare settings in the calculator. No device data or logs are read.
Check memory pressure
- Check RAM and swapping in Task Manager or Activity Monitor.
- Close unused memory-heavy applications. Use ollama ps to inspect loaded models.
- Compare a smaller model and shorter context in the calculator. Change one setting at a time.
Your setup
Three choices are enough for a first estimate.
Exact hardware valuesOnly needed when no preset fits.
The common local starting point for chat, coding, and agents.
Your result
Planning estimate — not a device benchmark.
Chat and coding should feel responsive.
How this estimate is calculated7B/8B Chat & Coding · Q4
Hardware guidance
Only relevant when choosing or upgrading hardware.
Data date & methodology
Good to know
What the calculator covers, where it stops, and which sources it uses.
VRAM, unified memory, RAM, download, and KV cache estimate for local LLMs
No real benchmark and no tokens-per-second guarantee
Can my PC run local AI?
Reddit questions about local LLMs almost always start with GPU, VRAM, quantization, and “will this run on my machine?” The checker turns those technical levers into a quick status.
Model memory first
For smooth local models, the first question is whether weights, KV cache, and runtime fit in dedicated VRAM or usable unified memory. Once offload is needed, speed often drops sharply.
Quantization explained
Q4, Q5, Q8, and FP16 change memory demand and quality. The checker shows which level looks realistic for your setup.
Context costs memory
Long context windows increase KV cache. This is exactly what many hardware discussions underestimate.