Can it run?

Provider

GPU

Count:

12 GB VRAM + 32 GB RAM (assumed)

Base model

Quant

Context

8k tokens
Offload capacity32 GB RAM (assumed)

System RAM beyond your GPU's VRAM. Used to estimate whether a model that doesn't fully fit in VRAM can still run by offloading some layers to RAM.

GB

Partial offload · slower

14.4 GB needed of 12 GB VRAM + 32 GB RAM (assumed) (incl. 8k-context KV cache)

Fits VRAM · fast
Partial offload · slower
CPU only · slowest

No comparable reports yet — the fit verdict above is pure memory math. Speed numbers only appear when real reports back them.

weights 11.3 GB + KV cache 3.1 GB = 14.4 GB needed

See all reports for this combo →