Can it run?
Provider
GPU
Count:
12 GB VRAM + 32 GB RAM (assumed)
Base model
Quant
Context
8k tokensOffload capacity32 GB RAM (assumed)
System RAM beyond your GPU's VRAM. Used to estimate whether a model that doesn't fully fit in VRAM can still run by offloading some layers to RAM.
GB
Fits VRAM · fast
Partial offload · slower
CPU only · slowest
weights 11.3 GB + KV cache 3.1 GB = 14.4 GB needed
See all reports for this combo →