Will this LLM fit my GPU?
Model
Size (billions of parameters)
Quantization
GPU or Mac
Context
This calculator needs JavaScript.
Open the full Vramwise app
.
Estimate: weights at the quant's average bits + FP16 KV cache + ~1 GB runtime. Model shapes from each model's published config.json.
by CyberMax