CyberMax
Home › Apps

Know if an LLM fits your GPU before you download it

For: Local-LLM hobbyists choosing a model and quantization, ML engineers sizing hardware, and anyone deciding which GPU or Mac to buy for local AI.

Paste a Hugging Face model id or URL and pick a GPU, a Mac or a custom machine. Vramwise reads the model's parameter counts and config.json live from the Hugging Face API and shows, for every quantization from BF16 down to 4-bit and below, the weights, the KV cache, the total memory, whether it fits, the longest context that fits and an estimated tokens-per-second range. It handles mixture-of-experts models, sliding-window layers and DeepSeek's MLA cache, and knows Apple silicon memory limits per chip.

Vramwise result: Qwen3-8B runs on RTX 4090 24GB at BF16, table of quantizations with memory and speed
Real screenshot: Qwen3-8B on an RTX 4090 at BF16

Qwen3-8B on an RTX 4090 24GB at 8K context: 18.6 GB of 24.0 GB usable at BF16

QuantizationWeightsKV cacheTotalFits?Max contextSpeed (tok/s)
BF16 / FP16 (16 bits)16.4 GB1.21 GB18.6 GBFits40K29–46
Q8_0 / FP8 / MLX 8-bit (8.5 bits)8.70 GB1.21 GB10.9 GBFits40K51–81
Q6_K (6.56 bits)6.72 GB1.21 GB8.92 GBFits40K64–102
Q5_K_M (5.69 bits)5.83 GB1.21 GB8.03 GBFits40K72–115
Q4_K_M / MLX 4-bit (4.85 bits)4.97 GB1.21 GB7.17 GBFits40K82–131
IQ4_XS / AWQ / GPTQ 4-bit4.35 GB1.21 GB6.56 GBFits40K91–145

Source: Rows transcribed from the published Vramwise screenshot https://cybermaxtools.com/store/img/vramwise-1.webp (live app, Qwen3-8B on RTX 4090 24GB). Estimates, not measurements.

Vramwise vs the alternatives

Published prices, each checked on the date shown; prices change, so confirm on each site.

ProductPriceFreeChecked
Vramwisefree; Pro $5 once (Markdown tables)full fit check, share links and CSVlive
ApX Machine Learning VRAM & Performance CalculatorBasic $0/month (single-node VRAM breakdown); Pro $19/month; Pro+ $59/month (multi-GPU topologies)Basic plan, free forever2026-10-10

Why teams pick Vramwise

How it works

  1. Open Vramwise and paste a Hugging Face model id (for example Qwen/Qwen3-8B) or URL.
  2. Pick your GPU, Mac or a custom machine, and the context length you need.
  3. Read the table: the best quantization that fits, its longest context and the speed range; share the link or export CSV.

FAQ

How much VRAM do I need to run an LLM?

Roughly the weights (parameters times bits per weight) plus the KV cache for your context plus about 1 GB of overhead. Vramwise does this per quantization from the model's real config; Qwen3-8B needs 18.6 GB at BF16 and 7.17 GB at Q4_K_M with 8K context.

Why not ApX's VRAM calculator?

ApX's free Basic plan gives a single-node VRAM breakdown and charges $19/month for Pro sizing. Vramwise is free for any model on Hugging Face, adds Apple silicon limits and MoE-aware speed, and its only paid extra is $5 once for Markdown tables.

How accurate is the speed?

It is an estimate: 50 to 80% of the memory-bandwidth ceiling for one user. Real speed depends on the runtime, drivers and batch size.

Does it work with gated models like Llama or Gemma?

Yes; the architecture is read from a public mirror. For GGUF repos, parameters come from the GGUF metadata.

Works well with

Leafmeltdocuments to Markdown for the modelDepmoordependency securityCyberMax Developer KitMCP configs and recipes

Related

Vramwise in the CyberMax StorePlans, checkout, FAQAll CyberMax web appsBuyer guides with pricesAPI alternativesPublished prices side by side
Product names of other companies are trademarks of their owners and are used only to compare published prices; no affiliation is implied.