How much memory Qwen, Gemma, and Llama weights take
The download counter on Hugging Face, read on 10 October 2026, is led by small Qwen checkpoints. The cards state parameters and context. Most of them never state the gigabytes a program needs.
| Parameters on the card | Context on the card | Weight files | Memory the card states | |
|---|---|---|---|---|
| Qwen3-0.6B | 0.6B | 32,768 | 1.503 GB | |
| Qwen3-4B | 4.0B | 32,768 native | 8.045 GB | |
| Qwen3-8B | 8.2B | 32,768 native | 16.382 GB | |
| Qwen2.5-7B-Instruct | 7.61B | 131,072 on the card | 15.231 GB | |
| Qwen3.5-9B | 9B | 262,144 | 19.306 GB | |
| Llama 3.1 8B Instruct | 8B | 128k | 16.061 GB | |
| gpt-oss-20b | 21B, 3.6B active | 131,072 in config.json | 13.761 GB | within 16GB |
| Gemma 4 12B | 11.95B | 256K | 23.920 GB |
The question readers type is how many gigabytes a model needs before it will run. The pages that can answer it are the vendor's model card and the weight files sitting in that same repository. A calculator site that multiplies a parameter count by a rule of thumb is not one of those pages, and none of its products appear below.
On 10 October 2026 the Hugging Face API, asked for text-generation models sorted by downloads, returned the counts in this brief. Qwen/Qwen3-0.6B stood at 30,768,858. Qwen/Qwen3-4B stood at 10,600,642. Qwen/Qwen3-8B stood at 9,821,408. Qwen/Qwen3.5-9B stood at 8,319,798. Qwen/Qwen2.5-7B-Instruct stood at 8,077,595. meta-llama/Llama-3.2-1B-Instruct stood at 6,976,807. meta-llama/Llama-3.1-8B-Instruct stood at 6,015,326. openai/gpt-oss-20b stood at 5,950,578. google/gemma-4-12b-it stood at 1,603,246.
A count on that list is a fetch of files in the repository. Mirrors, installers, and test jobs all increment it. The second row of the same list was openai-community/gpt2, at 15,521,105 downloads. An internal test checkpoint, trl-internal-testing/tiny-Qwen2ForCausalLM-2.5, sat in the top twelve at 6,239,622. The counter is a traffic figure for the hub. It is not a census of people sitting down to chat.
A separate search for repositories with GGUF in the name, the format local loaders actually open, put mudler/locate-anything.cpp-gguf first at 12,330,970 downloads. unsloth/Qwen3.8-27B-GGUF was at 6,449,020, and unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF was at 5,134,275. Those repositories hold repackaged files whose gigabyte sizes are not on the vendor cards, so this brief leaves them out of the table.
The Qwen3-0.6B card says the number of parameters is 0.6B and the context length is 32,768 tokens. The repository holds one safetensors file of 1,503,300,328 bytes. Counting a gigabyte as one billion bytes, that file is 1.503 gigabytes. The card never prints a memory requirement, so the chart has no memory bar for it.
The Qwen3-4B card says 4.0B parameters. Context is 32,768 tokens natively, and 131,072 tokens if the reader turns on the YaRN scaling the card describes. Three safetensors files add to 8,044,982,000 bytes, which is 8.045 gigabytes on the same counting. No gigabyte figure for running the model is written on the card.
The Qwen3-8B card says 8.2B parameters, with the same native context line and the same YaRN extension as the 4B card. Five safetensors files add to 16,381,516,776 bytes, which is 16.382 gigabytes. The config file sets torch_dtype to bfloat16 and max_position_embeddings to 40,960. The card says that 40,960 allocation reserves 32,768 tokens for the output and 8,192 tokens for a typical prompt. It still does not say how many gigabytes the program will hold while it runs.
The Qwen2.5-7B-Instruct card says 7.61B parameters. It says the full context is 131,072 tokens and generation is 8,192 tokens, and it says the config file as shipped is set for 32,768 tokens. Four safetensors files add to 15,231,271,888 bytes, which is 15.231 gigabytes. The card sends the reader to a separate speed-benchmark page for GPU memory and throughput. That other page is not quoted here, because the gigabyte figure is not printed on the card itself.
The Qwen3.5-9B card says 9B parameters and a native context of 262,144 tokens, with a note that the window can be extended to 1,010,000 tokens. Four safetensors files add to 19,306,310,880 bytes, which is 19.306 gigabytes. The card tells a reader who runs out of memory to shorten the context, and it advises keeping at least 128K tokens when the work needs the model's longer reasoning. It does not state a gigabyte budget.
The Llama 3.1 card describes a collection of text models in 8B, 70B, and 405B sizes. On the 8B column of that card the context length is 128k. Four safetensors files in the instruct repository add to 16,060,556,376 bytes, which is 16.061 gigabytes. The card does not state how much memory a running copy needs. Hugging Face's safetensors index counts 8,030,261,248 bfloat16 parameters inside those files. That number is an index of the tensors. It is not a measurement of the memory a chat program will allocate.
The gpt-oss-20b card says the model has 21B parameters, with 3.6B of them active. It says the expert weights were post-trained in MXFP4, and that this model runs within 16GB of memory. The same paragraph says the 120B sibling fits on a single 80GB GPU, and it gives that sibling 117B parameters with 5.1B active. Three safetensors files on the 20B repository add to 13,761,316,904 bytes, which is 13.761 gigabytes. The config file sets max_position_embeddings to 131,072. The README does not restate that context figure, so the brief marks it as a config value. The card does not split the 16GB into weights, cache, and runtime. The files are smaller than 16GB because they are the quantized weights the card describes.
The Gemma 4 12B card, in the column headed 12B Unified, lists 11.95B total parameters and a context of 256K tokens. One safetensors file is 23,919,549,408 bytes, which is 23.920 gigabytes. The card says this size is aimed at consumer GPUs and workstations. That is a class of machine, not a memory quantity, so the chart does not add a memory bar. The card's prose also says the 12B model accepts audio, while the parameter table leaves the 12B audio-encoder cell unmarked. This brief does not choose between those two lines.
The weight file is the floor, not the bill. A running program also keeps a cache of earlier tokens, the runtime, and sometimes a second copy of the weights while they load. Only the gpt-oss card, of the cards read for this brief, states a memory total. A 4-bit file published by a repackager is a different object, and its size is not filled in from a formula.
Where the repository ships bfloat16 weights, the file bytes land near two bytes a parameter. Dividing the Qwen3-8B file sum, 16,381,516,776, by the safetensors index count of 8,190,735,360 parameters, lands near two. The division is arithmetic we did, and it is not a sentence on the card. The config file does say the stored dtype is bfloat16, which is why the neighborhood is unsurprising. The gpt-oss files do not sit in that neighborhood, and the card explains why: the expert weights are MXFP4.
The download leader is the 0.6B checkpoint, whose single file is 1.503 gigabytes. The memory question starts where the vendor's own bfloat16 files already occupy 15 to 24 gigabytes before a long context is added. The one card in this set that answers in gigabytes is gpt-oss-20b, at 16GB. Anyone updating the table should re-read the card and re-sum the safetensors files. The download counts move every day. The byte sums move only when the repository replaces those files.
After Hugging Face. We did not report this. The pictures, if any, are theirs. Also filed by Qwen, Meta, OpenAI, Google.


