Downloading models
NobodyWho can either load a model from a path on disk or download it for you on first use, caching it for subsequent runs. This page covers the available model path formats, how to observe a download in progress, how to access gated/private models, and how to inspect what's already in the local cache.
Supported model path formats
The model_path argument to Chat, download_model, and friends accepts:
| Form | Example | Notes |
|---|---|---|
| HuggingFace reference | hf:owner/repo/file.gguf | Downloaded and cached on first use |
| llama.cpp-style reference | owner/repo:quantization | Downloaded and cached on first use |
| HTTPS URL | https://example.com/model.gguf | Downloaded and cached on first use |
| Local path | ./model.gguf | Used as-is |
The HuggingFace prefix is case-insensitive and the // is optional — hf:, hf://, huggingface:, and huggingface:// all mean the same thing. Remote models are downloaded to the platform cache directory on first load and re-used on subsequent runs.
The llama.cpp-style reference takes no prefix, and names a HuggingFace repo whose name must end in -GGUF — the convention llama.cpp relies on to work out the filename. NobodyWho resolves it the same way, so ggml-org/gemma-3-1b-it-GGUF:Q8_0 fetches gemma-3-1b-it-Q8_0.gguf from the ggml-org/gemma-3-1b-it-GGUF repo. Both the -GGUF suffix and the quantization are required; without them the string is read as a local path rather than a download. Note: llama.cpp will take the first model in the repo if there is no exact match for the quantization, but NobodyWho will fail with an error if the quantization is not found.
Tracking download progress
When loading a remote model, pass an on_download_progress callback to observe the download. It receives (downloaded_bytes, total_bytes) and is not called for cached or local files. If you don't pass anything, NobodyWho prints a default terminal progress bar.
from nobodywho import download_model
model_path = download_model(
'huggingface:NobodyWho/Qwen_Qwen3-0.6B-GGUF/Qwen_Qwen3-0.6B-Q4_K_M.gguf',
on_download_progress=lambda downloaded, total: print(f"{downloaded}/{total} bytes"),
)
Downloading a gated model
Some HuggingFace models are private or gated by a license you need to accept. In both cases you need to be authorized to download the model weights.
You can manually download the GGUF file via your web browser and then point Chat at the local path:
from nobodywho import Chat
chat = Chat('./model.gguf')
Or use download_model with an Authorization header:
from nobodywho import Chat, download_model
model_path = download_model(
'huggingface:NobodyWho/Qwen_Qwen3-0.6B-GGUF/Qwen_Qwen3-0.6B-Q4_K_M.gguf',
headers={ "Authorization": "Bearer your_hf_token" }
)
chat = Chat(model_path)
You can generate a HuggingFace token in your account settings.
Inspecting the model cache
get_cached_models returns every .gguf model that lives in NobodyWho's cache directory, paired with its size in bytes. This is the same cache used by download_model and by Chat's huggingface: paths.
from nobodywho import get_cached_models
for path, size in get_cached_models():
print(f"{path}: {size / 1024 / 1024:.1f} MiB")
- Paths are absolute.
- Sizes are in bytes.
- The list is empty if nothing has been downloaded yet.
- Raises
RuntimeErrorif the cache directory cannot be read.