Before you start
What happened
On Windows the engine sizes its pools as if nothing else held any of the card. get_free_memory reads torch.cuda.mem_get_info, and under WDDM cudaMemGetInfo leaves out what other processes hold - dwm and every desktop application with a GPU surface. With my desktop holding 2.8 GiB, it reported 23,313 MiB free where NVML and nvidia-smi reported 21,286.
The expert cache, the KV pages and the --vram-reserve-mb fit are all sized against that figure, so every start beside a desktop overcommits the card by what the desktop holds. WDDM does not fail the allocation; it pages it into shared memory, and the load runs there until the reserve fit shrinks the cache. With a VRAM hog standing in for a busier desktop, the engine read 22.75 GiB free every time:
free before the start (nvidia-smi) |
read by the engine |
slots sized, kept |
the fit's first prefill |
load |
| 22,198-22,890 MiB |
22.75 GiB |
2,900, 2,317 |
6.8 s |
132-148 s |
| 21,660 MiB |
22.75 GiB |
2,900, 1,820 |
45 s in shared memory |
215 s |
| 21,112 MiB |
22.75 GiB |
2,900, 1,432 |
62 s in shared memory |
230 s |
| 20,719 MiB |
22.75 GiB |
2,900, 1,290 |
228 s in shared memory |
384 s |
How did you install FreeToken
Built from source
FreeToken version
0d652e7 (get_free_memory is unchanged since v0.1.3, where I hit it)
OS
Windows 11
OS details
Windows 11 Pro 26200, Python 3.13.3, torch 2.11.0+cu130. The engine built with the Windows PRs linked from #384.
GPU and driver
RTX 3090 Ti 24 GB, driver 617.14
CPU and system RAM
AMD Ryzen 9 9950X3D, 256 GB
Checkpoint
RadixArk/Qwen3.8-Flash-Next-NVFP4
Command
python -c "import torch, subprocess; print(torch.cuda.mem_get_info(0)[0] >> 20, 'MiB free per CUDA'); print(subprocess.run(['nvidia-smi', '--query-gpu=memory.free', '--format=csv,noheader'], capture_output=True, text=True).stdout.strip(), 'free per nvidia-smi')"
ft serve --model RadixArk/Qwen3.8-Flash-Next-NVFP4 --memory-ratio 0.85 --vram-reserve-mb 2000
Full log
23313 MiB free per CUDA
21286 MiB free per nvidia-smi
INFO Free memory before loading model: 22.75 GiB
Anything else
NVML counts every process, and gpu_select.py already loads it. I have a fix with tests and will open a PR linked here.
Before you start
mainwhen building from source.What happened
On Windows the engine sizes its pools as if nothing else held any of the card.
get_free_memoryreadstorch.cuda.mem_get_info, and under WDDMcudaMemGetInfoleaves out what other processes hold -dwmand every desktop application with a GPU surface. With my desktop holding 2.8 GiB, it reported 23,313 MiB free where NVML andnvidia-smireported 21,286.The expert cache, the KV pages and the
--vram-reserve-mbfit are all sized against that figure, so every start beside a desktop overcommits the card by what the desktop holds. WDDM does not fail the allocation; it pages it into shared memory, and the load runs there until the reserve fit shrinks the cache. With a VRAM hog standing in for a busier desktop, the engine read 22.75 GiB free every time:nvidia-smi)How did you install FreeToken
Built from source
FreeToken version
0d652e7(get_free_memoryis unchanged since v0.1.3, where I hit it)OS
Windows 11
OS details
Windows 11 Pro 26200, Python 3.13.3, torch 2.11.0+cu130. The engine built with the Windows PRs linked from #384.
GPU and driver
RTX 3090 Ti 24 GB, driver 617.14
CPU and system RAM
AMD Ryzen 9 9950X3D, 256 GB
Checkpoint
RadixArk/Qwen3.8-Flash-Next-NVFP4
Command
python -c "import torch, subprocess; print(torch.cuda.mem_get_info(0)[0] >> 20, 'MiB free per CUDA'); print(subprocess.run(['nvidia-smi', '--query-gpu=memory.free', '--format=csv,noheader'], capture_output=True, text=True).stdout.strip(), 'free per nvidia-smi')" ft serve --model RadixArk/Qwen3.8-Flash-Next-NVFP4 --memory-ratio 0.85 --vram-reserve-mb 2000Full log
Anything else
NVML counts every process, and
gpu_select.pyalready loads it. I have a fix with tests and will open a PR linked here.