Skip to content

On Windows the engine sizes its pools as if no other process held any of the card #570

Description

@YevheniiKotyrlo

Before you start

  • I have read the FAQ and my problem is not answered there.
  • I have read the Roadmap and this is not already planned there.
  • I have searched existing issues and found no duplicate.
  • I am on the latest release, or on a freshly rebuilt main when building from source.

What happened

On Windows the engine sizes its pools as if nothing else held any of the card. get_free_memory reads torch.cuda.mem_get_info, and under WDDM cudaMemGetInfo leaves out what other processes hold - dwm and every desktop application with a GPU surface. With my desktop holding 2.8 GiB, it reported 23,313 MiB free where NVML and nvidia-smi reported 21,286.

The expert cache, the KV pages and the --vram-reserve-mb fit are all sized against that figure, so every start beside a desktop overcommits the card by what the desktop holds. WDDM does not fail the allocation; it pages it into shared memory, and the load runs there until the reserve fit shrinks the cache. With a VRAM hog standing in for a busier desktop, the engine read 22.75 GiB free every time:

free before the start (nvidia-smi) read by the engine slots sized, kept the fit's first prefill load
22,198-22,890 MiB 22.75 GiB 2,900, 2,317 6.8 s 132-148 s
21,660 MiB 22.75 GiB 2,900, 1,820 45 s in shared memory 215 s
21,112 MiB 22.75 GiB 2,900, 1,432 62 s in shared memory 230 s
20,719 MiB 22.75 GiB 2,900, 1,290 228 s in shared memory 384 s

How did you install FreeToken

Built from source

FreeToken version

0d652e7 (get_free_memory is unchanged since v0.1.3, where I hit it)

OS

Windows 11

OS details

Windows 11 Pro 26200, Python 3.13.3, torch 2.11.0+cu130. The engine built with the Windows PRs linked from #384.

GPU and driver

RTX 3090 Ti 24 GB, driver 617.14

CPU and system RAM

AMD Ryzen 9 9950X3D, 256 GB

Checkpoint

RadixArk/Qwen3.8-Flash-Next-NVFP4

Command

python -c "import torch, subprocess; print(torch.cuda.mem_get_info(0)[0] >> 20, 'MiB free per CUDA'); print(subprocess.run(['nvidia-smi', '--query-gpu=memory.free', '--format=csv,noheader'], capture_output=True, text=True).stdout.strip(), 'free per nvidia-smi')"
ft serve --model RadixArk/Qwen3.8-Flash-Next-NVFP4 --memory-ratio 0.85 --vram-reserve-mb 2000

Full log

23313 MiB free per CUDA
21286 MiB free per nvidia-smi
INFO     Free memory before loading model: 22.75 GiB

Anything else

NVML counts every process, and gpu_select.py already loads it. I have a fix with tests and will open a PR linked here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions