Skip to content

[Windows][Desktop] WeightLoadError: invalid python storage with Qwen3.6 NVFP4 and gpt-oss-20b MXFP4 #597

Description

@Shahrad-gh

Before you start

  • I have read the FAQ and my problem is not answered there.
  • I have read the Roadmap and this is not already planned there.
  • I have searched existing issues and found no duplicate.
  • I have restarted the Desktop app to pick up the latest update and the problem still happens.

What happened

FreeToken Desktop fails while loading model weights before inference starts.

I can reproduce the same error with two unrelated supported models and quantization formats:

  1. Qwen3.6-35B-A3B NVFP4
  2. gpt-oss-20b MXFP4

Both fail with:

WeightLoadError:
RuntimeError: Attempted to access the data pointer on an invalid python storage.

The failure happens during weight materialization/loading, before inference starts.

Qwen3.6 fails in the qwen3_5_moe weight loader.
gpt-oss-20b fails independently in the gpt_oss weight loader / merged tensor path.

This does not appear to be VRAM OOM:

  • GPU is detected normally by PyTorch
  • CUDA is available
  • FreeToken reports ~6.94 GiB free VRAM before loading Qwen
  • the same storage error occurs across NVFP4 and MXFP4 models

System:

  • OS: Windows
  • GPU: NVIDIA GeForce RTX 3070 8 GB
  • RAM: 32 GB
  • Compute capability: 8.6
  • Python: 3.12.14
  • PyTorch: 2.11.0+cu130
  • CUDA runtime: 13.0
  • FreeToken engine: 0.1.3+gc8ed699cb

FreeToken auto-selects Triton attention and hybrid MoE strategy.

Expected:
The supported models should load and start serving.

Actual:
The scheduler dies during weight loading with:
RuntimeError: Attempted to access the data pointer on an invalid python storage.

The same error reproducing with both Qwen3.6 NVFP4 and gpt-oss-20b MXFP4 makes this look like a Windows/Desktop weight-loader or storage-materialization issue rather than a checkpoint-specific problem.

Full server log attached below.

Desktop app version

v0.2.0-beta.23

OS

Windows 11

OS details

Windows 11 Pro 26H1, build 28000.1

GPU and driver

NVIDIA GeForce RTX 3070 8GB (8192 MiB), NVIDIA driver 616.92, WDDM mode, CUDA UMD 13.4

CPU and system RAM

AMD Ryzen 7 7800X3D 8-Core Processor, 32GB RAM DDR5

Checkpoint

nvidia/Qwen3.6-35B-A3B-NVFP4

Model settings

Default Desktop app settings.

FreeToken command:
ft serve --model C:\Users\Admin.freetoken\models\Qwen3.6-35B-A3B-NVFP4 --port 1919 --moe-backend auto --max-running-requests 4 --memory-ratio 0.85 --host 127.0.0.1

Resolved by FreeToken:

  • dtype: bfloat16
  • attention backend: triton
  • MoE strategy: hybrid
  • MoE cache: auto
  • expert loading: auto
  • memory ratio: 0.85
  • max running requests: 4
  • KV reserve tokens: 8192
  • max extend tokens: 8192
  • cache type: hybrid_radix
  • page size: 1
  • no manual context-length override
  • no manual quant backend override
  • multimodal encoders enabled by default

Engine log

[07:37:40] cmd/INFO  detecting CPU/PCIe bandwidth to pick the MoE backend (ft bench bw)…
[07:38:04] cmd/INFO  ft serve --model C:\Users\Admin\.freetoken\models\Qwen3.6-35B-A3B-NVFP4 --port 1919 --moe-backend auto --max-running-requests 4 --memory-ratio 0.85 --host 127.0.0.1 --cors-origins tauri://localhost,http://tauri.localhost,http://localhost:1420
[07:38:04] stdout/INFO  serve started (pid=23456 model=C:\Users\Admin\.freetoken\models\Qwen3.6-35B-A3B-NVFP4 port=1919)
[07:38:07] stdout/WARN  [2026-10-04|07:38:07] WARNING  --moe-backend is deprecated; use --moe-strategy
[07:38:07] stdout/INFO  [2026-10-04|07:38:07] INFO     Parsed arguments:
[07:38:07] stdout/INFO  ServerArgs(model_path='C:\\Users\\Admin\\.freetoken\\models\\Qwen3.6-35B-A3B-NVFP4', tp_info=DistributedInfo(rank=0, size=1), dtype=torch.bfloat16, max_running_req=4, attention_backend='auto', moe_strategy='auto', quant_backend=None, ple_backend='disk', expert_load='auto', moe_cache_size=0, moe_cache_rate=None, moe_cache_auto=False, kv_reserve_tokens=8192, moe_cache_policy='lru', moe_prefill_overlap=True, moe_prefill_hit_d2d=False, moe_collect_stats=False, moe_cpu_threads=0, moe_cpu_layers=None, moe_hybrid_max_fetch=-1, cuda_graph_bs=None, cuda_graph_max_bs=None, page_size=1, memory_ratio=0.85, linear_state_cache_ratio=2.0, swa_full_tokens_ratio=0.2, swa_num_pages_override=None, distributed_timeout=60.0, use_dummy_weight=False, use_pynccl=True, max_seq_len_override=None, num_page_override=None, num_token_override=None, mm=MultimodalConfig(disabled_encoders=frozenset(), embed_cache_device='cpu', encoder_weights='host', image_min_tokens=None, image_max_tokens=None, processor_kwargs={}), max_extend_tokens=8192, cache_type='radix', offline_mode=False, decode_log_interval=40, special_token_ckpt=False, _unique_suffix='.pid=22904', _ipc_base_port=63615, server_host='127.0.0.1', server_port=1919, num_tokenizer=0, silent_output=False, shell_mode=False, served_model_name='Qwen3.6-35B-A3B-NVFP4', tool_call_parser='qwen3_coder', reasoning_parser='qwen3', sampling_defaults='model', max_output_tokens=None, enable_cache_report=False, allowed_media_domains='', allowed_local_media_path='', cors_origins='tauri://localhost,http://tauri.localhost,http://localhost:1420', gpu=(), gpu_assigned=None)
[07:38:07] stdout/INFO  [2026-10-04|07:38:07|FrontendAPI] INFO     Default sampling config (source=model): temperature=1.0, top_k=20, top_p=0.95
[07:38:08] stdout/INFO  INFO:     Started server process [22904]
[07:38:08] stdout/INFO  INFO:     Waiting for application startup.
[07:38:08] stdout/INFO  INFO:     Application startup complete.
[07:38:08] stdout/INFO  INFO:     Uvicorn running on http://127.0.0.1:1919 (Press CTRL+C to quit)
[07:38:10] stdout/INFO  C:\Users\Admin\AppData\Local\FreeToken\venv\Lib\site-packages\freetoken\server\launch.py:81: FutureWarning: torch.cuda._set_allocator_settings is deprecated. Use torch._C._accelerator_setAllocatorSettings instead.
[07:38:10] stdout/INFO    scheduler = Scheduler(args)
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     Enabled expandable_segments (override via PYTORCH_ALLOC_CONF)
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     Auto-selected attention backend: triton
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     benchbw profile recommends hybrid for 'nvfp4' experts on this GPU
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     Auto-selected MoE strategy: hybrid
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     No MoE cache sizing flag given; defaulting to --moe-cache-auto for auto-selected strategy 'hybrid'
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     expert banks exceed the pin budget: --moe-cpu-layers defaults to 'auto' on native Windows
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     Resolved config: moe_strategy='hybrid', attention_backend='triton', cache_type='hybrid_radix', page_size=1
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     Free memory before loading model: 6.94 GiB
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     kernel triton selected; skipped torch: fp8 tensor cores need sm_89+, got sm_86  (×7)
[07:38:10] stdout/INFO  C:\Users\Admin\AppData\Local\FreeToken\venv\Lib\site-packages\torch\utils\_device.py:116: UserWarning: expandable_segments not supported on this platform (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\pytorch\c10/cuda/CUDAAllocatorConfig.h:39.)
[07:38:10] stdout/INFO    return func(*args, **kwargs)
[07:38:10] stdout/INFO  [2026-10-04|07:38:10|core|rank=0] INFO     kernel triton selected; skipped torch: fp8 tensor cores need sm_89+, got sm_86  (×73)
[07:38:10] stdout/INFO  Loading weights:   0%|          | 0/3 [00:00<?, ?it/s]
[07:38:10] stdout/INFO  Process freetoken-TP0-scheduler:
[07:38:10] stdout/ERROR  [2026-10-04|07:38:10|FrontendAPI] ERROR    Backend supervisor: WeightLoadError: RuntimeError: Attempted to access the data pointer on an invalid python storage.
[07:38:10] stdout/INFO  Traceback (most recent call last):
[07:38:10] stdout/INFO    File "python/freetoken/engine/engine.py", line 346, in freetoken.engine.engine._weight_load_context
[07:38:11] stdout/INFO    File "python/freetoken/engine/engine.py", line 600, in freetoken.engine.engine.Engine._load_weights
[07:38:11] stdout/INFO    File "python/freetoken/engine/engine.py", line 610, in freetoken.engine.engine.Engine._load_weight_state_dict
[07:38:11] stdout/INFO    File "python/freetoken/engine/engine.py", line 363, in freetoken.engine.engine._materialize_loaded_weight_state_dict
[07:38:11] stdout/INFO    File "python/freetoken/models/weight.py", line 240, in load_weight
[07:38:11] stdout/INFO    File "python/freetoken/models/qwen3_5_moe/weight.py", line 255, in iter_weights
[07:38:11] stdout/INFO    File "python/freetoken/models/qwen3_5_moe/weight.py", line 262, in _iter_shards
[07:38:11] stdout/INFO    File "python/freetoken/models/qwen3_5_moe/weight.py", line 275, in freetoken.models.qwen3_5_moe.weight._iter_shards
[07:38:11] stdout/INFO  RuntimeError: Attempted to access the data pointer on an invalid python storage.
[07:38:11] stdout/INFO  
[07:38:11] stdout/INFO  The above exception was the direct cause of the following exception:
[07:38:11] stdout/INFO  
[07:38:11] stdout/INFO  Traceback (most recent call last):
[07:38:11] stdout/INFO    File "C:\Users\Admin\AppData\Roaming\uv\python\cpython-3.12-windows-x86_64-none\Lib\multiprocessing\process.py", line 314, in _bootstrap
[07:38:11] stdout/INFO      self.run()
[07:38:11] stdout/INFO    File "C:\Users\Admin\AppData\Roaming\uv\python\cpython-3.12-windows-x86_64-none\Lib\multiprocessing\process.py", line 108, in run
[07:38:11] stdout/INFO      self._target(*self._args, **self._kwargs)
[07:38:11] stdout/INFO    File "C:\Users\Admin\AppData\Local\FreeToken\venv\Lib\site-packages\freetoken\server\launch.py", line 81, in _run_scheduler
[07:38:11] stdout/INFO      scheduler = Scheduler(args)
[07:38:11] stdout/INFO                  ^^^^^^^^^^^^^^^
[07:38:11] stdout/INFO    File "python/freetoken/scheduler/scheduler.py", line 68, in freetoken.scheduler.scheduler.Scheduler.__init__
[07:38:11] stdout/INFO    File "python/freetoken/engine/engine.py", line 419, in freetoken.engine.engine.Engine.__init__
[07:38:11] stdout/INFO    File "python/freetoken/engine/engine.py", line 599, in freetoken.engine.engine.Engine._load_weights
[07:38:11] stdout/INFO    File "C:\Users\Admin\AppData\Roaming\uv\python\cpython-3.12-windows-x86_64-none\Lib\contextlib.py", line 158, in __exit__
[07:38:11] stdout/INFO      self.gen.throw(value)
[07:38:11] stdout/INFO    File "python/freetoken/engine/engine.py", line 353, in _weight_load_context
[07:38:11] stdout/INFO  freetoken.engine.engine.WeightLoadError: RuntimeError: Attempted to access the data pointer on an invalid python storage.
[07:38:12] health/ERROR  WeightLoadError: RuntimeError: Attempted to access the data pointer on an invalid python storage.

Anything else

The same WeightLoadError also reproduces with gpt-oss-20b MXFP4 using the same default Desktop settings.

Activity

  1. added
    bugSomething isn't working
    Desktopproblem related to FreeToken Desktop
    on Oct 4, 2026
  2. jason-fxz commented on Oct 5, 2026

    @jason-fxz
    Collaborator

    This is Windows running out of RAM + page file while loading the weights, not VRAM or a broken checkpoint. On Windows, safetensors memory-mapped each shard twice, and Windows counts every mapped copy in full against RAM + page file, so a large shard briefly needed twice its size. Engine 0.1.3+gfa0e9537d reads the shards without mapping them. To get it, restart FreeToken Desktop and click Update engine in the top banner (or Settings -> Engine -> Reinstall engine). If it still fails, please paste the full server log.

  3. pablosilva87 commented on Oct 5, 2026

    @pablosilva87

    Adding two more models that reproduce this on the same engine build, plus a regression window that may help pinpoint it.

    Environment: Windows 11 Pro, RTX 5070 12 GB (sm_120), driver 616.56, 31.1 GB RAM, Python 3.12, torch 2.11.0+cu130, engine 0.1.3+gc8ed699cb.

    Additional repros on gc8ed699cb (same WeightLoadError: RuntimeError: Attempted to access the data pointer on an invalid python storage.):

    1. deepseek-ai/DeepSeek-V4-Flash-0731 BF16 (48 shards, 167 GB) — fails in the deepseek_v4 weight loader.
    2. RedHatAI/GLM-5.3-Flash-NVFP4 — fails via both load paths we tried (hf-cache repo id and plain local folder), at --kv-reserve-tokens 8192 and 65536, so it is not KV-allocation related.

    Not file corruption. Two independent checks:

    • A plain-python safetensors mass-probe over all 48 DeepSeek BF16 shards read 50,306 tensors successfully in 244 s and only then hit the same error at tensor ~50,307 (layers.32.hc_attn_base.*), reproducibly. The failure looks cumulative (storage materialization/poisoning), not a bad shard.
    • The GLM-5.3-Flash-NVFP4 files were later SHA256-verified against HF LFS hashes — all 13 files clean — and still failed to load on this build.

    Regression window. GLM-5.3-Flash-NVFP4 benched 6/6 at 60.9 tok/s at 23:56 on Oct 3 under the previously installed build. An upgrade landed at 23:59 Oct 3 (gc8ed699cb) and from that moment GLM failed with this exact error on every load path. We no longer have the previous wheel (the release tag is rolling), but the 3-minute window pins the regression to that upgrade.

    Status on 0.1.3+gfa0e9537d: Qwen3.6-35B-A3B NVFP4 loads and serves cleanly across several restarts; GLM-5.3-Flash now gets past weight reading (it fails later during host expert-bank building — see #600); DeepSeek-V4-Flash NVFP4 fails with a different error (KeyError: 'layers.0.ffn.experts.0.w1.scale'), which I'll file separately. So for the paths we could exercise, the storage-materialization failure mode looks resolved in gfa0e9537d.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Desktopproblem related to FreeToken DesktopbugSomething isn't workingwindows

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions