Before you start
What happened
FreeToken Desktop fails while loading model weights before inference starts.
I can reproduce the same error with two unrelated supported models and quantization formats:
- Qwen3.6-35B-A3B NVFP4
- gpt-oss-20b MXFP4
Both fail with:
WeightLoadError:
RuntimeError: Attempted to access the data pointer on an invalid python storage.
The failure happens during weight materialization/loading, before inference starts.
Qwen3.6 fails in the qwen3_5_moe weight loader.
gpt-oss-20b fails independently in the gpt_oss weight loader / merged tensor path.
This does not appear to be VRAM OOM:
- GPU is detected normally by PyTorch
- CUDA is available
- FreeToken reports ~6.94 GiB free VRAM before loading Qwen
- the same storage error occurs across NVFP4 and MXFP4 models
System:
- OS: Windows
- GPU: NVIDIA GeForce RTX 3070 8 GB
- RAM: 32 GB
- Compute capability: 8.6
- Python: 3.12.14
- PyTorch: 2.11.0+cu130
- CUDA runtime: 13.0
- FreeToken engine: 0.1.3+gc8ed699cb
FreeToken auto-selects Triton attention and hybrid MoE strategy.
Expected:
The supported models should load and start serving.
Actual:
The scheduler dies during weight loading with:
RuntimeError: Attempted to access the data pointer on an invalid python storage.
The same error reproducing with both Qwen3.6 NVFP4 and gpt-oss-20b MXFP4 makes this look like a Windows/Desktop weight-loader or storage-materialization issue rather than a checkpoint-specific problem.
Full server log attached below.
Desktop app version
v0.2.0-beta.23
OS
Windows 11
OS details
Windows 11 Pro 26H1, build 28000.1
GPU and driver
NVIDIA GeForce RTX 3070 8GB (8192 MiB), NVIDIA driver 616.92, WDDM mode, CUDA UMD 13.4
CPU and system RAM
AMD Ryzen 7 7800X3D 8-Core Processor, 32GB RAM DDR5
Checkpoint
nvidia/Qwen3.6-35B-A3B-NVFP4
Model settings
Default Desktop app settings.
FreeToken command:
ft serve --model C:\Users\Admin.freetoken\models\Qwen3.6-35B-A3B-NVFP4 --port 1919 --moe-backend auto --max-running-requests 4 --memory-ratio 0.85 --host 127.0.0.1
Resolved by FreeToken:
- dtype: bfloat16
- attention backend: triton
- MoE strategy: hybrid
- MoE cache: auto
- expert loading: auto
- memory ratio: 0.85
- max running requests: 4
- KV reserve tokens: 8192
- max extend tokens: 8192
- cache type: hybrid_radix
- page size: 1
- no manual context-length override
- no manual quant backend override
- multimodal encoders enabled by default
Engine log
[07:37:40] cmd/INFO detecting CPU/PCIe bandwidth to pick the MoE backend (ft bench bw)…
[07:38:04] cmd/INFO ft serve --model C:\Users\Admin\.freetoken\models\Qwen3.6-35B-A3B-NVFP4 --port 1919 --moe-backend auto --max-running-requests 4 --memory-ratio 0.85 --host 127.0.0.1 --cors-origins tauri://localhost,http://tauri.localhost,http://localhost:1420
[07:38:04] stdout/INFO serve started (pid=23456 model=C:\Users\Admin\.freetoken\models\Qwen3.6-35B-A3B-NVFP4 port=1919)
[07:38:07] stdout/WARN [2026-10-04|07:38:07] WARNING --moe-backend is deprecated; use --moe-strategy
[07:38:07] stdout/INFO [2026-10-04|07:38:07] INFO Parsed arguments:
[07:38:07] stdout/INFO ServerArgs(model_path='C:\\Users\\Admin\\.freetoken\\models\\Qwen3.6-35B-A3B-NVFP4', tp_info=DistributedInfo(rank=0, size=1), dtype=torch.bfloat16, max_running_req=4, attention_backend='auto', moe_strategy='auto', quant_backend=None, ple_backend='disk', expert_load='auto', moe_cache_size=0, moe_cache_rate=None, moe_cache_auto=False, kv_reserve_tokens=8192, moe_cache_policy='lru', moe_prefill_overlap=True, moe_prefill_hit_d2d=False, moe_collect_stats=False, moe_cpu_threads=0, moe_cpu_layers=None, moe_hybrid_max_fetch=-1, cuda_graph_bs=None, cuda_graph_max_bs=None, page_size=1, memory_ratio=0.85, linear_state_cache_ratio=2.0, swa_full_tokens_ratio=0.2, swa_num_pages_override=None, distributed_timeout=60.0, use_dummy_weight=False, use_pynccl=True, max_seq_len_override=None, num_page_override=None, num_token_override=None, mm=MultimodalConfig(disabled_encoders=frozenset(), embed_cache_device='cpu', encoder_weights='host', image_min_tokens=None, image_max_tokens=None, processor_kwargs={}), max_extend_tokens=8192, cache_type='radix', offline_mode=False, decode_log_interval=40, special_token_ckpt=False, _unique_suffix='.pid=22904', _ipc_base_port=63615, server_host='127.0.0.1', server_port=1919, num_tokenizer=0, silent_output=False, shell_mode=False, served_model_name='Qwen3.6-35B-A3B-NVFP4', tool_call_parser='qwen3_coder', reasoning_parser='qwen3', sampling_defaults='model', max_output_tokens=None, enable_cache_report=False, allowed_media_domains='', allowed_local_media_path='', cors_origins='tauri://localhost,http://tauri.localhost,http://localhost:1420', gpu=(), gpu_assigned=None)
[07:38:07] stdout/INFO [2026-10-04|07:38:07|FrontendAPI] INFO Default sampling config (source=model): temperature=1.0, top_k=20, top_p=0.95
[07:38:08] stdout/INFO INFO: Started server process [22904]
[07:38:08] stdout/INFO INFO: Waiting for application startup.
[07:38:08] stdout/INFO INFO: Application startup complete.
[07:38:08] stdout/INFO INFO: Uvicorn running on http://127.0.0.1:1919 (Press CTRL+C to quit)
[07:38:10] stdout/INFO C:\Users\Admin\AppData\Local\FreeToken\venv\Lib\site-packages\freetoken\server\launch.py:81: FutureWarning: torch.cuda._set_allocator_settings is deprecated. Use torch._C._accelerator_setAllocatorSettings instead.
[07:38:10] stdout/INFO scheduler = Scheduler(args)
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO Enabled expandable_segments (override via PYTORCH_ALLOC_CONF)
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO Auto-selected attention backend: triton
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO benchbw profile recommends hybrid for 'nvfp4' experts on this GPU
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO Auto-selected MoE strategy: hybrid
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO No MoE cache sizing flag given; defaulting to --moe-cache-auto for auto-selected strategy 'hybrid'
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO expert banks exceed the pin budget: --moe-cpu-layers defaults to 'auto' on native Windows
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO Resolved config: moe_strategy='hybrid', attention_backend='triton', cache_type='hybrid_radix', page_size=1
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO Free memory before loading model: 6.94 GiB
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO kernel triton selected; skipped torch: fp8 tensor cores need sm_89+, got sm_86 (×7)
[07:38:10] stdout/INFO C:\Users\Admin\AppData\Local\FreeToken\venv\Lib\site-packages\torch\utils\_device.py:116: UserWarning: expandable_segments not supported on this platform (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\pytorch\c10/cuda/CUDAAllocatorConfig.h:39.)
[07:38:10] stdout/INFO return func(*args, **kwargs)
[07:38:10] stdout/INFO [2026-10-04|07:38:10|core|rank=0] INFO kernel triton selected; skipped torch: fp8 tensor cores need sm_89+, got sm_86 (×73)
[07:38:10] stdout/INFO Loading weights: 0%| | 0/3 [00:00<?, ?it/s]
[07:38:10] stdout/INFO Process freetoken-TP0-scheduler:
[07:38:10] stdout/ERROR [2026-10-04|07:38:10|FrontendAPI] ERROR Backend supervisor: WeightLoadError: RuntimeError: Attempted to access the data pointer on an invalid python storage.
[07:38:10] stdout/INFO Traceback (most recent call last):
[07:38:10] stdout/INFO File "python/freetoken/engine/engine.py", line 346, in freetoken.engine.engine._weight_load_context
[07:38:11] stdout/INFO File "python/freetoken/engine/engine.py", line 600, in freetoken.engine.engine.Engine._load_weights
[07:38:11] stdout/INFO File "python/freetoken/engine/engine.py", line 610, in freetoken.engine.engine.Engine._load_weight_state_dict
[07:38:11] stdout/INFO File "python/freetoken/engine/engine.py", line 363, in freetoken.engine.engine._materialize_loaded_weight_state_dict
[07:38:11] stdout/INFO File "python/freetoken/models/weight.py", line 240, in load_weight
[07:38:11] stdout/INFO File "python/freetoken/models/qwen3_5_moe/weight.py", line 255, in iter_weights
[07:38:11] stdout/INFO File "python/freetoken/models/qwen3_5_moe/weight.py", line 262, in _iter_shards
[07:38:11] stdout/INFO File "python/freetoken/models/qwen3_5_moe/weight.py", line 275, in freetoken.models.qwen3_5_moe.weight._iter_shards
[07:38:11] stdout/INFO RuntimeError: Attempted to access the data pointer on an invalid python storage.
[07:38:11] stdout/INFO
[07:38:11] stdout/INFO The above exception was the direct cause of the following exception:
[07:38:11] stdout/INFO
[07:38:11] stdout/INFO Traceback (most recent call last):
[07:38:11] stdout/INFO File "C:\Users\Admin\AppData\Roaming\uv\python\cpython-3.12-windows-x86_64-none\Lib\multiprocessing\process.py", line 314, in _bootstrap
[07:38:11] stdout/INFO self.run()
[07:38:11] stdout/INFO File "C:\Users\Admin\AppData\Roaming\uv\python\cpython-3.12-windows-x86_64-none\Lib\multiprocessing\process.py", line 108, in run
[07:38:11] stdout/INFO self._target(*self._args, **self._kwargs)
[07:38:11] stdout/INFO File "C:\Users\Admin\AppData\Local\FreeToken\venv\Lib\site-packages\freetoken\server\launch.py", line 81, in _run_scheduler
[07:38:11] stdout/INFO scheduler = Scheduler(args)
[07:38:11] stdout/INFO ^^^^^^^^^^^^^^^
[07:38:11] stdout/INFO File "python/freetoken/scheduler/scheduler.py", line 68, in freetoken.scheduler.scheduler.Scheduler.__init__
[07:38:11] stdout/INFO File "python/freetoken/engine/engine.py", line 419, in freetoken.engine.engine.Engine.__init__
[07:38:11] stdout/INFO File "python/freetoken/engine/engine.py", line 599, in freetoken.engine.engine.Engine._load_weights
[07:38:11] stdout/INFO File "C:\Users\Admin\AppData\Roaming\uv\python\cpython-3.12-windows-x86_64-none\Lib\contextlib.py", line 158, in __exit__
[07:38:11] stdout/INFO self.gen.throw(value)
[07:38:11] stdout/INFO File "python/freetoken/engine/engine.py", line 353, in _weight_load_context
[07:38:11] stdout/INFO freetoken.engine.engine.WeightLoadError: RuntimeError: Attempted to access the data pointer on an invalid python storage.
[07:38:12] health/ERROR WeightLoadError: RuntimeError: Attempted to access the data pointer on an invalid python storage.
Anything else
The same WeightLoadError also reproduces with gpt-oss-20b MXFP4 using the same default Desktop settings.
Before you start
What happened
FreeToken Desktop fails while loading model weights before inference starts.
I can reproduce the same error with two unrelated supported models and quantization formats:
Both fail with:
WeightLoadError:
RuntimeError: Attempted to access the data pointer on an invalid python storage.
The failure happens during weight materialization/loading, before inference starts.
Qwen3.6 fails in the qwen3_5_moe weight loader.
gpt-oss-20b fails independently in the gpt_oss weight loader / merged tensor path.
This does not appear to be VRAM OOM:
System:
FreeToken auto-selects Triton attention and hybrid MoE strategy.
Expected:
The supported models should load and start serving.
Actual:
The scheduler dies during weight loading with:
RuntimeError: Attempted to access the data pointer on an invalid python storage.
The same error reproducing with both Qwen3.6 NVFP4 and gpt-oss-20b MXFP4 makes this look like a Windows/Desktop weight-loader or storage-materialization issue rather than a checkpoint-specific problem.
Full server log attached below.
Desktop app version
v0.2.0-beta.23
OS
Windows 11
OS details
Windows 11 Pro 26H1, build 28000.1
GPU and driver
NVIDIA GeForce RTX 3070 8GB (8192 MiB), NVIDIA driver 616.92, WDDM mode, CUDA UMD 13.4
CPU and system RAM
AMD Ryzen 7 7800X3D 8-Core Processor, 32GB RAM DDR5
Checkpoint
nvidia/Qwen3.6-35B-A3B-NVFP4
Model settings
Default Desktop app settings.
FreeToken command:
ft serve --model C:\Users\Admin.freetoken\models\Qwen3.6-35B-A3B-NVFP4 --port 1919 --moe-backend auto --max-running-requests 4 --memory-ratio 0.85 --host 127.0.0.1
Resolved by FreeToken:
Engine log
Anything else
The same WeightLoadError also reproduces with gpt-oss-20b MXFP4 using the same default Desktop settings.