Skip to content
HansimovPublic

About

Flash H3 inference with PyTorch and Triton

Resources

Stars

1 star

Watchers

1 watching

Forks

Repository files navigation

Vflash

Native MiniMax H3 inference with explicit GPU profiles. Vflash uses PyTorch and Triton, with its own denoising runtime and support for pinned LightX2V Turbo LoRAs.

Documentation · Get started · Release notes · 中文

0.6.13. H100 SXM and PRO 6000 Server/Workstation I2VA preview, CPU hybrid preparation and partition-aware discovery. Other new cards/MIG modes remain unqualified; existing SM89 defaults are unchanged.

0.6.12. Explicit single-SM89 hybrid/Veda supports qualified fifteen-second and nine-image requests within joint resource limits.

0.6.9. Optional FlashVSR upscaling preserves complete video and original audio; H3 defaults remain unchanged.

0.6.10. Explicit Veda sparse attention for original v0.1 and the hybrid model on one SM89, with a pinned standalone core installer, protected connections, and actual kernel/coverage receipts. Dense remains the automatic choice for these profiles.

0.6.8. Explicit hybrid references share the original FL-v0.1 backbone and add only Ref block modulation tables. One Python session serves real reference images and keyframes, with distinct conditioning identities. Original LightX v0.1 four-step keyframes have independent pinned profiles and complete SM89 media evidence. Dense remains their default; optional Sol was slower in three matched controls. The separate VDN/SelfLift pipeline, official Base16, attention-adapter and VAE interfaces remain available.

Generate synchronized video and audio through a complete Python or container pipeline. Turbo profiles create five-second MP4s from text, one to three images, or a short reference video. Official Base16 keyframe profiles accept a first frame, a last frame, or both, at integer durations from five through ten seconds. The optional 544p keyframe Turbo pairs also accept first-only, last-only and both endpoints with bounded five/ten-second evidence. One prepared keyframe pipeline switches among I2VA, L2VA and FL2VA without reloading weights, but not across adapter versions.

The native core targets RTX 3080 20 GB (SM86) and RTX 4090 48 GB (SM89). Native Base16 defaults to approximate Sol attention on single-SM89 official Base16. SM86, cooperating pairs and Turbo profiles default to dense attention. Original LightX v0.1 keyframes additionally permit explicit Sol on one SM89; other Turbo profiles do not. Python, CLI and the native HTTP service share this policy; --attention-backend torch-flash explicitly retains dense execution. Standard Docker builds include the pinned Sol dependency. A Python install needs python -m vflash.install_sol after its GPU extras. Missing dependencies fail clearly, never silently select another backend. Cross-step caches and quantized communication remain outside the defaults. Full media-quality equivalence is not claimed. See what was adopted, deferred or rejected.

The ten-second Base16 boundary has bounded complete-request evidence on the listed hardware, but it is not a guarantee for every canvas or prompt. A matching SM89 pair is an explicit single-request latency option; two independent workers remain the throughput default. Ref/T2 and reference-video profiles retain their separate five-second contracts. Read the qualification boundary.

Check your setup

Python 3.11 or newer is required. The base install does not download model weights or PyTorch.

git clone --branch v0.6.12 --depth 1 https://github.com/Hansimov/vflash.git
cd vflash
python -m venv .venv
source .venv/bin/activate
python -m pip install -e .

vflash doctor
vflash profiles
vflash plan ref2va-turbo4-exact-sm89 --gpu 0
GPU configuration Released profiles Weight placement
One RTX 4090 48 GB Ref2VA Turbo4 / Turbo8; T2VA Turbo4; Base16 and 544p Turbo keyframes Sol uses block streaming; dense native execution supports residency
One RTX 3080 20 GB Ref2VA Turbo4 Streamed from host RAM
Two RTX 3080 20 GB GPUs Ref2VA Turbo4; T2VA Turbo4; Base16 I2VA/L2VA/FL2VA Shared host weights; cooperative sequence-head
Two RTX 4090 48 GB GPUs Base16 I2VA/L2VA/FL2VA Opt-in cooperative latency path; not the throughput default

For a cooperating 3080 pair, select the peer explicitly with --peer-gpu 1. T2VA requires sequence-head; native Ref4 also offers tensor. The engine never selects another device automatically.

For the native denoiser, allow 64 GiB or more of available system memory per worker for the tested workload, plus headroom. The complete pipeline also keeps encoders and decoders in host memory and needs a larger budget. Larger inputs need separate capacity checks. See hardware, adapters and quality limits.

Build with Vflash

Turbo4 and Turbo8 are distilled adapters. Exact attention is not a base-model quality guarantee, and different GPU or parallel configurations need not produce bitwise-identical results. Complete generation uses explicit fixed-profile assets. Dual-SM86 Ref4 requires sequence-head; single-SM86 and other native-only Turbo profiles retain their latent interface. W8 and arbitrary adapter or mode switching are outside the supported profiles. Read the validation scope.

Contribute

python -m pip install -e '.[dev,server]'
pre-commit install
pre-commit run --all-files
pytest

Build the bilingual docs with npm ci and npm run docs:build. See the contribution guide and contributor map for evidence, privacy and release requirements.

License

Vflash source is Apache-2.0. Models and adapters retain their own licenses and terms; this repository contains no model weights. The LightX2V inference framework is not a runtime dependency.

Thanks to MiniMax H3, LightX2V, and the PyTorch, Triton and CUDA communities. Sources and acknowledgements.

About

Flash H3 inference with PyTorch and Triton

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages