lfm2-350m-webgpu โ€” GPTQ-int4 weights

Custom GPTQ-int4 quants (group 32, byte-sliced nibbles, dedicated q4 lm_head, q8 embedding gather) of LiquidAI/LFM2-350M and LiquidAI/LFM2.5-350M, in the wire format of the lfm2-350m-webgpu browser runtime (hand-written WGSL kernels, no ONNX / transformers.js / GGUF).

  • 269 MB per model; ~390 tok/s single-stream decode and ~1640 tok/s batched (64 streams) on an Apple M4 (10-core GPU) in Chrome.
  • LFM2: ppl +3.2% vs f32 on held-out text.
  • LFM2.5: GPTQ calibrated on chat-formatted passages (raw-prose calibration degrades an instruct checkpoint badly): ppl +4.9% vs f32 in its native chat format.
  • The runtime is verified bit-exact against a PyTorch dequantized-weights reference, greedy token-for-token.

f16/q8 blobs are reproducible with the repo's tools/ scripts.

Try it: https://huggingface.co/spaces/borkiss/lfm2-350m-webgpu-demo ยท run the benchmark on your GPU and share the .log from the /web/bench.html page.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for borkiss/lfm2-350m-webgpu

Finetuned
(63)
this model