bstnxbt (Open Source)

2026bstnxbt (Open Source)

dflash-mlx vlatest

Paid
Apple Silicon
MLX

dflash-mlx brings lossless DFlash speculative decoding to Apple Silicon via MLX, achieving up to 4.37x token generation speedup while guaranteeing every emitted token is verified against the target model. Uses a small ~1B draft model to generate 16 tokens in parallel via block diffusion, then verifies them in a single forward pass. Ideal for local LLM inference on Mac hardware.

Up to 4.37x speedup on Qwen3.5-9B with 86-91% acceptance rates across all tested models
Lossless output: every token is verified via greedy acceptance before being committed — no hallucinated tokens
OpenAI-compatible server (dflash serve) supporting streaming, tool calls, and chat templates; works with OpenCode, aider, Continue, Open WebUI
PricingFree (Open Source)
Versionlatest

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode