Models · 46Browse all models
Open weights, from day one.
Pick a model from the list to read its card here. You only leave for Hugging Face when you go to download.
- 46
- open-weight models
- 506k
- all-time downloads
- 40k
- downloads this month
Speech
Languages the industry underserves
Dialect-aware recognition trained in-house. lemura-arabic-asr-lite is a ~115M-parameter FastConformer-CTC model fine-tuned on ~2,900 hours spanning MSA and the Gulf, Egyptian, Levantine and Maghrebi groups — leaderboard-tier accuracy at roughly 18× fewer parameters, running in real time on CPU.
Expert pruning
Frontier models, pruned down
Our own expert-pruning work on MiniMax-M2: lower latency, smaller memory footprint and higher throughput, as a drop-in replacement in most inference stacks. A 50% cut and a pruned Kimi K2 Thinking are in development, with a paper and expanded evaluation set to follow.
Refusal ablation
Refusal-ablation research, measured
Refusal directions removed surgically and reported with numbers rather than adjectives — the Gemma-4-12B build drops refusals from 99/100 to 12/100 at a KL divergence of 0.053 from the original, well under the damage threshold. It is also the artifact that runs refusal-free vision, audio and video today.
Quantization
Builds that fit the hardware you have
OptiQ, MLX, GGUF and mxfp4 builds of frontier models, cut to the bit-widths that make them serveable on consumer and Apple Silicon machines.
Coding
Coding models for local and air-gapped work
Coding-focused builds aimed at teams that cannot send source to a third-party API.
Most adopted
What the community actually runs
The releases with the deepest uptake — uncensored Llama and Gemma builds in GGUF, downloaded tens of thousands of times.