Lemura AI Labs
Models · 46Browse all models

Open weights, from day one.

Pick a model from the list to read its card here. You only leave for Hugging Face when you go to download.

46
open-weight models
506k
all-time downloads
40k
downloads this month

Speech

Languages the industry underserves

Dialect-aware recognition trained in-house. lemura-arabic-asr-lite is a ~115M-parameter FastConformer-CTC model fine-tuned on ~2,900 hours spanning MSA and the Gulf, Egyptian, Levantine and Maghrebi groups — leaderboard-tier accuracy at roughly 18× fewer parameters, running in real time on CPU.

Expert pruning

Frontier models, pruned down

Our own expert-pruning work on MiniMax-M2: lower latency, smaller memory footprint and higher throughput, as a drop-in replacement in most inference stacks. A 50% cut and a pruned Kimi K2 Thinking are in development, with a paper and expanded evaluation set to follow.

Refusal ablation

Refusal-ablation research, measured

Refusal directions removed surgically and reported with numbers rather than adjectives — the Gemma-4-12B build drops refusals from 99/100 to 12/100 at a KL divergence of 0.053 from the original, well under the damage threshold. It is also the artifact that runs refusal-free vision, audio and video today.

Quantization

Builds that fit the hardware you have

OptiQ, MLX, GGUF and mxfp4 builds of frontier models, cut to the bit-widths that make them serveable on consumer and Apple Silicon machines.

Coding

Coding models for local and air-gapped work

Coding-focused builds aimed at teams that cannot send source to a third-party API.

Most adopted

What the community actually runs

The releases with the deepest uptake — uncensored Llama and Gemma builds in GGUF, downloaded tens of thousands of times.

See the org on Hugging Face