Lemura AI Labs
Models · 46lemura-arabic-asr-lite

lemura-arabic-asr-lite

automatic-speech-recognitioncc-by-4.0
Download on Hugging Face
Downloads
140
This month
113
Updated
1mo ago

lemura-arabic-asr-lite

Lemura Labs' Arabic-first speech-recognition model — leaderboard-grade at a fraction of the size.

Arabic accuracy, without the GPU bill.

License Task Format Params Open-AR-ASR Efficiency Runtime

Overview · Benchmarks · Efficiency · Inference · Live Demo · Sibling model


What it is

lemura-arabic-asr-lite is a compact multi-dialect Arabic ASR model built for accuracy and efficiency. It is a FastConformer-CTC acoustic model adapted in-house from the NVIDIA FastConformer foundation — the contribution is the adaptation (the Arabic data curriculum and dialect coverage), not the foundation.

  • Small and fast — ~115M parameters; runs comfortably on CPU and in real time, no GPU required.
  • Dialect-aware — fine-tuned on ~2,900 hours of Arabic spanning MSA and the Gulf, Egyptian, Levantine and Maghrebi dialect groups, not MSA-only.
  • Robust on real audio — strongest on broadcast, conversational and Gulf/MSA speech.
  • Open and simple — a single .nemo file, loadable in a few lines with NVIDIA NeMo.

The result is leaderboard-grade Arabic ASR at edge scale#2 of 36 systems on the Open Universal Arabic ASR Leaderboard, ahead of every 2B–30B audio-LLM evaluated, at ~115M parameters.

Model summary

Modellemura-arabic-asr-lite — compact multi-dialect Arabic ASR
TaskAutomatic speech recognition (audio → text)
ApproachDiscriminative ASR — FastConformer encoder + CTC decoder
Trainingadapted from the NVIDIA FastConformer foundation; ~2,900 h Arabic fine-tuning across 5 dialect groups
Total parameters~115,000,000 (~0.12B)
Checkpoint size~0.4 GB (single .nemo file)
Audio input16 kHz mono (auto-resampled)
LanguagesArabic — MSA + Gulf / Egyptian / Levantine / Maghrebi
RuntimeNVIDIA NeMo — CPU · GPU · real-time
LicenseCC-BY-4.0

Benchmarks

Arabic dialectal ASR is hard — heavily dialectal, conversational, code-switched speech is the frontier for every system. On the Open Universal Arabic ASR Leaderboard, lemura-arabic-asr-lite ranks #2 of 36 systems with an average WER of 25.08 % — ahead of every model evaluated except one, including systems 17× to 250× larger.

Open Universal Arabic ASR Leaderboard — full standings

Per-dataset WER % across all six leaderboard test sets. Lower is better; Avg WER is the ranking metric. Our rows were produced with the official leaderboard code (same normalizer, same WER metric). Ours in bold.

#ModelParamsAvg WERSADACV-18MASC-cleanMASC-noisyMGB-2Casablanca
1lemuralabs/lemura-arabic-asr-lite (Ours)0.12B25.0837.289.747.2723.6514.3358.24
2CohereLabs/cohere-transcribe-arabic-07-2026~2B25.8737.475.8219.6027.0715.5449.71
3omnilingual-asr/omniASR_LLM_7B7B28.3241.618.7519.6929.2914.1356.46
4omnilingual-asr/omniASR_LLM_3B3B29.9646.189.1519.9030.0314.2260.27
5omnilingual-asr/omniASR_LLM_1B1B29.9643.849.5520.0330.2615.3460.68
6CohereLabs/cohere-transcribe-03-2026~2B30.6760.118.178.6619.0125.3362.71
7Qwen/Qwen3-Omni-30B-A3B-Instruct30B30.7144.8211.4621.4730.8513.0962.55
8nvidia-conformer-ctc-large-arabic (lm)0.6B32.9144.528.8023.7434.2917.2068.90
9omnilingual-asr/omniASR_LLM_300M0.3B32.9651.3812.0320.6632.4516.5864.64
10google/gemma-4-E4B-it4B32.9843.4019.6524.8633.5917.7258.63
11Qwen/Qwen3-ASR-1.7B1.7B33.3645.5316.9024.3734.2916.5764.47
12mistralai/Voxtral-Small-24B-250724B34.4750.8215.2523.9634.4316.0366.30
13nvidia-conformer-ctc-large-arabic (greedy)0.6B34.7447.2610.6024.1235.6419.6971.13
14google/gemma-4-E2B-it2B35.8746.2323.7627.4736.1520.7260.87
15openai/whisper-large-v31.5B36.8655.9617.8324.6634.6316.2671.81
16omnilingual-asr/omniASR_CTC_3B3B37.7869.8514.1921.4834.6018.9667.58
17omnilingual-asr/omniASR_CTC_7B7B38.1272.6912.4721.0835.0420.4367.02
18facebook/seamless-m4t-v2-large2.3B38.1662.5221.7025.0433.2420.2366.25
19omnilingual-asr/omniASR_CTC_1B1B39.2971.4217.5522.7635.7319.9668.32
20openai/whisper-large-v3-turbo0.8B40.0560.3625.7325.5137.1617.7573.79
21openai/whisper-large-v21.5B40.2057.4621.7727.2538.5525.1771.01
22Qwen/Qwen3-ASR-0.6B0.6B42.1953.7528.2831.3442.6325.4571.68
23openai/whisper-large1.5B42.5763.2426.0428.8940.7924.2872.18
24mistralai/Voxtral-Mini-3B-25073B42.5863.6522.1228.3741.2722.5677.52
25asafaya/hubert-large-arabic-transcribe0.3B45.5067.828.0132.9450.1637.5176.53
26openai/whisper-medium0.8B45.5767.7128.0729.9942.9129.3275.44
27nvidia-Parakeet-ctc-1.1b-concat1.1B46.5470.7026.3430.4945.9524.9480.80
28omnilingual-asr/omniASR_CTC_300M0.3B46.6578.1127.9028.4043.2626.8575.35
29nvidia-Parakeet-ctc-1.1b-universal1.1B51.9673.5840.0136.1650.0330.6881.30
30microsoft/VibeVoice-ASR52.9969.8344.2532.9552.4325.1093.37
31facebook/mms-1b-all1B54.5477.4826.5238.8257.3339.1687.95
32openai/whisper-small0.24B55.1378.0224.1835.9356.3648.6487.64
33whitefox123/w2v-bert-2.0-arabic-40.6B58.1387.3441.7937.8253.2840.6687.88
34jonatasgrosman/wav2vec2-large-xlsr-53-arabic0.3B60.9886.8223.0042.7564.2756.2992.72
35speechbrain/asr-wav2vec2-commonvoice-14-ar0.1B65.7488.5429.1749.1069.5764.3793.68

Bold = our model and our best-in-column results. lemura-arabic-asr-lite is the single best system on MASC-clean (7.27) and MASC-noisy (23.65) of every model evaluated, and the highest-ranked system below 1B parameters by a wide margin. Casablanca (Moroccan Darija) is the hardest set for every system.

Competitor rows are the published Open Universal Arabic ASR Leaderboard standings, reproduced for context; they are not our measurements.

Efficiency

The systems around us on the leaderboard are large generative audio-LLMs (2–30B parameters) that need GPUs. lemura-arabic-asr-lite reaches the same accuracy tier with a ~115M-parameter CTC model:

lemura-arabic-asr-liteTypical top-5 systems
Parameters~115M2B – 30B
HardwareCPU or GPUGPU
LatencyReal-timeSeconds / clip
Footprint~0.4 GB4 – 60 GB
Avg WER25.0824.78 – 30.71

That makes it practical for on-device, low-cost and high-throughput Arabic transcription where the larger models are impractical.

Inference

import nemo.collections.asr as nemo_asr

model = nemo_asr.models.ASRModel.restore_from("asr_final.nemo")
print(model.transcribe(["audio.wav"]))   # 16 kHz mono

Try it live, no install: lemura-arabic-asr demo

Languages, dialects and tasks

LanguagesArabic
DialectsMSA, Gulf/Khaleeji, Egyptian, Levantine, Maghrebi
TasksTranscription (audio → text)
Audio16 kHz mono, auto-resampled

Intended use and limitations

Intended: transcribing Arabic speech across dialects — call centres, voice notes, media captioning, voice interfaces, and any deployment where CPU-only or real-time operation matters.

Limitations:

  • Maghrebi/Darija (Casablanca, 58.24 WER) remains hard, as it does for every system on the leaderboard.
  • Very noisy or far-field audio degrades accuracy.
  • CTC output has no built-in punctuation or diacritic restoration.
  • May reflect biases present in the training corpora.

For a larger, generative alternative with broader dialect fine-tuning, see lemura-arabic-asr-qwen3 (1.7B, audio-LLM). Note its published WER is an in-domain figure and is not directly comparable to the zero-shot leaderboard numbers above.

License

CC-BY-4.0.

Citation

@misc{lemura_arabic_asr_lite_2026,
  title  = {lemura-arabic-asr-lite: Compact Multi-Dialect Arabic Speech Recognition},
  author = {Lemura Labs},
  year   = {2026},
  url    = {https://huggingface.co/lemuralabs/lemura-arabic-asr-lite}
}

About Lemura Labs

Arabic-first, efficiency-first speech intelligence.

Lemura Labs builds compact, deployable speech and language models — accuracy at a size and cost that works outside the datacentre.