Lemura AI Labs

Roadmap

Where this goes.

One timeline rather than a set of promises. It starts with what already exists, because that is the part you can check.

Now

Shipped

Published, downloadable, and measured. This is what exists today.

  1. Open weights on the Hub46Every model we train goes out with open weights and a readable card, in the formats people actually serve with.Browse the models
  2. Expert pruningFrontier mixture-of-experts models cut down until they fit ordinary hardware, as a drop-in replacement.Read more
  3. Speech for underserved languagesDialect-aware Arabic recognition at roughly 18× fewer parameters than the systems it sits beside, and Tamil alongside it.Read more
  4. Quantization across formatsGGUF, MLX and mxfp4 builds at a range of bit-widths, so the trade is legible before you download.Read more
  5. Measured refusal ablationRed-team artifacts published with both numbers — refusal rate and the distributional cost of removing it.Read more

Next

Actively researching

Underway, with nothing published yet. Status is stated rather than implied.

  1. Detection models that run inlineClassifiers small enough to read every prompt, retrieved document and tool call without the latency making them unaffordable.Read more
  2. More languagesThe same approach pointed at other languages with real speakers, scarce labelled audio and no efficient open model already serving them.Read more
  3. A genomics benchmark for the GulfQuantifying, on public cohorts, how far existing genomic models fall short on Arab and South Asian samples. The measurement comes before the model.Read more
  4. The pruning method, written upAn expanded evaluation set and a paper, so how much capability survives each cut is shown rather than asserted.Read more

Later

Direction

Where the work points. Honest about being further out than the rest.

  1. Onto reconfigurable siliconCompiling quantized models onto FPGAs, where streamlined low-precision logic suits a classifier far better than a GPU does.Read more
  2. A deployable applianceInspection that racks like a network device and takes new models the way a firewall takes new signatures.Read more
  3. A population model for the GulfThe model that does not exist yet — built where the data lives, because national genomic programmes do not export it.Read more

The through-line is the same at every stage

A model small enough to run where the data already sits. That is what makes inline inspection affordable, what lets a hospital or a national genomics programme keep its data in the building, and what makes compiling onto silicon worth attempting at all.

Talk to us about a deployment