Lemura AI Labs
All research

Quantization · Shipped · open weights

Quantization & efficient serving

Bit-width work that turns large models into something a laptop or a single consumer GPU can hold.

Quantization & efficient serving — the thread in short

Formats, not favourites

The same model goes out in the formats people actually serve with, at a range of effective bit-widths. Which one is right depends on the machine in front of you, not on our preference — so we publish several rather than picking one.

Each build states its scheme and effective bits per weight, so the trade is legible before anyone downloads tens of gigabytes.

Deliberately unglamorous

This is the largest part of the catalogue by count and the least novel. Its value is simple: someone can run a 27B or 122B model at all, on hardware they already have.

The toolchain is shared with the pruning work, and both point at the same destination — models that stay useful as they get cheaper to run.

Models from this thread