Genomics · Actively researching
Population genomics for the Gulf
No genomic foundation model has been built for Arab or Emirati populations. That absence is the gap, and it is the one we are working on.
The gap
More than 80% of participants in genomic studies are of European ancestry. Every model built on that base — variant effect prediction, disease risk, drug response — carries the assumption forward, and there is no equivalent built for Arab or Emirati populations. Not a weaker one. None.
That is not an abstract fairness point once it reaches therapeutics. Which variants are pathogenic, which targets are worth pursuing, and how a patient metabolises a drug all vary by ancestry. A pipeline calibrated on European cohorts will rank the wrong candidates for a Gulf population, and nothing in the pipeline will flag it.
Why the Gulf specifically
The data is now among the best in the world. The Emirati programme has sequenced roughly 700,000 whole genomes toward its national population, and Qatar has built a reference panel aimed squarely at improving Middle Eastern ancestry prediction. These are national-scale cohorts, not convenience samples.
Gulf populations also have high rates of consanguinity, which enriches for rare recessive variants. That makes the region genuinely distinct scientifically — a different target space, not just a different sample of the same one.
Two populations, not one
Arab and South Asian are distinct ancestry groups and we treat them as separate programmes. Gulf populations sit in the Middle Eastern group; India, Pakistan, Bangladesh and Sri Lanka sit in the South Asian one. Their allele frequencies differ, and reference panels built for one do not transfer to the other.
Conflating them is a common and costly error — a model tuned on Indian cohorts and marketed for Emirati patients would reproduce exactly the mismatch this work exists to correct. Separate data, separate evaluation, separate claims.
Where we are
Early, and open about it. The first job is measurement: showing on public cohorts how far the available models fall short on Arab and South Asian samples, because that number is what makes the gap arguable rather than assertable.
The harder constraint is that national genomic programmes do not export their data, and should not. That points at the architecture we already build for — models compact enough to run inside the institution that holds the data, so the model travels and the genomes stay put.