How we work: sources and methods

Zolamir is built to be checked. Every number on the site comes from a public source that we name, is refreshed automatically, and is shown with the method behind it. This page explains how.

Where the data comes from

  • Models and prices: the public OpenRouter API, which lists the models offered by dozens of providers with their prices per million tokens, context windows and capabilities. Prices and uptime per provider come from the same API.
  • Quality scores: independent test results published by Epoch AI under the CC BY 4.0 licence: the Epoch Capabilities Index, SWE-bench Verified, GPQA Diamond, AIME, FrontierMath and SimpleQA Verified. We do not use scores announced by the AI companies themselves.
  • News: the public RSS feeds of AI labs and of the tech press in four languages, plus the most discussed AI stories on Hacker News. We only show a headline, a short excerpt and a link: the article always stays on the publisher’s site.
  • Research: papers highlighted by the Hugging Face community, and the open models trending on Hugging Face.

How often it is updated

An automatic job runs every four hours. It downloads the new data, records new models and price changes in the change log, translates new headlines and rebuilds the site. Quality scores are refreshed every day and our token measurements every month. Each week is summed up in This week in AI.

How the rankings work

Rankings sort models by one independent test at a time. When a model was tested with several settings, we keep its best result. New models appear once they have been tested, usually a few weeks after release. The "best for" pages combine these scores with current prices using simple, stated rules (for example, "capable" means an overall index of 140 or more).

How we measure the cost of each language

We wrote the same four texts in 15 languages and counted the tokens produced by each company’s public tokenizer. The ratio to English gives the extra cost of each language for each model family. Companies that do not publish their tokenizer get the average of the measured families, marked with ≈.

Translations

Headlines and summaries published in other languages are translated automatically by open-source models (Opus-MT and NLLB) that we run ourselves. They are always labelled "Translated from…" and link to the original article. Interface text, guides and explanations are written by us in each language.

Limits and corrections

Automatic data can be wrong or out of date: a provider may change a price before our next update, and any test favours some kinds of models. Always check the provider’s official page before a purchase decision. If you spot an error, tell us and we will correct it.