AI glossary

What is a token, a context window, a reasoning model, RAG or fine-tuning? 36 AI terms explained in plain language.

Large language model (LLM)
An AI model trained on huge amounts of text to predict the next piece of text. By doing this very well, it learns to answer questions, write, translate, summarise and write code. Claude, GPT and Gemini are LLMs.
Token
The unit a model reads and writes. A token is a word or a piece of a word: in English, 1,000 tokens is about 750 words. French and especially Arabic text usually needs more tokens for the same meaning. API prices are counted in tokens.
Context window
The maximum amount of text, counted in tokens, that a model can take into account at once: your instructions, the conversation so far, any documents you attach and its own answer. A 200K context window holds roughly a 500-page book.
Input and output price
API providers charge separately for the tokens you send (input) and the tokens the model writes (output), usually per million tokens. Output is typically 3 to 8 times more expensive than input, because generating text takes more computing.
Reasoning model
A model that can work through a problem step by step before giving its final answer, sometimes called "thinking". It is slower and uses more tokens, but is noticeably better at maths, logic, planning and hard coding tasks.
Multimodal
A model that works with more than one type of data: for example, it can read images, PDFs, audio or video in addition to text, or generate images as well as text.
Open weights
A model whose trained parameters are published so anyone can download it and run it on their own computer or server. Llama, Gemma, Qwen, DeepSeek and many Mistral models are open-weight. "Open weights" is not always the same as open source: the training data and code may stay private, and licences vary.
Parameters
The internal numbers a model learns during training. Model sizes are often given in billions of parameters (for example "70B"). More parameters usually means more capability but also more memory and cost to run.
Mixture of experts (MoE)
A model design where only part of the network, a few "experts", is used for each token. A model can have hundreds of billions of parameters in total while using only a fraction of them per step, which makes it cheaper and faster to run.
Prompt
The text you give a model: your question, instructions and any examples or documents. Clear prompts with context and examples usually get much better answers.
System prompt
Instructions set by the developer of an app that apply to the whole conversation, such as the assistant’s role, tone and rules. Users usually don’t see it.
Hallucination
When a model states something false with confidence, such as an invented quote, reference or number. Newer models hallucinate less, but important facts should still be checked, ideally against a source.
RAG (retrieval-augmented generation)
A technique where an app first searches your documents or the web for relevant passages, then gives them to the model with the question. The model answers from real sources, which reduces hallucinations and lets it use private or recent information.
Embeddings
Lists of numbers that represent the meaning of a text, so that similar texts get similar numbers. They power semantic search, recommendations and RAG.
Fine-tuning
Training an existing model further on your own examples so it adopts a specific style, format or expertise. It is often less needed today: good prompts and RAG solve most problems more cheaply.
AI agent
A model that works towards a goal over several steps on its own: it plans, uses tools (search, code, a browser, other apps), checks the results and continues until the task is done.
Tool use (function calling)
The ability of a model to ask an app to run a function, such as searching the web, querying a database or sending an email, and then use the result in its answer. It is the building block of AI agents.
MCP (Model Context Protocol)
An open standard, introduced by Anthropic, for connecting AI assistants to tools and data sources in a uniform way. An app that supports MCP can plug into any MCP server, such as a calendar, a code repository or a database.
Prompt caching
A discount for re-sending the same beginning of a prompt, such as a long document or system prompt. The provider stores it and charges a much lower "cached input" price the next times, often around a tenth of the normal price.
Benchmark
A standard test used to compare models, such as a set of maths problems, coding tasks or exam questions. Useful as a rough guide, but results can be gamed and don’t always match real-world use: testing on your own tasks is best.
Knowledge cutoff
The date after which a model has not seen training data. It won’t know about later events unless the app gives it search or documents.
Temperature
A setting that controls how random a model’s answers are. Low temperature gives focused, repeatable answers; high temperature gives more varied and creative ones.
API
A way for software to use a model directly, without a chat interface. Developers send a request with the prompt and receive the answer, and pay per token used.
Distillation
Training a smaller, cheaper model to imitate a larger one. Many "mini", "flash" or "lite" models are made this way: they keep much of the quality at a fraction of the price.
Quantization
Storing a model’s numbers with less precision (for example 4 bits instead of 16) so it needs less memory and runs on smaller hardware, at the cost of a small drop in quality. Common for running open models on a laptop.
Max output
The longest answer a model can write in one go, counted in tokens. A long report or a big piece of code needs a high limit; short replies do not.
Structured output
The model can answer in an exact format that you define, such as JSON with fixed fields, so that a program can read the answer without errors.
Free tier
A free version of the model offered through the API, usually with low limits on how many requests you can send. Good for trying it out, rarely enough for a product.
Uptime
The share of requests a provider answered successfully over the last 24 hours. 99% means about one failed request in a hundred.
Tokenizer
The part of a model that cuts text into tokens before the model reads it. Each model family has its own, which is why the same text can count a different number of tokens from one model to another.

Tests and scores

The independent tests behind the scores on this site. Each one checks a different skill; the overall index combines many of them.

Overall index
Epoch Capabilities Index: one score that combines dozens of benchmarks, so models tested on different tests can still be compared. Higher is better; it is an index, not a percentage.
GPQA Diamond
Very hard multiple-choice questions in biology, physics and chemistry, written by PhD experts. Random guessing scores 25%.
SWE-bench Verified
Can the model fix real bugs in real open-source software projects? Share of tasks solved.
AIME math
Competition-level math problems in the style of the American Invitational Mathematics Examination. Share solved.
FrontierMath
Original, unpublished math problems written by professional mathematicians, from advanced to research level. Share solved.
SimpleQA Verified
Short factual questions. Measures how often the model gets facts right instead of making something up.