Home › AI Explained
Hub · AI Explained
How AI works: plain-English guides
Six short guides on large language models, training, tokens, inference, GPUs and open versus closed models. Each ends with a one-sentence summary and links to its sources.
Last reviewed Oct 3, 2026
Last reviewed Oct 3, 2026. Plain-English guides; terms link to the glossary.
Where to start. The short version of how a chatbot works is on How AI works. These guides go one level deeper, one idea at a time. Each ends with a one-sentence summary, and terms link to the glossary.
1. What is a large language model?
Modern general-purpose AI is built with deep learning: instead of hand-written rules, a model learns patterns from very large amounts of data. Its inner workings are layers of connected units loosely inspired by brain cells; the strength of each connection is called a weight (the model’s parameters), and training adjusts them. Most leading systems use an architecture called the transformer, whose “attention” mechanism helps the model focus on the most relevant parts of its input.1
A large language model (LLM) applies this to text. OpenAI describes the first stage of training as predicting the next word in huge amounts of text.2 A chatbot such as ChatGPT is a system: one or more models plus a chat interface, content filters, web access and tools.1
In one sentence: an LLM is a very large set of learned numbers that predicts what text should come next, and a chatbot is that model wrapped in a product.
2. How models are trained
The International AI Safety Report 2026 lays out six stages in building a general-purpose AI system.1
| Stage | What happens |
|---|---|
| 1. Collect and clean data | Developers and data workers collect, clean and filter training data; this can be labor intensive. |
| 2. Pre-training | The model sees billions or trillions of examples and produces a “base model”. It takes weeks or months on tens or hundreds of thousands of GPUs or TPUs. |
| 3. Post-training and fine-tuning | The base model is refined for tasks. Methods include supervised fine-tuning and reinforcement learning, such as reinforcement learning from human feedback. |
| 4. System integration | Models are combined with other parts (filters, tools, a chat interface) into an AI system. |
| 5. Deployment | The system is released through apps or an API (a way for other software to use it). |
| 6. Monitoring and updates | Developers gather feedback and make improvements without repeating full pre-training. |
Cost: the report says building a leading system from scratch now costs hundreds of millions of US dollars, and that current frontier training runs cost about $500 million in computing alone, with next-generation models projected at $1 billion to $10 billion.1 It also notes that most recent gains come from post-training and from extra computing at the time of use, not from making models bigger alone.1
A technique called distillation trains a smaller “student” model on the outputs of a larger “teacher” model; the report cites DeepSeek, which reportedly fine-tuned a model for about $10,000 this way, though its pre-training costs were not reported.1 See also training and fine-tuning in the glossary.
In one sentence: training is a multi-stage process, with a huge first stage on raw data and later stages that shape the model into a useful assistant.
3. Tokens
Models do not read words; they read tokens. OpenAI says a token can be a character, part of a word, a whole word or punctuation, and that a token count is not a word count. As rough English estimates it gives: 1 token ≈ 4 characters, 1 token ≈ three-quarters of a word, 100 tokens ≈ 75 words. Counts differ by model and language.3
Tokens also set limits and prices. Input (prompt) tokens and output tokens are counted separately, and reasoning models use extra hidden “reasoning tokens” that are billed as output even though you do not see them.3 The context window is how many tokens a model can consider at once: the model pages we read list 1.05 million tokens for OpenAI’s GPT-6 family and 1 million for most current Claude models.4,5
Worked example (our arithmetic). Developers pay per million tokens (MTok). Anthropic lists Claude Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens.5 A request with 1,500 input tokens (about 1,100 words) and 500 output tokens costs 1,500 × $2 ÷ 1,000,000 = $0.003 plus 500 × $10 ÷ 1,000,000 = $0.005, or $0.008 in total. Prices change often; app subscriptions are priced separately (see Which AI?).
In one sentence: a token is a piece of text about three-quarters of a word long, and models are limited and priced by how many tokens go in and come out.
4. Inference
NVIDIA defines inference as the process where a trained model generates new outputs by applying what it learned to new data, and notes that inference cannot happen without training.6 Every time you press Enter in a chatbot, that is inference.
Since 2025, a major change has been “reasoning” models, which spend more computing at the time of use to produce intermediate steps (“chains of thought”) before a final answer. The safety report says this “inference-time scaling” has led to large gains in mathematics, software engineering and science, and costs more computing per answer.1 See inference and reasoning model. For how much electricity AI answers use, see Data Centers and the water myth.
In one sentence: training builds the model once; inference is using it, again and again, and reasoning models use more of it per answer.
5. GPUs and chips
AI training and inference need enormous numbers of calculations done at the same time. The safety report describes GPUs (graphics processing units) and TPUs (tensor processing units) as specialized computer chips designed to rapidly process many such calculations, and says pre-training uses tens or hundreds of thousands of them.1 “Compute” is the shorthand for these chips and the infrastructure that runs them.
The report says the compute used to train the most compute-intensive models has grown about 5 times a year, and that the largest training runs have likely exceeded 1026 floating-point operations.1 The AI Index 2026 puts AI data-center power capacity at 29.6 GW.7 See GPU, TPU, accelerator, and the Data Centers hub for the buildings that house them.
In one sentence: GPUs are chips built to do many calculations in parallel, which is what AI needs, and they are the main reason AI data centers use so much power.
6. Open vs closed models
The safety report distinguishes open-weight releases (the trained numbers can be downloaded) from closed-weight ones (access only through an API or app).1 Open-weight does not automatically mean “open source”. The Open Source Initiative’s definition of Open Source AI requires the freedoms to use, study, modify and share a system, with access to the preferred form for making modifications, which includes detailed information about the training data and the code as well as the weights.8 Terms are in the glossary: open-weight model and open-source AI.
Examples, from official pages
- OpenAI released two open-weight models, gpt-oss-120b and gpt-oss-20b, under the Apache 2.0 license on Aug 5, 2025.9
- Mistral’s docs list several of its models, including Mistral Large 3, as Apache 2.0 open-weight.10
- Meta’s official Hugging Face page for Llama says you must accept the license terms and acceptable use policy to access its models.11
The trade-offs the sources describe
- The safety report says open-weight models bring significant research and commercial benefits, particularly for less-resourced groups, but cannot be recalled after release and have safeguards that are easier to remove.1
- The AI Index 2026 reports that the average score on the Foundation Model Transparency Index fell to 40 points from 58 a year earlier, and that the most capable models often disclose the least.7
In one sentence: open-weight models can be downloaded and run by anyone, closed models are reached only through the maker’s service, and “open source” is a stricter claim than either.
Keep reading
- How AI works (the short version)
- AI Glossary
- Companies & Models
- Data Centers
- AI Safety & Security
- Prompting
- Try it
What we left out. We did not give parameter counts or training-compute figures for specific commercial models, because the companies do not publish them on the pages we read (the AI Index notes that frontier developers are increasingly withholding them). We did not explain model internals beyond what the safety report states. See How we verify; report an error through the corrections page.
Sources for this page
- International AI Safety Report 2026. International AI Safety Report (chair: Yoshua Bengio; secretariat: UK AI Security Institute).
- Why language models hallucinate. OpenAI.
- Understanding and counting tokens. OpenAI Help Center.
- Models | OpenAI API. OpenAI.
- Models overview. Anthropic (Claude Platform Docs).
- What’s the Difference Between Deep Learning Training and Inference?. NVIDIA Blog.
- Inside the AI Index: 12 Takeaways from the 2026 Report. Stanford Institute for Human-Centered AI.
- The Open Source AI Definition 1.0. Open Source Initiative.
- Introducing gpt-oss. OpenAI.
- Models Overview. Mistral AI (docs).
- Meta Llama (official Hugging Face organization). Meta.