Technology

How large language models work, in plain English

A plain-English guide to what happens inside ChatGPT-style AI tools: tokens, next-word prediction, transformers, training, context windows, and why running them costs real money.

Illustrative cover: How large language models work, in plain English
Illustration: Pointales

A large language model (LLM) is a computer program trained on enormous amounts of text to do one thing very well: predict what piece of text should come next. Chatbots such as ChatGPT, Claude, Gemini and Llama-based assistants are built on LLMs. When you ask one a question, it does not look up an answer in a database. It generates a reply one small piece at a time, each time choosing a likely continuation of everything written so far. Almost everything an LLM does well, and almost every mistake it makes, follows from that one idea.

This guide walks through the machinery in plain English, from the smallest unit of text the model sees to the reason every reply costs someone money.

Step one: text becomes tokens

LLMs do not read letters or whole words. They read tokens: chunks of text that may be a whole word, part of a word, a punctuation mark or a space. A common word such as “the” is usually one token; a rarer word may be split into several.

Google’s developer documentation for its Gemini models gives a useful rule of thumb: a token is about four characters, and 100 tokens are roughly 60 to 80 English words. Other developers’ tokenizers differ in detail, but the ballpark is similar for English.

Tokens matter for three practical reasons:

  • Cost. Commercial AI services usually charge per token, separately for what you send in and what the model writes back. Anthropic’s documentation, for example, prices requests by input and output tokens.
  • Limits. Every model can only consider a fixed number of tokens at once (more on this under context windows below).
  • Language. Tokenizers trained mostly on English text split other languages into more pieces. A July 2026 preprint (not yet peer-reviewed) measured this for 14 Indian languages: under the tokenizer used by GPT-3.5 and GPT-4, the same content needed on average 8 times as many tokens as English, and 13 times for Malayalam. Newer multilingual tokenizers cut that gap by about 73%, the author found. For Indian users and companies, this can mean higher bills and less room in the model’s working memory for the same content.

Step two: predicting the next token

At its core, an LLM is a prediction engine. Given a sequence of tokens, it calculates a probability for every token in its vocabulary being the next one. Then the system picks one, adds it to the sequence, and repeats.

The GPT-3 research paper (OpenAI, 2020) describes its model in exactly these terms: “an autoregressive language model”, meaning one that generates each new token based on the tokens before it. The US National Institute of Standards and Technology (NIST) puts it just as plainly in its Generative AI risk profile (July 2024): “LLMs predict the next token or word in a sentence or phrase.”

It sounds too simple to produce essays, code and translations. The trick is scale. To predict the next word in a legal contract, a cricket report or a Python program well, a model has to capture a great deal about grammar, facts, style and reasoning patterns. Prediction at scale forces the model to learn those patterns.

Most chat systems do not simply pick the single most likely token every time. They add some controlled randomness (often adjusted by a setting called “temperature”), which is why the same question can get differently worded answers.

Step three: learning from training data

Before it can predict anything, the model has to learn from examples. This first stage is called pre-training. The model is shown huge amounts of text, asked to predict the next token, and adjusted slightly every time it is wrong. Repeat this trillions of times and the predictions get very good.

The scale is hard to picture. Meta’s Llama 3 paper (July 2024), one of the few that discloses such details, says the models were pre-trained on “about 15T multilingual tokens”, that is, roughly 15 trillion tokens of text, and that the largest version was trained on up to 16,000 Nvidia H100 graphics processors.

Two consequences follow:

  1. The model’s knowledge has a cut-off date. It only “knows” what was in its training data. Developers acknowledge this openly: Anthropic, for instance, describes its web search tool as a way to find information beyond the knowledge cutoff.
  2. The model absorbs the data’s gaps and errors. If a topic was rare, contested or wrong in the training text, the model’s predictions about it will be weaker.

Step four: the transformer and “attention”

Nearly all of today’s LLMs use a design called the transformer, introduced by Google researchers in the 2017 paper “Attention Is All You Need”. Its central idea is a mechanism called attention.

Here is the intuition. Take the sentence: “The bank refused the loan because it was short of capital.” To understand “it”, you need to connect it back to “the bank”, not “the loan”. Attention lets the model, for every token, weigh how relevant every other token in the passage is, and blend that information in. The word “it” can “pay attention” strongly to “bank” and weakly to everything else.

A transformer stacks many layers of this. Early layers might pick up simple relationships such as which words go together; later layers build up more abstract patterns. Nobody hand-codes these patterns. They emerge from training.

The paper’s other practical advantage was speed. Earlier designs read text one word after another, which was slow to train. The transformer processes a whole passage at once, which made it, in the authors’ words, “more parallelizable and requiring significantly less time to train”. That property is what made training on trillions of tokens feasible.

Step five: parameters, the model’s adjustable dials

When people say a model has “175 billion parameters”, they mean it has 175 billion numbers that are tuned during training. Each parameter is like a tiny dial. Together they encode everything the model has learned. GPT-3 had 175 billion parameters; the largest Llama 3 model has 405 billion.

More parameters generally means more capacity to learn, but also more computing power to train and run. Size is not everything, though, as the next section shows. Many developers no longer disclose parameter counts for their flagship models, so treat any number you see quoted without a source with caution.

Step six: from raw predictor to helpful assistant

A pre-trained model is a powerful text predictor, but not yet a useful assistant. Ask it a question and it might continue with more questions, because that is what often follows a question on the internet. Turning it into a chatbot takes further training, usually in two stages.

Fine-tuning means continuing training on a smaller, carefully chosen dataset. For an assistant, that dataset is typically examples of good answers to instructions, often written by people.

Reinforcement learning from human feedback (RLHF) goes one step further. People compare several model answers to the same prompt and rank them. Those rankings are used to train the model to produce the kind of answers people prefer.

The best-known description of this process is OpenAI’s 2022 InstructGPT paper. Its headline result shows how much this stage matters: “outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters.” The authors also reported improvements in truthfulness and reductions in toxic output.

StageWhat the model learns fromWhat it adds
Pre-trainingTrillions of tokens of general textLanguage, facts, patterns, broad knowledge
Fine-tuningSmaller sets of example instructions and good answersFollows instructions, answers in a useful format
RLHF (or similar)Human rankings of alternative answersAnswers people judge more helpful, honest and safe

Post-training also shapes a model’s weaknesses. A September 2025 research paper published by OpenAI, “Why Language Models Hallucinate”, argues that common training and evaluation methods reward confident guessing over admitting uncertainty. We explain what that means for you in why AI chatbots “hallucinate”, and how to catch it.

Step seven: the context window, the model’s working memory

The context window is the amount of text a model can consider at once when writing a reply. Anthropic’s documentation defines it as “all the text a language model can reference when generating a response, including the response itself”, and describes it as the model’s “working memory”, distinct from the training data.

Everything counts against this limit: your question, any documents you paste in, the conversation so far, hidden instructions from the app, and the reply being written. Context windows have grown fast. As of October 2026, Anthropic lists windows of up to 1 million tokens for some of its models, and Meta’s Llama 3 paper described a 128,000-token window in 2024.

Bigger is not automatically better. The same Anthropic page notes that as the token count grows, “accuracy and recall degrade”, a problem it calls “context rot”. In practice, a model can lose track of a detail buried in a very long document. Giving it only what it needs often works better than pasting in everything.

A model also does not remember you between conversations unless the app deliberately stores and re-sends earlier material. Each reply is generated from what is inside the context window at that moment.

Step eight: inference, and why every answer costs money

Using a trained model to generate text is called inference. Every token in a reply requires running the input through the entire model, with billions of calculations. That is why LLMs run on graphics processing units (GPUs), specialised chips that perform huge numbers of calculations in parallel.

That hardware is expensive, scarce and power-hungry, which is why it has become a matter of national policy. The IndiaAI Mission, approved with an outlay of about ₹10,372 crore, had expanded shared compute capacity to more than 45,000 GPUs as of June 2026, according to a PIB factsheet dated 13 August 2026. A February 2026 PIB backgrounder said the IndiaAI Compute Portal offered GPU access at subsidised rates of under ₹100 per hour, compared with global rates above ₹200 per hour. For the chip side of this story, see India’s semiconductor push, explained.

The cost structure explains several things you may notice as a user:

  • Free tiers have usage caps, and longer answers or bigger documents use them up faster.
  • Long conversations get more expensive to run, because the whole history is processed again with each reply.
  • Faster or cheaper “mini” versions of models exist because smaller models need less computation per token.

What LLMs are good at, and where they fail

Because LLMs are pattern predictors trained on vast text, their strengths and weaknesses are fairly predictable.

Generally strong at:

  • Drafting, rewriting, summarising and changing the tone of text.
  • Explaining well-documented concepts in simpler terms.
  • Translation between widely used languages.
  • Writing and explaining common kinds of code.
  • Working with material you supply in the context window, such as a document to summarise.

Generally weak at, or risky for:

  • Facts they were not trained on, including anything after their cut-off date, unless the app adds a search tool.
  • Precise recall of rare facts, such as a specific person’s birthday, an exact citation or a niche statistic. Here the model may produce a fluent, confident and wrong answer. NIST calls this “confabulation” and notes it is “a natural result of the way generative models are designed”.
  • Explaining their own reasoning reliably. NIST also warns that models can produce “confabulated logic or citations” that seem to justify an answer even when the answer is wrong.
  • Taking actions in the world. A plain LLM only produces text. Systems that let a model use tools, browse or run tasks over many steps are called agents, and they bring a different set of risks; see what is an AI agent?

A useful rule: treat an LLM’s output as a strong first draft from a well-read assistant who never checks sources. Use it for speed, and check anything that matters.

Glossary

TermPlain-English meaning
Large language model (LLM)A program trained on huge amounts of text to predict the next piece of text
TokenThe chunk of text a model reads and writes; often part of a word
Pre-trainingThe first, largest training stage, on general text
ParametersThe billions of internal numbers adjusted during training
TransformerThe model design behind most LLMs, introduced in 2017
AttentionThe mechanism that lets each token weigh the relevance of every other token
Fine-tuningFurther training on a smaller, targeted dataset
RLHFReinforcement learning from human feedback: training on human rankings of answers
Context windowHow much text the model can consider at once, including its reply
InferenceRunning a trained model to generate output
GPUA chip that performs many calculations in parallel; the workhorse of AI
Hallucination / confabulationConfident output that is false or unsupported
AgentA system in which a model uses tools in a loop to complete a goal

The point: An LLM is a very large next-token predictor, shaped first by trillions of tokens of text and then by human feedback into a helpful assistant. That design makes it fluent and broadly useful, but it also means it can produce confident errors and knows nothing beyond its training and what you give it. Use it as a fast drafter and explainer, and verify anything you will act on.

Sources

Chander Prakash

Chander Prakash

Chander Prakash is the founder and editor of Pointales. He reviews every story before it is published and sets the publication's editorial standards, with a focus on clear, well-sourced explanations of business and technology.