A large language model (LLM) is a computer program trained on enormous amounts of text to do one thing very well: predict what piece of text should come next. Chatbots such as ChatGPT, Claude, Gemini and Llama-based assistants are built on LLMs. When you ask one a question, it does not look up an answer in a database. It generates a reply one small piece at a time, each time choosing a likely continuation of everything written so far. Almost everything an LLM does well, and almost every mistake it makes, follows from that one idea.
This guide walks through the machinery in plain English, from the smallest unit of text the model sees to the reason every reply costs someone money.
Step one: text becomes tokens
LLMs do not read letters or whole words. They read tokens: chunks of text that may be a whole word, part of a word, a punctuation mark or a space. A common word such as “the” is usually one token; a rarer word may be split into several.
Google’s developer documentation for its Gemini models gives a useful rule of thumb: a token is about four characters, and 100 tokens are roughly 60 to 80 English words. Other developers’ tokenizers differ in detail, but the ballpark is similar for English.
Tokens matter for three practical reasons:
- Cost. Commercial AI services usually charge per token, separately for what you send in and what the model writes back. Anthropic’s documentation, for example, prices requests by input and output tokens.
- Limits. Every model can only consider a fixed number of tokens at once (more on this under context windows below).
- Language. Tokenizers trained mostly on English text split other languages into more pieces. A July 2026 preprint (not yet peer-reviewed) measured this for 14 Indian languages: under the tokenizer used by GPT-3.5 and GPT-4, the same content needed on average 8 times as many tokens as English, and 13 times for Malayalam. Newer multilingual tokenizers cut that gap by about 73%, the author found. For Indian users and companies, this can mean higher bills and less room in the model’s working memory for the same content.
Step two: predicting the next token
At its core, an LLM is a prediction engine. Given a sequence of tokens, it calculates a probability for every token in its vocabulary being the next one. Then the system picks one, adds it to the sequence, and repeats.
The GPT-3 research paper (OpenAI, 2020) describes its model in exactly these terms: “an autoregressive language model”, meaning one that generates each new token based on the tokens before it. The US National Institute of Standards and Technology (NIST) puts it just as plainly in its Generative AI risk profile (July 2024): “LLMs predict the next token or word in a sentence or phrase.”
It sounds too simple to produce essays, code and translations. The trick is scale. To predict the next word in a legal contract, a cricket report or a Python program well, a model has to capture a great deal about grammar, facts, style and reasoning patterns. Prediction at scale forces the model to learn those patterns.
Most chat systems do not simply pick the single most likely token every time. They add some controlled randomness (often adjusted by a setting called “temperature”), which is why the same question can get differently worded answers.
Step three: learning from training data
Before it can predict anything, the model has to learn from examples. This first stage is called pre-training. The model is shown huge amounts of text, asked to predict the next token, and adjusted slightly every time it is wrong. Repeat this trillions of times and the predictions get very good.
The scale is hard to picture. Meta’s Llama 3 paper (July 2024), one of the few that discloses such details, says the models were pre-trained on “about 15T multilingual tokens”, that is, roughly 15 trillion tokens of text, and that the largest version was trained on up to 16,000 Nvidia H100 graphics processors.
Two consequences follow:
- The model’s knowledge has a cut-off date. It only “knows” what was in its training data. Developers acknowledge this openly: Anthropic, for instance, describes its web search tool as a way to find information beyond the knowledge cutoff.
- The model absorbs the data’s gaps and errors. If a topic was rare, contested or wrong in the training text, the model’s predictions about it will be weaker.
Step four: the transformer and “attention”
Nearly all of today’s LLMs use a design called the transformer, introduced by Google researchers in the 2017 paper “Attention Is All You Need”. Its central idea is a mechanism called attention.
Here is the intuition. Take the sentence: “The bank refused the loan because it was short of capital.” To understand “it”, you need to connect it back to “the bank”, not “the loan”. Attention lets the model, for every token, weigh how relevant every other token in the passage is, and blend that information in. The word “it” can “pay attention” strongly to “bank” and weakly to everything else.
A transformer stacks many layers of this. Early layers might pick up simple relationships such as which words go together; later layers build up more abstract patterns. Nobody hand-codes these patterns. They emerge from training.
The paper’s other practical advantage was speed. Earlier designs read text one word after another, which was slow to train. The transformer processes a whole passage at once, which made it, in the authors’ words, “more parallelizable and requiring significantly less time to train”. That property is what made training on trillions of tokens feasible.
Step five: parameters, the model’s adjustable dials
When people say a model has “175 billion parameters”, they mean it has 175 billion numbers that are tuned during training. Each parameter is like a tiny dial. Together they encode everything the model has learned. GPT-3 had 175 billion parameters; the largest Llama 3 model has 405 billion.
More parameters generally means more capacity to learn, but also more computing power to train and run. Size is not everything, though, as the next section shows. Many developers no longer disclose parameter counts for their flagship models, so treat any number you see quoted without a source with caution.
Step six: from raw predictor to helpful assistant
A pre-trained model is a powerful text predictor, but not yet a useful assistant. Ask it a question and it might continue with more questions, because that is what often follows a question on the internet. Turning it into a chatbot takes further training, usually in two stages.
Fine-tuning means continuing training on a smaller, carefully chosen dataset. For an assistant, that dataset is typically examples of good answers to instructions, often written by people.
Reinforcement learning from human feedback (RLHF) goes one step further. People compare several model answers to the same prompt and rank them. Those rankings are used to train the model to produce the kind of answers people prefer.
The best-known description of this process is OpenAI’s 2022 InstructGPT paper. Its headline result shows how much this stage matters: “outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters.” The authors also reported improvements in truthfulness and reductions in toxic output.
| Stage | What the model learns from | What it adds |
|---|---|---|
| Pre-training | Trillions of tokens of general text | Language, facts, patterns, broad knowledge |
| Fine-tuning | Smaller sets of example instructions and good answers | Follows instructions, answers in a useful format |
| RLHF (or similar) | Human rankings of alternative answers | Answers people judge more helpful, honest and safe |
Post-training also shapes a model’s weaknesses. A September 2025 research paper published by OpenAI, “Why Language Models Hallucinate”, argues that common training and evaluation methods reward confident guessing over admitting uncertainty. We explain what that means for you in why AI chatbots “hallucinate”, and how to catch it.
Step seven: the context window, the model’s working memory
The context window is the amount of text a model can consider at once when writing a reply. Anthropic’s documentation defines it as “all the text a language model can reference when generating a response, including the response itself”, and describes it as the model’s “working memory”, distinct from the training data.
Everything counts against this limit: your question, any documents you paste in, the conversation so far, hidden instructions from the app, and the reply being written. Context windows have grown fast. As of October 2026, Anthropic lists windows of up to 1 million tokens for some of its models, and Meta’s Llama 3 paper described a 128,000-token window in 2024.
Bigger is not automatically better. The same Anthropic page notes that as the token count grows, “accuracy and recall degrade”, a problem it calls “context rot”. In practice, a model can lose track of a detail buried in a very long document. Giving it only what it needs often works better than pasting in everything.
A model also does not remember you between conversations unless the app deliberately stores and re-sends earlier material. Each reply is generated from what is inside the context window at that moment.
Step eight: inference, and why every answer costs money
Using a trained model to generate text is called inference. Every token in a reply requires running the input through the entire model, with billions of calculations. That is why LLMs run on graphics processing units (GPUs), specialised chips that perform huge numbers of calculations in parallel.
That hardware is expensive, scarce and power-hungry, which is why it has become a matter of national policy. The IndiaAI Mission, approved with an outlay of about ₹10,372 crore, had expanded shared compute capacity to more than 45,000 GPUs as of June 2026, according to a PIB factsheet dated 13 August 2026. A February 2026 PIB backgrounder said the IndiaAI Compute Portal offered GPU access at subsidised rates of under ₹100 per hour, compared with global rates above ₹200 per hour. For the chip side of this story, see India’s semiconductor push, explained.
The cost structure explains several things you may notice as a user:
- Free tiers have usage caps, and longer answers or bigger documents use them up faster.
- Long conversations get more expensive to run, because the whole history is processed again with each reply.
- Faster or cheaper “mini” versions of models exist because smaller models need less computation per token.
What LLMs are good at, and where they fail
Because LLMs are pattern predictors trained on vast text, their strengths and weaknesses are fairly predictable.
Generally strong at:
- Drafting, rewriting, summarising and changing the tone of text.
- Explaining well-documented concepts in simpler terms.
- Translation between widely used languages.
- Writing and explaining common kinds of code.
- Working with material you supply in the context window, such as a document to summarise.
Generally weak at, or risky for:
- Facts they were not trained on, including anything after their cut-off date, unless the app adds a search tool.
- Precise recall of rare facts, such as a specific person’s birthday, an exact citation or a niche statistic. Here the model may produce a fluent, confident and wrong answer. NIST calls this “confabulation” and notes it is “a natural result of the way generative models are designed”.
- Explaining their own reasoning reliably. NIST also warns that models can produce “confabulated logic or citations” that seem to justify an answer even when the answer is wrong.
- Taking actions in the world. A plain LLM only produces text. Systems that let a model use tools, browse or run tasks over many steps are called agents, and they bring a different set of risks; see what is an AI agent?
A useful rule: treat an LLM’s output as a strong first draft from a well-read assistant who never checks sources. Use it for speed, and check anything that matters.
Glossary
| Term | Plain-English meaning |
|---|---|
| Large language model (LLM) | A program trained on huge amounts of text to predict the next piece of text |
| Token | The chunk of text a model reads and writes; often part of a word |
| Pre-training | The first, largest training stage, on general text |
| Parameters | The billions of internal numbers adjusted during training |
| Transformer | The model design behind most LLMs, introduced in 2017 |
| Attention | The mechanism that lets each token weigh the relevance of every other token |
| Fine-tuning | Further training on a smaller, targeted dataset |
| RLHF | Reinforcement learning from human feedback: training on human rankings of answers |
| Context window | How much text the model can consider at once, including its reply |
| Inference | Running a trained model to generate output |
| GPU | A chip that performs many calculations in parallel; the workhorse of AI |
| Hallucination / confabulation | Confident output that is false or unsupported |
| Agent | A system in which a model uses tools in a loop to complete a goal |
The point: An LLM is a very large next-token predictor, shaped first by trillions of tokens of text and then by human feedback into a helpful assistant. That design makes it fluent and broadly useful, but it also means it can produce confident errors and knows nothing beyond its training and what you give it. Use it as a fast drafter and explainer, and verify anything you will act on.
Sources
- Attention Is All You Need, Vaswani et al., arXiv / NeurIPS 2017
- Language Models are Few-Shot Learners, Brown et al. (OpenAI), arXiv 2020
- Training language models to follow instructions with human feedback, Ouyang et al. (OpenAI), arXiv 2022
- The Llama 3 Herd of Models, Llama Team (Meta), arXiv 2024
- Why Language Models Hallucinate, Kalai et al., arXiv 2025
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1), NIST, July 2024
- Context windows, Anthropic developer documentation (accessed October 2026)
- Tool use with Claude, Anthropic developer documentation (accessed October 2026)
- Understand and count tokens, Google AI for Developers (accessed October 2026)
- The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages, Srivastava, arXiv preprint, July 2026
- Factsheet: Semiconductor and AI Revolution, Press Information Bureau, 13 August 2026
- India AI Stack: Powering Intelligence at Scale, PIB Backgrounder, 4 February 2026