A large language model, or LLM, is a neural network trained on a very large collection of text to predict the next token in a sequence. A token is a word or a piece of a word. Given "The capital of France is", the model assigns a probability to every possible next token and "Paris" comes out on top. Asked to keep going, it predicts the token after that, and the one after that, until it produces a stop token.
Everything an LLM does in practice, from answering questions to writing the paragraph at the top of an AI search result, is built on that one prediction task. The model has no database of facts and no search engine inside it. What it has is a very detailed statistical picture of how text tends to continue, learned from the text it was trained on.
How is a large language model trained?
Training happens in stages, and each stage changes what the model is good at.
Pretraining is the expensive part. The model reads a large corpus of web pages, books, code and other text, and adjusts its internal weights so that its next-token predictions get better. No one labels the data. The text itself is the supervision: the model guesses the next token, compares the guess with the real one, and corrects itself. Over billions of examples this produces a model that has absorbed grammar, facts, reasoning patterns and the style of many kinds of writing.
Fine-tuning teaches the pretrained model to behave like an assistant. It is trained on examples of questions and good answers, so it learns to respond to instructions instead of continuing the prompt as if it were a document. Fine-tuning can also specialize a model for a domain or a task.
Preference training adjusts the model with human or automated feedback on which of two answers is better. This is where a model learns to decline harmful requests, to admit uncertainty and to format answers the way people expect.
The weights the model ends up with are its parameters. A larger model has more of them, which usually means better predictions at a higher cost per token. Model sizes are measured in billions of parameters.
What are tokens and the context window?
Models do not read characters or whole words. They read tokens, which are pieces of text chosen so that common words are a single token and rare words split into several. English prose averages about three quarters of a word per token.
The context window is the maximum number of tokens the model can take in at once, counting the system instructions, the conversation so far, any retrieved documents and the answer it is writing. Text outside the window does not exist for the model. In AI search this limit decides how many retrieved passages can be shown to the model for one answer, which is one reason retrieval systems select a few good passages instead of whole pages.
Why can a model be fluent and still be wrong?
The model is trained to produce likely text, and likely text is usually true text, because most of the training data is written by people describing the world. But the two are not the same. When the model has not seen reliable information about a topic, it still produces a fluent continuation, and that continuation can be a confident, well-formed sentence with a wrong fact in it. This is an AI hallucination.
Two properties of training make the problem predictable. The model's knowledge stops at its knowledge cutoff, the date of the last text in the training set, so anything after it is unknown. And rare topics are covered by little training text, so the model's picture of them is blurry. A model knows a lot about Paris and very little about a small company that launched last year.
The standard fix is to give the model the facts at the moment of the question instead of relying on what it absorbed during training. That is what retrieval-augmented generation does, and it is the design behind every AI search product.
How do AI search engines use language models?
AI search uses a language model in at least three places. It writes the search queries that go to the index, turning a vague question into specific queries. It reads the retrieved passages and decides which ones answer the question. And it writes the final answer, attaching citations to the passages it used.
In all three roles the model works from text it was handed during the request. That is why the pages a search engine can find, and the passages inside them that stand on their own, decide what the answer says and which brands it names.