Explainer · By Arthur Shafer · Published

What is an LLM?

A language model can explain a concept, summarize a document, or help write code. Understanding what happens during a model call makes its strengths and its mistakes easier to recognize.

What the name means

An LLM is a large language model. It works with sequences of tokens, which can be whole words or smaller pieces of text. During generation, it estimates what token should come next in the context it has been given. Repeating that process produces a sentence, a paragraph, or a longer answer. The word large points to the scale of its training and model, rather than to one precise threshold every LLM must meet.

It is tempting to think of the model as a library containing every page it read. Training does not give it a dependable catalog of documents and citations. It adjusts numerical weights across many examples. Those weights help it recognize patterns in language, code, and other structured material, but they do not provide an audit trail for each statement it produces.

How a model learns

In an initial training stage, a model learns by predicting missing or subsequent tokens in a large collection of material. Many models receive further training to follow instructions and respond more usefully in conversation. Different model families use different methods, but the result is a system that can apply learned patterns to a prompt it has not seen before. Google's Machine Learning Crash Course provides a more technical introduction to tokens, training, and transformers.

I find it useful to separate training from use. Training changes the model's weights. A prompt supplies temporary context for a particular response. Adding a document to a prompt can help the model answer a question about that document, but that interaction does not itself rewrite the model's trained weights. Whether a particular service saves conversations or later uses them for training is a separate product policy question.

What happens when you ask

An application assembles the input for a model call. That input may include a system instruction, the user's request, earlier conversation, retrieved passages, and tool results. The model generates an output from what it was given and what its training enables it to do. The application may display that output directly or use it as one step in a longer workflow.

This is also where the idea of a single mind behind every chat can mislead. A service can run many requests using the same model. A chat interface may preserve history and present it to a later request, making the exchange feel continuous. That continuity is managed by the application and its stored state. Anthropic's Messages API, for example, explicitly asks the caller to provide conversation history for each request.

Why it can do so many things

The same model can summarize a report, compare two passages, draft an email, explain code, or propose a plan. Those tasks all involve interpreting context and producing language or another structured response. Broad training gives the model a flexible starting point, while instructions and examples narrow the task in front of it.

The surrounding application determines what information the model can use. It may retrieve current records, call a calculator, or place a person's documents into context. That is useful when the answer depends on material outside the model's training. It also means the quality of the answer depends on which records were selected, how permissions were enforced, and whether the tool results are accurate.

Why fluent is not the same as right

A plausible sentence can still be wrong. A model may fill a gap with a detail that sounds right, misunderstand a source, or state a time-sensitive fact as if it were current. It can attach a citation to a passage that is related to the answer without actually supporting it. The writing is often smooth enough that the error takes work to find. OpenAI's research on hallucinations examines why models may guess instead of expressing uncertainty.

I ran into a narrower version of this problem with the assistant on this site. I use code to select approved records before the model writes from them. The system checks the response structure and its citation references, then falls back to record summaries if that check fails. I still have to evaluate whether each cited sentence is a fair reading of its source. A citation number alone cannot answer that question.

Where the model fits

The model is one part of an AI system. Code can decide which records it sees, which tools it may request, what format the response must take, and what happens if it fails a check. When a system repeatedly asks the model to choose an action, carries out permitted actions, and returns the results for another decision, it starts to resemble an AI agent.

I do not think understanding the model requires treating it as either a simple autocomplete feature or a complete mind. The practical questions are what it was trained to do, what context it received for this call, and which parts of the application can verify or correct its output. Those questions remain useful as models improve.

Arthur's portfolio assistant

Beyond the resume.

Get to know the experience behind the resume.

    Ask about my experience, the systems I've built, or how I lead delivery.

    Public experience · Saved in this browser