LLM Glossary: The Short Version
Forty-seven LLM terms in one or two lines each: models, running them, prompts, knowledge search and agents.
The short version. One or two lines per term, no maths. When a word in a note stops you, look it up here.
The Model Words
Alignment
The work of making a model helpful, honest and safe, usually through human feedback training.
Attention
The mechanism that lets a model look back at every earlier token and decide which ones matter for the next one.
Distillation
Training a small model to imitate a bigger one, keeping much of the skill at a fraction of the size.
Fine-tuning
Extra training on a smaller, focused dataset to change how a model behaves or what it knows.
LLM
A large language model: a neural network trained on huge amounts of text to predict the next token.
LoRA
A cheap fine-tuning method that trains small add-on layers instead of the whole model.
Mixture of Experts
A model split into many expert blocks where only a few run per token. Large on disk, fast to run.
Multimodal
A model that handles more than text, such as images, audio or video.
Open weights
A model whose weights you can download and run yourself, even if the training data stays private.
Pretraining
The first, expensive phase where a model learns language by reading a massive text corpus.
Reasoning model
A model trained to think step by step before answering. Slower and costlier, better on hard problems.
Transformer
The architecture behind almost every modern LLM. It uses attention to weigh how words relate to each other.
Weights
The numbers a model learned during training. Downloading a model means downloading its weights.
The Running Words
Batching
Processing several requests at once so the GPU stays busy. Key to serving many users.
Inference
Running a trained model to get an answer. Training builds the model; inference uses it.
K-quants
The llama.cpp quantization family with names like Q4_K_M. Lower number means smaller and rougher.
KV cache
Memory that stores attention results for earlier tokens so the model does not recompute them. It grows with context.
Latency
How long you wait for the first token of the answer.
Offloading
Moving part of a model from GPU to system RAM when VRAM runs out. It works, but it is much slower.
OpenAI-compatible API
An API shape most local runtimes copy, so tools built for OpenAI can talk to your own server.
Safetensors
A safe, fast file format for model weights, the standard on Hugging Face.
Speculative decoding
A small draft model guesses ahead and the big model checks the guesses, speeding up output.
Throughput
How many tokens per second a setup produces, for one user or across many.
Tokenizer
The component that splits text into tokens and back. Each model family has its own.
The Prompt Words
Chain of thought
Asking the model to reason in steps before answering, which often improves accuracy.
Few-shot
Showing the model a few worked examples in the prompt so it copies the pattern.
Guardrails
Checks around a model that block unsafe input, output or actions.
Hallucination
A confident answer that is simply wrong or made up.
Prompt
The input you give a model: instructions, context and the question itself.
Prompt injection
Hidden instructions in a document or web page that try to hijack what the model does.
Structured output
Forcing the model to answer in a fixed format such as JSON, so code can read it reliably.
System prompt
Hidden instructions set before the conversation that shape the model's role and rules.
Temperature
A setting for randomness. Low gives focused, repeatable answers; high gives more varied ones.
Top-p
A sampling limit that only lets the model pick from the most likely tokens adding up to probability p.
The Knowledge Words
Chunking
Cutting documents into small pieces before embedding them, so search returns focused passages.
Cosine similarity
The usual score for how close two embeddings are. Closer to 1 means more similar meaning.
Embedding model
A model that turns text into vectors for search, not for chat.
Grounding
Making the model answer from supplied sources instead of from memory, ideally with citations.
Hybrid search
Combining keyword search with vector search. Catches exact names and fuzzy meaning at once.
Reranker
A second model that re-scores search results so the best passages reach the LLM first.
Vector database
A database built to store embeddings and find the closest ones fast.
The Agent Words
Agent
An LLM that can plan, call tools and act over several steps toward a goal.
Benchmark
A public, standardized eval used to compare models. Useful, but easy to over-trust.
Eval
A repeatable test set that measures whether a model or prompt change made things better or worse.
Human in the loop
A setup where a person approves or corrects the agent before risky actions run.
MCP
Model Context Protocol: an open standard for plugging tools and data sources into AI apps.
Tool calling
When a model asks to run a function, like a search or an API call, and uses the result.