Workbench / Glossary
Glossary/
The vocabulary of self-hosted AI, in plain words. Every term links to the note it comes from.
A
- Agent
- An LLM that can plan, call tools and act over several steps toward a goal. From: LLM Glossary: The Short Version →
- Alignment
- The work of making a model helpful, honest and safe, usually through human feedback training. From: LLM Glossary: The Short Version →
- Attention
- The mechanism that lets a model look back at every earlier token and decide which ones matter for the next one. From: LLM Glossary: The Short Version →
B
- Base
- Base means it is closer to the raw pretrained model and usually less friendly for chat. From: Terms That Unlock Self-hosted LLMs →
- Batching
- Processing several requests at once so the GPU stays busy. Key to serving many users. From: LLM Glossary: The Short Version →
- Benchmark
- A public, standardized eval used to compare models. Useful, but easy to over-trust. From: LLM Glossary: The Short Version →
C
- Chain of thought
- Asking the model to reason in steps before answering, which often improves accuracy. From: LLM Glossary: The Short Version →
- Chunking
- Cutting documents into small pieces before embedding them, so search returns focused passages. From: LLM Glossary: The Short Version →
- Context window
- Context window is how much text the model can consider at once. From: Terms That Unlock Self-hosted LLMs →
- Cosine similarity
- The usual score for how close two embeddings are. Closer to 1 means more similar meaning. From: LLM Glossary: The Short Version →
D
- Distillation
- Training a small model to imitate a bigger one, keeping much of the skill at a fraction of the size. From: LLM Glossary: The Short Version →
E
- Embedding
- Embedding is a numeric representation of meaning. From: Terms That Unlock Self-hosted LLMs →
- Embedding model
- A model that turns text into vectors for search, not for chat. From: LLM Glossary: The Short Version →
- Embedding Vector
- When people talk about vectors in LLMs, it can sound technical at first, but the idea is surprisingly close to the word vector from graphic design. An embedding vector is… From: Embedding Vector →
- Eval
- A repeatable test set that measures whether a model or prompt change made things better or worse. From: LLM Glossary: The Short Version →
F
- Few-shot
- Showing the model a few worked examples in the prompt so it copies the pattern. From: LLM Glossary: The Short Version →
- Fine-tuning
- Extra training on a smaller, focused dataset to change how a model behaves or what it knows. From: LLM Glossary: The Short Version →
G
- GGUF
- A local-friendly model file format commonly used with llama.cpp-based tools. From: Terms That Unlock Self-hosted LLMs →
- Grounding
- Making the model answer from supplied sources instead of from memory, ideally with citations. From: LLM Glossary: The Short Version →
- Guardrails
- Checks around a model that block unsafe input, output or actions. From: LLM Glossary: The Short Version →
H
- Hallucination
- A confident answer that is simply wrong or made up. From: LLM Glossary: The Short Version →
- Human in the loop
- A setup where a person approves or corrects the agent before risky actions run. From: LLM Glossary: The Short Version →
- Hybrid search
- Combining keyword search with vector search. Catches exact names and fuzzy meaning at once. From: LLM Glossary: The Short Version →
I
- Inference
- Running a trained model to get an answer. Training builds the model; inference uses it. From: LLM Glossary: The Short Version →
- Instruct
- Instruct usually means the model has been tuned to follow directions. From: Terms That Unlock Self-hosted LLMs →
K
- K-quants
- The llama.cpp quantization family with names like Q4_K_M. Lower number means smaller and rougher. From: LLM Glossary: The Short Version →
- KV cache
- Memory that stores attention results for earlier tokens so the model does not recompute them. It grows with context. From: LLM Glossary: The Short Version →
L
- Latency
- How long you wait for the first token of the answer. From: LLM Glossary: The Short Version →
- LLM
- A large language model: a neural network trained on huge amounts of text to predict the next token. From: LLM Glossary: The Short Version →
- LoRA
- A cheap fine-tuning method that trains small add-on layers instead of the whole model. From: LLM Glossary: The Short Version →
M
- MCP
- Model Context Protocol: an open standard for plugging tools and data sources into AI apps. From: LLM Glossary: The Short Version →
- Mixture of Experts
- A model split into many expert blocks where only a few run per token. Large on disk, fast to run. From: LLM Glossary: The Short Version →
- Multimodal
- A model that handles more than text, such as images, audio or video. From: LLM Glossary: The Short Version →
O
- Offloading
- Moving part of a model from GPU to system RAM when VRAM runs out. It works, but it is much slower. From: LLM Glossary: The Short Version →
- Open weights
- A model whose weights you can download and run yourself, even if the training data stays private. From: LLM Glossary: The Short Version →
- OpenAI-compatible API
- An API shape most local runtimes copy, so tools built for OpenAI can talk to your own server. From: LLM Glossary: The Short Version →
P
- Parameters
- Parameters are the learned weights inside the model. From: Terms That Unlock Self-hosted LLMs →
- Pretraining
- The first, expensive phase where a model learns language by reading a massive text corpus. From: LLM Glossary: The Short Version →
- Prompt
- The input you give a model: instructions, context and the question itself. From: LLM Glossary: The Short Version →
- Prompt injection
- Hidden instructions in a document or web page that try to hijack what the model does. From: LLM Glossary: The Short Version →
Q
- Quantization
- A way to shrink model weights so larger models can fit on smaller hardware. From: Terms That Unlock Self-hosted LLMs →
R
- RAG
- RAG is the pattern where your documents are searched first, then relevant pieces are sent to the model. From: Terms That Unlock Self-hosted LLMs →
- Reasoning model
- A model trained to think step by step before answering. Slower and costlier, better on hard problems. From: LLM Glossary: The Short Version →
- Reranker
- A second model that re-scores search results so the best passages reach the LLM first. From: LLM Glossary: The Short Version →
- Runtime
- Runtime is the software that runs the model. From: Terms That Unlock Self-hosted LLMs →
S
- Safetensors
- A safe, fast file format for model weights, the standard on Hugging Face. From: LLM Glossary: The Short Version →
- Speculative decoding
- A small draft model guesses ahead and the big model checks the guesses, speeding up output. From: LLM Glossary: The Short Version →
- Structured output
- Forcing the model to answer in a fixed format such as JSON, so code can read it reliably. From: LLM Glossary: The Short Version →
- System prompt
- Hidden instructions set before the conversation that shape the model's role and rules. From: LLM Glossary: The Short Version →
T
- Temperature
- A setting for randomness. Low gives focused, repeatable answers; high gives more varied ones. From: LLM Glossary: The Short Version →
- Throughput
- How many tokens per second a setup produces, for one user or across many. From: LLM Glossary: The Short Version →
- Tokenizer
- The component that splits text into tokens and back. Each model family has its own. From: LLM Glossary: The Short Version →
- Tokens
- Tokens are the pieces of text the model reads and writes. From: Terms That Unlock Self-hosted LLMs →
- Tool calling
- When a model asks to run a function, like a search or an API call, and uses the result. From: LLM Glossary: The Short Version →
- Top-p
- A sampling limit that only lets the model pick from the most likely tokens adding up to probability p. From: LLM Glossary: The Short Version →
- Transformer
- The architecture behind almost every modern LLM. It uses attention to weigh how words relate to each other. From: LLM Glossary: The Short Version →
V
- Vector database
- A database built to store embeddings and find the closest ones fast. From: LLM Glossary: The Short Version →
- VRAM
- VRAM is GPU memory, and it becomes one of the main limits when running models locally. From: Terms That Unlock Self-hosted LLMs →
W
- Weights
- The numbers a model learned during training. Downloading a model means downloading its weights. From: LLM Glossary: The Short Version →