Skip to content
Workbench

Workbench / Glossary

Glossary/

The vocabulary of self-hosted AI, in plain words. Every term links to the note it comes from.

A

Agent
An LLM that can plan, call tools and act over several steps toward a goal. From: LLM Glossary: The Short Version →
Alignment
The work of making a model helpful, honest and safe, usually through human feedback training. From: LLM Glossary: The Short Version →
Attention
The mechanism that lets a model look back at every earlier token and decide which ones matter for the next one. From: LLM Glossary: The Short Version →

B

Base
Base means it is closer to the raw pretrained model and usually less friendly for chat. From: Terms That Unlock Self-hosted LLMs →
Batching
Processing several requests at once so the GPU stays busy. Key to serving many users. From: LLM Glossary: The Short Version →
Benchmark
A public, standardized eval used to compare models. Useful, but easy to over-trust. From: LLM Glossary: The Short Version →

C

Chain of thought
Asking the model to reason in steps before answering, which often improves accuracy. From: LLM Glossary: The Short Version →
Chunking
Cutting documents into small pieces before embedding them, so search returns focused passages. From: LLM Glossary: The Short Version →
Context window
Context window is how much text the model can consider at once. From: Terms That Unlock Self-hosted LLMs →
Cosine similarity
The usual score for how close two embeddings are. Closer to 1 means more similar meaning. From: LLM Glossary: The Short Version →

D

Distillation
Training a small model to imitate a bigger one, keeping much of the skill at a fraction of the size. From: LLM Glossary: The Short Version →

E

Embedding
Embedding is a numeric representation of meaning. From: Terms That Unlock Self-hosted LLMs →
Embedding model
A model that turns text into vectors for search, not for chat. From: LLM Glossary: The Short Version →
Embedding Vector
When people talk about vectors in LLMs, it can sound technical at first, but the idea is surprisingly close to the word vector from graphic design. An embedding vector is… From: Embedding Vector →
Eval
A repeatable test set that measures whether a model or prompt change made things better or worse. From: LLM Glossary: The Short Version →

F

Few-shot
Showing the model a few worked examples in the prompt so it copies the pattern. From: LLM Glossary: The Short Version →
Fine-tuning
Extra training on a smaller, focused dataset to change how a model behaves or what it knows. From: LLM Glossary: The Short Version →

G

GGUF
A local-friendly model file format commonly used with llama.cpp-based tools. From: Terms That Unlock Self-hosted LLMs →
Grounding
Making the model answer from supplied sources instead of from memory, ideally with citations. From: LLM Glossary: The Short Version →
Guardrails
Checks around a model that block unsafe input, output or actions. From: LLM Glossary: The Short Version →

H

Hallucination
A confident answer that is simply wrong or made up. From: LLM Glossary: The Short Version →
Human in the loop
A setup where a person approves or corrects the agent before risky actions run. From: LLM Glossary: The Short Version →

I

Inference
Running a trained model to get an answer. Training builds the model; inference uses it. From: LLM Glossary: The Short Version →
Instruct
Instruct usually means the model has been tuned to follow directions. From: Terms That Unlock Self-hosted LLMs →

K

K-quants
The llama.cpp quantization family with names like Q4_K_M. Lower number means smaller and rougher. From: LLM Glossary: The Short Version →
KV cache
Memory that stores attention results for earlier tokens so the model does not recompute them. It grows with context. From: LLM Glossary: The Short Version →

L

Latency
How long you wait for the first token of the answer. From: LLM Glossary: The Short Version →
LLM
A large language model: a neural network trained on huge amounts of text to predict the next token. From: LLM Glossary: The Short Version →
LoRA
A cheap fine-tuning method that trains small add-on layers instead of the whole model. From: LLM Glossary: The Short Version →

M

MCP
Model Context Protocol: an open standard for plugging tools and data sources into AI apps. From: LLM Glossary: The Short Version →
Mixture of Experts
A model split into many expert blocks where only a few run per token. Large on disk, fast to run. From: LLM Glossary: The Short Version →
Multimodal
A model that handles more than text, such as images, audio or video. From: LLM Glossary: The Short Version →

O

Offloading
Moving part of a model from GPU to system RAM when VRAM runs out. It works, but it is much slower. From: LLM Glossary: The Short Version →
Open weights
A model whose weights you can download and run yourself, even if the training data stays private. From: LLM Glossary: The Short Version →
OpenAI-compatible API
An API shape most local runtimes copy, so tools built for OpenAI can talk to your own server. From: LLM Glossary: The Short Version →

P

Parameters
Parameters are the learned weights inside the model. From: Terms That Unlock Self-hosted LLMs →
Pretraining
The first, expensive phase where a model learns language by reading a massive text corpus. From: LLM Glossary: The Short Version →
Prompt
The input you give a model: instructions, context and the question itself. From: LLM Glossary: The Short Version →
Prompt injection
Hidden instructions in a document or web page that try to hijack what the model does. From: LLM Glossary: The Short Version →

Q

Quantization
A way to shrink model weights so larger models can fit on smaller hardware. From: Terms That Unlock Self-hosted LLMs →

R

RAG
RAG is the pattern where your documents are searched first, then relevant pieces are sent to the model. From: Terms That Unlock Self-hosted LLMs →
Reasoning model
A model trained to think step by step before answering. Slower and costlier, better on hard problems. From: LLM Glossary: The Short Version →
Reranker
A second model that re-scores search results so the best passages reach the LLM first. From: LLM Glossary: The Short Version →
Runtime
Runtime is the software that runs the model. From: Terms That Unlock Self-hosted LLMs →

S

Safetensors
A safe, fast file format for model weights, the standard on Hugging Face. From: LLM Glossary: The Short Version →
Speculative decoding
A small draft model guesses ahead and the big model checks the guesses, speeding up output. From: LLM Glossary: The Short Version →
Structured output
Forcing the model to answer in a fixed format such as JSON, so code can read it reliably. From: LLM Glossary: The Short Version →
System prompt
Hidden instructions set before the conversation that shape the model's role and rules. From: LLM Glossary: The Short Version →

T

Temperature
A setting for randomness. Low gives focused, repeatable answers; high gives more varied ones. From: LLM Glossary: The Short Version →
Throughput
How many tokens per second a setup produces, for one user or across many. From: LLM Glossary: The Short Version →
Tokenizer
The component that splits text into tokens and back. Each model family has its own. From: LLM Glossary: The Short Version →
Tokens
Tokens are the pieces of text the model reads and writes. From: Terms That Unlock Self-hosted LLMs →
Tool calling
When a model asks to run a function, like a search or an API call, and uses the result. From: LLM Glossary: The Short Version →
Top-p
A sampling limit that only lets the model pick from the most likely tokens adding up to probability p. From: LLM Glossary: The Short Version →
Transformer
The architecture behind almost every modern LLM. It uses attention to weigh how words relate to each other. From: LLM Glossary: The Short Version →

V

Vector database
A database built to store embeddings and find the closest ones fast. From: LLM Glossary: The Short Version →
VRAM
VRAM is GPU memory, and it becomes one of the main limits when running models locally. From: Terms That Unlock Self-hosted LLMs →

W

Weights
The numbers a model learned during training. Downloading a model means downloading its weights. From: LLM Glossary: The Short Version →