Skip to content
Workbench

Workbench / Methods

LLM Glossary: The Short Version

Forty-seven LLM terms in one or two lines each: models, running them, prompts, knowledge search and agents.

The short version. One or two lines per term, no maths. When a word in a note stops you, look it up here.

The Model Words

Alignment

The work of making a model helpful, honest and safe, usually through human feedback training.

Attention

The mechanism that lets a model look back at every earlier token and decide which ones matter for the next one.

Distillation

Training a small model to imitate a bigger one, keeping much of the skill at a fraction of the size.

Fine-tuning

Extra training on a smaller, focused dataset to change how a model behaves or what it knows.

LLM

A large language model: a neural network trained on huge amounts of text to predict the next token.

LoRA

A cheap fine-tuning method that trains small add-on layers instead of the whole model.

Mixture of Experts

A model split into many expert blocks where only a few run per token. Large on disk, fast to run.

Multimodal

A model that handles more than text, such as images, audio or video.

Open weights

A model whose weights you can download and run yourself, even if the training data stays private.

Pretraining

The first, expensive phase where a model learns language by reading a massive text corpus.

Reasoning model

A model trained to think step by step before answering. Slower and costlier, better on hard problems.

Transformer

The architecture behind almost every modern LLM. It uses attention to weigh how words relate to each other.

Weights

The numbers a model learned during training. Downloading a model means downloading its weights.

The Running Words

Batching

Processing several requests at once so the GPU stays busy. Key to serving many users.

Inference

Running a trained model to get an answer. Training builds the model; inference uses it.

K-quants

The llama.cpp quantization family with names like Q4_K_M. Lower number means smaller and rougher.

KV cache

Memory that stores attention results for earlier tokens so the model does not recompute them. It grows with context.

Latency

How long you wait for the first token of the answer.

Offloading

Moving part of a model from GPU to system RAM when VRAM runs out. It works, but it is much slower.

OpenAI-compatible API

An API shape most local runtimes copy, so tools built for OpenAI can talk to your own server.

Safetensors

A safe, fast file format for model weights, the standard on Hugging Face.

Speculative decoding

A small draft model guesses ahead and the big model checks the guesses, speeding up output.

Throughput

How many tokens per second a setup produces, for one user or across many.

Tokenizer

The component that splits text into tokens and back. Each model family has its own.

The Prompt Words

Chain of thought

Asking the model to reason in steps before answering, which often improves accuracy.

Few-shot

Showing the model a few worked examples in the prompt so it copies the pattern.

Guardrails

Checks around a model that block unsafe input, output or actions.

Hallucination

A confident answer that is simply wrong or made up.

Prompt

The input you give a model: instructions, context and the question itself.

Prompt injection

Hidden instructions in a document or web page that try to hijack what the model does.

Structured output

Forcing the model to answer in a fixed format such as JSON, so code can read it reliably.

System prompt

Hidden instructions set before the conversation that shape the model's role and rules.

Temperature

A setting for randomness. Low gives focused, repeatable answers; high gives more varied ones.

Top-p

A sampling limit that only lets the model pick from the most likely tokens adding up to probability p.

The Knowledge Words

Chunking

Cutting documents into small pieces before embedding them, so search returns focused passages.

Cosine similarity

The usual score for how close two embeddings are. Closer to 1 means more similar meaning.

Embedding model

A model that turns text into vectors for search, not for chat.

Grounding

Making the model answer from supplied sources instead of from memory, ideally with citations.

Hybrid search

Combining keyword search with vector search. Catches exact names and fuzzy meaning at once.

Reranker

A second model that re-scores search results so the best passages reach the LLM first.

Vector database

A database built to store embeddings and find the closest ones fast.

The Agent Words

Agent

An LLM that can plan, call tools and act over several steps toward a goal.

Benchmark

A public, standardized eval used to compare models. Useful, but easy to over-trust.

Eval

A repeatable test set that measures whether a model or prompt change made things better or worse.

Human in the loop

A setup where a person approves or corrects the agent before risky actions run.

MCP

Model Context Protocol: an open standard for plugging tools and data sources into AI apps.

Tool calling

When a model asks to run a function, like a search or an API call, and uses the result.