Building a Local AI Server: The Hardware, the Models, and What It Does All Day

One workstation-class machine in a normal office room runs open-weight language models, a document search stack, an image and video studio, and a small crew of AI agents that keep the whole thing running. Nothing is exposed to the public internet. This is the technical walkthrough: what is in the box, what runs on it, and what it actually does all day.

Why local at all?

Cloud models are excellent, and we still use them where they make sense. But three things pushed us toward running our own hardware.

  • Data stays home. Product data sheets, safety sheets, internal documents and our own agent memory never leave the building. Retrieval, embeddings and most chat run entirely on the box.
  • Flat cost. A local model answers ten thousand questions for the same electricity as one. Once the hardware is paid for, experimentation is free, and that changes how much you experiment.
  • It is genuinely fun. The open-weight model scene moves weekly. Having a 32 GB GPU on the LAN means you can pull the new thing on Tuesday and have an opinion about it by Wednesday.

The design principle from day one: RAG, not fine-tuning. We do not train models. We give good general-purpose models access to our documents and tools, and we let them retrieve what they need.

The hardware

It is one machine, built as a workstation rather than a rack server, and it lives in a normal office room.

PartChoiceWhy
CPUAMD Ryzen 9 9950X (16 cores)Plenty of headroom for containers, document conversion and OCR while the GPU is busy
GPUNVIDIA RTX 5090, 32 GB VRAMThe single most important part. VRAM decides which models you can run at all
RAM128 GBLets the Linux side, databases and vector store breathe; also allows CPU offload for models that spill over VRAM
Primary storage4 TB PCIe 5.0 NVMeModels, databases and container volumes. Model loading is I/O bound; fast flash matters
Archive storage8 TB HDD (SMR)Cold archive and backups only. SMR drives hate random writes, so nothing live runs from it
PSU1350 WA 5090 under sustained load is not shy

The one number that dictates everything is 32 GB of VRAM. Several of the models we like are 18 to 25 GB on their own. Two of them loaded at once will overcommit the card, and we have watched it happen.

The working rule is simple: one chat model resident at a time, plus the embedding model, and nothing else. The model runner unloads idle models automatically, so switching costs a few seconds of load time rather than a crash.

The software layers

The box runs Windows as the host operating system with Ubuntu inside WSL2. That sounds like a compromise, and in some ways it is, but it gives us a useful split.

Native Windows

Owns the GPU-heavy processes: the model runner and the image generation engine. GPU passthrough to WSL and to containers works fine, but keeping the two hungriest processes native avoids a layer of indirection and makes driver updates boring.

Ubuntu in WSL2

Runs Docker with a dedicated daemon (not Docker Desktop) and hosts everything else: the chat front end, databases, the vector store, the document converter, workflow automation, the agent framework and a handful of in-house apps.

A mesh VPN is the only way in. Every service binds to localhost and is published to team members through identity-aware routes on the VPN. There is no port forwarding to the internet for any of the AI services.

Almost every confusing symptom is one service being visible in one layer and invisible in another. Before concluding something is down, ask which layer you are asking.

That is the practical lesson from this layout. We keep an internal handbook chapter just on that.

The model runner: Ollama

All local language models run through Ollama. We evaluated a desktop alternative early on and dropped it. One runner, one API, one place to look when something is slow. Models live on the NVMe, the runner listens on the LAN so containers can reach it, and everything else in the stack talks to that single endpoint.

The current inventory, roughly grouped by job:

ModelParamsOn diskJob
qwen3-embedding:8b7.6B4.7 GBAgent memory embeddings (4096 dimensions)
qwen3-embedding:4b4.0B2.5 GBDocument search embeddings (2560 dimensions)
nemotron-3.5-lightning32.9B25.4 GBLargest general chat model on the box
qwen3.827.3B17.7 GBDefault chat and the product-knowledge assistant
gemma4:26b / gemma4:e4b25.8B / 8.0B18.0 / 9.6 GBGeneral chat, multimodal
qwen3-coder:30b30.5B (MoE)18.6 GBCode. Mixture-of-experts, so it is fast: around 300 tokens/s
devstral:24b23.6B14.3 GBCode, dense. Slower at around 95 tokens/s but a useful second opinion
deepseek-r1:8b8.2B8.9 GBReasoning experiments
qwen2.5:3b3.1B1.9 GBTiny and instant; smoke tests and trivial classification

Picking models for a 32 GB card

  • Mixture-of-experts models punch above their weight. A 30B MoE model with roughly 3B active parameters runs three times faster than a 24B dense model and fits comfortably with a 32K context. For code assistance that speed difference is the whole experience.
  • Turn thinking off for production writing. Reasoning models are great at puzzles and slow at prose. For product descriptions and translations, the “think first” mode mostly adds latency.
  • Context length eats VRAM too. A 30B model at 32K context sat at about 22 GB. There is room for 64K, but not for a second model alongside it.
  • Embedding models are cheap and worth keeping loaded. The 4B embedder uses under 4 GB and is what makes search feel instant.

Cloud models are still in the mix. The agent framework can route to hosted frontier models when a task needs them, but retrieval and embeddings never leave the box regardless of which model answers.

Document search: the retrieval stack

The single most useful thing the server does is answer questions about our own documents. We are a company that sells physical products, and every product has a technical data sheet, a safety data sheet, often a manual and a certificate, in Icelandic or English or both. That is thousands of PDFs, many of them scanned.

  1. Ingest. A nightly job picks up new or changed files from a shared folder. It never deletes anything.
  2. Convert. Docling turns PDFs into structured text. OCR runs only on pages that are actually scanned images; born-digital PDFs go straight through.
  3. Chunk and embed. Each page becomes a chunk. We store a dense vector from qwen3-embedding:4b and a sparse BM25 vector with accent-stripped forms, so a query for “þynnir” and “thynnir” land in the same place.
  4. Store. Everything goes into Qdrant, one collection for the product corpus, with metadata for document type, brand, language and page number.
  5. Retrieve. Queries run dense and sparse search in parallel and fuse the results with reciprocal rank fusion. Results link straight to the original PDF at the cited page.

Current size: roughly 118,000 page-chunks from about 5,150 documents. A nightly run that finds nothing new takes a couple of minutes.

Two front doors sit on top of that: a plain search page for quick lookups, and a chat assistant inside Open WebUI that answers with numbered citations. The assistant runs on qwen3.8 locally. It also rewrites sloppy queries before searching and shows you what it searched for, which turned out to matter more than we expected for trust.

What did not work: we tested a dedicated reranker model on top of the fused results, expecting the usual boost. It made retrieval slightly worse overall and noticeably worse for Icelandic queries against English documents. It is off.

Measure before you add layers.

Two memories that are not the same thing

This confused us for a while, so it is worth spelling out. There are two separate retrieval systems on the box, and they must not be merged.

Document RAG

The Qdrant corpus above. It answers “what is the flash point of product X”. Anyone on the team can query it.

Agent memory

A small per-agent SQLite store that indexes markdown notes only. It is how our AI admin agent remembers decisions, past incidents and preferences between sessions. Embedded with the larger 8B model and private to each agent.

Rule of thumb: the agent thinks or remembers, that is agent memory. The agent searches a corpus, that is Qdrant. We deliberately keep the internal operations handbook out of the document RAG, because the audience for a vector collection is wider than the audience for a handbook full of port numbers.

Image and video: the creative studio

The same GPU doubles as an image and short-video workstation. ComfyUI runs natively on Windows, and we built a small React front end with a Python backend on top of it so that non-technical colleagues get a normal-looking app instead of a node graph.

  • Images: FLUX.2 Klein for generation and editing.
  • Video: LTX 2.5 for short clips.
  • Product photos: a fixed pipeline that frames a product shot to 1200 × 1200, cleans the background and upscales with Real-ESRGAN, so a phone photo becomes a catalogue image.

Image jobs are queued rather than run in parallel with chat. On a single card that is the only sane policy.

The agents

The part that makes the box feel alive is an agent framework running a small number of persistent AI agents, each with its own workspace, memory, tools and persona. The main one is an AI admin that looks after the server itself and our externally hosted websites. Two specialists handle web content and creative work.

What an agent can actually do here:

  • Read and edit files, run shell commands in WSL and on Windows, and drive a headless browser for QA.
  • Talk to the online shops through Model Context Protocol tool servers: search products, update descriptions and SEO fields, compare category trees between two stores, pull analytics.
  • Query a self-hosted SEO toolkit, again over MCP.
  • Reach us on a chat channel and in a web UI, and be scheduled to run on its own.

The agents keep a written handbook of the machine, chapter by chapter, and are instructed to read the relevant chapter before touching anything. Every fact in it was observed live on the box, and when a fact turns out wrong the agent fixes the source and re-imports it.

That handbook is the single most effective thing we have done for reliability. An agent that starts each session fresh but can look up “how does this box actually work” behaves very differently from one guessing from training data.

Automation: deciding what belongs where

Three different schedulers exist on this box, and the rule for which one owns a job is simple.

If the job can be drawn as a flowchart with no judgement calls, it is an n8n workflow. If it needs judgement, it is an agent automation that n8n can trigger. If it needs root, or must survive the Linux side dying, it is a Windows scheduled task.

Concretely:

  • n8n handles deterministic glue: a 30-minute internal health check of the containers, webhook plumbing from the shops, and future content pipelines. No model in the loop, zero token cost.
  • Agent automations run the things that need reading and judgement: a 30-minute heartbeat where the admin agent decides whether anything needs attention, a nightly memory consolidation pass, a nightly re-index of agent memory, and a weekly update sweep that checks every layer (Windows apps, WSL packages, both container engines, git repositories) and reports what is behind. It never installs anything. Applying an update is a human decision, taken deliberately, verified afterwards.
  • Windows tasks own backups, the nightly document sync, and, crucially, uptime monitoring.

A monitor has to sit outside its own failure domain. We proved this the hard way. A container-based monitor reported green through an entire outage because the container network was fine and the host network was not.

The route monitor now runs on the Windows side, every five minutes, and pushes alerts to a chat channel. A second monitor on a different machine is the next step.

Backups

Nightly: both Postgres databases are dumped and compressed, the vector store takes a full snapshot, configuration and the in-house apps’ source and data are copied, all to the archive drive with two weeks of retention. Restores have been tested against a throwaway database, which is the only way to know a backup is real. Selected projects also push to private off-site git repositories on meaningful changes.

What it is used for today

  • Product knowledge. Staff ask the assistant about products and get cited answers from the actual data sheets in seconds instead of opening a dozen PDFs.
  • Web shop maintenance. The admin agent fixes bugs, writes and translates product copy in Icelandic and English, tunes SEO metadata and audits indexing, all through tool servers rather than by hand.
  • A digital asset manager. A small in-house Node app for marketing assets and ad bookings, running in a container, maintained entirely by the agent.
  • Creative work. Catalogue images, ad variations, short product clips.
  • Operations. The server looks after itself: health checks, update sweeps, backup verification, and a written record of every incident.

What is next

  • Generated technical brochures. The retrieval corpus is good enough that per-product summaries could be drafted automatically and reviewed by a human. Parked for now in favour of getting search right first.
  • Event-driven agents. Deterministic checks waking an agent only when something is actually wrong, instead of on a timer.
  • A second, independent uptime monitor on a different machine.
  • Role-based access on the VPN before more staff get accounts, so that the search and chat front ends are open to everyone and the admin surfaces are not.
  • Keep chasing models. The gap between a 27B local model and a hosted frontier model has narrowed a lot in a year. Every few months we re-run the same document questions against the new releases and move the default when one wins.

If you are building something similar

  1. Buy VRAM. Everything else is secondary.
  2. Run one model runner and one vector store. Resist the urge to have two of each.
  3. Keep GPU-heavy processes native and everything else in containers.
  4. Never expose anything to the public internet. A mesh VPN with identity is cheap and removes a whole category of worry.
  5. Pin your container images. Floating tags will update themselves on the next pull and take your evening with them.
  6. Put your monitoring somewhere the outage cannot reach.
  7. Write down how the box works, in a place both humans and agents read. Then keep it true.

None of this needed a rack, a cluster, or a cloud contract. It needed one card with enough memory, one runner, one vector store, a private network, and a handbook that stays true. The rest is maintenance, and the box now does most of that itself.

Built and maintained in-house.