Skip to content
Lindner Tech

NewLocal AI & RAG Systems

Local AI for Your Business: What Hardware Do You Really Need?

Local AI without a data centre: what the GPU, video memory, RAM and CPU do, pilot vs. full operation, own server or EU data centre.

Markus Lindner · 7 min read

"Do we need a data centre in the basement for local AI?" I hear this question a lot when the topic is GDPR-friendly AI on company documents. The honest answer: usually not. A RAG system that searches your contracts, manuals and quotes and writes answers with sources runs on a single server with a suitable graphics card for a small team. Here I explain, component by component, what each part does, when a pilot is enough and when a rented server in a German data centre is the better choice.

Why local at all: data protection in one paragraph

With cloud AI, your questions and documents go to an external provider, often based in the US. If the language model runs on your own hardware or on a dedicated server in the EU instead, the circle of parties stays small: no external AI provider, no transfer to a third country, and with a rented server a data processing agreement with the data centre. You still need a legal basis, an access concept and deletion rules. The details are in the article on GDPR-compliant RAG systems. You make the legal assessment together with your data protection officer.

Want to tackle this in your company?

Send me a short message or give me a call. I'll get back to you within 24 hours on weekdays.

Free consultation0162 8036 905

5.0 on Google · 50 reviews

15 minutes, no obligation · Reply within 24 hours on weekdays

Graphics card and video memory: the key component

A language model writes answers word by word, and to do that it should sit entirely in the graphics card's fast memory. That is why video memory (VRAM) matters most for local AI, more than raw computing power. If the model does not fit completely, tools like Ollama move part of it into normal system memory. That works, but answers become noticeably slower.

How much memory a model needs can be estimated roughly. At full 16-bit precision, every billion parameters takes about 2 GB. Quantisation, meaning the weights are stored more coarsely, cuts that considerably: a model with 8 billion parameters is just under 5 GB in the common 4-bit version, a model with 70 billion parameters a little over 40 GB. Quality often drops only slightly at 4 bits, but that should be tested with your own questions.

On top of that comes memory for the context: the question, the passages found and the answer being written. With RAG this context is larger than in a simple chat, because several document sections are sent along every time. The rule of thumb is therefore: model size plus headroom for context and simultaneous users.

System memory, storage and processor

The other components are less critical, but they still matter:

  • System memory (RAM): it holds the operating system, the search index, the database and the processing of new documents. Size it generously so the model has room while loading and somewhere to fall back to.
  • Storage: fast SSDs shorten model loading and index searches. Models take up anything from a few to many gigabytes depending on size, the index grows with your documents, and model versions kept for comparison need extra space.
  • Processor (CPU): it splits up PDFs and Office files, reads scanned documents with text recognition and runs the workflow. A current server processor is usually enough, because the graphics card does the writing.
  • Embedding model: it turns your documents into number vectors for the search and is tiny in comparison. A widely used open-source model for this is smaller than 300 MB and runs alongside the language model.

Pilot or full operation: what changes

In a pilot, two or three people test one use case with real documents. Questions rarely arrive at the same time, and short waits bother nobody. Existing hardware or a smaller model is often enough, and that is the point: you see on a real case whether the answers are useful before any hardware is bought.

In full operation two things change. First, several people ask at the same time, and every parallel request needs its own memory for its context. With Ollama, for example, memory needs grow with the number of parallel requests times the context length. For many simultaneous users there is specialised server software such as vLLM that manages this memory more efficiently. Second, expectations rise: answers should come quickly, the server should run reliably, and access should be logged.

How to find the right first use case is covered in the article AI in small businesses: where it really saves time.

Sizing by users and use case

A table along the lines of "this many employees, this much video memory" would not be serious, because three factors interact:

  • How many people ask at the same time? What counts is not the number of accounts but the peaks, for example in the morning or before deadlines.
  • How long are questions and answers? A contract question with a short answer needs less context than a quote draft spanning several pages.
  • Which language and what quality? Many smaller models write decent German, but for demanding texts a larger model can make the difference.

What you can reuse

That is why I measure during the pilot: how fast do answers come, how good are they, how much memory is in use? This determines the setup for operation, with headroom for the next use case. At BSU-Holding, a locally run RAG system creates draft quotes from contracts, service specifications and past quotes, and they only leave the company after approval. Long drafts like these place different demands on the hardware than a short answer from a manual.

Often more can be reused than expected: an existing server or workstation with a graphics card for the pilot, the network drive, SharePoint or DMS as the document source, the existing backup and the user accounts with their permissions. Frequently the only new piece is the AI server itself.

In your own office or in an EU data centre?

An AI server in the office fits if you have a server room or at least a well-ventilated rack and the data should not leave the building at all. Don't underestimate the side effects: a powerful graphics card under load draws a lot of power and produces heat and fan noise. Next to a desk or in a storeroom without ventilation it is not a good idea. An uninterruptible power supply and a clean connection to the company network belong with it, which is part of my network and server setup.

A dedicated GPU server in a German or EU data centre fits if there is no suitable room, several sites need access or you don't want to look after hardware. The data centre takes care of power, cooling and replacing hardware. You need a data processing agreement and an encrypted connection. The word "dedicated" matters: the server is yours alone, and the language model runs on it as your installation, not as an AI provider's service.

The software is the same in both cases. A pilot on rented hardware followed by your own server in-house, or the other way round, is therefore not a fresh start.

Operation: updates, backups and oversight

An AI server is a server like any other and needs the same care: security updates for the operating system and software, graphics driver updates that are tested before they go in, and monitoring that reports when memory or disk space runs low. I don't install new model versions blindly but test them first against a set of real questions.

For backups, the model matters less, since it can be downloaded again at any time, than everything created on your side: configuration, access rules, logs and the index. The same 3-2-1 rule applies as for the rest of your IT, including tested restores.

And regardless of the hardware: AI can be wrong, local AI included. RAG makes answers checkable because every statement names its source. A person still has to check them, especially before anything leaves the company.

Local AI does not need its own data centre, but it does need equipment that fits your use case. My advice: start small, with one use case and a pilot on real documents, and decide on the hardware only after measuring. Whether your existing equipment is enough or an EU server makes more sense is something we clarify in a free consultation. After that you get a fixed price for setting up your local AI and an honest assessment of the hardware you really need.

Local AI

Using AI in your business?

I'll show you on a real use case what local AI can do for you.

Free consultation

5.0 on Google · 50 reviews

15 minutes, no obligation · Reply within 24 hours on weekdays