LLM integration: AI reads technical documentation, circuit diagrams and manuals

What RAG Is — and Why the Model Does Not Learn Your Data

Retrieval-augmented generation in one paragraph

With RAG (retrieval-augmented generation) the language model is not trained on your data. Instead, your documents are split into sections, stored as vectors in a search database, and for every question the most relevant sections are retrieved. Only these sections are handed to the model together with the question — and it formulates an answer with a source reference.

This has three advantages: New documents are searchable immediately, without training. Every answer can be traced back to its source. And access rights are preserved — anyone who may not see a document does not get an answer from it either.

How We Proceed: Five Steps to a Knowledge Base

From the first conversation to operation

StepWhat happensDuration
1. Review sourcesWhich systems (file server, SharePoint, Confluence, ERP, ticket system), which formats (PDF, Word, Excel, e-mail, scans), which permissions. We tell you honestly what is worthwhile and what is not.1 week
2. PrototypeA subset of your data, a local model, a simple chat interface. You ask real questions from daily work and see how good the answers are.2–4 weeks
3. Integration of sourcesAutomatic import from your systems, permission mapping (e.g. from Active Directory), processing of tables, scans and images, updates on change.3–8 weeks
4. Integration into your toolsChat in the intranet, in Teams, in the ERP or as an interface for your own software. Voice control or connection to exhibits on request.2–4 weeks
5. OperationModel updates, quality measurement based on real questions, extension to new sources. Runs on your hardware, maintained by us.ongoing
AI on your own hardware: from industrial PC to workstation

Local, Cloud or Hybrid?

An honest comparison

Local open-source models (Qwen, Llama, Mistral, Gemma — run via Ollama or vLLM) operate entirely on your hardware. Your data does not leave the building, there are no API costs per request, and the solution works without an internet connection. Answer quality is below the large cloud models but sufficient for document search, summaries and internal assistants in most cases — above all because with RAG the quality of the retrieval matters more than the quality of the model.

Cloud models (OpenAI, Anthropic, Google) deliver the best language quality and are quick to integrate, but send every request including document excerpts to the provider — feasible with a data processing agreement, yet a deal-breaker for many clients.

Hybrid means: retrieval and sensitive data local, only non-critical phrasing tasks in the cloud. We build all three variants and recommend according to your protection needs, not according to fashion.

What Hardware Does a Local LLM Need?

Reference values for operation in a company

UseHardwareModelsHardware price (approx.)
Prototype / up to 5 usersWorkstation with one GPU (16–24 GB VRAM)7B–14B models, e.g. Qwen 2.5 14B, Llama 3.1 8B€3,000 – 6,000
Department / up to 50 usersServer with one GPU (48–80 GB VRAM)30B–70B models, quantised€10,000 – 25,000
Company-wideServer with 2–4 GPUs, redundant70B models at full quality, several models in parallel€30,000 – 80,000
Search only, no chatexisting server, CPU is enoughEmbedding model without LLMusually €0
The search database itself is undemanding — a PostgreSQL with pgvector on existing infrastructure is often enough. Costs only rise with concurrent users and long answers, not with the amount of data.
Data protection and EU AI Act for local AI solutions

What a Knowledge Base Costs

And why projects fail in practice

A feasibility prototype with a subset of your data starts at at² under €3,000. A productive knowledge base connected to two or three source systems, with permission mapping and intranet integration, typically ranges from €15,000 to €50,000, plus the hardware from the table above. Running costs: electricity, maintenance, model updates — no fees per request.

What projects fail on in practice is rarely the model: It is outdated documents that contradict current ones, scans without text, tables stored as images, and permissions nobody maintains any more. That is why every project with us starts with a review of the sources — and an honest statement about which of them need cleaning up.

Frequently Asked Questions About AI Knowledge Bases With Local LLMs

What clients ask us about RAG, Ollama and data protection

How can I make my company's internal knowledge searchable with AI? +
With a RAG knowledge base: Your documents from file servers, SharePoint, wikis or ERP are transferred into a vector search database. A language model answers questions in natural language and cites the source. The model does not need to be trained for this; new documents are available immediately. We implement this with locally hosted models so that no data leaves the company.
Can LLMs really be run without the cloud? +
Yes. Open-source models such as Qwen, Llama, Mistral or Gemma run entirely on your own hardware via Ollama or vLLM — from a workstation to a GPU server. There is no external connection, no cost per request and no data processing agreement to arrange.
How good are local models compared to ChatGPT? +
For free text generation the large cloud models are ahead. For a knowledge base, however, what matters most is whether the right document sections are found — and that is done by the retrieval, not the model. With a good 14B to 70B model you get answers for document search, summarisation and internal assistants that are perfectly adequate in daily use. We show you the difference in the prototype using your own questions.
Which sources can you connect? +
File servers and network drives, SharePoint and OneDrive, Confluence and other wikis, e-mail mailboxes, ticket systems, ERP and CRM databases, websites and manuals — as PDF, Office documents, scans (with text recognition), tables and images. Whatever exists as a database or interface, we can connect.
Are access rights preserved? +
Yes, that is mandatory. We take over the permissions from your source systems (e.g. Active Directory, SharePoint permissions) and filter with every question before the model sees any documents. An employee only receives answers from documents they would be allowed to open directly.
How much does a RAG knowledge base cost? +
The feasibility prototype starts under €3,000. A productive solution with connection to two or three source systems, permission mapping and intranet integration typically ranges from €15,000 to €50,000, plus hardware between €3,000 (workstation) and €25,000 (departmental server). After the free initial consultation we give you a specific figure.
How long until the first result? +
A prototype with a subset of your data is ready in 2 to 4 weeks. The productive solution including source integration takes 2 to 4 months depending on the number of source systems.
Is a local AI knowledge base GDPR-compliant? +
It is the most privacy-friendly variant: no transfer to third parties, no data processing agreement, data minimisation through permission filters and logging of access. For the EU AI Act we classify your project into its risk class — internal knowledge assistants generally fall into the lowest.

The Free AI Check for Your Knowledge Base

In 30 minutes you know whether it pays off for you

Tell us where your knowledge lives and who is looking for it — we tell you whether a local knowledge base makes sense, which hardware you need and what the prototype costs. More about our AI development.

Software development since 1999 - Made in Germany

Made in Germany

Software development since 1999

Developed 100 % in Germany — at our Nuremberg and Kempten (Allgäu) locations. Over 1,170 completed projects, with 82 % of our revenue coming from repeat customers. If you're nearby: stop by for a coffee.