Skip to content
Iterus

On-Premise LLM Deployment

On-premise LLM deployment means queries and data stay on your own hardware — nothing leaves the company network. We deploy open-source models on your own server or workstation; on-premise deployment starts from 350,000 CZK (hardware not included). The architecture is the same as for a cloud API integration — only the model runs on your premises.

Who is it for?

Companies where sensitive data — health records, contracts, internal documents — must not leave the network, yet want AI for search, categorisation or answering questions over their own documents. If the data may leave the network and price and answer quality matter most, a Claude or OpenAI API integration is usually the better fit — we'll say so on the call.

What can we show?

We don't run an on-premise deployment for a client yet, and we don't claim otherwise. What we have is the architecture it stands on: layered conversation memory and switching between model providers run in production in Innea; in a local deployment, what changes is the model provider and where it runs, not the application built around it.

Frequently asked questions

What hardware do we need?

For a smaller model, one server or workstation with a capable GPU; a larger model or more concurrent users means more GPUs or more machines. We size the model to the task, not the other way round — on the call we first establish what the model must do and derive the hardware from that, which you buy yourself.

Can a local model do what ChatGPT does?

Not for general knowledge and open conversation — a smaller open-source model only keeps up with the cloud where it has a narrowly defined task. Document search, categorisation, answers grounded in your data: those it handles well, without the data leaving the network. Where exactly the line runs for your task we verify on a sample of your data as part of discovery, before anything is bought.

Who updates the model and who watches it?

We do — monitoring and an update plan are part of the project, not an add-on: you get the same automated tests and runtime watching we have on our own products. Ongoing support after handover is agreed separately — from "call us when something happens" to regular maintenance.

Does it work without an internet connection?

Yes, the whole thing can run disconnected from the internet — that is usually the main reason companies want a local deployment. Just expect to bring model and application updates in by hand; we design a procedure that manages that without opening the network.