Support and sales assistants
A chatbot that answers from your documentation and product data, hands over to a person when it should, and logs the questions it left open.
AI integration
Assistants, document search and automation built into what you already run. Cloud models where they fit, EU-hosted models where residency matters, or a private LLM on your own servers with Kipper AI.
What we build
A chatbot that answers from your documentation and product data, hands over to a person when it should, and logs the questions it left open.
Retrieval over contracts, manuals, tickets or wikis, with the source shown for every answer. Runs against your files where they already are.
Classification, extraction, summaries and drafting built into the tools your team already uses, with a human in the loop where it matters.
Kipper AI installs Ollama and a chat UI on your own servers with one command. You pay for the server, keep your own keys, and your data stays inside the cluster.
Model and data residency
The choice is made once, in writing, before any code. It follows your data-protection requirements and your budget, in that order.
| Models | Where it runs | When it fits |
|---|---|---|
| Claude, OpenAI, Gemini | Provider clouds | Strongest models, fastest to ship. Data leaves your environment under the provider's terms. |
| Mistral | EU-hosted | Good models with European hosting and contracts. The default when EU residency is a requirement. |
| Kipper AI (Ollama) | Your own servers | Open models running inside your cluster. Usable from 16 GiB RAM, fast with a GPU. Everything stays on the server. |
From the work
Our own HR product ships assistants for onboarding and document drafting on EU-hosted models, with the data staying in German data centers. The same pattern is available for your software.
See PersoHRBefore you call
Wherever you decide. Provider clouds, an EU-hosted model, or a model running inside your own cluster. We have shipped set-ups where all data stays inside the EU and set-ups where all data stays on one server.
The installer refuses to run with less than 8 GiB of free memory on a node. For anything beyond a demo, plan on 16 GiB RAM and 4 vCPUs; a GPU with 16 GiB or more of VRAM makes chat fast and lets you run larger models.
Yes, so it shows its sources, stays inside the documents you give it, and hands over to a person for anything it cannot back up. We measure answer quality before and after launch rather than promising accuracy.
A short assessment to pick the use case and the model, then milestones like any other software work. Running costs are either the provider's per-token bill or your own server, and we show both before you choose.
Talk to an engineer
Thirty minutes with a senior engineer. Reply within one business day.