Our stack — what actually runs
No PowerPoint architecture, no wall of logos. This is running right now on our server in Germany — the same stack we use to build and operate our own products.
Anyone advising companies on should be able to show that systems work in continuous operation, not just in a demo. So we open up our own operation: the same requirements we meet for clients — availability, data protection, cost and control.
Every number here is a snapshot of the running system, not a target state. And where something did not work, that is on this page too.
Figures as of September 2026
The server
More than 50 separate services run on our own server in a German data centre — each in its own container, each updatable on its own, all reachable encrypted behind a .
Three machines, cleanly separated
- Server in a German data centre — runs the client applications around the clock.
- Mac Studio on our own premises — carries the load: language models, image analysis, company memory.
- Development machine — where things are built and checked before anything ships.
The separation is the actual point: whatever needs computing power runs on our own premises, whatever needs to be reachable sits in the data centre. Which software runs in which version and where is deliberately not on this page — a build manual for attackers does not belong in marketing.
Self-hosted AI without the cloud
The models run on our own machine, not in someone else's cloud. Day to day we use a model with 27.3 billion parameters that also understands images. For code and research there is one with 79.7 billion, and search across our own company memory runs on a model. The holds around 262,000 tokens — enough for long documents in one piece.
What the models do for us
Blog pipeline
Social media content
Legal-text preparation
Why we moved from Gemma to Qwen
- Until May 2026 a model with 26 billion parameters ran on the server, purely on the CPU and without a graphics card. For standard German copy that was good enough.
- The limit showed up in operation, not in testing: invented details in a production run, and weak behaviour as soon as a task ran across several steps and tools.
- Moving to the Qwen family brought both — considerably more parameters and a that holds a long document in one piece. Plus the move from the CPU to the graphics unit of the Studio machine: minutes became seconds.
- This is not a closed verdict. We keep checking whether a newer model does the work better, and we switch when the answer is yes. What runs today is the current state of a series of decisions, not their end.
What deliberately does NOT run on the local model
Legal research
Software development
Client data
That dividing line is the real point: is strong at producing text at volume and weak wherever verifiability matters. Anyone selling it the other way round has not run it in production.
Automation — 34 workflows
is our nervous system: 34 active workflows handle everything that has to run without being asked.
Content and marketing
Deployment
Integrations
Business processes
Operations
What makes the difference
- Every has error handling. Failures get reported, not swallowed.
- Deployments are triggered through Git, not started by hand.
- Routine changes from our agents are merged and shipped automatically once the test run is green.
- All credentials encrypted and — no cloud subscription.
AI agents in our development process
Our products are built with the help of agents. Several specialised instances divide the work across issues, pull requests and automated tests — under graduated autonomy rules.
Graduated autonomy
Semantic company memory
Human approval where it counts
When something breaks, we know within 60 seconds
Six containers work purely on observation — server metrics, container health, log analysis, realtime measurement and an external uptime check. and form the backbone.
On top of that, the layer reports for itself: every error and every release announces itself.
What this means for you
We run this stack for our own products. That is why we know from experience rather than from brochures:
- how to keep dozens of containers stable, with updates, backups and uptime;
- how to run language models yourself — on your own hardware, without a cloud connection, and what switching models looks like in live operation;
- how to automate processes that used to be manual work;
- how to set up monitoring that spots problems before clients do;
- and where the limits are — what self-hosting does well and what it does not.
When we build an solution for you, it runs on infrastructure that has proven itself in daily use, not on a whiteboard. What comes out of it is open to inspection too: our consent tool “KaaTai Consent Manager” is listed in the official WordPress directory at wordpress.org/plugins/kaatai-consent-manager and has been downloaded more than 880 times since April 2026.
Ready for the next step?
Free intro call, no strings attached. In 30 minutes you'll know whether and how AI can help your business.