Book a call
Our own operation

Our stack — what actually runs

No PowerPoint architecture, no wall of logos. This is running right now on our server in Germany — the same stack we use to build and operate our own products.

Anyone advising companies on should be able to show that systems work in continuous operation, not just in a demo. So we open up our own operation: the same requirements we meet for clients — availability, data protection, cost and control.

Every number here is a snapshot of the running system, not a target state. And where something did not work, that is on this page too.

Figures as of September 2026

The server

More than 50 separate services run on our own server in a German data centre — each in its own container, each updatable on its own, all reachable encrypted behind a .

Three machines, cleanly separated

  • Server in a German data centre — runs the client applications around the clock.
  • Mac Studio on our own premises — carries the load: language models, image analysis, company memory.
  • Development machine — where things are built and checked before anything ships.

The separation is the actual point: whatever needs computing power runs on our own premises, whatever needs to be reachable sits in the data centre. Which software runs in which version and where is deliberately not on this page — a build manual for attackers does not belong in marketing.

Self-hosted AI without the cloud

The models run on our own machine, not in someone else's cloud. Day to day we use a model with 27.3 billion parameters that also understands images. For code and research there is one with 79.7 billion, and search across our own company memory runs on a model. The holds around 262,000 tokens — enough for long documents in one piece.

What the models do for us

Blog pipeline

The pipeline exists and has run: specialist articles in a two-pass process — outline first, then full text — image added, published without a manual step. In a six-week continuous run it produced 47 articles. We set up the same pipeline for your topics.

Social media content

Copy for Facebook, Instagram and TikTok is produced the same way; images and videos come from connected services. No post goes out before a human has approved it.

Legal-text preparation

Researched legal content is formatted into finished blog posts — the research itself deliberately does not run on the local model.

Why we moved from Gemma to Qwen

  • Until May 2026 a model with 26 billion parameters ran on the server, purely on the CPU and without a graphics card. For standard German copy that was good enough.
  • The limit showed up in operation, not in testing: invented details in a production run, and weak behaviour as soon as a task ran across several steps and tools.
  • Moving to the Qwen family brought both — considerably more parameters and a that holds a long document in one piece. Plus the move from the CPU to the graphics unit of the Studio machine: minutes became seconds.
  • This is not a closed verdict. We keep checking whether a newer model does the work better, and we switch when the answer is yes. What runs today is the current state of a series of decisions, not their end.

What deliberately does NOT run on the local model

Legal research

Too high a risk of invented case references. For that we use a service that finds and cites real sources.

Software development

For our codebase the local model is too slow and too imprecise. We use a instead.

Client data

Goes through the database, never through a language model.

That dividing line is the real point: is strong at producing text at volume and weak wherever verifiability matters. Anyone selling it the other way round has not run it in production.

Automation — 34 workflows

is our nervous system: 34 active workflows handle everything that has to run without being asked.

Content and marketing

Blog publishing, post and story generation, scheduled publishing across several channels, drafts for approval and a regular legal-text monitor across six sources.

Deployment

A push to the main branch triggers build, restart and notification — nobody has to ask for a deploy.

Integrations

Access tokens renew themselves; media uploads and syncing into the company memory run automatically.

Business processes

Client onboarding with master data and strategy, competitor monitoring, video production for hospitality clients.

Operations

One central error handler plus fallbacks — nothing dies quietly.

What makes the difference

  • Every has error handling. Failures get reported, not swallowed.
  • Deployments are triggered through Git, not started by hand.
  • Routine changes from our agents are merged and shipped automatically once the test run is green.
  • All credentials encrypted and — no cloud subscription.

AI agents in our development process

Our products are built with the help of agents. Several specialised instances divide the work across issues, pull requests and automated tests — under graduated autonomy rules.

Graduated autonomy

Routine changes run through fully automatically. Changes to the core need human sign-off. The level depends on the matter, not on the mood.

Semantic company memory

A knowledge base with vector search and a that every agent can reach across machines. Decisions and operational knowledge stay retrievable instead of vanishing into chat logs.

Human approval where it counts

Content for client channels only goes out after review. The machine produces, the human decides.

When something breaks, we know within 60 seconds

Six containers work purely on observation — server metrics, container health, log analysis, realtime measurement and an external uptime check. and form the backbone.

On top of that, the layer reports for itself: every error and every release announces itself.

What this means for you

We run this stack for our own products. That is why we know from experience rather than from brochures:

  • how to keep dozens of containers stable, with updates, backups and uptime;
  • how to run language models yourself — on your own hardware, without a cloud connection, and what switching models looks like in live operation;
  • how to automate processes that used to be manual work;
  • how to set up monitoring that spots problems before clients do;
  • and where the limits are — what self-hosting does well and what it does not.

When we build an solution for you, it runs on infrastructure that has proven itself in daily use, not on a whiteboard. What comes out of it is open to inspection too: our consent tool “KaaTai Consent Manager” is listed in the official WordPress directory at wordpress.org/plugins/kaatai-consent-manager and has been downloaded more than 880 times since April 2026.

Ready for the next step?

Free intro call, no strings attached. In 30 minutes you'll know whether and how AI can help your business.

Book a callBAFA funding