Skip to content
← All posts

How Piloti works

Build logEntry 02By Jonathan Uhlemann, Matthias Bigl and Ferdinand Rubenbauer

Risograph: six sheets floating one above another as an exploded drawing, from sources through preparation, index, search and check to an answer with citations.

A typical Piloti question looks harmless. “Which escape route widths apply to this kindergarten?” Behind it sit an OIB guideline, one of nine state building codes, the project brief, and possibly a decree that changed last spring. The answer is only worth something if it names the clause it came from, quotes it correctly, and reflects the version that is actually in force.

Getting that right is not a prompt. It is a pipeline with six layers, and this post walks through all of them, including the parts that did not work on the first try.

Every Piloti answer falls through this stack. Each layer only trusts what the layer above already cleaned up.

Reading the stack top to bottom: sources are the documents as your office keeps them. Ingestion turns those documents into machine-searchable pieces. The index stores every piece twice, once as a semantic fingerprint and once as plain searchable text. Retrieval asks both representations at the same time for every question. Verification checks the evidence before it is allowed to appear. And the answer is what you see: a response that carries its sources with it.

The two dashed boxes feed the answer without being layers themselves: a project memory that keeps later conversations from starting at zero, and the lessons distilled from thumbs-down reports across the whole platform. Both get their own section below; the lesson loop even gets its own post.

Where knowledge lives

Early on we made a decision that shaped everything after it: we do not ask offices to reorganize their documents for us. Knowledge is created wherever it is convenient. Building law lives in official PDFs, office standards live in folders that grew over fifteen years, project knowledge lives in briefs and email attachments. Forcing all of that into one new structure is the kind of migration that gets planned twice and finished never.

Two structural facts organize what Piloti ingests. The first is what a source is: building law, office knowledge, project document, or web. Every source in an answer carries a chip naming its kind, so you always see what class of evidence you are looking at.

The second fact is the important one, and it is a hierarchy: where a source lives. Knowledge sits on four nested levels. At the outermost level, the building law everyone shares. Inside it, your office’s own archive, which no other office can see. Inside that, each project with its documents, visible to its members. Innermost, the single conversation with whatever you attached to it.

A question searches outward from where it is asked: its conversation, its project, your office, the shared law. It never searches inward from outside, and never sideways into another office.

The nesting is what makes the answer trustworthy in both directions. Outward, it means one question reaches everything relevant to it at once, from the sketch you attached a minute ago to the OIB guideline everyone shares. Inward, it means the boundary holds: your office’s knowledge improves your office’s answers and nobody else’s.

One rule sits above all the automation: when a person classifies a document, that classification wins. A filename can suggest what a document is. It cannot outvote the architect who knows.

Most building law arrives as PDF, and PDF fights back. This is the ingestion layer, and it extracts page by page. Text is pulled with watermarks stripped, tables are extracted separately so a load table stays a table, and plan pages are rendered and described by a vision model, which turns “the site plan on page 12” into text a search can actually find. The result is split into chunks of roughly a thousand tokens with overlap, so each piece is small enough to match precisely and large enough to keep its context.

Each chunk is then stored twice in the index layer. Once as an embedding, a numeric fingerprint of its meaning that lets the system find a clause even when the question uses entirely different words. And once as plain text in a full-text index, so an exact term can be found exactly. Keeping both is deliberate, and the next section is about why.

The official guidelines get one extra layer of care. Piloti syncs OIB documents against a registry of content hashes, so a revised guideline replaces its predecessor instead of quietly coexisting with it. Two versions of the same fire protection rule in one index is how you end up citing the wrong one with full confidence.

This is also where we part ways with a ranking trick that is popular in systems built on chat archives: letting newer content outrank older content, because chatter goes stale. Law does not go stale, it gets replaced. A newer text is not a more in-force text, so nothing in our document ranking rewards age. Versions are resolved at ingestion, where the revision replaces its predecessor outright, and what remains in the index is current by construction. The same rule covers your own uploads: a document re-uploaded under the same name replaces its earlier version in the index instead of competing with it.

Why vector search alone was not enough

Our first retrieval was pure embedding search. It was impressive for a week. Embeddings catch paraphrase, which matters in a field where the person asking and the clause answering rarely share vocabulary. Someone asks about wheelchair access, the clause says “stufenlos erreichbar”, and vector similarity connects the two anyway.

Then the failure cases arrived. When a planner types “OIB-Richtlinie 4, Punkt 2.1.1”, semantic similarity has no business deciding what comes back. That question has exactly one right document, and an exact match on the reference is the strongest evidence there is. So the retrieval layer runs every question through two channels at once, a semantic one for paraphrase and a lexical one for norm references, section numbers, and defined terms, and the candidates are merged.

Each channel covers the other's blind spot: embeddings find paraphrase, exact matching finds section numbers. The cap keeps a 200 page guideline from crowding out a two page decree.

One more lesson from production: long documents are bullies. A 200 page guideline can fill the entire candidate list on its own and crowd out the two page decree that actually decides the case. We cap how many chunks a single document may contribute to one answer.

An answer you can check

Retrieval finds candidates. It does not make them true. The verification layer is where evidence earns the right to appear: every Piloti answer cites its sources down to the document and page, and quoted passages are checked against the source text before you see them. If a quote cannot be verified, it does not ship.

This is the part we are least willing to compromise on. In building law, “trust me” is not a feature. An answer you cannot trace is an opinion, and our users can get opinions for free.

Memory that earns its place

Answering one question well is table stakes. The more interesting problem is the fifth conversation about the same project, which should not start from zero.

After every answer, a reflection step runs in the background and writes down what mattered: decisions, constraints, corrections. It runs after the reply has already reached you, so it never costs you a second of waiting, and it checks its finding against what is already remembered, so the memory does not fill up with restatements of itself.

When the next question comes in, the stored notes compete for a small budget of space, scored on three things at once: how relevant they are to this question, how important they were judged to be when written, and how recently they last helped. Notes that keep proving useful stay vivid. Notes that never help sink down the ranking, though irrelevance alone never deletes them, because the note you have not needed for a year is sometimes the one that decides a dispute. What does retire a note is being corrected: when a new finding contradicts an old one, the old one is superseded by it, with the record of what replaced it. Your thumbs-down reaches the memory as well: a complaint that sits semantically next to a stored note costs that note standing and confidence, inside your own project and nowhere else. And this is the one place in Piloti where recency gets a vote at all, deliberately the smallest of the three.

Two more rules keep the memory honest. It never outranks what you type in the moment; the live turn always wins. And it is never a citation source; memory shapes an answer, the cited clause still has to come from a document.

We did not invent this scoring. The shape comes from Stanford’s Generative Agents work, the decay curve from the MemoryBank paper. Borrowing a published, tested design beat a weekend of our own intuition, and we suspect that is true more often than engineers like to admit.

And when the answer is still wrong

Every layer above exists to make a bad answer rare. It will not make bad answers impossible, and pretending otherwise would be the least trustworthy sentence on this page. What we can promise is what happens next: under every answer sits a thumbs-down button, and behind it a pipeline that turns your report into an anonymized, generalized lesson that every future answer knows about, for every office on the platform.

That loop deserves more than a paragraph, so we gave it its own post. The short version: your report stays private, the improvement does not, and a failure that one user already paid for should never be paid for twice.

What holds it together

The knowledge base works because it meets documents where they already live, because every answer can be checked against its source, and because a human correction is treated as a bug to be fixed once, not as feedback to be averaged. None of it is finished. The retrieval channels can get smarter, the memory can get sharper, and the lesson loop has only begun to accumulate.