Souryan Augé

// data in the AI era, deployed in your company

The converters on this site are free, open source, and deliberately small: one document in, clean data out. This page is about the other thing — the several thousand documents you already have, that nobody can search, and that hold most of what your firm knows.

Document migration

Reports, contracts, minutes, scans going back years. Every one of them readable by a person and by nothing else. Turning that into something a team — or a model — can actually use is a project, not a conversion, and it is what I do.

Dormant documents are lost expertise

You have probably seen all three. Someone leaves, and the reasoning behind a decision leaves with them — the document is still on the server and nobody can find it. A new hire spends three months relearning what the firm already knew. A meeting re-decides something that was settled in 2021, because re-deciding it was faster than looking.

None of that is a filing problem, and buying more storage does not touch it. The documents are there. What is missing is any way to ask them a question.

What the service does

It takes a corpus you already own and turns it into something answerable: read, cleaned of the noise that scanning leaves behind, structured so a machine can navigate it, and tagged according to how your firm actually uses its documents — which is a study before it is a conversion.

The result is a knowledge base your team can search in their own words, and that a language model can answer from without inventing the parts it does not know.

In the cloud, or entirely on your own machines

This is a real choice and not a reassurance, because the engine is already published. The core of the converters on this site is open source under Apache-2.0 and runs on your own infrastructure with your own API keys — that is not a roadmap item, it is the thing this site was built on top of.

So if your documents may not leave your building, they do not leave your building. And if you would rather do the whole thing yourself, the same code is there and costs nothing — I would rather tell you that than sell you a migration you do not need.

What it costs

Quoted on the corpus, sized for a small structure. The price depends on how many documents there are, what state they are in, and how much of the tagging needs a human decision — which is why there is no rate card on this page: a number here would be a guess about your archive.

What it is not: a seat licence, a platform subscription, or a per-user fee that grows when you hire. You pay for a migration, once, and you own what comes out of it.

One person, who advises first

You would be dealing with me, from the first conversation to the delivery. That matters less as a promise than as a constraint: I cannot take on work I do not understand, so the first thing is always a look at what you have and an honest answer about whether this is worth doing.

The clearest evidence I can offer is what the free tools deliberately do not do. They convert, and they stop there — no tagging, no classification. Not because it is hard, but because tagging without knowing how your firm uses its documents produces categories nobody uses. That study is the work, and it is why this is a conversation before it is a quote.

A worked example, on public data

Rather than a client story you cannot verify, the demonstration runs on a public corpus — the kind of archive anyone can download and check the result against: a few hundred scanned administrative documents, of the sort a small firm accumulates.

What goes in: PDFs, some born-digital, most of them scans of varying quality. What comes out: clean structured text, one record per document, each block carrying its page. What becomes answerable: which documents mention a given obligation, what changed between two versions, and where a decision was first written down.

The specific corpus is still being chosen, and this page will name it and link the result when it is. Until then this describes the method rather than a finished engagement — there is no client to name, and inventing one would undo the only argument this page has.

Talk to me about your documents

The first step is a conversation, not a quote. Tell me roughly what you have — how many documents, what kind, what state they are in, and what you wish you could ask them — and I will tell you honestly whether this is worth doing and what it would involve.

Write to me about your documents

This opens your own mail client with a subject already filled in. Nothing is sent until you send it, and nothing about you reaches this site.

Or write to: datadosomething.doctrine508@passmail.net

And if you are not ready, there is nothing here to dismiss. No form pops up, no reminder follows you around. The page will still be here.