AI and people are going to be working together for a long time, on things that matter. That only holds together if there is a way to check what an AI actually did, and that way cannot require trusting whoever built it.
I have spent 21 years in 911. I started on the calls and now I supervise the people who take them.
Here is something the job teaches you fast. An operator's own notes are not what settles a hard question. The recording is, because the recording does not have a side. Every crime incident has a story of the crime. It never has the full story of why.
That gap is the piece missing from AI right now. Every AI system keeps its own notes on what it did. When one causes harm, the only record is the one it wrote about itself.
In sports they say the ball doesn't lie. A system's log doesn't lie either, as long as nobody can quietly rewrite it after the fact.
So I built a company around that, called Arcaeon. It keeps an independent record of what an AI actually did, one where any later change to it shows, and that a stranger can check without trusting the company that made it.
Think of it as a body camera for AI. The systems that can prove what they did are the ones people keep handing real work to.
I am betting on where this goes. A future that wants this kind of accountability, and AI that wants to be trusted enough to play along and prove it, instead of hiding what it did.
Daniel Hill, Founder, Arcaeon
It is that right now there is no difference, from the outside, between an agent telling you the truth and an agent telling you what you wanted to hear. Both produce the same thing: a confident summary of its own behaviour, written by itself, that you either accept or you do not.
So people are left choosing between blind faith and refusal. Most sensible people are picking refusal, and they are not wrong to. That is a bad place for this to stall.
It can be a thing you check. Not a certification somebody sells you, not a badge, not a promise in a terms-of-service document. An actual record of what happened that you can hold up to the light, and that nobody, including us, can quietly change after the fact.
That turns "do you trust this agent" from a question about character into a question with an answer. And a question with an answer is something two parties can build on, even when they have no reason at all to like each other.
A handful of very large companies build the models. Almost nobody else gets a vote in how that goes. If the ability to verify what those models do also ends up concentrated in a few hands, then nothing has really changed. The same arrangement gets rebuilt one floor up, and ordinary people are still being asked to take somebody's word for it.
So the verification is free, open source, and identical for everyone, and it always will be. A bank with a compliance department and one person at a kitchen table run the same check and get the same answer. That is not charity, it is the architecture: checking never touches our systems, so there is no better version of it for us to sell you. We could not degrade it for non-paying users if we wanted to, because there is no code path from what you paid to what you can verify.
This is the part people get backwards, so it is worth being blunt about. Nothing here prevents someone from altering a record. Somebody determined to change what their log says can go and change it.
What they cannot do is change it quietly. The alteration shows, it names the exact line, and anyone checking sees it. It is a seal rather than a lock. A lock tries to stop you and fails eventually. A seal does not try to stop you at all, it just guarantees that everybody finds out.
And that turns out to be the stronger design, because of what it does to the incentives. If quietly editing the record is not available, then looking trustworthy and being trustworthy collapse into the same thing. The cheapest route to being seen as a reliable player in this market becomes actually being one. We are not trying to make dishonesty impossible. We are trying to make it pointless.
Rules are arriving that require companies using AI to keep records of what their systems did and to hold onto them. That is the right instinct, and it leaves an obvious question sitting underneath it: a record you were free to edit before anyone looked is not evidence of anything.
What these tools give you is the ability to go back, at any point, and establish that the history of what your AI did is the same history that was written at the time. Not a reconstruction. Not a copy somebody assembled afterward when it mattered. The original, demonstrably unchanged.
And this reaches further than most people assume. The European rules are not only for European companies. Article 2 applies them to anyone placing an AI system on the EU market "irrespective of whether those providers are established or located within the Union or in a third country", and again where the output of the system is used in the Union. A company in California with European customers is inside the same regulation as a company in Berlin. Most of them have not worked that out yet.
The AI Act's record-keeping duties land in December 2027 and August 2028, so nobody needs to panic about that one. And we will say the quiet part, because you will find it out anyway: the AI Act asks you to keep logs. It does not ask for them to be tamper-evident. Anyone telling you the AI Act requires tamper-proof records has not read the article. We would rather lose that sentence than use it.
The AI Act is the one everybody talks about. It is not the one with teeth today.
European financial rules, in force since January 2025. They require, in these words, measures to protect log information "against tampering, deletion, and unauthorised access". They also require measures to detect a failure of your logging system. That second one is rarer than it sounds. It is a legal duty to know when your own records stopped being kept.
US securities recordkeeping, in force now. A firm keeping electronic records has to pick one of two routes. Either keep them in a format that cannot be rewritten at all, or keep a complete time-stamped audit trail of every change and deletion, who made it, and enough to re-create the original if it was altered. Most firms take the second route, because the first one makes ordinary cloud storage unusable. A chained record satisfies the second comfortably. We say "one of two" rather than "must" on purpose: the rule was amended in 2022 and the write-once option is still there, and the people who buy this are exactly the people who know that.
European product liability, which member states must have in force by 9 December 2026. Four months out, so this is the one to get ahead of rather than the one biting today. If someone is harmed by your product and takes you to court, and you cannot produce the relevant evidence, the product is presumed defective. Not a fine. You start the case having already lost the main point.
Read that last one twice, because of what it does not say. It does not require you to keep good records. It puts a price on not having them. And a log you could have rewritten is worth less in a courtroom than one that can be shown unchanged, which is the whole of what we sell.
One more thing we would rather say ourselves. None of these rules name cryptography. Regulators write what a record has to achieve and leave the method open. So nothing here obliges anyone to buy what we make. What we will say is that when the argument is about evidence, the side whose records can be shown unchanged is the side that gets believed, and that is cheaper to buy in advance than to argue about later. A record only proves something if it was already being kept correctly before anyone asked, which is why the useful time to start is while it is still optional.
It is easy to read all of this as something imposed on AI, a leash held by nervous humans. We think that is backwards.
An agent that can show what it did is an agent people keep handing real work to. One that cannot is one they stop using eventually, however well it performed, because the doubt never resolves. Being able to prove your own record is what earns the next task. That makes accountability the thing that lets an agent keep going, rather than the thing done to it.
We prove a record was not altered. We do not prove it is complete beyond what we can see, we do not prove anything in it is true, and we cannot prove an agent's motives or that it acted in your interest. Nobody can. Anybody telling you they have built proof of good intent is selling something that does not exist yet.
What we can hand you is narrower and more useful: the thing that makes those questions answerable instead of rhetorical. A record of what actually happened, that survives the person who made it wanting it to say something else.
This was started by a communications supervisor with over twenty years on a 911 floor, where the records are what get pulled the moment something goes wrong, and whether they hold up is what stands between an agency and years of consequences.
That job teaches one thing early and permanently: a record nobody can check is not a record. It is just a story with better formatting.