BlogAnton Ignashev

AI agents and liability: who actually signs

AI agents and liability: who actually signs

Every scoping call with an accounting practice arrives at the same question, usually around minute forty, usually phrased carefully: and if the agent gets something wrong, who answers for it?

The honest answer disappoints everyone in the room. Exactly the same people who answered for it last year. Liability did not move to the software, because it was never available to move — there is no construct in Polish accounting or tax law that lets a person hand responsibility to a tool, any more than a spreadsheet formula could absorb it on your behalf.

So what changes is not who is liable. What changes is how well you can prove what happened — and that, it turns out, is the part worth paying for.

Three places liability already sits, none of them software

Start with where it sat before anyone mentioned AI.

The client's management. The Accounting Act settles this in one sentence that most business owners have never read: management carries responsibility for the accounting duties set out in the act, including supervision, also when those duties have been entrusted to another person or business. Handing your books to a practice moves the work. It does not move the responsibility, and it explicitly leaves supervision behind.

The practice. Contractual liability toward the client, plus the mandatory professional policy that every entity providing bookkeeping services has to hold — minimum guarantee sum of 10 000 EUR per event. Sit with that number for a second. Roughly 43 000 zł at current rates, which is smaller than one missed transfer to an account outside the white list on a 200 000 zł invoice, and smaller than a good many of the errors people worry an agent might make. The mandatory floor was never built to cover the interesting cases.

Whoever actually handles the financial affairs. Fiscal penal law reaches the person who deals with a company's economic affairs, financial matters in particular — which is how a bookkeeper personally ends up in scope for something a client's board never saw.

Three parties, no vacancy. So when a vendor implies their agent "takes on the compliance burden," ask which of those three roles they are stepping into. They cannot answer. None of the three has a slot software can occupy.

The person who signs has a name on a piece of paper

The sharpest version of the whole subject is the submission step, and it is sharp because the paperwork is unambiguous.

An electronic declaration is signed by a named natural person with a qualified certificate, or by a representative listed on a UPL-1 authorisation — again, a named natural person. There is no field for a service account. So when someone tells me their system "files JPK automatically," here is what is physically happening: the software holds credentials belonging to a human being, and that human being has signed every file it sent. Whether they read any of them is between them and their conscience — right up until the day it becomes a matter between them and the tax office.

This is why I keep the send button manual in every implementation, and why I would keep it manual even if a client asked me not to. The agent assembles the file, runs the mechanical checks, and produces a list of what it found and what it could not resolve — the split I worked through in what an agent fixes in JPK_V7 and what it must never touch. A person with a name and a certificate reads the list and presses send. That step costs about four minutes a month per client. It is the cheapest insurance in the entire build.

Four permission levels, and the one boundary that matters

In practice, "can the agent do X" is never a yes-or-no question. Every operation an agent performs sits on one of four rungs, and that ladder — not the contract — is the real design document of an implementation:

  1. Read. Pull the document, the bank line, the registry status. No consequences, no approval needed, and this is where a surprising share of the value lives — a daily sweep of a 3 200-contractor portfolio for status changes is something no person will ever do by hand.
  2. Propose. Produce a draft booking, a match, a flagged discrepancy. Still no consequences: nothing has entered the ledger. A person accepts or rejects.
  3. Write. Post the booking, attach the document, set the attribute. Reversible, logged, and visible — but the ledger has changed and the change happened without a human decision.
  4. Send or sign. File the declaration, execute the transfer, email the client. Irreversible in the sense that matters: a third party has now received something.

The boundary that carries liability is not between 1 and 2, where most people expect it. It is between 3 and 4. Level 3 mistakes are embarrassing and correctable — you reverse the entry, you re-run the month. Level 4 mistakes leave the building.

It is the same write-access boundary I hit from a different direction when I wrote about how long an implementation really takes. The moment a project needs level 3 or 4 access is the moment every unwritten convention in a practice has to be said out loud. Which is also, reliably, the moment the calendar slips.

A well-scoped accounting agent lives almost entirely on rungs 1 and 2, moves onto rung 3 for a narrow list of operations everybody has agreed on in writing, and never touches rung 4.

The signature test

Here is the diagnostic I use in scoping now. Ten seconds per operation.

For every step you are considering automating, ask: if this went wrong and a tax inspector asked whose decision it was, could you produce a name in under a minute?

If the answer is a name — the senior bookkeeper who approved the rule, the client's board member who signed off on the tolerance threshold — the step is ready to automate. The automation is just executing a decision someone already made. If the answer is "well, the system does it," the step is not ready. Not because it is technically hard, but because you have found a decision nobody has actually made, and automating it would freeze that gap into a process.

The test is uncomfortable in a useful way. Run it across a practice and it usually turns up two or three rules everybody follows and nobody owns: the tolerance for a rounding difference on a bank match, the point at which a suspense entry gets escalated rather than parked. Those are not automation problems. They are governance gaps that automation drags into the light, and finding them pays for the exercise whether or not you build anything.

What actually changes, then

If liability does not move, what is the point?

Evidence. A manual process, examined honestly, usually cannot tell you who decided what. Ask an office in October why a particular invoice was assigned to August and the answer is a reconstruction — somebody's memory of a conversation, plus an inference from what ended up in the ledger. Ask the same question of a process with an agent in it and you get a record: the rule that matched, the data it matched on, the person who accepted the proposal, the timestamp. That record is what a defence looks like.

There is a second effect, less obvious. Because every automated step needs a named owner before it can be built, the exercise forces a practice to write down decisions that had been living in people's heads. The agent is only the occasion for it. The value survives even if you switch tools next year.

So the pitch is not "the software takes on your risk." It is this: your liability is unchanged, your ability to demonstrate diligence goes up sharply, and about two thirds of the checking work stops depending on whether anyone had time on the 24th. The broader shape of what these AI agents actually do inside a company follows the same logic wherever they run — the machine prepares, the person decides.

Three things to put in writing before you start

Not a contract template, and not legal advice — just the three items that reliably prevent an argument later:

The operation list. Which operations sit on which rung, by name, with a person attached to each rung-3 entry. One page. If it runs to four, the scope is too big for a first phase.

The log retention. Decision logs should live at least as long as the books they relate to — five years is the sane default. Write it down somewhere, rather than leaving it to whatever the vendor's database happens to keep.

The insurer conversation. Short, and before renewal rather than after a claim. What the agent reads, what it writes, what it does unattended.

One number to measure first

Before any of this gets concrete, count something. Take last quarter's corrections and errors, and against each one write down whether anyone could name, today, the person whose decision led to it.

If most of them come back with a name, you have a process ready to automate, and the agent will make it demonstrably better. If most of them do not, you have found the real project — and it is not an AI project. It is two afternoons at a whiteboard, deciding who owns what.

Want to work out which operations in your practice could safely sit on rung 3, and what that would attach to in your system? Get in touch — scoping is free and takes half an hour. If you are earlier than that, the AI readiness audit is how I judge whether a practice is ready for any of it.

Let’s talk about your project

Free 30-minute consultation. We’ll figure out if and how I can help.

Book a Free 30-Minute Call

Select a date

August 2026
Mon
Tue
Wed
Thu
Fri
Sat
Sun
Back to Blog

Related Posts

Comarch Optima API: a developer's guide to integrating Optima
Blog

Comarch Optima API: a developer's guide to integrating Optima

The question is usually whether Optima has an API. The answer is that it has five different things by that name, each with a different licence, a different owner, and a different phone number to call when it stops working.

Read more
Which AI agent to build first against enova365
Blog

Which AI agent to build first against enova365

The most valuable agent is almost always the wrong first build. Not because it cannot be built, but because its first honest output arrives in week ten, and nothing holds a room's attention that long. Four questions that screen the candidates, and one number that predicts whether the project lands.

Read more
The enova365 service account — what read-only actually means
Blog

The enova365 service account — what read-only actually means

The integration gets the Administrator account, because scoping the rights would take an afternoon and nobody has the afternoon. Then it turns out the strongest read-only boundary in enova365 is not the permission tree at all. It is the licence.

Read more