Which AI agent to build first against enova365
Six posts in this enova365 series ended the same way: here is the first build I would do. A nightly three-check job on the KSeF submission state. An audit of what customers actually type into payment titles. A fill-rate inventory of custom fields. The ninety-day line for ulga na złe długi. A warning at seventy percent of a credit limit. Every one of them was honest about its own scope. Not one answered the question clients actually ask me: fine, but which of these do I do first?
There is a better answer than the one worth the most money. And the gap between those two answers is where most of these projects quietly fall apart.
The most valuable agent is usually the wrong first build
Ask an accounting office which agent would help most and you get the answer in about fifteen seconds: the one that books documents. That is the correct answer to the question asked. Posting is where the hours are, and I have written up how that architecture works in enough detail that nobody can accuse me of arguing it cannot be built.
It can. This is an argument about order, not feasibility.
The posting agent writes. Writing means the widest permission scope of anything on the list, which drags the licence and operator negotiation onto the critical path in week one — the conversation I have watched run five weeks on its own, because an access request is usually a procurement request wearing a disguise. On top of that it needs contractor mapping, account classification and an exception queue in place before it emits anything a human would bother reading. Add it up: the first honest output lands somewhere between week six and week ten.
Ten weeks is a long time to hold a room's attention on something nobody has seen work. Long enough for the sponsor to change roles, for the budget to get a second look, or for a quiet consensus to form that this is not really going anywhere. I have lost projects that were working. Not failing — working, and too slow to prove it.
Four questions that screen a first build
So I stopped ranking candidates by value and started running them through these four.
Does it write anything? If yes, it is not first. Not because writes are dangerous in principle — a draft sitting in buffer status is genuinely safe — but because writes pull the entire access negotiation forward and park it in front of the work.
How fast does it produce a number somebody argues about? This is the question that actually predicts whether the project lands. It gets its own section below.
Does it fail loudly? A job that quietly returns nothing looks exactly like a quiet week — the failure mode from the post on the scheduler. A first build has to report what it examined, not only what it found, because at that stage nobody has a baseline for what normal looks like. Zero exceptions on Tuesday is either excellent news or a broken connection, and for the first month you genuinely cannot tell which.
Does it survive being wrong? Every first build is wrong somewhere. The question is what the wrongness costs. A misjudged line in an internal report costs a comment in a meeting. A misjudged line in an email to a customer costs the relationship — chasing somebody who already paid, the mistake I called unrecoverable in the receivables post. That is why customer contact never belongs in a first build, however good it looks in a demo.
The shortlist, scored
Six candidates this series has already produced, put through the four questions.
| Candidate | Writes? | Days to a disputed number | Fails loudly? | Cost of being wrong |
|---|---|---|---|---|
| KSeF submission-state check | no | 3–5 | yes, with a state count | a re-check |
| Bank description-field audit | no | 1–2 | one-off, n/a | none |
| Custom-field fill-rate inventory | no | 1–2 | one-off, n/a | none |
| Ulga na złe długi ninety-day line | no | 5–10 | yes | an unnecessary review |
| Credit limit warning at seventy percent | no | 7–14 | yes | an ignored alert |
| Posting agent, drafts only | yes | 40–70 | partly | a wrong draft, caught |
Two entries on that list are not agents at all. The description-field audit and the fill-rate inventory are an afternoon of SQL and a spreadsheet. They belong here precisely because they are the cheapest way to find out whether the expensive thing is worth building at all. The description-field audit settles the ceiling on automatic settlement before anyone signs anything: if parseable invoice references sit at thirty-five percent of incoming lines, no agent and no ERP raises that number, and the honest recommendation is to go and fix the payment titles first. Nobody enjoys that conversation. It is still much cheaper in August than in December.
The boring one wins, and that is the point
Among the actual agents, the one I recommend most often is the least impressive in a demo: a nightly read-only job that reconciles three things and reports the exceptions. It books nothing, emails no one, and its entire output is a list of these did not match, and here is why.
It wins for a reason that has very little to do with technology. On day three somebody looks at line eleven of the exception list and says that cannot be right. Then they check. And it is right. That moment does more for the project than any demo, because it is the first time the system has told them something they did not already know — and it moves the agent out of the category of purchase and into the category of colleague. Once that has happened, every scope conversation for the rest of the engagement gets easier.
It also delivers something nobody budgets for. Six weeks of real exceptions is the taxonomy the posting agent will need anyway, and it is a specification you cannot write in a workshop — the interesting cases are exactly the ones nobody thinks to mention when you ask them to describe their own process.
The metric: time-to-first-disputed-number
That day-three moment is measurable, so measure it. Time-to-first-disputed-number: days from kickoff to the first time a person at the client challenges an output and turns out to be wrong.
Both halves carry weight. If nobody ever challenges it, nobody is reading it, and you have built a report that gets filed. If they challenge it and they are correct, you have a bug — normal, expected, but a completely different signal. What you want is a human being surprised by their own data.
Under ten days, the project lands. I cannot think of a counterexample. Past thirty, something structural is wrong, and in my experience it is almost always the same thing: the pilot is running against a copy of the database with last year's data in it. That is the most comfortable way to build one of these and the least informative. Nobody argues with last year's numbers. Nothing rests on them, nobody is accountable for them, and a system that produces them is not being tested — it is being watched politely.
Which is the honest framing of the read-only phase. It is not a dress rehearsal for the real thing. It is the phase where you find out whether the data supports what you sold. And if it does not, this is a cheap place to learn that.
How you graduate
The four permission levels — read, propose, write, send — are the ladder, and the real boundary sits between three and four. The common mistake is climbing it on a calendar. Read-only for a month, then we turn on writes is a plan resting on nothing but the passage of time.
Graduate on a measured number instead. The override rate — the share of proposed actions a person cancels — tells you where you actually are. Above roughly thirty percent after a few weeks, the rules are still wrong, and switching on writes at that point just industrialises the error at speed. Under ten percent, the human step has quietly become a rubber stamp, and you are paying somebody to click confirm on things they no longer read. That second failure is the more dangerous one. It looks like success on every dashboard you have.
Where I would start on Monday
Pick the process that generates the most exception emails inside your own office. Not the most hours — the most emails. Email volume is the visible trace of a filter living in one person's head, and a filter nobody has written down is the most reliable sign that there is something worth automating underneath it.
Build the read-only version of that, run it nightly, and have it report what it examined and not only what it found. Then, on day one and in parallel, start the access conversation for the write phase you will want in month three. It is the long pole, it does not depend on anything you are building, and starting it late is the single most common reason these timelines slip. For the wider picture of what these agents do once they are running, there is the agents overview.
Send me your shortlist and I will tell you which candidate reaches a disputed number fastest — and which one you should not build first, however much you want to. Half an hour, no charge — get in touch.
Let’s talk about your project
Free 30-minute consultation. We’ll figure out if and how I can help.



