All articles

Show me the workings

Firms judge AI by whether the demo output looks right. The profession holds its own staff to a harder standard, and should hold software to the same one.

Watch a partner review a grad’s first company tax return. The refund matches the estimate and the return balances, and none of that earns a signature. The partner opens the workpapers and starts asking the questions a total cannot answer. Did the bank rec reconcile, or was the difference plugged? Was the franking account reconciled, or rolled forward from last year because that was easier?

The partner is not being difficult. A right answer reached the wrong way is a liability that has not found its client yet. The grad who plugs a small reconciliation difference on a clean file will plug a bigger one on a messy file, and the messy file is where it costs you.

Public practice built its whole quality apparatus on that idea. The review hierarchy, the workings kept on file, the documentation standards and the sign-off before lodgement all exist because an answer is only as trustworthy as the process behind it. Nobody lodges a return with a note saying the total looked right. The file has to show how it got there, because sooner or later someone else has to verify the work, and the file is all they will have.

Then the demo arrives

Put an AI tool in front of the same partner. The vendor runs a prepared file, a correct-looking BAS appears within minutes, and the conversation moves to pricing. The software has just been judged by a standard the profession refuses to apply to people. A grad who handed over a finished BAS with nothing behind it would get it handed straight back. Software that does the same thing gets called a successful demo.

The demo file is also the cleanest file that tool will ever see. The ledger is reconciled and every source document is already in the folder. Your clients are not like that. The files that cost you money have shoebox records and a loan account nobody has examined since 2019, and the demo tells you nothing about what happens there.

One file is not a track record

A correct output is one data point about a system you cannot see inside. Whether the same quality arrives next quarter, on a client with worse records, or after the vendor ships an update that quietly changes how the tool behaves, is invisible from the one file you watched it produce. So the reviewer starts from zero on every file and re-checks everything against source, which is exactly the burden the tool was supposed to remove. An answer you have to re-derive before you can trust it has saved you nothing.

Machine learning research reached the profession’s position from the other direction. In 2023 OpenAI published a paper, Let’s Verify Step by Step, which found that a model rewarded for each correct step of a maths solution outperformed a model rewarded only for the final answer. Check how the work was done and quality holds on the next problem. Check only the answer and the first warning you get is the failure.

Ask for the workings

Refuse AI output that arrives without its reasoning attached, exactly as you would refuse a file with no workpapers behind it.

The demand is not exotic. It is the same two things you already require of staff: a record of the decisions made, written so a reviewer can confirm each one in seconds rather than reconstruct it, and evidence behind every figure. A tool that cannot produce either is asking you to redo its work and calling that review.

The two standards lead to different places. Judged by output, you get one correct-looking file on the cleanest data the tool will ever see, no view of how the figures were reached, and a review that starts from zero each time. Trust expires with every file. Judged by process, you get a reasoning log for every decision, the assumptions flagged for a professional’s call, and every figure traced to its source. Trust carries to the next file.

The standard we hold ourselves to

A MindLedger job comes back as a work packet, not a bare return. The packet carries the finished work together with a reasoning log for every decision and a register of every assumption: what was assumed, why, and what the reviewer should check. Every figure traces to the source document that produced it. We have written before about what goes into a work packet; the point here is that the packet exists to be checked.

Reviewers typically clear a standard company tax return packet in 10 to 20 minutes. That number is only possible because the process arrives with the answer, so the reviewer confirms decisions instead of reconstructing them. Where the reviewer disagrees, the comment goes on the exact cell in the workpapers. The rework comes back with that change recorded, and it is included in the job. Judgement stays with the firm’s qualified accountants, and so does the signature.

The same discipline covers the data the process runs on. Each firm runs isolated. Nothing is pooled across firms and nothing is used to train a model, ours or anyone else’s. The detail is on our security page.

The next demo you sit through will end the way they all end, with a correct return on the screen. Ask the question you would ask a grad: show me the workings. If the tool cannot, what exactly are you being asked to sign?