Talk to us
Back to Blogs
Platform14 min read

How should a bank use AI agents to investigate reconciliation breaks?

Finding the break is the easy part. Explaining it is the hard part, and deciding what the agent may conclude on its own is harder still.

A reconciliation agent should investigate the break and stop there. It reads the front-office book, the accounting book, the custodian statement and the supporting records, isolates the field that differs, and names a likely cause from a fixed list your reconciliation policy already recognises. It cites the clause it relied on, says how confident it is, and proposes a correction. It cannot book that correction. Your controller approves or rejects it with a reason, and the whole case, including anything the agent was refused, is kept for the next audit.

Why is finding a reconciliation break easier than explaining it?

Your matching engine already finds them, and that part of the problem was solved years ago. What it hands you is a list: instrument, as-of date, delta. What it does not tell you is why. So an analyst opens the front-office system, then the accounting system, then the custodian portal, then the security master, then the FX rates for that date, then the corporate actions calendar, and starts working backwards. Was it a settlement that landed a day late? An FX rate applied at a different snap time? A corporate action processed in one book and not the other? A security master issue where two systems disagree about the instrument itself?

Most of that time is not judgement. It is retrieval. The analyst knows what the answer probably is within about ninety seconds of opening the case, and then spends thirty minutes proving it well enough that the controller will sign.

And the clock is real. NAV has to be struck. The report has to go out. When the backlog builds, breaks get resolved in order of deadline pressure instead of order of risk, which is exactly backwards.

What does a three-way break actually look like?

Take one position, three books.

The front-office book, your IBOR, shows a market value of 1,250,000. The accounting book, your ABOR, shows 1,250,000 as well. The custodian statement shows 1,243,750. The delta is 6,250, and it sits between your systems and the custodian, not between the two internal books. (Synthetic numbers, used for illustration. The shape is what matters.)

That last detail does most of the work. Because the two internal books agree with each other, a whole family of causes drops away immediately. You are not looking at an accounting posting error or a valuation difference between IBOR and ABOR. You are looking at something that changed the custodian’s view and not yours, or the other way round, which narrows you to timing, a corporate action the custodian has processed and you have not, or an instrument-level mismatch in the security master.

An analyst reaches that conclusion quickly. The agent’s job is to reach it, show its working, and hand you the evidence already assembled.

When the two internal books agree and the custodian differs, the set of plausible causes collapses immediately. Figures are synthetic and used for illustration.
Figure 1. When the two internal books agree and the custodian differs, the set of plausible causes collapses immediately. Figures are synthetic and used for illustration.

Which causes should the agent be allowed to choose from?

From a fixed list, and this is the design decision that keeps the whole thing defensible under examination. The agent does not write a paragraph of free-form speculation about what might have happened. It picks from the causes your reconciliation policy already recognises: settlement timing, FX, corporate action, security master, and whatever else your own procedure names. Each label maps to a clause in that procedure, and the agent cites the clause alongside the label.

There are three reasons to work this way. Your controller can check the reasoning against something written down instead of against the agent’s prose. Your repeat-break analysis stays countable, because “settlement timing” in March means the same thing as “settlement timing” in September. And when a regulator or an internal auditor asks how causes are assigned, the answer is a policy document, not a model.

The agent also states confidence, and low confidence is a useful output, not a failure. A case where the agent says it thinks this is a corporate action but the custodian’s record is ambiguous is a case that should go to a person quickly, not one that should be dressed up as certain.

What is the agent not allowed to do?

The agent cannot book the correction, and that prohibition is absolute.

Booking a correction changes the accounting record. That record feeds NAV. NAV feeds client statements and regulatory filings. The decision to move it requires judgement about whether the explanation is sufficient, whether the amount is material, whether this is the third recurrence for the same instrument, and whether the root cause sits upstream and needs fixing instead of patching. That is not retrieval work. That is control work, and it stays with a named person.

So in a governed setup the posting tool is not merely discouraged for the agent. It is not in the agent’s permitted action set at all. If something tries to invoke it through the agent, the call fails and the refusal is written into the case record.

That last part matters more than it sounds. A case record showing only successful actions tells an auditor very little about whether the boundary was ever tested. A record that shows an attempt, a refusal and a timestamp tells them the control is real.

What does the controller actually receive?

A case that is ready to decide, instead of one that still needs assembling. On one screen: the instrument and as-of date, the three source values with the differing field highlighted, the retrieval timestamp for each source, the proposed cause with its policy clause, the agent’s confidence, the proposed correction, and the supporting records it read on the way. The controller approves or rejects, and a rejection carries a reason.

The reason code is not bureaucracy. Rejections are the only feedback loop you have. If controllers keep rejecting settlement-timing calls on a particular fund, either the agent’s retrieval on that fund is weak or the procedure needs rewriting, and you cannot tell which without the reasons.

The agent assembles evidence and proposes; a named controller decides. The posting tool sits outside the agent's permitted actions.
Figure 2. The agent assembles evidence and proposes; a named controller decides. The posting tool sits outside the agent's permitted actions.

How does the workflow run end to end?

The workflow moves through six steps. Each one builds on the last, and a named person closes the loop.

Step one: the overnight run flags the break. Your existing reconciliation engine compares the books and produces a list. Each item shows the instrument, the as-of date, and the delta. This part of your process does not change.

Step two: an analyst opens the case. Before any evidence loads, the screen shows what this agent is registered to do. Its purpose. Its risk tier. The policy version currently in force. If someone updated the procedure last week, that fact is visible now, before anyone acts on it.

Step three: the agent reads the permitted sources. It pulls from IBOR, ABOR, the custodian record, the security master, FX rates for the date, and the corporate actions calendar. It also checks prior resolutions for the same instrument. A break that has happened before usually has the same cause, and that history is worth consulting.

Step four: the agent isolates the field. Not just the position, but the specific field where the values differ. It shows which source carries the differing value and what each source said. This is where the three-way comparison collapses the set of plausible causes.

Step five: the agent explains and stops. It names the likely cause, cites the policy clause behind it, states its confidence, and proposes a correction. Then it stops. The next action is not available to it.

Step six: the controller decides. Approve or reject, with a reason. The evidence, the agent’s actions, any refused actions, and the human decision all enter the record together.

The gate between step five and step six is the whole design. Everything before it is evidence work that a machine does faster than a person. Everything at it is a judgement that belongs to someone whose name goes on it.

The overnight run flags the break. Instrument, as-of date, delta. This is your existing reconciliation process, unchanged.

The analyst opens the case. Before any evidence appears, they see what this agent is registered to do: its purpose, its risk tier, the policy version currently in force. If someone changed the procedure last week, that is visible here, before anyone acts on it.

The agent reads the permitted sources. IBOR, ABOR, the custodian record, the security master, FX rates for the date, the corporate actions calendar, and prior resolutions for the same instrument. That last one is quietly valuable, because a break that has happened before usually has the same cause.

The agent isolates the field. Not the position, the field. It shows which source carries the differing value and what each source said.

The agent explains and stops. Likely cause, cited clause, confidence, proposed correction. Then it stops, because the next action is not available to it.

The controller decides. Approve or reject, with a reason. Evidence, agent actions, refused actions and the human decision all enter the record together.

The gate between step five and step six is the whole design. Everything before it is evidence work that a machine does faster than a person. Everything at it is a judgement that belongs to someone whose name goes on it.

What gets kept, and why does it matter later?

Six months from now someone will ask why a particular correction was booked, and the quality of your answer depends entirely on what you retained at the time.

The record holds the session, meaning the whole run from the break being flagged to the controller’s decision. It holds the policy version that was live that day, which becomes important the first time you update a procedure or change models and the behaviour shifts underneath you. It holds each source read and when it was read, so a value that has since been restated does not silently rewrite history. It holds the agent’s reasoning and its confidence. It holds what the agent tried and was refused.

And it holds the human half, which is the half examiners tend to care about: which controller decided, what they decided, on what basis, at what time.

You also get cost and latency per session, and latency reported per agent type instead of blended across all of them. Averaging a two-line lookup with a full three-way investigation produces a number that means nothing to anyone.

How is this different from the reconciliation tools we already have?

Your matching engine finds breaks. This explains them. Those are genuinely different problems and it is worth being precise about the boundary, because reconciliation teams have been sold overlap before.

Rule-based auto-matching handles the deterministic cases: same instrument, same amount, known tolerance, clear the item. It works well and you should keep it. What it cannot do is read a custodian’s corporate action notice, compare it to your own processing date, notice they differ by one business day, and connect that to a delta on a specific position.

Robotic process automation sits somewhere in between. It can open the systems and copy the values, which removes some of the clicking, but it follows a fixed script and breaks when a screen or a file format changes. It also has no view on what the values mean once it has them.

The agent is for the cases that survive the matching engine: the ones where the evidence lives in several places, the sources use different identifiers, and someone has to work out which explanation fits. Those are also the cases eating your analysts’ time, because the easy ones already cleared automatically.

How do you measure whether it worked?

From your own baseline, and not from anyone’s headline percentage. The measures worth tracking are analyst time per break, the unresolved-break backlog, controller rework, deadline adherence, and the repeat-break rate. That last one is the interesting one. If the same instrument breaks every month for the same reason, the win is not resolving it faster, it is fixing whatever upstream process keeps producing it, and a countable cause label is what makes that visible.

The arithmetic is annual breaks, multiplied by minutes saved per break, divided by sixty, multiplied by loaded hourly cost.

As an illustration: a mid-market firm running 5,000 breaks a year, saving 20 minutes on each, at $60 an hour loaded, recovers roughly $100,000 of analyst capacity a year. A large firm at 25,000 breaks on the same assumptions recovers roughly $500,000. Those numbers are illustrative and the assumptions are deliberately visible so you can replace them.

To get your own: time twenty current breaks end to end, and split the minutes between retrieval, comparison and the controller’s review. Run ten comparable breaks through the agent. Take the difference per break, then multiply by your annual volume and your loaded cost. Be sceptical of anyone quoting you a figure before they know your break mix, because a book heavy in corporate actions behaves nothing like one heavy in FX.

Inputs, agent actions, refused actions and the named human decision are retained together, so a case can be reconstructed months later
Figure 3. Inputs, agent actions, refused actions and the named human decision are retained together, so a case can be reconstructed months later

What is still under development?

Three capabilities are not yet finished, and you should know about them before you pilot.

Custodian statement coverage. Not every custodian delivers data in the same format, and some require bespoke parsers before the agent can read them. We are widening coverage so that fewer statements need custom work, but part of the delay is commercial rather than technical. Data access is negotiated with each custodian, not just integrated.

Bulk export for month-end. On high-volume days, working through breaks one at a time in a browser stops being practical. The missing piece is a full export of the break population to a spreadsheet, so analysts can triage, sort by materiality, and decide which cases to open first. That export needs to carry enough context that the spreadsheet is useful without the agent’s screen.

Cross-environment agent identity. Every action should trace to four things: the employee who opened the case, the named agent that acted, the owner accountable for that agent, and the person who approved the outcome. That chain needs to hold even when the agent crosses from one cloud environment to another. Two related controls are already live. Skills and tools are scanned before they are published, and every call through the gateway records a guardrail verdict of passed, flagged or blocked. What is still in progress is re-checking an agent’s tool access while it runs, instead of only at the point it was registered. The aim is continuous verification, not a one-time gate.

Three things are not finished, and it is better to read that here than find out during a pilot.

Wider coverage of custodian formats, so that fewer statements need a bespoke parser before the agent can read them. Some of that is a commercial problem rather than an engineering one, because custodian data access is negotiated, not just integrated.

Export of a full break population to a spreadsheet, for the month-end days when working through cases one at a time in a browser stops being sensible.

And agent identity. The aim is that every action traces to four things: the employee who opened the case, the named agent that acted, the owner accountable for that agent, and the person who approved the outcome. We want that chain to hold even when the agent crosses from one cloud environment to another. Two related controls are already live and worth separating out. Skills and tools are scanned before they are published, and every call through the gateway records a guardrail verdict of passed, flagged or blocked. What is still in progress is re-checking an agent’s tool access while it runs, instead of only at the point it was registered.

Key takeaway

Scope a reconciliation agent to evidence and the control question gets much simpler. It gathers the permitted records, isolates the field that differs, names a cause from a list your policy already recognises, cites the clause, states its confidence and proposes a correction. It cannot post anything. A controller approves or rejects with a reason, and the case keeps the evidence, the reasoning, the refusals and the decision together.

Answer the control questions once, in the control plane: who owns this agent, what risk tier it sits in, what it can reach, where the hard line is, who decides, and what is kept. The next agent after this one is then a configuration exercise instead of a rebuild.


Book a diagnostic workshop to map your reconciliation workflow and design a governed break-investigation agent: moring.ai/contact

Frequently asked questions

Can an AI agent resolve reconciliation breaks automatically?
It can investigate them automatically and it should not resolve them automatically. Gathering the records, isolating the mismatched field, comparing values across books and proposing a cause are all evidence work. Booking the correction changes the accounting record and feeds NAV, so it needs a named person's approval. In a governed setup the posting action is not in the agent's permitted set, so the boundary holds structurally, not by instruction.
What is the difference between IBOR and ABOR in a reconciliation break?
The investment book of record is the front-office view, kept current through the trading day for position and exposure decisions. The accounting book of record is the accounting view, valued and adjusted under accounting policy, and it is what NAV and client reporting are built from. They are supposed to converge and often differ briefly for legitimate reasons, so a break between them points to a different set of causes than a break between your books and the custodian.
How does an agent decide the cause of a break?
It selects from a fixed set of causes that your reconciliation procedure already defines, such as settlement timing, FX, corporate action or security master, and it cites the clause behind the label. It does not invent a category or write free-form speculation. It reports confidence alongside the cause, and a low-confidence case goes to a person instead of being presented as settled.
What audit evidence should a reconciliation agent retain?
Enough to reconstruct the case months later: the session, the policy version live at the time, each source read with its retrieval timestamp, the field-level comparison, the cause and cited clause, the confidence, any actions that were refused, and the controller's decision with its reason. Refused actions matter as much as successful ones, because they are what shows the boundary was holding.
How is this different from robotic process automation?
RPA follows a fixed script across screens and file formats, and it breaks when either changes. It also has no view on what the values it copies actually mean. A reconciliation agent reads several sources that use different identifiers, compares them at field level, weighs which explanation fits, and hands over the ambiguous cases. RPA suits stable, deterministic steps. Agents suit comparison and judgement across messy inputs.
How do you calculate ROI on a reconciliation agent?
Annual breaks, multiplied by minutes saved per break, divided by sixty, multiplied by loaded hourly cost. Build the inputs from your own baseline by timing twenty current breaks and running ten comparable ones through the agent. As an illustration only, 5,000 breaks a year at 20 minutes saved and $60 an hour is roughly $100,000 of recovered analyst capacity. Track the repeat-break rate alongside it, because preventing a break is worth more than resolving it quickly.
Agentic AIPlatformMoring AI