A pattern is only visible in context
A single invoice, read on its own, almost never looks obviously wrong. The amount is a plausible number, the supplier name is real, the format looks like every other invoice from that vendor. What makes it an outlier only becomes visible next to the invoices that came before it — and that comparison is exactly the part that gets skipped when a team is reviewing invoices against a deadline, one at a time, from memory.
This page describes how that comparison happens systematically instead of by chance — every transaction read, a baseline built from real history behind each supplier and account, and a flag raised the moment something breaks it, with the exact comparison shown rather than buried inside an opaque score nobody can explain to an auditor later.
Why a single rule isn't enough
The problem isn't applying one fixed rule to one obvious case — flagging anything over a huge round number is simple. The problem is that normal looks completely different from one supplier to the next, and from one account to the next. A logistics supplier might invoice a different amount every single time; a software subscription invoices the same figure for years. A single global threshold either buries real anomalies in noise or drowns a team in irrelevant flags, often both at once.
Add seasonal suppliers, ones that genuinely change price partway through a year, and new relationships with no track record yet, and a fixed rule starts producing results that are technically consistent but practically useless — not because the rule is wrong, but because it's being applied uniformly to relationships that behave nothing alike.
What gets read from each transaction
| Field | What it's used for |
|---|---|
| Counterparty identifier | Groups every transaction into the right baseline |
| Amount | Compared against the historical range for that counterparty |
| Date | Checked against the established timing pattern |
| Bank account details | Where present, compared against previously used details |
| Reference or document number | For traceability back to the source |
Five fields, read per transaction — never a single figure judged in isolation, disconnected from the history that gives it meaning.
Where this shows up in practice
| Situation | What changes |
|---|---|
| A supplier's invoice arrives at triple its usual amount | Flagged as an amount deviation, with the historical range shown |
| A recurring supplier's payment lands two weeks early | Flagged as a timing deviation from the established cadence |
| A supplier's bank account changes without notice | Flagged as a bank-detail deviation, the highest-priority signal |
| Several invoices land just under an approval threshold | Flagged as a threshold-proximity pattern across multiple documents |
Four common situations, each caught without a separate configuration for the specific case — the same underlying comparison applies to every one, because it's checking the transaction against real history rather than a fixed list of known scenarios prepared in advance.
From a read transaction to a flag
Once read, a transaction is compared against the historical range established for that specific counterparty — typical amount, typical timing, and the bank details normally on file. A transaction well inside that range passes without any flag. One that sits outside it is flagged, with the specific signal that triggered the flag attached.
A transaction close to the edge of the historical range — not dramatically outside it, but noticeably different from the recent pattern — is flagged as a milder deviation rather than being silently absorbed as normal or escalated as if it were a dramatic outlier. The severity of the flag reflects how far the transaction actually sits from what history would predict.
Confidence and what gets flagged
Not every flag carries the same certainty, and treating them identically would hide exactly the ones that need the closest look. Each flag carries a confidence level based on how much history backs the baseline it's being compared against — a deviation from three years of consistent data is a stronger signal than one from a relationship that's only sent four invoices.
This distinction matters most for genuinely new relationships, where there simply isn't enough history yet to say confidently what “normal” even looks like. Rather than forcing a flag either way, those cases are marked as low-confidence-baseline, so a reviewer understands why a large early invoice from a brand-new supplier isn't treated with the same certainty as a deviation from a five-year relationship.
How it works
Upload invoices and statements
PDF, scan or photo — individually or in a batch.
Every transaction is read
Counterparty, amount, date and bank details, per document.
Compared against its own history
Each new transaction checked against the baseline built from prior transactions.
Export with flags attached
Every deviation exported with its signal, comparison and confidence as separate fields.
A month of transactions, flagged
A mid-sized business, one month, 340 invoice and statement transactions across 52 counterparties.
| Outcome | Count |
|---|---|
| Matched established pattern | 333 |
| Flagged — amount deviation | 4 |
| Flagged — timing deviation | 2 |
| Flagged — bank detail change | 1 |
Seven flags out of 340 transactions took under twenty minutes to review — six turned out to be explainable business changes, confirmed with the comparison data already attached. The seventh, a bank-detail change on a mid-size recurring supplier, needed a verification call before payment, which is exactly the outcome the flag exists to produce.
By hand against automatic
By hand
A reviewer relies on memory or a quick mental estimate of what a supplier “usually” charges, without a documented record of the actual historical range, and with no consistency across who happens to be reviewing that week.
Automatic
Every transaction is checked against the same documented baseline every time, with the exact comparison shown instead of relying on whoever happens to remember what “normal” looks like for that supplier.
From one account to a whole business
Comparing one transaction against recent memory is a quick mental task. Doing it for every transaction across dozens or hundreds of counterparties, every month, changes what's actually practical — not because any single comparison is hard, but because the number of comparisons required grows with every new supplier and every new account.
At that volume, the value isn't only coverage — it's that every counterparty gets checked with the same rigor, instead of the handful of large, familiar suppliers getting careful attention while a smaller, less frequent one slips through with barely a glance.
Who uses it
Finance teams
A documented, consistent comparison behind every approval, not a remembered estimate.
Accounts payable
Deviations flagged before a payment run, not discovered afterward.
Controllers and finance directors
A defensible, exportable record of what was flagged and why, for audit or investigation.
Accountants and bookkeepers
The same consistent scrutiny applied across every client's transaction history.
Across all four groups, the underlying need is the same: a documented, repeatable comparison that doesn't depend on any one person's memory of what a given supplier or account usually looks like — useful precisely because that memory is exactly what tends to disappear when the person holding it moves on.
Edge cases worth knowing about
A supplier that genuinely doubles its invoice amount because of an agreed price increase is a common case that looks alarming but isn't — it's flagged the same as any other amount deviation, with the historical range shown, so confirming it as a known, agreed change takes seconds rather than requiring a separate exception process to be configured in advance.
A seasonal supplier — one that invoices heavily in certain months and barely at all in others — is read the same way every year. Once enough seasonal history accumulates, the baseline itself starts to reflect that seasonality, so a large invoice in the supplier's known busy month stops being flagged as an anomaly, while the same amount in its known quiet month still would be.
A brand-new counterparty with an unusually large first invoice is never silently accepted or silently blocked. It's flagged as a low-confidence-baseline case — genuinely worth a look precisely because there's no track record yet to confirm the amount is normal for that relationship.
What happens after export
Once exported, flagged transactions become ordinary data in your spreadsheet or system — filterable by signal type, sortable by confidence, groupable by counterparty however your workflow requires. That flexibility doesn't require returning to the original comparison each time; it's already built into the structure of the export itself.
Many finance teams build a running flag log on top of this export — one sheet showing, for the current period, how many transactions were flagged, by which signal, and how each was ultimately resolved, refreshed every time a new batch is read.
That log becomes especially useful heading into an audit or an insurance claim — a controller can hand over not just the final transactions but the full pattern of what was flagged and reviewed across the period, showing that the process itself was applied consistently, not just that the final numbers happen to look reasonable.
Why a flagged transaction beats a hidden score
A system that outputs a single risk score per transaction, without showing the comparison behind it, is convenient to glance at but nearly useless the moment someone asks why a specific transaction scored the way it did. Reconstructing that reasoning after the fact — especially months later, during an investigation — means effectively rebuilding the analysis from scratch.
Showing the comparison from the start — the exact historical range, the actual value, and which signal triggered the flag — moves that reconstruction work from a future emergency into a detail that's simply already present in the data. The difference matters most at precisely the wrong moment to discover it missing: during an investigation, when time to respond is already short.
When normal itself changes
A business that genuinely grows — new pricing, larger orders, a supplier relationship that scales up over a year — will see its own baseline shift, and that shift itself shouldn't be treated as a permanent stream of false positives. As real transactions accumulate at the new normal, the baseline moves with them, rather than staying anchored to an outdated historical range indefinitely.
That gradual drift is different from a sudden, single deviation, and the two are treated differently by design — a pattern that shifts consistently over several transactions reads as the business genuinely changing, while a single transaction that jumps sharply away from an otherwise stable pattern reads as the kind of deviation worth a closer look.
Reviewing the baseline over time
A baseline built from real history doesn't need manual retuning the way a fixed rule does, but it's still worth a periodic look — not to adjust the underlying logic, but to review what's been flagged over the past quarter and confirm the pattern of resolutions still makes sense for how the business actually operates.
That periodic review is a natural fit for a broader quarterly finance review, alongside other operational metrics — a controller can see, in one glance, how many flags came up, how they resolved, and whether the sensitivity setting still fits the volume the team can comfortably absorb.
That same review is a natural moment to check whether any supplier relationship has enough history now to warrant tighter scrutiny than it originally received — a supplier that started thin and low-confidence a year ago may, by the time of the review, have accumulated enough consistent history to be treated with the same confidence as a much older relationship.
What this doesn't do
Doesn't decide something is fraud
It flags a deviation from an established pattern. Confirming whether it's actually a problem remains a human judgment call.
Doesn't work well with zero history
A brand-new counterparty has no baseline yet — flagging accuracy improves as real transactions accumulate.
Doesn't block payments
It reads and flags. Stopping or releasing a payment stays inside your existing approval workflow.
Doesn't replace a documented fraud policy
It's a detection layer that feeds your existing process, not a substitute for having one.
These limits share a common thread: separating what is reading and comparing documents from what is a judgment about fraud. The comparison is this feature's job; the judgment stays with you.
