The grouping that is not on the document
Every extracted row already knows a great deal about itself: what it cost, when, from whom, and which file it came out of. All of that is printed somewhere and can be read.
What is not printed anywhere is the thing you most want to group by. Which project consumed this. Which grant funded it. Which matter it is recoverable against. Which site incurred it. Those are facts about your organisation, and no supplier has any reason to know them.
So they have to be applied. The question is whether they get applied once, as a rule that persists, or re-decided every month by whoever is doing the coding — which is the difference between a dataset that gets more useful over time and one that gets less comparable.
Not the same thing as entity tagging
The two are easy to confuse and behave differently, so it is worth separating them before anything else.
| Entity tagging | Dimension tagging | |
|---|---|---|
| What it records | Where the row came from | What the row is for |
| Source | The file and account | A reference, a rule, or you |
| Automatic? | Yes — structural | No — it is a judgement |
| Can it be wrong? | Only if the file was | Yes, and that matters |
| Splittable? | No, a row has one origin | Yes, and often must be |
| Changes over time? | Never | Only if you change the rules |
The fourth and fifth rows are where the practical difference lies. A row’s origin cannot be argued with; its dimension can, and a single row can legitimately belong partly to two of them. Everything else on this page follows from those two facts.
They work together rather than competing: entity taggingkeeps the row traceable, and the dimension makes it groupable. A dataset with both can answer “what did this project cost” and “prove it” in the same breath.
What counts as a dimension
Anything you would want a subtotal for that is not a category of spending. Four cover almost every case, and they behave differently enough to be worth naming.
Project. Has a start, an end and a budget. The dominant case in agencies and professional services, and the one where late-arriving cost causes the most trouble.
Fund or grant. Money with strings attached, where the point of the tag is proving that restricted money was spent on what it was restricted to.
Matter or engagement. Client work where some costs are recoverable and some are not, so the tag decides what can be billed on.
Site or location. Permanent, comparable to each other, and where the interesting question is relative performance rather than a total.
Department is a fifth that behaves like site — permanent and comparable — and is covered from the reporting side on our cost centre spend report page.
Where the value actually comes from
Three sources, in descending order of reliability and ascending order of effort.
A reference on the document. A project code in the line description, a purchase order number, a delivery address. The best case, because it is evidence rather than assertion — and it is more common than people expect once suppliers are asked.
A rule against a supplier or payee.This supplier only ever works on one site; this subscription belongs to one team. Recorded once and applied thereafter, which is what turns the first month’s work into later months’ routine.
A person deciding. Necessary for the rest, and it should be the residue rather than the norm. If most rows need a human every month, the rules are not being recorded — or the suppliers have never been asked for a reference.
Worth tracking the proportions once a quarter. A rising share of manual decisions is the earliest sign that the tagging is drifting back into a monthly chore.
Splitting one document across dimensions
The case that makes dimension tagging different from ordinary categorisation, and the one most home-made systems handle by pretending it does not happen.
A supplier bills monthly and their month covered three of your projects. A landlord invoices one building used by two departments. A software subscription is used by four teams. In every case one document becomes several allocated rows.
Where the lines themselves carry the answer, the split follows the document and needs no judgement. Where they do not, it needs a basis — and the basis should be recorded against that supplier rather than reconstructed each month.
The reason to record it is not tidiness. A split that is re-decided monthly drifts, and drift in an allocation is invisible until someone compares two years and finds a trend that is really a change of method.
And where a cost genuinely is shared with no meaningful basis, say so. Forcing it onto dimensions by an arbitrary rule produces numbers that look precise and are not, which is worse than an honest overhead line.
Stability beats precision
The principle the whole feature is built on, and it runs against the instinct to improve the tagging whenever you notice something imperfect.
A crude rule applied identically for two years produces comparable numbers. A refined rule introduced in month seven produces two half-datasets that cannot be compared to each other — and the improvement is invisible while the discontinuity is permanent.
That does not mean never improving. It means improving at a boundary, and restating the prior period when you do, so that the change shows up as a change of method rather than as a change in the business.
In practice: review the rules once a year, at the point the budget is rebuilt. Between times, handle awkward rows with notes rather than with new rules. The same discipline governs period comparison, for exactly the same reason.
How it works
1 · Extract the documents
Invoices, statements, receipts — up to 100 files in a batch, scans through OCR first.
2 · Read any reference present
A project code, purchase order or site in the line detail is picked up as a candidate value rather than assumed.
3 · Apply your rules
Recorded against supplier or payee, so they run without anyone re-deciding.
4 · Split where needed
One line becomes several allocated rows, on the basis you recorded, with the basis visible.
5 · Review the residue
What no rule and no reference covered is listed for a decision — not filled in with a guess.
6 · Export with the tags
Dimension columns travel alongside source file and page, so the allocation stays checkable.
Four dimensions, four different problems
The mechanism is the same; what changes is which failure hurts most, and it is worth knowing which one you are exposed to.
| Dimension | What it decides | Worst failure |
|---|---|---|
| Project | Whether the work made money | Late cost lands after the project closed |
| Fund or grant | Whether restricted money was used correctly | A cost attributed to the wrong funder |
| Matter | What can be billed on to a client | A recoverable cost never recovered |
| Site | How locations compare | An allocation basis that flatters one site |
The fund row is the one with the sharpest consequence. A misallocation between projects is an accuracy problem; a misallocation between restricted funds is a reporting problem with a funder, and those are corrected in public.
The matter row is the one that quietly costs money rather than accuracy: a recoverable cost that never got tagged is simply never billed, and nothing anywhere reports its absence.
More than one dimension at a time
A cost can be both a project and a department, or both a site and a fund. Supporting two is genuinely useful; supporting five is where these systems collapse.
The cost is not technical, it is human. Every additional dimension multiplies the number of decisions per row, and rows that need three judgements do not get three judgements — they get one and two defaults.
Two is the practical ceiling for most organisations. If a third feels necessary, it is often because one of the first two is being used for something it does not describe well, and redefining it is cheaper than adding.
When you do use two, make one primary. The primary one gets the attention and the rules; the second is derived where possible — a project usually implies a department, so deriving it costs nothing and asks nobody.
The unassigned bucket, and why it should exist
There will always be rows that no rule covers and no reference explains. What happens to them determines whether the whole dataset can be trusted.
The tempting move is to force them somewhere — spread them proportionally, or default them to the largest project. It produces a report with no awkward gaps, and it is the single fastest way to make the numbers untrue.
An explicit unassigned total is far better. It is honest, it is visible, and it is actionable: a large unassigned figure tells you exactly where to spend the next hour, and a shrinking one is the clearest evidence that the tagging is improving.
Track it as a percentage of value rather than of rows. Fifty small rows unassigned matters much less than one large one, and counting rows hides that completely.
What you get back
The extracted row, plus the columns that make the allocation visible and checkable.
| Column | Why it is there |
|---|---|
| Dimension value | The project, fund, matter or site |
| Allocated amount | This row's share, after any split |
| Allocation basis | How the split was decided, in words |
| Source of the tag | Reference, rule or manual |
| Source file and page | So the row stays traceable |
| Original line total | The undivided amount, for checking |
The third and fourth rows are the ones people do not ask for and later depend on. Knowing that an allocation came from a rule rather than from someone’s judgement changes how much you trust it — and knowing the basis in words means it can be questioned without an archaeology project.
The last row is the arithmetic safety net. Allocated amounts should sum back to the original line total, and when they do not, a split has gone wrong in a way no total-level check would catch.
Naming your dimensions, which matters more than it sounds
A dimension value is a label that has to survive being typed by several people over several years. Most of the pain in tagged data comes from labels that were obvious to whoever invented them and ambiguous to everyone since.
Prefer codes to names.“PRJ-114” has exactly one spelling. “Northside rebrand” has at least six, and a supplier will use a seventh. Codes are ugly and they are the only thing that matches reliably across a document, a spreadsheet and a ledger.
Never reuse a code. When a project ends, its code retires with it. Reusing it two years later means any historical query silently mixes two unrelated bodies of work, and nothing will ever flag that it happened.
Keep the client and the project separate.Encoding one into the other — “ACME-REBRAND” — feels efficient and prevents you from ever asking a question about the client across all their projects without string surgery.
Decide what happens to a value that no longer applies. A closed project, a discontinued cost centre, a client who left. Marking it inactive keeps history intact; deleting it breaks every historical report that referenced it.
None of this is exotic, and all of it is cheaper to decide at the start than to repair after two years of data. The one rule that matters most is the second: reused codes are the only mistake on this list that cannot be cleaned up afterwards.
When the structure changes underneath you
It will. Departments merge, a client is acquired, two projects become one, a cost centre is split in half. The question is what happens to the two years of tagged history that used the old structure.
The instinct is to rewrite history so everything is consistent. It is usually the wrong instinct. Rewritten history no longer matches the reports that were issued at the time, and the first time someone compares a current report against a printed board pack from eighteen months ago, the numbers disagree and the data loses its credibility permanently.
The better pattern is to keep the original tag and add a mapping. History stays as it was recorded, current data uses the current structure, and a lookup connects them for anyone who needs a continuous series. It is slightly more work and it is the only version that survives being audited.
A related case is worth flagging because it catches people out: a project that was tagged to one client and is later transferred to another. Both the costs and the reason for the move need recording, or the receiving client appears to have incurred costs before the relationship existed — which is the kind of anomaly that consumes an afternoon to explain and thirty seconds to have documented at the time.
What it will not do
It will not guess. Where there is no reference and no rule, the row is presented as needing a decision rather than filled with the most likely value — because an allocation you did not make is one you cannot defend.
It will not decide your allocation policy. Whether a shared cost is spread by headcount, by revenue or not at all is a management judgement, and it stays one.
It will not know your projects, funds or matters. Those come from you or from a reference on the document; there is no source of truth for them inside a supplier’s invoice.
And it will not repair tagging that was inconsistent when it was applied. It can show that a rule changed and when, which is most of the value — but restating the earlier period remains a decision someone has to take.
