FlowParse
Feature 10 August 2026 15 min read

Policy field extraction

Insured, policy number, period, sums insured by section, excess, premium, tax. All of it stated plainly on the schedule, and none of it available to anything except a person reading the page — which is why most organisations do not know what their own policies say without opening them.

FlowParse
flowparse.io

Stated on the page, unavailable to everything else

A policy schedule is an unusually well-organised document. Someone has already done the work of setting out the facts: who is insured, for what periods, up to what amounts, with what excess.

And then it arrives as a PDF, which means all of that is text on a page rather than data. To find out what your excess is on one section, a person opens a file and reads. To find out across eight policies, a person opens eight files.

The consequence is quiet and widespread: organisations routinely do not know their own aggregate position. Not because anyone was careless, but because assembling it means an afternoon with a highlighter every time the question comes up.

Extraction turns those documents into rows. What it does not do — deliberately, and this is most of the page — is tell you what any of it means.

Reading against interpreting

This distinction is the whole design, and it is worth being blunt about because plenty of products in this area blur it.

QuestionKind of questionAnswered here
What is the stated excess?ReadingYes
What is the sum insured on stock?ReadingYes
When does cover start and end?ReadingYes
Is this event covered?InterpretationNo
Am I underinsured?AssessmentNo
Should I accept this renewal?AdviceNo
Does this endorsement hurt me?InterpretationNo

The line falls where a document stops stating and starts requiring judgement. Everything in the first group is written down; everything in the second depends on wording, circumstances and case law, and a wrong answer there has consequences that no accuracy figure covers.

What extraction does for the second group is make the inputs available quickly, so the person who can answer spends their time on the judgement rather than on finding the numbers.

FlowParse
flowparse.io

The field set

FieldWhy it is wantedWhat complicates it
Insured nameMatching to the right entityTrading name against legal name
Policy numberIdentity, correspondenceSits beside quote and account numbers
InsurerWho to contactBroker's name printed more prominently
Period of coverWhether an event falls insideMid-term adjustments
Sums insured by sectionThe limits that matterSeveral sections, several bases
Excess or deductibleWhat you bearDifferent per section, sometimes per event
PremiumCost, comparisonInstalments against annual
Insurance taxSeparating cost from taxSometimes only a gross figure
Endorsement referencesWhat has been variedListed as codes, wording elsewhere
File and pageGetting back to the sourceNever present — added

Notice how many complications in the right-hand column are about something sitting next tothe field you want. Schedules are dense, and a rule as simple as “the long number near the top” picks up the quote reference about as often as the policy number.

The last row is the only field that does not come from the document, and it is the one that makes the rest usable a year later.

FlowParse
flowparse.io

Four fields that are genuinely hard

Sums insured. Rarely one number. A commercial schedule carries separate limits for buildings, contents, stock, plant, and often business interruption on a different basis again. Returning a single total would be tidy and useless — you cannot check a total against anything.

Excess. Frequently different per section, sometimes per event rather than per claim, sometimes expressed as a percentage. Each variant is stated clearly and none of them is stated the same way twice across insurers.

Endorsements. Usually a list of codes on the schedule with the wording somewhere else entirely. The references come back; the wording they point to is a separate document, and what it changes is interpretation.

The insured entity. Sounds trivial and is a recurring source of trouble in groups: the schedule may name a trading name, a holding company or a list of subsidiaries, and matching that to your own records is not always mechanical.

On all four the principle is the same as everywhere else: what is stated comes back as stated, and where the document does not say, the field stays empty rather than being filled with the most likely answer.

Periods and dates

A policy has more dates on it than people expect, and they are not interchangeable: inception, expiry, renewal, the date the schedule was issued, and the effective date of any mid-term change.

They come back separately for a practical reason. The question that actually matters — whether something falls inside a period — depends on which pair you use, and a schedule reissued mid-term can show an issue date months after inception.

Date formats get the same treatment as anywhere in the product: ambiguous forms are resolved from the rest of the document rather than guessed, and a two-digit year is completed from context.

One thing worth stating plainly: whether cover applied at a particular moment is not a date-arithmetic question. Notification requirements, continuity of cover and the effect of a mid-term adjustment are all matters for a broker. What extraction supplies is the dates, correctly read.

FlowParse
flowparse.io

Comparing this year against last

The single most useful thing extraction enables, and the one nobody does by hand: putting two schedules side by side.

Renewals rarely arrive with a summary of what changed. The premium is visible because it is the number everyone looks at. An excess that moved from one figure to another, or a sum insured that stayed flat while your stock doubled, is visible only to someone comparing two documents line by line — which almost nobody does at renewal.

Extracted into rows, the comparison takes a minute. Same fields, two columns, and the differences are simply visible.

Being precise about what that gives you: it shows the changes. Whether a change matters is a judgement, and one worth taking to your broker with the differences already in hand rather than asking them to find them.

The same mechanism answers the question that comes up after a loss — what the position was at a particular renewal — provided the schedules were extracted at the time rather than being dug out afterwards.

FlowParse
flowparse.io

Endorsements, and why the schedule alone is not the position

A schedule describes the policy as it stood on the day it was issued. Anything that changed afterwards arrives as an endorsement — a separate document, often a single page, frequently filed somewhere other than the schedule it amends.

This is a quiet problem because nothing about it looks wrong. The schedule is genuine, complete and current-looking. It simply no longer describes what you are covered for, and the difference only surfaces when the difference matters.

Endorsements are read into the same fields as the schedule, with the policy number as the key that joins them. The result is a table where a location can appear twice — once as originally scheduled and once as amended — and the amendment is visible rather than buried in a folder.

The pattern to watch for is a location added mid-term with a sum insured that was set in a hurry, or a cover removed to save premium with nobody downstream told. Both are ordinary business decisions; both become expensive when nobody can see them a year later.

A practical habit that costs nothing: extract each endorsement as it arrives rather than accumulating them for renewal. One document, thirty seconds, and the table stays true. The alternative is an annual reconciliation exercise that nobody enjoys and most people skip.

When there is more than one policy

Most organisations of any size have several: property, liability, motor, cyber, professional indemnity, sometimes one per site or per entity. They renew on different dates, sit with different insurers and are filed wherever they arrived.

Three questions become answerable once they are all rows rather than files. What renews next? — a sort by date rather than a diary reconstruction. What do we pay in total? — including the policies nobody remembers. What is the stated limit on a given exposure across the portfolio?

The third is the one that produces surprises, usually of the form “that site has been on the old sum insured since we bought the extension”. It is not a subtle problem; it is just invisible while the answer lives in eight PDFs.

A hundred documents can be processed in one pass, so a portfolio across several entities is an afternoon once and a filter thereafter.

FlowParse
flowparse.io

What uncertainty means here

Every field carries an indication of how confident the reading is, and on insurance documents that matters more than on most.

A misread invoice line is caught by arithmetic: the lines stop summing to the total. A policy schedule has almost no arithmetic to check against — an excess is just a number, and a wrong one looks exactly like a right one.

So the confidence indication is doing more of the work. A figure over a stamp, a poorly scanned page, a field where the label was ambiguous: those are flagged, and they are where a person should look.

The practical rule that follows: on policy documents, check the flagged fields against the original before relying on any of it for a decision. It is a handful of fields, and the alternative — treating extracted policy data as authoritative — is exactly the overreach this page exists to avoid.

FlowParse
flowparse.io

Scans, and documents that are simply long

Insurance documents have two properties that make them awkward beyond the usual. They are long — a commercial policy with wording can run to a hundred pages — and they are frequently scanned rather than born digital, because somebody printed, signed and re-scanned them.

Length matters because the fields you want are concentrated in a few pages and the rest is wording. Reading is directed at the schedule sections rather than at everything, which is why a hundred-page document does not take a hundred times longer than a one-page one.

Scan quality matters because it lowers certainty, and here that is reported rather than hidden. A schedule scanned at an angle with a stamp over the excess produces a flagged field, not a confident wrong number.

Worth saying explicitly: the policy wording is not extracted into fields, because wording is not a field. What comes back is the schedule data, and the wording stays where it is, as the document that governs.

FlowParse
flowparse.io

What it will never do

Interpret the wording

What a clause means, and how it applies to circumstances, is legal interpretation. No extraction is a substitute for a broker or solicitor reading it.

Tell you whether cover responds

The most consequential question in insurance, and the one where a confident wrong answer does the most damage.

Assess adequacy of sums insured

It can place a stated sum next to a figure you supply. Whether that is enough depends on basis of settlement, indexation and averaging — all judgements.

Compare policies for quality

Two schedules with identical numbers can be very different contracts. Comparison of cover is not comparison of fields.

Replace the document

Extracted fields are a working layer. The schedule and wording remain what governs, and every row points back at them.

The fourth deserves emphasis because it is the most tempting misuse. Reducing two policies to a row each makes them look comparable, and the fields are the part of a policy that varies least. The differences that matter are usually in the wording.

The day something happens

Everything above is administration until the morning a site floods, and then two numbers matter within the hour: the sum insured for that specific location, and the excess that applies to that specific cover.

Those two figures are on the schedule. Whether they are findable is a different question, and the answer depends entirely on work done before there was any reason to do it. A forty-page PDF with the locations in the order the insurer chose, plus three endorsements filed elsewhere, is a twenty-minute problem on a morning when twenty minutes is not available.

The same table answers the questions that follow over the next fortnight: which other covers might respond, what the business interruption sum insured is, whether there is a condition attached to this location that somebody should have been complying with, and which insurer the notification actually goes to when the programme is split.

None of these are extraction questions once the data is rows. All of them are document-hunting questions while it is not. The distinction only becomes visible on the one day a year it costs something, which is precisely why it never feels urgent beforehand.

What happens after that first hour — assembling the evidence for the claim itself — is a different body of work, covered in how to evidence a business interruption claimand, from the policyholder’s side, in claims evidence for policyholders.

When to extract, and why timing changes the value

The same operation is worth very different amounts depending on when it happens, and the pattern is consistent enough to plan around.

At renewal, on receipt. The best moment. The comparison against last year is available while there is still time to ask about a change, and the schedule is in your hand rather than in an archive.

When a portfolio is inherited. A new finance role, an acquisition, a broker change. Extracting everything once establishes what exists — which is otherwise a question answered gradually over a year, usually by discovering gaps.

Before a material change. New site, new stock levels, new activity. The stated sums insured are the starting point for the conversation, and having them ready shortens it considerably.

After a loss. Necessary and the least useful timing, because at that point the schedule tells you what the position was rather than letting you change it. Everything about extraction is more valuable before it is needed.

The pattern in all four: the value is in having the facts available at the moment a decision is being made. Extracted a year in advance, they cost nothing to keep and are simply there.

A small habit makes all four work: extract each schedule when it arrives rather than in batches. It takes a minute per document at the point where the document is already open, and it means the portfolio table is never out of date — which is the state in which anyone actually consults it.

What you get back

One row per policy for the headline facts, or one row per section where limits and excesses differ — which on commercial schedules is the more useful shape.

Formats are the same as everywhere: Excel for working by hand, CSV for importing, JSON for anything programmatic. Every row keeps the file and page it came from, so checking a figure is opening a document rather than searching a folder.

What this is good for, concretely: a renewal calendar, a total of what you actually pay, a comparison against last year, and a starting point for the conversation with your broker. What it is not good for is deciding anything about cover on your own.

The claim-side use of the same extraction — turning loss documents into a schedule — is on insurance claim document processing, and the premium and instalment side on premium schedule to Excel.

FlowParse
flowparse.io

Frequently asked questions

Try it on two renewals

This year’s schedule and last year’s, extracted and put side by side. It takes a few minutes and it is the comparison almost nobody makes.

Related