The number on the slide is not your number
Every document-extraction vendor pitching a broking firm quotes an accuracy figure, and it is almost always a single number in the high nineties. It is also almost always useless for the decision you are making, because it answers a question you did not ask. It is one averaged figure, computed on the vendor's own test set, which is cleaner than your inbox, weighted differently from your work, and measured on fields you may not even consume.
The corpus already has plenty on what extraction can do: reading policy schedules for advisor tooling, pulling loss runs for renewals, handling the edge cases in commercial claims. This post is about the question none of those answers: when a vendor says its tool does this, how do you test whether that is true on your documents before you sign.
The short version is that you test it the way you would test any claim about the physical world: with your own materials, on the fields you actually use, against a truth you established yourself, including the hard cases the vendor's demo quietly avoided. This playbook works through building that test, measuring it in a way that means something, stressing it where your real documents are hardest, and then reading the pricing, the data-handling and the contract with the same scepticism. The firms that do this buy tools that survive contact with their mailbox. The firms that buy on the slide find out the accuracy figure was real and irrelevant.
Build a ground-truth set from your own documents
The foundation of any honest evaluation is a test set the vendor did not build and does not see. You construct it from your own documents, and its whole value is that it looks like your real work rather than a curated sample.
Start by sampling across the two dimensions that make your document population varied: type and source. The types are the documents you actually process, policy schedules, endorsement documents, commission statements, claim forms, loss runs, proposal forms. The sources are the insurers you place with, because a policy schedule from one insurer looks nothing like another's, and a tool that reads two insurers well can fail on the third. A test set drawn from one insurer's clean schedules will tell you nothing about the account that places across fifteen.
Then do the unglamorous part: hand-label the fields you consume. For each document in the set, a person records the correct value of every field you would want the tool to extract, the policy number, the insured, the premium, the sum insured, the dates, the insurer, the peril, whatever your workflow actually uses. This hand-labelled truth is the yardstick, and it is worth the effort because everything downstream is measured against it. Keep the set held out: the vendor runs its tool on your documents and returns its extractions, and you score those against your labels, so the vendor never gets to tune to your test.
Size the set for coverage, not volume. A few hundred documents spread deliberately across your insurers, document types and quality levels tells you far more than a few thousand from one clean source. Include the messy ones on purpose, because the whole point is to find out where the tool breaks on the work you actually do, not to confirm it handles the easy cases.
Measure per field, not per document
With a labelled set in hand, the measurement decision is where most evaluations go soft. A document-level accuracy figure, the vendor's favourite, averages away exactly the information you need. What matters is how the tool does on each field, weighted by what a wrong value costs you.
Measure accuracy field by field. The policy number, the sum insured, the premium, the inception and expiry dates, the insurer name, each gets its own score, because they fail at different rates and matter to different degrees. A tool can read insured names almost perfectly and still misread sums insured often enough to be unusable, and a document-level average hides that behind a comfortable headline. For each field, look at both how often it is right when the field is present and how often it invents a value that is not there, because a confidently wrong number is worse than a blank.
The metric that vendors least like to discuss, and that you should insist on, is the silent-error rate: how often the tool returns a confident value that is wrong, with nothing to signal the problem. A tool that flags its uncertainty lets you route the doubtful cases to a human; a tool that is confidently wrong pushes the error straight into your workflow undetected. Measure the confidence the tool reports against whether it was actually right, because a confidence score that does not track accuracy is worse than none, since it invites you to trust the wrong extractions.
Stress the cases your inbox actually contains
Vendor demos run on clean, native-text PDFs, and your inbox does not. The evaluation has to include the document conditions that make real extraction hard, because those are where tools diverge and where the vendor's headline number was never tested.
Four stress cases matter most for an Indian broking firm.
- Scanned and faxed documents. A schedule printed, signed, and rescanned has no text layer, so the tool must read it as an image through optical character recognition. Poor scans, skew, shadows and low resolution are where accuracy collapses, and a tool that is excellent on native PDFs can be poor here.
- Vernacular and mixed-language documents. Documents carrying regional-language text, or a mix of English and a local language, are common, and a tool trained mainly on English can drop or garble the vernacular content. Government efforts such as Bhashini have pushed vernacular document handling forward, but the tool you are buying either handles your languages or it does not, and only your documents will tell you.
- Multi-policy and multi-document PDFs. A single PDF that contains several policies, or a schedule bundled with endorsements and a covering letter, tests whether the tool can segment the file correctly before extracting, and many stumble by merging fields across policy boundaries.
- Tables, annotations and non-standard layouts. Schedules of property, vehicle lists, and hand-annotated documents test structure handling, and a tool that reads a clean key-value layout can fail on a dense table.
Run each stress case as its own scored slice, not folded into the average, so you can see where the tool is strong and where it needs a human safety net. The answer is rarely that a tool is uniformly good or bad; it is that it is reliable on some conditions and not others, and knowing which is what lets you decide whether the failure modes are ones your workflow can absorb.
Read the pricing model, not just the price
Two vendors quoting a similar headline price can cost very different amounts on your actual volume, because the pricing model interacts with your document mix in ways the sticker does not show. The model matters as much as the number.
The common structures are per-page, per-document, per-field, and subscription or committed-volume. Each has a shape that favours some workloads and punishes others. Per-page pricing penalises long documents, so a firm whose schedules run to many pages of property or vehicle lists pays disproportionately, while a per-document model favours exactly that firm. Per-field pricing can look cheap until you count how many fields you actually pull from each document. A subscription flattens the cost but only pays off above a volume you should verify you will hit.
The costs that break the business case are usually the ones outside the headline rate. Ask how re-processing is charged when a document fails and has to be run again, whether human-in-the-loop review of low-confidence extractions is included or billed separately, and what happens on overage above a committed volume. A tool priced attractively per page can become expensive once the re-runs on your scanned documents and the review time on your silent errors are counted. Model the total cost on your real volume mix, not the demo.
The honest way to compare is to take your own annual volume, split by document type and page count, and run it through each vendor's model to get a real total cost, including the review effort each tool's error rate will impose on your team. A tool with a slightly higher unit price but a lower silent-error rate can be cheaper overall, because it saves the human hours that a cheaper, less reliable tool spends. Price the outcome, the correctly extracted document with the errors caught, not the raw extraction.
Data residency, DPDP and where your documents go
Insurance documents are full of personal and commercial data, and where a vendor processes them is a compliance question a broking firm cannot skip. Under the Digital Personal Data Protection Act, 2023 (DPDP), a firm handling client personal data carries obligations as the entity that determines how that data is processed, and pushing documents through an extraction vendor does not transfer those obligations away.
The questions to ask are specific. Where are the documents processed, in India or abroad, and if abroad, does the arrangement meet your cross-border-transfer position under the Act. Does the vendor use a sub-processor, in particular a foreign large-language-model API, to do the actual extraction, because that means your client documents pass through a third party you must then diligence in turn. What is the retention and deletion regime: how long does the vendor hold your documents after processing, and can you require deletion. And, the question vendors are least forthcoming about, are your documents used to train the vendor's models, and can you opt out, because a firm's client documents becoming training data for a tool sold to competitors is a confidentiality problem as much as a data-protection one.
These are not procurement formalities. A broking firm is trusted with its clients' commercial and personal information, and the choice of an extraction vendor is a choice about who else touches that information. The data-handling terms belong in the evaluation with the same weight as the accuracy, because a tool that is accurate and careless with data is not a tool a firm can responsibly use.
Integration reality and the contract that holds
A tool that scores well on the test set can still fail in production if it does not fit the workflow, so the last two checks are about fit and about what happens when things go wrong.
The integration reality check asks whether the tool fits how you actually work. Does it offer an interface your systems can call, or does it force manual upload and download that adds a step rather than removing one. What is its throughput and latency at your volumes, since a tool that is accurate but slow can bottleneck a renewal cycle. How do low-confidence extractions surface for human review, because the review interface is where your team will spend its time and a bad one erases the time the tool saved. And does it make your team change their process to suit the tool, or does it slot into the process they have. A tool that is technically good but operationally awkward gets abandoned.
The contract is the final defence, and three terms carry the weight. An accuracy service level has to be defined on something measurable, on which fields, on what document set, measured how, because an accuracy promise with no defined basis is unenforceable marketing. An exit clause has to let you leave without penalty if the tool underperforms against that defined standard. And the data-return-and-deletion terms have to guarantee that on exit your documents are returned and deleted and your hand-labelled test set and any accumulated corrections are yours to take, not locked into the vendor, because that labelled data is an asset you built and will need for the next evaluation.
The through-line of the whole playbook is that extraction accuracy is a claim to be tested on your own materials, not a number to be accepted. That same scepticism applies to what sits behind the extracted fields: knowing a tool pulled the sum insured is not the same as knowing whether the policy-wording behind it actually responds to the client's risk. Sarvada gives brokers searchable access to insurer wordings, so the extracted data connects to the cover it describes and a broker can check what a policy actually says rather than trusting a field in isolation. If you are evaluating extraction tooling and want the wording layer behind the data, Request Access to Sarvada.