AI & Insurtech

Buying a One-Year-Old AI Startup: Why Training-Data Reps Are the Hardest Thing to Insure

Adobe's acquisition of Rilo, an AI marketing startup founded in 2025, is a template for a wave of fast Indian AI exits. Warranty and indemnity insurance is now common on mid-market deals, but AI targets break the IP, non-infringement and model-behaviour warranties underwriters are asked to cover.

Sarvada Editorial TeamInsurance Intelligence
11 min read

Listen to this article

Audio version • 11 min read

warranty and indemnitym&aai liabilityip risktransactional risk

Last reviewed: September 2026

A company founded in 2025 just got bought in 2026

On 3 September 2026 the StartupTalky Daily Indian Funding Roundup recorded that Adobe had acquired Rilo, an AI marketing startup founded in 2025, in Adobe's second India acquisition. Financial terms were undisclosed. Strip out the buyer's name and the shape of that deal is the one Indian dealmakers should expect to see repeatedly over the next two years: a target barely twelve months old, a small team, a product built on models the company did not train from scratch, and a strategic acquirer paying for capability rather than for revenue.

That shape breaks the diligence habits built around conventional mid-market M&A. A twelve-month-old company has no audited history worth arguing about, almost no customer concentration to unpick, and a balance sheet that fits on one page. What it does have is a training corpus, a set of model weights, a fine-tuning pipeline, third-party API dependencies, and a product whose behaviour nobody can fully specify. Nearly all of the purchase price sits in those assets, and nearly all of the risk sits in how they were assembled.

Warranty and indemnity insurance is now common on Indian mid-market deals, and buyers reach for it by reflex. The product works by insuring the buyer against loss from a breach of the warranties the seller gives in the share purchase agreement, so the seller can exit with a low cap and no escrow while the buyer keeps recourse against an insurer. On an AI target that reflex meets a problem. The warranties a buyer most needs are the ones an underwriter is least willing to write, for reasons that have nothing to do with the seller's honesty.

The three warranties that do not survive contact with an AI target

Almost every dispute about insuring an AI acquisition reduces to three warranties in the share purchase agreement.

Clean title to the intellectual property. The seller warrants it owns or is validly licensed to use the IP being sold. On a software company this is a documentation exercise: assignment deeds from founders and contractors, employment agreements with IP clauses, open-source licence inventory. On an AI company the same warranty extends to a training corpus that may have been scraped, bought from a data broker with thin provenance records, derived from customer data used beyond the terms the customer agreed to, or built on base weights released under a licence that restricts commercial redistribution. Ownership of the model is not the same question as the right to have used what went into it.

Non-infringement. The seller warrants that the business as conducted does not infringe third-party rights. Indian law gives no comfortable answer here. The Copyright Act, 1957 provides an enumerated fair dealing exception rather than an open-ended fair use test, and India has no statutory text and data mining exception of the kind some other jurisdictions have introduced. Indian proceedings on whether training a model on copyrighted material infringes remain undecided. A seller giving an unqualified non-infringement warranty on a model trained partly on scraped material is warranting the outcome of an open legal question.

Model behaviour. Buyers increasingly ask for warranties that the product performs to stated accuracy, that outputs do not reproduce third-party content, that no prohibited data sits in the training set. A customer contract can be diligenced by reading it. A model cannot. Evaluation results describe behaviour on a test set, not behaviour in general, so a warranty about model output is a warranty about a system whose failure modes are discovered rather than enumerated.

What underwriters are actually excluding

The exclusions arriving on AI targets are not improvised deal by deal. They follow a market-wide retreat from unpriced AI exposure. Fenwick's 2026 analysis The End of Silent AI describes AI-related exposures being carved out of conventional covers, with insurers introducing AI exclusions and coverage fragmenting across cyber, technology errors and omissions, and general liability wordings. Transactional risk underwriters sit downstream of the same reinsurance appetite, and the caution shows up in W&I policy wording in four recurring forms.

  • Training-data provenance exclusions. The policy carves out loss arising from the sources, licensing or lawfulness of data used to train, fine-tune or evaluate any model.
  • Output infringement exclusions. Claims that model output reproduces or is derivative of third-party copyrighted material are excluded, sometimes alongside a broader IP infringement carve-out.
  • Open-source and model-licence exclusions. Loss from breach of licence terms attaching to base weights or open-source components, including copyleft contamination of proprietary code.
  • Regulatory and data-protection carve-outs. Exposure under the Digital Personal Data Protection Act, 2023 where personal data entered a training set without a lawful basis, and fines that are uninsurable as a matter of public policy in any event.

A fifth exclusion is structural rather than AI-specific and catches buyers out more often. W&I never covers known issues. Anything identified in diligence, disclosed in the data room, or flagged in the disclosure letter is excluded by definition, because insurance prices uncertainty and a known problem is not uncertain. On an AI target, thorough diligence therefore has an awkward property: the better the buyer understands the provenance gaps, the more of them move from insured warranty risk into the disclosed and excluded column. That is not an argument for diligencing less. It is an argument for deciding early which risks the policy is expected to carry and which need a different instrument.

Diligence built to keep cover available

Underwriters do not exclude AI risk because it is AI. They exclude it because the submission gives them nothing to underwrite. The counter is a provenance record assembled before the risk goes to market, structured so an underwriter can trace each asset to a right to use it.

The provenance ledger

One row per dataset, with the source, the acquisition date, the legal basis for use, and the evidence supporting it. Scraped sources need the terms of service and robots directives in force at the time of collection, not today's. Purchased datasets need the vendor contract, the vendor's own warranties, and any indemnity flowing from the vendor to the target, which is an asset the buyer inherits and should value. Customer-derived data needs the specific contractual clause permitting training use, and the answer is frequently that no such clause exists.

The model lineage record

Base weights and their licence terms, every fine-tuning run and the data it consumed, and the boundary between what the target built and what it merely calls through an API. A product that is a thin layer over a third-party foundation model carries a different risk profile from one trained in house, and pretending otherwise in the submission costs credibility.

Two further records sit alongside those. Whether personal data entered training sets, on what lawful basis under the DPDP Act, 2023, and whether deletion requests can be honoured against a trained model. Alongside that, founder and contractor IP assignments, which on a company that went from incorporation to acquisition in about a year are frequently incomplete.

This workstream is a variant of the discipline set out in insurance-linked due diligence for Indian M&A, applied to assets that have no title register. Start it in the first week of diligence. A provenance pack assembled in the final fortnight before signing will not change an underwriter's mind, because there is no time left to verify it.

Pricing a specific indemnity for training-data risk

When the general W&I policy will not carry training-data exposure, the risk has to be allocated somewhere else. Four routes are used in practice, and they are not alternatives so much as a ladder the parties climb until someone accepts the risk.

  1. A seller specific indemnity, uninsured. The seller indemnifies the buyer for training-data and IP infringement claims on a rupee-for-rupee basis, usually with a longer survival period than the general warranties and often outside the general cap. On a founder-led target this is only as good as the founders' post-completion solvency, which after a small all-cash exit is rarely reassuring.
  2. A ring-fenced escrow or holdback. A portion of consideration held back specifically against IP and training-data claims, released on a longer clock than any other escrow. Certain, funded, and expensive in seller goodwill, because it is exactly the clean break the seller wanted W&I to deliver.
  3. Affirmative cover bought back into the W&I tower. Some underwriters will remove or narrow the training-data exclusion where the provenance pack is strong, the corpus is licensed or synthetic, and the buyer accepts a higher retention on that head of loss. Pricing is quoted as an additional rate on line on the layer, and it sits well above the base W&I rate because the underwriter is pricing an unresolved legal question rather than a diligenced fact.
  4. A standalone contingent legal risk or IP policy. Where a specific identified exposure is capable of being scoped, for example one dataset with a defective licence, contingent risk underwriters will sometimes price it separately from the deal policy.

The honest description of the market is that route three is available on a minority of AI targets and route four on fewer still. What moves an underwriter is evidence: a corpus that is licensed, synthetic or first-party; upstream vendor indemnities that survive change of control; a documented deletion and retraining capability; and a target that can show what it did not train on. What loses the argument is a provenance record reconstructed from memory after the term sheet is signed.

The post-completion exposure W&I was never going to cover

Even a perfectly negotiated deal policy leaves the acquirer with a live operating risk from the day it takes control. W&I looks backwards at facts as they stood at completion. From completion onwards, the buyer owns a model that generates outputs, and those outputs create third-party exposure that belongs to a different set of policies entirely.

That market has begun to organise. The Insurer reported on 24 April 2026 that a standalone AI liability market has started to form, including Vanguard AI, a coordinated structure launched by Chaucer and Armilla AI on 10 February 2026. It combines cyber and technology errors and omissions cover with a standalone AI liability policy responding to inaccurate outputs, hallucinations and autonomous agent failures, with AI aggregate limits of USD 25 million or more per organisation.

The reason a coordinated structure matters to an acquirer is that the acquired product's failures otherwise land unevenly across policies the group already holds. An inaccurate output that harms a customer could be argued into technology E&O, into general liability, or out of both, and the Fenwick analysis of exclusion drift makes the third outcome more likely each renewal. Mapping the acquired exposure across the buyer's existing tower is a distinct exercise from the deal policy, and one worth doing before the first post-completion renewal rather than after the first claim. The gap map across cyber, PI, D&O and general liability sets out where those seams usually open.

For Indian acquirers there is a placement question underneath all of this. Standalone AI capacity today sits largely with Lloyd's syndicates and international carriers, so cover for an Indian subsidiary's AI exposure typically arrives through a parent programme, a non-admitted placement, or an IFSCA GIFT City structure. Settle how an Indian claim would actually be paid before relying on the limit.

A sequence for the acquirer

Order matters here more than in a conventional deal, because the diligence output is the underwriting submission.

  1. Decide in week one whether the deal is insurable as structured. If the target trained on scraped data and cannot evidence licences, assume the training-data exclusion applies and plan the allocation around it rather than discovering it at the underwriting call.
  2. Run the provenance workstream in parallel with legal diligence, not after it. Datasets, model lineage, personal data, and founder and contractor IP assignments, each with documentary evidence.
  3. Separate warranty risk from operating risk on paper. Historical facts at completion belong in the W&I discussion. Future model behaviour belongs in the liability tower discussion. Conflating them produces a policy that disappoints on both.
  4. Take the exclusions list from the underwriter early and negotiate against it. The value of a W&I policy is what remains after the exclusions, not the headline limit, a point that holds on every deal and bites hardest on AI targets.
  5. Match survival periods to how the claims actually arrive. Long-tail IP and data-protection exposure needs a longer clock than the general warranty period, whether the protection comes from the policy, a specific indemnity, or escrow.
  6. Price the residual honestly and put it in the model. Where no insurer and no solvent seller will carry training-data risk, the buyer is carrying it. That belongs in the valuation, not in a footnote.

Working through this well means comparing what a transactional risk policy grants against what the buyer's cyber, technology E&O and liability wordings exclude, clause by clause, across several insurers. Sarvada gives commercial-insurance brokers structured, searchable access to insurer policy wordings, so a deal policy's exclusions and the surrounding liability tower can be read side by side rather than reassembled by hand. Brokers advising on AI acquisitions can Request Access to evaluate the platform for their transactional risk practice.

Frequently Asked Questions

Can warranty and indemnity insurance cover training-data risk on an AI acquisition at all?
Sometimes, and only where the evidence supports it. Many W&I policies now carry a training-data provenance exclusion by default. Underwriters will consider narrowing or removing it where the corpus is licensed, synthetic or first-party, upstream vendor indemnities survive change of control, and the buyer accepts a higher retention on that head of loss. Pricing comes as an additional rate on line above the base W&I rate, because the underwriter is pricing an unresolved legal question rather than a diligenced fact.
Why does thorough diligence sometimes reduce what the W&I policy covers?
W&I insures uncertainty, not known problems. Anything identified in diligence, disclosed in the data room or recorded in the disclosure letter is excluded by definition. On an AI target, the deeper the provenance review, the more gaps move from insured warranty risk into the disclosed and excluded column. The answer is not to diligence less. It is to decide early which risks the policy is expected to carry and to allocate the rest through a specific seller indemnity, a ring-fenced escrow, or a price adjustment.
Does a W&I policy protect the buyer if the acquired model produces a harmful output after completion?
No. W&I responds to loss from a breach of warranty about facts that existed at completion. Post-completion model failure is an operating exposure that belongs to cyber, technology errors and omissions, or standalone AI liability cover. The Insurer reported on 24 April 2026 that a standalone AI liability market has begun forming, including Vanguard AI from Chaucer and Armilla AI, which combines cyber and technology E&O with an AI liability policy covering inaccurate outputs, hallucinations and autonomous agent failures.
How long should a training-data indemnity survive after completion?
Longer than the general warranties. Copyright and data-protection claims surface years after the conduct that caused them, so an indemnity running on the same eighteen-month clock as the general warranty package will usually have lapsed before the exposure it was written for arrives. Match the survival period to how the claim actually arrives, and make sure the escrow or indemnity backing it runs on the same longer clock.
What should an Indian acquirer prepare before taking an AI target to the W&I market?
A provenance ledger with one row per dataset showing source, acquisition date, legal basis and supporting evidence; a model lineage record covering base weights and their licence terms, every fine-tuning run, and the boundary between owned and API-accessed capability; a record of whether personal data entered training sets and on what lawful basis under the DPDP Act, 2023; and complete founder and contractor IP assignments, which are frequently missing on a company that went from incorporation to exit in about a year.

Related Glossary Terms

Related Insurance Types

Related Industries

Related Articles

Sarvada Intelligence

Ready to see Sarvada in action?

Explore the platform workflow or start a product conversation with our underwriting automation team.

Explore the platform