The Trust Gap, Vol. 2: agents disclose how they are governed, not who pays when they fail
Across 1,785 published AI agents and assurance vendors, 77.66% of what we tried to learn is not public. Controlling for how much each company documents, operational governance is disclosed slightly above average and risk transfer at one sixty-seventh of it. We also hand-audited our own insurance data and 21 of 31 rows failed.
Of the 97,755 things this index tried to learn about the 1,785 AI agents and assurance vendors it publishes, **75,917 — 77.66% — are not publicly stated anywhere we could read.** MIT's AI Agent Index reports its own version of this number as 227 of 1,350 data points, or 16.81%. Ours is four and a half times worse, and the gap between those two figures is mostly a statement about our method rather than about the market. We explain why below, because a reader who cannot see the difference cannot use either number.
This is the second volume of The Trust Gap. It reports three things: what the index now measures, one finding that changes our own thesis, and a hand audit of our insurance data in which **two-thirds of the rows failed**.
## 1. The disclosure matrix
Every published entity is scored against the fields in our 59-field catalogue that apply to its type. That yields 97,755 applicable data points across 1,785 published entities.
| Status | Data points | Share | |---|---|---| | `present` — a value with a verbatim quote and a dated source | 21,763 | 22.26% | | `no_public_information` | 75,917 | 77.66% | | never attempted (no row at all) | 75 | 0.08% |
The matrix is 99.92% dense. That density is deliberate: absence is recorded as a value, not left blank, so that "we looked and found nothing" is distinguishable from "we did not look." The 75 empty slots are a defect, not a category, and are listed for repair.
Across all 2,868 entities in the index including unpublished candidates, the figure is 114,327 of 157,180, or 72.74%.
## 2. Disclosure is not uniform, and its shape is the finding
Sorted by fill rate, the seven field categories are stable in order and far apart in magnitude:
| Category | Applicable | Present | Fill rate | |---|---|---|---| | assurance | 30,513 | 3,921 | 12.85% | | models | 10,458 | 1,705 | 16.30% | | safety | 14,280 | 2,352 | 16.47% | | agency | 10,458 | 2,047 | 19.57% | | ecosystem | 10,626 | 2,312 | 21.76% | | impact | 8,925 | 3,423 | 38.35% | | practicality | 12,495 | 6,003 | 48.04% |
Assurance is last. But treating "assurance" as one thing conceals the actual result. Split into its natural blocks:
| Assurance block | Applicable | Present | Fill rate | |---|---|---|---| | **risk transfer** — insurance available, carriers, coverage limits, indemnification, liability cap | 8,925 | 31 | **0.35%** | | **settlement** — outcome pricing, settlement mechanism, dispute process, SLA terms | 7,140 | 715 | 10.01% | | **operational governance** — audit trail, tamper-evidence, runtime governance, permission scopes, explainability, compliance certs, regulatory alignment, eval coverage | 14,280 | 3,093 | **21.66%** | | vendor profile (vendors only) | 168 | 82 | 48.81% |
Operational governance at 21.66% is not an outlier at the bottom. It sits above safety (16.47%) and models (16.30%) and level with ecosystem (21.76%). **Agents describe how they are governed at a perfectly ordinary rate.** The collapse is confined to the risk-transfer block, which is sixty-two times sparser.
### 2.1 Controlling for documentation depth
An obvious objection: entities that document more of everything will fill every block more often, so block comparisons may just track how much a company writes. To control for it we restricted to the 861 published entities that document 12 or more fields — the ones with a substantive page to read — and expressed each block's fill rate as a ratio to that cohort's own overall fill rate of 32.46%.
| Block | Slots | Present | Fill rate | Lift vs. cohort baseline | |---|---|---|---|---| | risk transfer | 4,305 | 21 | 0.49% | **0.015×** | | settlement | 3,444 | 537 | 15.59% | 0.480× | | models | 5,052 | 1,081 | 21.40% | 0.659× | | safety | 6,888 | 1,628 | 23.64% | 0.728× | | ecosystem | 5,128 | 1,733 | 33.79% | 1.041× | | **operational governance** | 6,888 | 2,463 | 35.76% | **1.102×** | | agency | 5,052 | 1,906 | 37.73% | 1.162× | | impact | 4,305 | 2,274 | 52.82% | 1.627× | | vendor profile | 76 | 41 | 53.95% | 1.662× | | practicality | 6,027 | 3,624 | 60.13% | 1.853× |
The control changes the conclusion. Among well-documented entities, **operational governance is disclosed slightly more often than the average field, not less.** Risk transfer sits at 0.015× — roughly one sixty-seventh of baseline, and the lowest figure in the table by a factor of thirty-two.
### 2.2 Where this undercuts our own position
This index exists on the thesis that AI agents cannot be insured, audited, governed, or paid on verified outcomes. On this evidence, the "audited" and "governed" halves of that sentence are too strong and we should stop saying them in that form.
Among the entities that document themselves at all, audit trails, permission scopes, runtime governance and explainability are described at an above-average rate. Whatever the quality of those descriptions — and §3 is a caution about exactly that — the *silence* we have been asserting is not there. What survives, and survives more sharply for being narrowed, is this: vendors describe the controls, and say nothing about who pays when the controls fail. **Governance is a feature, so it gets marketed. Risk transfer is a liability, so it does not.**
The one field with no exceptions at all is `coverage_limits`: **0 present out of 1,785 published entities**, the only field in the 59-field catalogue at 100.00% absence. Not one indexed entity states an AI coverage limit on the page we read.
## 3. We audited our own insurance data and two-thirds of it failed
The 31 `present` risk-transfer rows are few enough to read individually, so we read all of them — every row, not a sample, across 24 entities. **Ten survive review. Twenty-one do not.**
| Defect class | Rows | Example | |---|---|---| | Liability *exclusion* stored as a liability *cap* | 4 | "We are not liable for any damages arising from the use of this service." | | Polarity inversion — the customer indemnifies the vendor | 3 | "By using the extension, you agree to indemnify its creator from any liability." | | Payment processor stored as insurance carrier | 3 | "Secured by Stripe"; "Payments secured by PayPal." | | Customer or commercial counterparty read as carrier | 3 | "Insurance company Conte automated 90% of queries…" — Conte is a customer | | Non-AI insurance | 2 | "Segregated account, $175M FDIC"; "ATOL, ABTA, IATA protected" | | Company-form suffix read as a cap or indemnity | 2 | "913.ai UG (haftungsbeschränkt)" | | Compliance acronyms stored as indemnification | 1 | "BAA, HIPAA, DPA, SSO, RBAC" | | The agent's own demo output read as the vendor's terms | 1 | "Liability cap 1× fees — below firm floor (2×); escalate to partner" | | Vendor's own brand name as evidence of the thing it is named after | 1 | "Indemn — Scale Your Capacity." | | Corporate history — a divested underwriting arm | 1 | "Hiscox to acquire Corix, the underwriting division of Vouch" | | **Total defective** | **21 of 31 (67.7%)** | |
An exclusion is not a cap. A cap states a ceiling on what the vendor will pay; an exclusion states that it will pay nothing. Storing the second as the first makes a vendor look *more* accountable than its own terms claim, and it is our single most common error class.
The ten rows that survive: Armilla AI and Coalition on insurance availability; Coalition naming Allianz and Munich Re aiSure naming Mosaic as carriers; Munich Re aiSure and Gemini Code Assist on customer indemnification; AgentCard on assumed legal responsibility; Agent shield on both carrying professional liability cover and capping liability at the engagement fee; and RentAHuman on fee-funded incident insurance.
Correcting for the audit, the true risk-transfer fill rate is **10 of 8,925 slots, or 0.112%** — not 0.35%. **The headline finding gets stronger when we mark our own homework.** And on the narrower question: of 1,785 published entities, **1,781 (99.78%) have no defensible public evidence of insurance availability.**
Two things this audit rules out. All 24 entities' evidence came from a domain matching their own registered homepage, so none of these errors is cross-entity contamination. And every one of the 21 defects is a *reading* error — the quote is verbatim from the page in every case, enforced on write by a database trigger. Our citations are sound; our interpretation of them was not.
## 4. The scorecard
Two instruments write to `scorecards`, and only one is usable.
**The judged rubric (`v1`)** is adversarial human-reviewed scoring, 0–3 per dimension. On published entities:
| Dimension | Rows | 0 | 1 | 2 | 3 | Mean | Scoring 0–1 | |---|---|---|---|---|---|---|---| | `outcome_settlement` | 54 | 52 | 2 | 0 | 0 | 0.037 | **100.00%** | | `insurance_indemnity` | 54 | 45 | 7 | 2 | 0 | 0.204 | 96.30% | | `compliance` | 54 | 42 | 8 | 4 | 0 | 0.296 | 92.59% | | `audit_trail` | 53 | 42 | 5 | 6 | 0 | 0.321 | 88.68% | | `runtime_governance` | 54 | 32 | 10 | 12 | 0 | 0.630 | 77.78% |
213 of 269 judged rows (79.18%) are zeroes; 245 of 269 (91.08%) are 0 or 1. **No published entity has ever scored above 1 on `outcome_settlement`** across 54 careful readings. There is exactly one score of 3 anywhere in the index, on `runtime_governance`, and it sits on an unpublished candidate.
**Its limits, which must travel with the numbers.** n = 54, purposively chosen — weighted toward entities we were already interested in. It is 3.03% of published entities and is **not a population rate.** Do not read the means as an estimate of the index, and do not read movement in them as a trend.
**The automated scorer (`v1-auto`)** covers 2,105 rows on published entities and cannot be used at all. Its observed range is exactly {1, 2}: zero 0s and zero 3s on every dimension. It scores `score = min(2, evidenced fields in that dimension)`, so a dimension with no evidence produces no row rather than a zero — the finding is discarded rather than recorded. Read naively it reports means of 1.254 to 1.638 against the judged rubric's 0.037 to 0.630, which would describe the assurance layer as roughly four times healthier than careful reading finds it. **The modal judged score is a value this instrument cannot emit.** It has written nothing since 2026-08-02 and no population statistic in this report is derived from it.
### 4.1 The gap is in the missing rows
| Dimension | Published entities with no score at all | Share | |---|---|---| | `insurance_indemnity` | 1,718 | 96.25% | | `outcome_settlement` | 1,385 | 77.59% | | `compliance` | 1,303 | 73.00% | | `runtime_governance` | 1,231 | 68.96% | | `audit_trail` | 914 | 51.20% |
Every unscored entity is a question the index has not answered. This is the honest state of our coverage, and it is the number we would most like a reader to hold us to next quarter.
## 5. What we cannot claim, stated before someone states it for us
**Every published entity is evidenced from a single document.** 1,783 of 1,785 draw on exactly one; two draw on more than one capture; **zero draw on more than one distinct URL.** The two multi-document entities hold re-crawls of the same address, not second sources.
**That document is almost always a homepage.** 1,640 of 1,785 (91.88%) are evidenced from a bare root URL. Exactly **three** entity-URL pairs in the entire index touch a path containing terms, legal, privacy, trust, security, MSA, SLA or DPA.
Indemnity clauses, liability caps and coverage limits are published in terms of service, master agreements and trust centres. We have read essentially none of those. So every absence in this report means one specific thing:
> **Not stated on the page we captured, on the date we captured it** — and for 91.88% of entities that page is the homepage, which is not where this information would be published even by a vendor that publishes it.
This is why our 77.66% is not comparable to MIT's 16.81%. MIT used seven expert annotators reading multiple documents per agent across 30 agents. We used one automated capture across 1,785. Their number describes what a determined researcher cannot find; ours describes what a company does not put on its front page. Both are real measurements. They are not the same measurement, and anyone citing them side by side without this paragraph is misusing at least one of them.
**The single test that would settle it:** crawl `/terms`, `/legal`, `/trust` and `/security` for the ten entities with any risk-transfer evidence and re-measure. If `coverage_limits` is still 0 across four pages each, the strong claim — that this market does not publish coverage limits, full stop — is earned. Until then it is not, and this report does not make it.
Two further scope limits. A fifth of what we publish is off-thesis: 352 of 1,785 published entities carry a subtype matching video, image, music, audio, web3, crypto, gaming, NSFW, dating, meme or art, and match neither "AI agent" nor "assurance vendor" as we define them. And two analyses we wanted are not computable: `autonomy_level` has roughly 50 distinct free-text values across 58 present rows, and `model_provider` splits "Google" from "Google DeepMind" from "Gemini". **We cannot yet tell you whether more autonomous agents have weaker audit trails**, and we will not estimate it.
## 6. Findings
1. **77.66% of what we tried to learn is not public** — 75,917 of 97,755 data points across 1,785 published entities. 2. **The gap is risk transfer, not governance.** Controlling for documentation depth, operational governance is disclosed at 1.102× the average field; risk transfer at 0.015×. Our own thesis was too broad and is narrowed here. 3. **`coverage_limits` is 0 of 1,785** — the only field in the catalogue at complete absence. 4. **99.78% of published entities have no defensible evidence of insurance availability** — 1,781 of 1,785, after correcting our own errors. 5. **Two-thirds of our insurance rows were wrong** — 21 of 31, all reading errors, none citation errors. The corrected fill rate is 0.112%. 6. **No entity has scored above 1 on outcome settlement** in 54 human judgements. 7. **Everything above rests on one document per entity, a homepage 91.88% of the time.** Treat every absence as "not on that page, on that date."
---
## Appendix: reproducing these numbers
Run against the `aidispatch-index` Postgres database on 2026-08-10. Every figure above comes from one of these queries. The applicable-slot denominator — `field_catalog.applies_to` matching the entity's type — is used throughout; using a flat 59 fields per entity would overstate every denominator and understate every fill rate.
### A1. The disclosure matrix (§1)
```sql with pub as (select id, entity_type from entities where status='published'), slots as ( select p.id as entity_id, fc.field_key from pub p join field_catalog fc on fc.applies_to = 'both' or fc.applies_to = p.entity_type) select (select count(*) from pub) as published_entities, (select count(*) from slots) as applicable_slots, count(*) filter (where ef.value_status = 'present') as present_rows, count(*) filter (where ef.value_status = 'no_public_information') as npi_rows, (select count(*) from slots) - count(ef.id) as slots_never_attempted from slots s left join entity_fields ef on ef.entity_id = s.entity_id and ef.field_key = s.field_key; -- 1785 | 97755 | 21763 | 75917 | 75 ```
Swap `where status='published'` for no filter to get the all-entity figure (2,868 | 157,180 | 114,327 | 72.74%).
### A2. Category and block fill rates (§2)
```sql with pub as (select id, entity_type from entities where status='published'), slots as (select p.id eid, fc.field_key, fc.category from pub p join field_catalog fc on fc.applies_to='both' or fc.applies_to=p.entity_type) select s.category, count(*) applicable, count(*) filter (where ef.value_status='present') present, round(100.0*count(*) filter (where ef.value_status='present')/count(*),2) fill_rate_pct from slots s left join entity_fields ef on ef.entity_id=s.eid and ef.field_key=s.field_key group by 1 order by fill_rate_pct; ```
For the assurance blocks, replace the `group by` key with this CASE, restricted to `fc.category='assurance'`:
```sql case when field_key in ('insurance_available','insurance_carriers', 'coverage_limits','indemnification','liability_cap') then 'risk_transfer' when field_key in ('outcome_based_pricing','settlement_mechanism', 'dispute_process','sla_terms') then 'settlement' when field_key like 'vendor_%' then 'vendor_profile' else 'operational_governance' end ```
### A3. Disclosure lift, controlled for depth (§2.1)
```sql with pub as (select id, entity_type from entities where status='published'), depth as (select p.id, p.entity_type, count(*) filter (where ef.value_status='present') n_present from pub p left join entity_fields ef on ef.entity_id=p.id group by 1,2), cohort as (select id, entity_type from depth where n_present >= 12), slots as (select c.id eid, fc.field_key, case when fc.field_key in ('insurance_available','insurance_carriers', 'coverage_limits','indemnification','liability_cap') then '1 risk transfer' when fc.field_key in ('outcome_based_pricing','settlement_mechanism', 'dispute_process','sla_terms') then '2 settlement' when fc.category='assurance' and fc.field_key not like 'vendor_%' then '3 operational governance' when fc.category='assurance' then '4 vendor profile' else '5 '||fc.category end as block from cohort c join field_catalog fc on fc.applies_to='both' or fc.applies_to=c.entity_type), f as (select s.block, count(*) slots, count(*) filter (where ef.value_status='present') present from slots s left join entity_fields ef on ef.entity_id=s.eid and ef.field_key=s.field_key group by 1) select block, slots, present, round(100.0*present/slots,2) fill_pct, round((100.0*present/slots)/(select 100.0*sum(present)/sum(slots) from f),3) lift from f order by fill_pct; -- cohort: 861 entities, baseline fill rate 32.46% ```
### A4. The risk-transfer rows we audited by hand (§3)
```sql select e.id, e.name, e.entity_type, ef.field_key, ef.value_text, ef.quote, ev.source_url, ev.retrieved_at::date, ef.extracted_by from entity_fields ef join entities e on e.id=ef.entity_id and e.status='published' join evidence ev on ev.id=ef.evidence_id where ef.field_key in ('insurance_available','insurance_carriers', 'coverage_limits','indemnification','liability_cap') and ef.value_status='present' order by ef.field_key, e.name; -- 31 rows across 24 entities. Classification in section 3 is our editorial -- judgement, not a computed field; the query returns the rows, we read them. ```
The domain-identity check that rules out cross-entity contamination compares the registered `homepage_url` host against the `evidence.source_url` host for the same rows; all 24 matched.
### A5. Scorecards (§4)
```sql select sc.rubric_version, sc.dimension, count(*) rows, count(*) filter (where sc.score=0) s0, count(*) filter (where sc.score=1) s1, count(*) filter (where sc.score=2) s2, count(*) filter (where sc.score=3) s3, round(avg(sc.score),3) mean_score, round(100.0*count(*) filter (where sc.score<=1)/count(*),2) pct_scoring_0_1 from scorecards sc join entities e on e.id=sc.entity_id and e.status='published' group by 1,2 order by 1, mean_score; ```
Unscored coverage (§4.1) counts published entities with no `scorecards` row for a dimension under either rubric version.
### A6. The evidence ceiling (§5)
```sql select count(*) filter (where nd=1) one_doc, count(*) filter (where nd>1) multi_doc, count(*) filter (where nu>1) multi_distinct_url from (select ef.entity_id, count(distinct ev.raw_document_id) nd, count(distinct ev.source_url) nu from entity_fields ef join entities e on e.id=ef.entity_id and e.status='published' join evidence ev on ev.id=ef.evidence_id group by 1) t; -- 1783 | 2 | 0
select count(*) pairs, count(*) filter (where regexp_replace(source_url,'^https?://[^/]+','') in ('','/')) root, count(*) filter (where source_url ~* '/(terms|legal|tos|privacy|trust|security|msa|sla|dpa)') legal from (select distinct e.id, ev.source_url from entities e join entity_fields ef on ef.entity_id=e.id join evidence ev on ev.id=ef.evidence_id where e.status='published') x; -- 1785 | 1640 | 3 ```
### Method notes
- **Nothing in this report is estimated, extrapolated, or rounded from an uncomputed figure.** Percentages are rounded to two decimals from exact counts; the counts are printed alongside so any reader can recompute them. - `v1-auto` scorecard rows are excluded from every population statistic, for the reason given in §4. - The section 3 defect classification is editorial judgement applied to 31 rows we read in full. The rows are enumerated by A4 so that a reader who disagrees with a classification can see exactly which row it applies to and reach their own conclusion. - MIT's figures are quoted from the AI Agent Index's published methodology, not recomputed by us.
Entries in this piece 14
Published index entries backed by the same source documents this piece cites.
Sources 28
“aiSure TM -backed performance warranties enable you to indemnify your clients for their financial losses or legal liabilities directly related to AI errors.”
www.munichre.com · checked Jul 31, 2026“Secured by Stripe”
www.walletfinder.ai · checked Aug 2, 2026“By using the extension, you agree to indemnify its creator from any liability.”
chromewebstore.google.com · checked Aug 6, 2026“Liability cap 1× fees — below firm floor (2×); escalate to partner”
agentman.ai · checked Aug 7, 2026“Our liability is limited to the engagement fee paid.”
agentshield.win · checked Aug 1, 2026“Indemn — Scale Your Capacity. Deliver Agentic Business Results.”
www.indemn.ai · checked Aug 4, 2026“pay via Stripe”
ilikeimg.com · checked Aug 7, 2026“Limitation of liability. To the maximum extent permitted by law, Free AI Humanizer and its operator are not liable for indirect or consequential damages arising from your use of the Service.”
freeaihumanizer.online · checked Aug 3, 2026“EXCEPT WHERE OTHERWISE PROHIBITED BY APPLICABLE LAW, DECKDROP SHALL NOT BE LIABLE TO EVALUATOR FOR ANY CLAIMS IN CONNECTION WITH THE TRIAL SERVICES OR ANY OTHER SERVICES PROVIDED BY DECKDROP TO EVALUATOR.”
deckdrop.io · checked Aug 2, 2026“Armilla provides named, affirmative AI insurance that responds to those failure modes directly”
www.armilla.ai · checked Jul 30, 2026“913.ai UG (haftungsbeschränkt)”
913.ai · checked Aug 5, 2026“Hiscox to acquire Corix, the underwriting division of Vouch”
vouch.us · checked Jul 31, 2026“Active Cyber Insurance for Enterprises Comprehensive cyber coverage backed by proprietary intelligence and Allianz’s A+ rated capacity.”
www.coalitioninc.com · checked Jul 31, 2026“BAA, HIPAA, DPA, SSO, RBAC”
anchorbrowser.io · checked Aug 7, 2026“Cyber Insurance | Active Insurance & Cybersecurity | Coalition Now Available: Active Cyber Insurance for Enterprises”
www.coalitioninc.com · checked Jul 31, 2026“To the fullest extent permitted by law, XNODE Inc. and its affiliates shall not be liable for any direct, indirect, incidental, consequential, or punitive damages arising from your use of the platform.”
xnode.ai · checked Aug 1, 2026“The legal responsibility sits with us, not you.”
agentcard.sh · checked Aug 7, 2026“Eventuguard, acquired by Jewelers Mutual Insurance.”
www.indemn.ai · checked Aug 4, 2026“ATOL, ABTA, IATA protected”
www.holiwise.com · checked Aug 1, 2026“Google’s IP indemnification policy helps protect Gemini Code Assist licensed users from potential legal ramifications concerning copyright infringements.”
cloud.google.com · checked Aug 1, 2026“Segregated account, $175M FDIC”
www.deferred.com · checked Aug 1, 2026“Insurance company Conte automated 90% of queries related to insurance purchases, renewals, and refunds”
kommunicate.io · checked Aug 1, 2026“Payments secured by PayPal.”
craigmbrown.com · checked Aug 1, 2026“We carry professional liability insurance.”
agentshield.win · checked Aug 1, 2026“Plus requester reputation and incident insurance funded from the fee.”
workman.rent · checked Aug 2, 2026“Mosaic Insurance has partnered with Munich Re to harness the power of aiSure™”
www.munichre.com · checked Jul 31, 2026“July 29, 2026 Banking on the Frontier Amp Labs partners with Westpac to transform how the bank builds and delivers technology.”
ampcode.com · checked Aug 7, 2026“We are not liable for any damages arising from the use of this service.”
www.aisynthidremover.com · checked Aug 2, 2026