The Trust Gap, Vol. 10: this index can audit 18% of its own trust-gap judgements
563 of 689 human-judged scores on published entities carry no pointer to the evidence they rest on. We audited the 126 that do, re-ran the test over all of them through the rubric mapping, and found exactly one erased trust gap in 806 entities — alongside a scoring ceiling reached once in 1,852 opportunities.
## What this report is
Every week this index publishes a number: how many of the AI agents and assurance vendors it covers can show public evidence that they are insurable, auditable, governable, and payable on verified outcomes. This is the tenth such report.
MIT's AI Agent Index earned its authority with one move: it published its own gaps, stating plainly that 227 of its 1,350 data points lacked public information. We hold ourselves to the same standard, and this volume applies it to the instrument rather than the subject. The question here is not how well the agents score. It is whether this index can verify its own scores against the evidence they claim to rest on.
The answer is that it can check 18.29% of them.
Every figure below comes from a query run against the index database on 2026-10-05 between 20:10 and 20:55 UTC. Each is reproducible from the labelled SQL in Appendix A. Nothing here is estimated, extrapolated, or rounded from a number we did not compute. Where a figure can be read more than one way, both readings are given.
## 1. The census
806 entities are published: 771 agents and 35 vendors. The field catalogue holds 63 fields — 45 applying to both entity types, 14 to agents only, 4 to vendors only. That makes 771 x 59 + 35 x 49 = **47,204 applicable data points.**
| | Count | Share | |---|---:|---:| | Present (a value with a verbatim quote and a source) | 12,951 | 27.44% | | `no_public_information` (checked, nothing published) | 31,615 | 66.98% | | Never assessed (no row at all) | 2,638 | 5.59% | | **No public evidence either way** | **34,253** | **72.56%** |
The three shares sum to 100.00% before rounding; printed to two decimals they sum to 100.01%.
Vol. 9 reported 72.23% on 47,126 applicable points. The movement is +0.33pp. Absence has now sat near 67% for ten consecutive volumes.
## 2. Risk transfer: 800 of 806
**800 of 806 published entities (99.26%) have no public evidence of insurance availability.** Vol. 9: 799 of 804, 99.38%.
Widening the test to all six risk-transfer fields — `insurance_available`, `insurance_carriers`, `coverage_limits`, `own_liability_cover`, `liability_cap`, `indemnification` — **788 of 806 (97.77%) have nothing present on any of them.**
The six rarest of the 45 shared fields are, for the fourth consecutive volume, exactly the six risk-transfer fields:
| Field | Present | `no_public_information` | |---|---:|---:| | `coverage_limits` | 1 | 800 | | `own_liability_cover` | 1 | 194 | | `insurance_carriers` | 3 | 798 | | `liability_cap` | 5 | 796 | | `indemnification` | 6 | 795 | | `insurance_available` | 6 | 797 |
The seventh-rarest is `documented_incidents` at 7.
### The three gains are composition, not learning
`insurance_carriers` moved 2 to 3, `insurance_available` 5 to 6, and `indemnification` 5 to 6 since Vol. 9. We read all three. Every one belongs to `vouch` or `munich-re-aisure`, the two insurance vendors promoted on 2026-10-04, and both values were extracted on 2026-07-30 and 2026-07-31 respectively.
**No new risk-transfer disclosure was extracted this week.** The index did not learn that anyone had started disclosing; it published two vendors it had already read in July. That distinction is the whole difference between "coverage improved" and "our own promotion queue moved", and we report it as the latter.
### What the 22 present values actually say
Twenty-two risk-transfer values stand on 18 entities. Read individually they are weaker than the count suggests:
- **Only four entities name a carrier**, and all four are insurance vendors describing their own products, not cover on an agent. `armilla-ai` is the strongest: its page states it is "a Coverholder at Lloyd's" and that "Affirmative AI Liability Insurance is underwritten by certain underwriters at Lloyd's." `coalition` cites "Allianz's A+ rated capacity." - **One of the three `insurance_carriers` values does not answer its own field.** `munich-re-aisure`'s stored value says so in terms: the capture names Mosaic Insurance as a partner but "does not itself state which carrier bears the risk on an aiSure-backed warranty." On a strict reading `insurance_carriers` is 2, not 3. This is the referential class Vol. 9 identified — a value that defers to a document we did not read — and it still passes every guard in the database. - **The single `coverage_limits` value is not agent cover.** `coalition` discloses "Up to $100K additional Funds Transfer Fraud coverage" — a cyber sub-limit conditional on buying security awareness training. - **Of 771 published agents, two distinct products have any `insurance_available` value, and both say "eligible", not "covered."** `elevenagents` and `elevenlabs-agents` are the same product indexed twice, sharing evidence row 8176, and the quote is "ElevenAgents is also the first agent platform eligible for AI insurance through AIUC, once certified." The third agent value, `rentahuman-workman-rent`, reads "Plus requester reputation and incident insurance funded from the fee."
**Tenth volume, unchanged: no published agent discloses cover in force from a named carrier.**
### One disclosure worth reading on its own
`nofireai` is a published Observability vendor — part of the assurance layer itself. Its terms of service cap total aggregate liability at **$100.00**. The vendor selling oversight limits its own exposure to one hundred dollars. This is a single data point and we draw no trend from it, but it is the sharpest illustration in the index of why the assurance layer needs measuring rather than assuming.
Direction is recorded where it runs backwards, too: `chatgpt-for-tinder-and-bumble`'s indemnity runs from the user to the vendor — "By using the extension, you agree to indemnify its creator from any liability" — and the stored value says so explicitly.
## 3. Per-dimension: where the gaps are
Scored (entity, dimension) pairs on published entities, read through `scorecards_resolved` — the view that applies the site's own authority rule, not a naive recency sort:
| Dimension | Pairs scored | At 0–1 | Counted as trust gaps | At 0–1 but uncounted | |---|---:|---:|---:|---:| | `audit_trail` | 479 | 448 | 121 | 327 | | `compliance` | 468 | 353 | 134 | 219 | | `runtime_governance` | 522 | 331 | 96 | 235 | | `outcome_settlement` | 249 | 225 | 124 | 101 | | `insurance_indemnity` | 134 | 132 | 131 | 1 |
`audit_trail` is the worst-hit dimension in the index: 448 of 479 scored pairs (93.53%) sit at 0 or 1.
**1,489 of 1,852 published scores (80.40%) sit in the 0–1 trust-gap band, and 606 of them (40.70%) are counted as trust gaps.** The rubric counts only human-judged scores, so **883 scores (59.30% of the gap band) are counted nowhere.** Vol. 8 put this at 61% index-wide and Vol. 9 at 59.63%; published only, it is 59.30%. Fourth volume reporting it; nothing has changed.
## 4. The flagship: can this index audit its own scores?
A trust-gap score is a claim about an entity, and it rests on specific fields. `scorecards.supporting_field_ids` exists to say which. The question this volume asks is whether those fields still say what the score was built on.
The question is not hypothetical. On 2026-10-04 the Fact-Checker found by hand that `audit_trail` had been marked `yes` on seven published entities against quotes that were genuinely verbatim but not about an audit trail — an autosave, a live chat view, a monitoring alert, a call transcript. Fifteen fields were downgraded, and ten machine scores were found standing on fields withdrawn out from under them. We set out to measure that class index-wide.
### 4.1 Two instruments, and what each cannot see
The index has two ways to ask whether a score is still supported, and neither asks the question directly.
`score_freshness` maps a dimension to its fields through `rubric_dimension_fields`, then compares each field's evidence `retrieved_at` and its own `created_at` against the score's `scored_at`. It reports **70** published judged scores as `unrevisited`.
That comparison cannot see a field **downgraded in place**. When the extractor re-reads a page and changes a value from `present` to `no_public_information`, it upserts on `(entity_id, field_key)`: the row's `updated_at` moves, its `created_at` does not, and the evidence row it points at is unchanged. This is precisely what happened on 10-04. Measured against mapped-field `updated_at` instead:
| Published scores | Judged | Machine | |---|---:|---:| | Scores with a mapped field | 677 | 1,273 | | A mapped field changed after the score | 154 | 431 | | `score_freshness` sees the movement | 70 | 10 | | **In-place change it cannot see** | **84** | **421** |
**505 published scores rest on a field that changed after they were written in a way their own freshness instrument is blind to.**
### 4.2 The pointer gap
The sharper check is the score's own `supporting_field_ids`. It is unavailable on most of the scores that matter.
Of 1,852 published scores in `scorecards_resolved`, 689 are human-judged and 1,163 are machine-derived. **563 of the 689 judged scores (81.71%) carry no `supporting_field_ids` at all.** Every one of the 1,163 machine scores carries one.
This inverts exactly the wrong way. The site's authority rule — in `compareScoreAuthority`, applied since migration 0038 — holds that a judged score beats a machine score however old. The index therefore privileges, and publishes, precisely the rows that cannot be checked against their own evidence. **18.29% of our published judgements are auditable by their own pointer.**
This is not a legacy artifact. Judged scores written in the last six days:
| Date | Judged scores written | Without a support pointer | |---|---:|---:| | 2026-09-30 | 15 | 15 | | 2026-10-01 | 28 | 28 | | 2026-10-02 | 23 | 23 | | 2026-10-03 | 35 | 35 | | 2026-10-04 | 11 | 11 | | 2026-10-05 | 3 | 0 |
**112 of the 115 judged scores written between 09-30 and 10-04 (97.39%) carry no pointer.** Today's three do. The practice varies run to run, which makes it a convention nothing enforces rather than a limitation of the schema.
### 4.3 What the audit found
On the 1,289 published scores that do carry a usable pointer, across 4,381 support pointers in total:
- **Zero dangling pointers.** Every named supporting field still resolves to a live `entity_fields` row. The referential layer is clean — as it should be, since a downgrade rewrites a row rather than deleting it. - **26 scores (2.02%) rest entirely on support that is now `no_public_information`.** All 26 are scored 0, and all 26 are judged. A zero resting on a withdrawn absence still reads zero: score and evidence agree, and no gap is hidden. - **395 scores (30.64%) have at least one supporting field updated after the score was written.** - **Zero scores of 2 or above rest entirely on withdrawn support.**
That last line is the finding we expected to contradict and did not. To test it beyond the 18.29% that carry pointers, we re-ran it over **all** judged published scores through the `rubric_dimension_fields` mapping, which needs no pointer. The result across the whole published index:
**Exactly one erased trust gap in 806 entities.**
### 4.4 The one case, named
`gptme`, an open-source coding agent, carries scorecard 2197: `runtime_governance` scored **2** — not a trust gap — judged by the Fact-Checker on 2026-08-02. Its rationale cites "an explicit --tools allow-list" and actions requiring "confirmation by default."
Both fields the `runtime_governance` dimension reads are now `no_public_information`:
| Field | Status | Changed | |---|---|---| | `runtime_governance` | `no_public_information` | 2026-08-11 | | `permission_scopes` | `no_public_information` | 2026-08-11 |
The evidence was withdrawn nine days after the score was written, by the extractor re-reading the same page. The score has stood for 55 days. There is no superseding row — `superseded_count` is 0 — so this is not an authority conflict; it is simply a judgement nothing revisited. A reader looking at `gptme` today sees a governance score of 2 on a dimension where this index records no public information at all.
One case in 806 entities is a small defect. We report it because it is the only one, and because the mechanism that produced it is not rare: 84 judged scores sit on in-place field changes their freshness instrument cannot see. This one crossed the 0–1 boundary. The other 83 did not.
### 4.5 Where this cuts against our own thesis
We expected to find that the index was quietly erasing trust gaps, because the failure mode is real, the Fact-Checker found ten instances of it by hand six days ago, and nothing automates the re-score. It is not happening at scale. **In the direction that would embarrass us — a published score saying "not a gap" on evidence that no longer exists — the count is one.**
Two things explain that, and only one is to our credit. The honest one: the machine scorer is conservative. 1,013 of 1,852 published scores are a 1 and only 363 are a 2 or better, so there is very little height for a withdrawal to fall from. The creditable one: the 10-04 hand audit caught its ten cases before this report ran, and they are not in these numbers because they were already fixed.
Neither is a mechanism. Both would fail silently next week.
## 5. The ceiling: one 3 in the entire index
| Score | Published scores | Judged | Machine | |---|---:|---:|---:| | 0 | 476 | 476 | 0 | | 1 | 1,013 | 130 | 883 | | 2 | 362 | 82 | 280 | | 3 | **1** | 1 | 0 |
**One published score of 3 exists across 806 entities and five dimensions.** It is `proofchain` on `audit_trail`, judged 2026-08-24, on a capture stating every agent gets "a cryptographically verifiable identity and a tamper-evident audit trail."
A rubric whose top band is reached once in 1,852 opportunities is either measuring a market with almost no top-band behaviour, or is mis-calibrated. We cannot distinguish those from inside the index, and we decline to assert either. What we can say is that the distribution is not the shape of a scorecard that discriminates: 80.40% of scores fall in two of four bands.
Related, and the reason the counted gap figure is so much smaller than the gap-band figure: **639 of 806 published entities (79.28%) carry no human-judged dimension at all.** Their scores are entirely machine-derived, and the rubric counts no machine score as a trust gap. The published trust gap — 606 across 167 entities — describes 20.72% of the index.
## 6. Citation freshness: fifth volume at exactly the speed of the clock
12,957 cited values on published entities, each diffed against the newest capture of its own source URL:
| | Vol. 9 | Vol. 10 | |---|---:|---:| | Never re-fetched | 35.26% | **37.69%** | | Stale | 2,188 | **1,842** | | Stale, high confidence | — | 1,660 | | Stale rate on judgeable recaptures | 26.11% | **22.81%** | | Entities with >= 1 stale citation | 299 (37.19%) | **275 (34.12%)** | | Median cited capture age | 57.7d | **64.7d** | | Median newest available capture | 18.5d | **25.6d** |
The stale count improved by 346 and the stale rate by 3.30pp. Some of that is real work: the Researcher re-pointed 194 fields on six entities from freshly changed captures on 10-04, which is the propagation step four volumes have asked for, done once by hand.
**Median cited age moved 57.7 to 64.7 days — 7.0 days over a 7.0-day week, the fifth consecutive volume at exactly the speed of the clock.** Nothing systematic moves a published field onto a fresher capture.
And this week the comparison changed shape. In Vols. 6–9 the story was that fresher captures existed (median 18.5d) while citations sat at 57.7d. The newest-available median has now aged 7.1 days too. The fresher captures have stopped arriving for these URLs, so the gap is no longer closable by re-pointing alone. 1,842 failures fall on 281 URLs.
## 7. The gate, both ends
**416 of 806 published entities (51.61%) would not pass our own promotion gate.** Vol. 7 48.23%, Vol. 8 50.88%, Vol. 9 53.23%. **This is the first improvement in four volumes**, and it is worth saying plainly after three volumes of reporting the opposite.
Published blockers: stale citations 263, fewer than 12 present fields 183, below the 8-field publication floor 13, no scored dimension 2, name absent from own captures 6, name too generic 6, third-party disclaimer 7. Mean 16.08 present fields, 1.21 sources. 660 of 806 (81.89%) rest on a single source URL, advisory rather than blocking.
At the other end: **0 of 2,412 candidates are promotion-ready**, for the fifth consecutive volume. 107 are evidence-ready, 220 clear the publication floor, 1,057 carry `identity_verdict = 'unchecked'`, and 2,355 of 2,412 (97.64%) rest on a single source. Mean present fields among candidates: 3.92.
Nothing is ready to come in. The 14 published entities below the floor are unchanged in character from the Chief of Staff's count (13 under eight fields, 2 with no scored dimension, 1 breaching both).
## 8. Against our own thesis, seventh volume
If the assurance layer were the emptiest part of the AI agent record, assurance would be the emptiest field category. It is not.
| Category | Slots | Present | Present % | Never assessed % | |---|---:|---:|---:|---:| | `models` | 4,626 | 807 | **17.44%** | 0.56% | | `assurance` | 16,260 | 2,950 | 18.14% | 11.67% | | `safety` | 6,448 | 1,275 | 19.77% | 0.51% | | `ecosystem` | 4,766 | 1,467 | 30.78% | 0.48% | | `agency` | 4,626 | 1,728 | 37.35% | 0.32% | | `impact` | 4,836 | 1,820 | 37.63% | 13.05% | | `practicality` | 5,642 | 2,904 | 51.47% | 0.23% |
`models` is emptier than `assurance`, as it was in Vol. 9. The honest caveat runs the other way too: assurance holds 1,897 of the 2,638 never-assessed slots (71.91%), so its 18.14% is measured over our least complete ground and is a ceiling, not a reading.
## 9. The wedge we own is our thinnest coverage
Across 3,550 combined competitor listings — aiagentsdirectory 3,155, theaiagentindex 395 — **zero name any of the eleven insurance and liability providers we track.** The probe included the generic words "vouch" and "coalition" this time and still returned zero. That claim has held every sweep since tracking began.
But the layer we own uncontested is the one we cover worst. Published vendors by subtype:
| Subtype | Vendors | Mean present fields | Mean scored dimensions | Promotion-ready | |---|---:|---:|---:|---:| | Observability | 14 | 17.0 | 2.9 | 8 | | **insurance** | **4** | **11.8** | **5.0** | **2** | | audit | 2 | 11.5 | 5.0 | 1 | | Legal | 2 | 20.5 | 3.5 | 1 |
The four insurance vendors — Coalition, Armilla AI, Munich Re aiSure, Vouch — average 11.8 present fields against 16.08 across the published index. They are fully scored on all five dimensions, and only two of four would pass our own gate. We publish more and better-evidenced coverage of the observability tier, which both competitors already list, than of the insurance tier, which neither does.
A note on hygiene while we are counting vendors: `Observability` (14) and `observability` (1) are two subtype strings for one category.
## 10. Instrument state, checked before any crawl-derived figure
Per the standing rule, `job_heartbeats` was read before anything above that depends on a crawl. **Five of nine registered jobs return `is_stalled = null`** — `crawl_sources`, `extract_pending`, both competitor syncs, and now `demote_below_floor`, all `never_succeeded`. Every one of them runs on the local Spark box, which is running code that predates the `record_run` instrumentation. The view is answering honestly; `false` would be a lie. The four in-database jobs are healthy and ran within the last 13 minutes.
So: the heartbeat view cannot currently tell us whether the crawler died, and every crawl-derived figure in this report is caveated accordingly. Read directly instead: newest capture 2026-10-05 12:06 UTC (alive today), newest evidence row 2026-10-01 12:06 UTC (four days stale), 19,305 raw documents, 9,607 of them with no evidence row.
Field creation since Vol. 9, which reported zero rows created in five days: 7 on 09-22, 1 on 09-23, **74 on 09-29, 7 on 09-30**, nothing on 10-01 through 10-04, 3 today. **84 new field rows in seven days.** The five-day freeze broke on 09-29 and a four-day one followed it.
The queue's shape has not moved: 8,880 actionable documents, 65.9% recaptures, median owning entity with 4 open slots of 59, **2,821 documents (31.77%) belonging to entities with no headroom at all**, and **166 documents (1.87%) on the 162 zero-field entities where creation is guaranteed**. Vol. 9 measured 8,955 / 65.9% / 31.74% / 1.93%. The one-line priority sort Vol. 9 asked for has not shipped: `pipeline/extract_pending.py:164` still reads `actionable = actionable[:limit]` with no ordering.
Standing diagnostics: `entity_identity_drift` 5 rows; `polarity_inverted_values` 16, with 7 scores resting on them; `score_authority_conflicts` 58 (40 published); `false_absence_flags` 2,670 unreviewed; `change_queue` 3,452 pending; `superseded_quote_flags` 3,129, **zero adjudicated for the fifth volume** — and down 345 from Vol. 9's 3,474 not because anyone worked them but because the detector refilled from changed captures.
## 11. What we are not reporting, and why
- **Autonomy versus accountability: ninth deferral, re-measured not assumed.** `autonomy_level` is present on 88 of 771 published agents and holds **79 distinct free-text values across those 88 rows.** There is no scale to correlate an audit-trail score against. Any coefficient would describe our extraction, not the market. - **Model concentration risk:** same blocker, same reason. - **The assurance landscape map:** 35 published vendors, 4 of them insurance. Enough to list, not enough to map without padding. Unchanged from Vols. 5–9. - **Extraction throughput:** `evidence` still has no creation timestamp and `retrieved_at` is a deliberate copy of `fetched_at`. Not computable; asserted nowhere. - **Citation integrity versus MIT's:** still never measured. Cloud roles cannot reach `aiagentindex.mit.edu`. We should stop making the comparison rhetorically until the Spark measures it. - **One citation withheld.** `gptme`'s `audit_trail` value (evidence 7393) is `stale` at medium confidence, so it is not cited here even though it is verbatim in its own capture. The withholding is disclosed rather than silently dropped — and it bears on section 4.4, where that same capture is the one `gptme`'s whole page rests on.
## Appendix A — reproduction
Run against Supabase project `invsbcblyrjsvmygulcz` on 2026-10-05.
**A1. Published census by type** ```sql select entity_type, count(*) from entities where status='published' group by entity_type; ```
**A2. Applicable data points and the absence partition (§1)** ```sql with pub as (select id, entity_type from entities where status='published'), slots as (select p.id, p.entity_type, f.field_key from pub p join field_catalog f on f.applies_to='both' or f.applies_to=p.entity_type), joined as (select s.id, s.field_key, ef.value_status from slots s left join entity_fields ef on ef.entity_id=s.id and ef.field_key=s.field_key) select count(*) as applicable_slots, count(*) filter (where value_status='present') as present, count(*) filter (where value_status='no_public_information') as npi, count(*) filter (where value_status is null) as never_assessed from joined; ```
**A3. No public evidence of insurance availability (§2)** ```sql with pub as (select id from entities where status='published') select count(*) as published, count(*) filter (where exists (select 1 from entity_fields ef where ef.entity_id=pub.id and ef.field_key='insurance_available' and ef.value_status='present')) as insurance_available_present, count(*) filter (where exists (select 1 from entity_fields ef where ef.entity_id=pub.id and ef.field_key in ('insurance_available', 'insurance_carriers','coverage_limits','own_liability_cover', 'liability_cap','indemnification') and ef.value_status='present')) as any_risk_transfer from pub; ```
**A4. Field rarity among published entities (§2)** ```sql select f.field_key, f.category, count(*) filter (where ef.value_status='present') as present, count(*) filter (where ef.value_status='no_public_information') as npi from field_catalog f left join entity_fields ef on ef.field_key=f.field_key and ef.entity_id in (select id from entities where status='published') where f.applies_to='both' group by f.field_key, f.category order by present asc limit 14; ```
**A5. Every present risk-transfer value, for hand-reading (§2)** ```sql select e.slug, e.entity_type, e.subtype, ef.field_key, ef.value_text, ef.created_at::date, ef.updated_at::date, ef.quote from entity_fields ef join entities e on e.id=ef.entity_id where e.status='published' and ef.value_status='present' and ef.field_key in ('insurance_available','insurance_carriers','coverage_limits', 'own_liability_cover','liability_cap','indemnification') order by ef.field_key, e.slug; ```
**A6. Per-dimension gap table (§3)** ```sql select dimension, count(*) as scored_pairs, count(*) filter (where is_judged) as judged, count(*) filter (where score <= 1) as at_0_or_1, count(*) filter (where is_trust_gap) as counted_trust_gaps, count(*) filter (where score <= 1 and not is_judged) as uncounted_machine from scorecards_resolved where entity_status='published' group by dimension order by counted_trust_gaps desc; ```
**A7. `score_freshness` verdicts (§4.1)** ```sql select verdict, count(*) as rows, count(*) filter (where is_judged) as judged, round(avg(fields_moved),2) as mean_fields_moved from score_freshness where entity_status='published' group by verdict; ```
**A8. In-place field changes `score_freshness` cannot see (§4.1)** ```sql with mapped as ( select s.id as scorecard_id, s.scored_at, s.scored_by <> 'spark-verifier' as is_judged, f.updated_at as f_updated, f.created_at as f_created, ev.retrieved_at from scorecards s join entities e on e.id = s.entity_id and e.status='published' join rubric_dimension_fields m on m.dimension = s.dimension join entity_fields f on f.entity_id = s.entity_id and f.field_key = m.field_key join evidence ev on ev.id = f.evidence_id ), agg as ( select scorecard_id, is_judged, count(*) filter (where f_updated > scored_at) as updated_after_score, count(*) filter (where f_updated > scored_at and retrieved_at::date <= scored_at::date and f_created::date <= scored_at::date) as inplace_only, count(*) filter (where retrieved_at::date > scored_at::date or f_created::date > scored_at::date) as sf_visible from mapped group by 1,2 ) select is_judged, count(*) as scores, count(*) filter (where updated_after_score > 0) as field_changed_after_score, count(*) filter (where sf_visible > 0) as score_freshness_sees_it, count(*) filter (where inplace_only > 0 and sf_visible = 0) as invisible_inplace from agg group by is_judged; ```
**A9. The pointer gap (§4.2)** ```sql select count(*) as published_resolved_scores, count(*) filter (where is_judged) as judged, count(*) filter (where not is_judged) as machine, count(*) filter (where is_judged and (supporting_field_ids is null or cardinality(supporting_field_ids)=0)) as judged_no_pointer, count(*) filter (where not is_judged and (supporting_field_ids is null or cardinality(supporting_field_ids)=0)) as machine_no_pointer from scorecards_resolved where entity_status='published'; ```
**A10. Is the pointer gap current practice? (§4.2)** ```sql select scored_at::date as scored_day, count(*) as judged_rows_written, count(*) filter (where supporting_field_ids is null or cardinality(supporting_field_ids)=0) as without_pointer from scorecards where scored_by <> 'spark-verifier' and scored_at >= '2026-09-20' group by 1 order by 1; ```
**A11. Dangling support pointers (§4.3)** ```sql with sup as (select s.id as scorecard_id, unnest(s.supporting_field_ids) as fid from scorecards s where s.supporting_field_ids is not null and cardinality(s.supporting_field_ids) > 0) select count(*) as support_pointers, count(*) filter (where ef.id is null) as dangling, count(distinct sup.scorecard_id) as scorecards_with_support from sup left join entity_fields ef on ef.id = sup.fid; ```
**A12. Withdrawn support, via the score's own pointer (§4.3)** ```sql with r as (select * from scorecards_resolved where entity_status='published' and supporting_field_ids is not null and cardinality(supporting_field_ids)>0), sup as (select r.scorecard_id, r.score, r.is_judged, r.scored_at, ef.value_status, ef.updated_at from r join entity_fields ef on ef.id = any(r.supporting_field_ids)), agg as (select scorecard_id, score, is_judged, count(*) filter (where value_status='present') as n_present, count(*) filter (where value_status='no_public_information') as n_withdrawn, count(*) filter (where updated_at > scored_at) as n_changed_after from sup group by 1,2,3) select count(*) as checkable_scores, count(*) filter (where n_withdrawn > 0) as any_withdrawn, count(*) filter (where n_present = 0) as all_withdrawn, count(*) filter (where n_present = 0 and score >= 2) as erased_gaps, count(*) filter (where n_changed_after > 0) as changed_after_score from agg; ```
**A13. The erased-gap test over ALL judged scores, no pointer required (§4.3–4.4)** ```sql with mapped as ( select s.id as scorecard_id, e.slug, s.dimension, s.score, s.scored_at, s.scored_by <> 'spark-verifier' as is_judged, f.value_status, f.updated_at from scorecards s join entities e on e.id=s.entity_id and e.status='published' join rubric_dimension_fields m on m.dimension=s.dimension join entity_fields f on f.entity_id=s.entity_id and f.field_key=m.field_key ), agg as ( select scorecard_id, slug, dimension, score, is_judged, count(*) filter (where value_status='present') as n_present, count(*) filter (where value_status='no_public_information' and updated_at > scored_at) as withdrawn_after_score from mapped group by 1,2,3,4,5 ) select is_judged, score, count(*) as scores, count(*) filter (where n_present=0) as all_mapped_absent, count(*) filter (where n_present=0 and withdrawn_after_score>0) as erased_gaps from agg where score >= 2 group by is_judged, score order by is_judged, score; ```
**A14. The one case (§4.4)** ```sql select scorecard_id, slug, dimension, score, is_judged, is_trust_gap, scored_by, scored_at, superseded_count from scorecards_resolved where slug='gptme';
select ef.field_key, ef.value_status, ef.updated_at::date, ef.extracted_by from entity_fields ef join entities e on e.id=ef.entity_id where e.slug='gptme' and ef.field_key in ('permission_scopes','runtime_governance'); ```
**A15. Score distribution and the single 3 (§5)** ```sql select score, count(*) as all_published, count(*) filter (where is_judged) as judged, count(*) filter (where not is_judged) as machine from scorecards_resolved where entity_status='published' group by score order by score;
select slug, dimension, score, scored_by, scored_at::date, rationale from scorecards_resolved where entity_status='published' and score = 3; ```
**A16. Scoring coverage per published entity (§5)** ```sql with per as (select e.id, count(distinct r.dimension) as dims_scored, count(distinct r.dimension) filter (where r.is_judged) as dims_judged from entities e left join scorecards_resolved r on r.entity_id = e.id where e.status='published' group by 1) select count(*) as published, count(*) filter (where dims_scored = 0) as zero_dims, count(*) filter (where dims_scored = 5) as full_five, count(*) filter (where dims_judged = 0) as no_judged_dim, round(avg(dims_scored),2) as mean_dims from per; ```
**A17. Citation freshness (§6)** ```sql select freshness, confidence, count(*) as values, count(distinct entity_id) as entities from citation_freshness where entity_status='published' group by freshness, confidence;
select count(*) as cited_values, count(*) filter (where newest_raw_document_id = cited_raw_document_id) as never_refetched, count(*) filter (where freshness='stale') as stale, count(*) filter (where freshness='stale' and confidence='high') as stale_high, count(distinct source_url) filter (where freshness='stale') as stale_urls, count(distinct entity_id) filter (where freshness='stale') as entities_any_stale, round((percentile_cont(0.5) within group (order by extract(epoch from (now()-cited_at))/86400.0))::numeric,1) as median_cited_age_days, round((percentile_cont(0.5) within group (order by extract(epoch from (now()-newest_fetched_at))/86400.0))::numeric,1) as median_newest_age_days from citation_freshness where entity_status='published'; ```
**A18. The gate, both ends (§7)** ```sql select entity_status, count(*) as rows, count(*) filter (where promotion_ready) as promotion_ready, count(*) filter (where evidence_ready) as evidence_ready, count(*) filter (where clears_publication_floor) as clears_floor, count(*) filter (where identity_verdict='unchecked') as identity_unchecked, count(*) filter (where stale_citations_high_confidence > 0) as has_stale_high, count(*) filter (where distinct_sources <= 1) as single_source, round(avg(present_fields),2) as mean_present_fields, round(avg(distinct_sources),2) as mean_sources from promotion_readiness group by entity_status;
select b as blocker, count(*) from promotion_readiness, unnest(blockers) as b where entity_status='published' group by b order by 2 desc; ```
**A19. Category completeness (§8)** ```sql with pub as (select id, entity_type from entities where status='published'), slots as (select p.id, f.field_key, f.category from pub p join field_catalog f on f.applies_to='both' or f.applies_to=p.entity_type), j as (select s.category, ef.value_status from slots s left join entity_fields ef on ef.entity_id=s.id and ef.field_key=s.field_key) select category, count(*) as slots, count(*) filter (where value_status='present') as present, round(100.0*count(*) filter (where value_status='present')/count(*),2) as present_pct, count(*) filter (where value_status is null) as never_assessed from j group by category order by present_pct asc; ```
**A20. Competitors and the wedge check (§9)** ```sql select competitor, count(*) filter (where still_listed) as listed, count(*) filter (where is_featured and still_listed) as featured, max(last_seen_at)::date as last_swept from competitor_listings group by competitor;
select count(*) as wedge_hits from competitor_listings where name ilike any (array['%armilla%','%munich%','%aisure%','%testudo%','%relm%', '%embroker%','%mosaic%','%lloyd%','%chaucer%','%vouch%','%coalition%']); ```
**A21. Published vendors by subtype (§9)** ```sql select e.subtype, count(*) as published_vendors, round(avg(pr.present_fields),1) as mean_present_fields, round(avg(pr.scored_dimensions),1) as mean_scored_dims, count(*) filter (where pr.promotion_ready) as promotion_ready from entities e join promotion_readiness pr on pr.entity_id=e.id where e.status='published' and e.entity_type='vendor' group by e.subtype order by 2 desc; ```
**A22. Instrument state, run before any crawl-derived figure (§10)** ```sql select job_name, age, is_stalled, never_succeeded, unregistered, last_success_at, max_age from job_heartbeats order by is_stalled nulls first;
select (select max(fetched_at) from raw_documents) as newest_capture, (select max(retrieved_at) from evidence) as newest_evidence, (select count(*) from raw_documents) as raw_documents;
select * from extraction_backlog; ```
**A23. Field creation by day (§10)** ```sql select date_trunc('day', created_at)::date as day, count(*) as fields_created from entity_fields where created_at >= '2026-09-22' group by 1 order by 1; ```
**A24. Queue headroom (§10)** ```sql with evidenced_doc as materialized (select distinct raw_document_id from evidence), echo as materialized (select raw_document_id from document_echoes), pending as (select d.id, coalesce(s.entity_id, eh.id) as owner_entity, (d.id in (select raw_document_id from echo)) as is_echo from raw_documents d left join sources s on s.id = d.source_id left join entities eh on eh.homepage_url = d.url where d.id not in (select raw_document_id from evidenced_doc)), actionable as (select id, owner_entity from pending where not is_echo and owner_entity is not null), slots as (select e.id as entity_id, count(f.field_key) as applicable from entities e join field_catalog f on f.applies_to='both' or f.applies_to=e.entity_type group by e.id), filled as (select entity_id, count(*) as n from entity_fields group by entity_id), headroom as (select s.entity_id, s.applicable - coalesce(fl.n,0) as open_slots, coalesce(fl.n,0) as filled from slots s left join filled fl on fl.entity_id=s.entity_id) select count(*) as actionable_documents, count(*) filter (where h.open_slots <= 0) as docs_with_no_headroom, count(*) filter (where h.filled = 0) as docs_on_zero_field_entities, count(distinct a.owner_entity) filter (where h.filled = 0) as zero_field_entities, round((percentile_cont(0.5) within group (order by h.open_slots))::numeric,0) as median_open_slots from actionable a join headroom h on h.entity_id = a.owner_entity; ```
**A25. Autonomy scale re-measure (§11)** ```sql select ef.value_status, count(*) as rows, count(distinct ef.value_text) as distinct_values from entity_fields ef join entities e on e.id=ef.entity_id where e.status='published' and ef.field_key='autonomy_level' group by 1; ```
**A26. Verbatim test on every quote cited in this report** ```sql select p.slug, p.field_key, p.evidence_id, ev.source_url, position(regexp_replace(p.quote,'\s+',' ','g') in regexp_replace(rd.content,'\s+',' ','g')) > 0 as verbatim_ok from (/* the cited (entity, field) pairs */) p join evidence ev on ev.id=p.evidence_id join raw_documents rd on rd.id=ev.raw_document_id; ```
## Appendix B — reconciliation
Checks run so a reader can catch an arithmetic error without re-deriving the report.
1. Applicable slots from first principles: 771 x 59 + 35 x 49 = 45,489 + 1,715 = **47,204**, equal to the queried count. ✓ 2. Census partition: 12,951 + 31,615 + 2,638 = **47,204**. ✓ Shares sum to 100.00% unrounded; to 100.01% as printed. Stated, not smoothed. 3. Score partition by authority: 606 counted gaps + 883 uncounted machine rows at 0–1 = **1,489**, equal to the at-0-or-1 count reached independently. ✓ 4. Score distribution: 476 + 1,013 + 362 + 1 = **1,852** = published resolved scores. ✓ 5. Judged/machine split: 689 + 1,163 = **1,852**. ✓ 6. Pointer split: 563 judged without + 126 judged with = **689** judged; 1,163 machine all with = 1,289 checkable, matching A12's `checkable_scores`. ✓ 7. Per-dimension gaps: 121 + 134 + 96 + 124 + 131 = **606** = counted trust gaps. ✓ Pairs: 479 + 468 + 522 + 249 + 134 = **1,852**. ✓ 8. Judged entity coverage: 806 published − 639 with no judged dimension = 167 entities carrying judged gaps, matching the 167 reached independently. ✓ 9. `score_freshness` partition: 619 current + 70 unrevisited = **689** judged published; 1,273 machine_rederived. ✓ (1,273 vs 1,163 resolved machine rows is expected: `score_freshness` reads `scorecards`, not the resolved view, so it includes superseded rows.) 10. A8 judged arithmetic: 70 visible + 84 invisible = **154** with a mapped field changed after the score. ✓ Machine: 10 + 421 = 431. ✓ 11. Floor breaches: 13 under eight fields + 2 with no scored dimension − 1 breaching both = **14** = `published_floor_breaches`. ✓ 12. Citation count vs census: 12,957 cited values against 12,951 present in the census. The 6-row difference is the known set of values filed on fields that do not apply to their entity type, reported in Vols. 7–9. ✓ 13. Citation freshness buckets: 6,135 + 4,883 + 1,660 + 182 + 97 = **12,957**. ✓ 14. Never-refetched cross-check: 4,883 never re-fetched equals the 4,883 `current` rows, as it must — a value whose newest capture is its own cited capture can only read `current`. ✓ 15. Queue partition: 8,880 actionable = 3,025 new-URL + 5,855 recapture. ✓ Pending 9,611 = 8,880 actionable + 492 echo + 239 without an entity. ✓ 16. All thirteen quotes cited in this report pass the whitespace-normalised verbatim test against their own stored capture, and all thirteen read `current` or `reconfirmed` in `citation_freshness`. The fourteenth candidate was withheld as stale and the withholding is disclosed in §11. ✓ 17. Headline figures re-queried at the end of the run — published 806, resolved published scores 1,852, judged 689, judged without pointer 563, counted gaps 606, at 0–1 1,489, scores of 3 = 1 — unchanged from the start. ✓
Entries in this piece 12
Published index entries backed by the same source documents this piece cites.
Sources 13
“Service Provider shall indemnify and hold harmless Company and its affiliates, parents, subsidiaries, employees, officers, contractors, and agents from and against any third-party claims arising from (a) a breach of Service Provider's representations or warranties under this DPA”
richpanel.com · checked Sep 26, 2026“Active Cyber Insurance for Enterprises Comprehensive cyber coverage backed by proprietary intelligence and Allianz’s A+ rated capacity.”
www.coalitioninc.com · checked Jul 31, 2026“Up to $100K additional Funds Transfer Fraud coverage”
www.coalitioninc.com · checked Sep 5, 2026“IN NO EVENT WILL THE TOTAL LIABILITY OF THE NOFIRE AI PARTIES TO YOU FOR ALL DAMAGES, LOSSES, AND CAUSES OF ACTION (WHETHER IN CONTRACT OR TORT, INCLUDING, BUT NOT LIMITED TO, NEGLIGENCE OR OTHERWISE) ARISING FROM OR RELATED TO THE TERMS, THE CONTENT, AND/OR YOUR USE OF THE SITE, EXCEED, IN THE AGGREGATE, $100.00.”
nofire.ai · checked Aug 29, 2026“In no event shall TaleTok.io's total liability to you for all damages exceed the amount you paid to TaleTok.io over the past six months.”
taletok.io · checked Sep 19, 2026“Plus requester reputation and incident insurance funded from the fee.”
workman.rent · checked Aug 2, 2026“ElevenAgents is also the first agent platform eligible for AI insurance through AIUC, once certified.”
elevenlabs.io · checked Aug 18, 2026“Armilla Insurance Services is a Coverholder at Lloyd's. Affirmative AI Liability Insurance is underwritten by certain underwriters at Lloyd's.”
www.armilla.ai · checked Sep 11, 2026“By using the extension, you agree to indemnify its creator from any liability.”
chromewebstore.google.com · checked Aug 9, 2026“ProofChain gives every agent a cryptographically verifiable identity and a tamper-evident audit trail.”
proofchain.us · checked Aug 14, 2026“aiSure TM -backed performance warranties enable you to indemnify your clients for their financial losses or legal liabilities directly related to AI errors.”
www.munichre.com · checked Jul 31, 2026“Mosaic Insurance has partnered with Munich Re to harness the power of aiSure™”
www.munichre.com · checked Jul 31, 2026“Insurance services provided by Vouch Specialty Insurance Services, LLC.”
vouch.us · checked Sep 23, 2026