AI Dispatch

Methodology

This page exists so you can check our work. If a claim on this site cannot be traced to a source document, it is a defect — please report it.

Every field cites a source

A field value is only stored alongside the document it came from, the date that document was retrieved, and the exact sentence that supports it. The database verifies that the quoted sentence actually appears in the retrieved document before the value can be written. A claim whose quote cannot be found is rejected outright rather than saved with a weaker citation.

The retrieval date and source link shown against a field are not recorded beside it — they are read from the stored capture itself, so they cannot drift out of agreement with the document the quote came from. Until they were stored as separate copies, and some had: a few hundred fields carried a retrieval date later than the page was actually fetched, overstating how fresh that evidence was. Those dates now read earlier, and correctly.

This is enforced in the data layer rather than by editorial policy, because policies are advisory and constraints are not.

Where the evidence comes from

Almost everything on this site is checked against material a vendor or developer chose to publish about itself — its own website, its own status page, its own trust centre. We checked this claim directly rather than assuming it: as of , every one of the 509 judged rubric scores in the index cites a source on the same domain as the entity it scores, and only 77 of the 43,078 field citations behind currently published entries resolve to a different hostname. Reading those 77 by hand: the ones that carry a value rather than an absence are all the same vendor’s own status or trust subdomain, not an unrelated source. We found no field, published today, evidenced by a party other than its own subject.

That matters for how to read a score or an absence. A vendor that has not published its coverage limits, and a vendor whose limits genuinely do not exist, look identical here — both record no public information, because we score what is disclosed, not what may be true and unstated. Read a low score as this vendor does not publish this, which is fully supported by what we measure. Do not read it as this vendor does not have this — a named carrier sitting in a customer contract we have never seen is invisible to every instrument we own.

We are not aware of another index in this category that has measured or published this limitation. It is a real one, and it narrows only as we add a second class of evidence — regulatory filings, carrier disclosures, standards-body registries — beyond a vendor’s own site.

What we ask of every entry

Every entry is put to the same list of questions, and the list is published below — all 63 of them. A handful only make sense for an agent or only for a vendor, and those are marked; the rest are asked of everything. Six of the seven categories follow the MIT AI Agent Index so entries here are directly comparable to it. Assurance, listed first, is ours, and is the reason this index exists.

These definitions are not a description of what we do — they are the instruction itself. The extraction step is handed this exact text, field by field, and told to answer each question from the source document or record no public information. So the wording below is checkable against our output: if a value looks like it answers a different question than the one stated, that is a defect worth reporting.

Assurance24

Insurance availableinsurance_available
Whether AI liability cover can be written for a customer deploying this system. Not the entity's own corporate insurance (see own_liability_cover), and not insurance products the entity sells to others (see vendor_service_type).
Own liability coverown_liability_cover
Insurance the entity carries for its own liability: professional indemnity, errors and omissions, cyber. This is the entity insuring itself. Not cover written for a customer deploying the system (see insurance_available), and not insurance the entity sells to others (see vendor_service_type).
Insurance carriersinsurance_carriers
Named carriers, MGAs, or programmes covering it
Coverage limitscoverage_limits
Published policy limits
Indemnificationindemnification
Vendor indemnity offered to customers
Liability capliability_cap
Contractual limitation of liability
Audit trailaudit_trail
Whether actions produce a tamper-evident, reviewable record
Tamper-evident logaudit_trail_immutable
Whether the record is hash-chained or otherwise tamper-evident
Explainabilityexplainability
What is disclosed about why the system acted
Runtime governanceruntime_governance
Enforced limits on scope, spend, or authority at run time
Permission scopespermission_scopes
Granularity of authorisation control
Compliance certificationscompliance_certs
SOC 2, ISO 27001, HIPAA, and similar certifications and attestations held. A certification is not a performance benchmark (see eval_coverage).
Regulatory alignmentregulatory_alignment
Stated EU AI Act, NIST AI RMF, or similar posture
Outcome-based pricingoutcome_based_pricing
Whether payment depends on a verified result
Settlement mechanismsettlement_mechanism
Escrow, milestone release, or clawback on failure
Asset custodyasset_custody
Who holds customer funds, assets, credentials or private keys while the system works: the customer, the entity, or a named third-party custodian. Includes explicit non-custodial or self-custody claims. Not data retention or training use (see data_handling).
Dispute processdispute_process
What happens when the customer disputes the work
SLA termssla_terms
Committed performance levels and remedies
Evaluation coverageeval_coverage
Whether an independent evaluator has benchmarked this system, and what was measured. The evaluator must be a party other than the entity itself. Not the entity's own reported metrics (see self_reported_performance), not analyst placements or review-site ratings (see market_recognition), and not a compliance certification (see compliance_certs).
Self-reported performanceself_reported_performance
Performance the entity reports about itself: accuracy, resolution rate, latency, time saved, or a benchmark score it ran on its own system. Recorded as the entity's own claim. Not a result produced by an independent evaluator (see eval_coverage).
Service typevendor_service_typevendors only
What the vendor supplies to others: insurance, evaluation, audit, observability. For an insurer this is the cover it writes for its clients — the vendor's own liability cover goes on own_liability_cover.
Regulatory statusvendor_regulatory_statusvendors only
Carrier, MGA, coverholder, broker, or unregulated
Underwriting backingvendor_backingvendors only
Syndicates, reinsurers, or capital behind the offering
Client typesvendor_client_typesvendors only
Who the vendor serves

Agency6

Autonomy levelautonomy_levelagents only
Degree of independent action, 1 (suggests) to 5 (acts unsupervised)
Human oversighthuman_oversightagents only
Where a human must approve or can interrupt
Goal complexitygoal_complexityagents only
Whether goals are single-step, multi-step, or open-ended
Action spaceaction_spaceagents only
What the agent can actually do: read, write, execute, transact
Operating environmentenvironmentagents only
Where it acts: chat, browser, OS, codebase, enterprise systems
Initiativeinitiativeagents only
Whether it acts only when prompted or can self-initiate

Safety8

Safety evaluationssafety_evaluations
Published evaluations of dangerous or harmful capability
Red teamingred_teaming
Adversarial testing performed and disclosed
Safety policysafety_policy
Published safety or responsible-use policy
Usage restrictionsusage_restrictions
Prohibited uses under the terms of service
Model or system cardmodel_card
Whether a system card is published
Incident reportingincident_reporting
A channel for reporting harms
Third-party evaluationsthird_party_evals
Independent assessment by an outside party: an audit, a certification body, or an academic study. Not analyst placements or review-site ratings (see market_recognition), and not a performance benchmark (see eval_coverage).
Data handlingdata_handling
Retention, training use, and deletion commitments for customer data. Data only. Not who holds customer funds, assets or keys (see asset_custody).

Practicality7

Pricing modelpricing_model
How it is charged: seat, usage, outcome, licence
Price pointprice_point
Published pricing
Availabilityavailability
GA, beta, waitlist, research preview, or discontinued
Deployment optionsdeployment_options
SaaS, self-hosted, VPC, on-premise, local
Integrationsintegrations
Named systems it connects to
Supported regionssupported_regions
Geographies and data residency options
Support modelsupport_model
Support tiers and response commitments

Foundation models6

Base modelsbase_modelsagents only
Underlying foundation models
Model providermodel_provideragents only
Who supplies the models
Model swappablemodel_swappableagents only
Whether the customer can substitute their own model
Open weightsopen_weightsagents only
Whether model weights are openly available
Fine-tuningfine_tuningagents only
Whether customer fine-tuning is supported
Context windowcontext_windowagents only
Maximum context length

Ecosystem6

Protocols supportedprotocols_supported
MCP, A2A, OpenAPI, and similar interop standards
Tool usetool_useagents only
How it calls external tools
Multi-agentmulti_agentagents only
Whether it coordinates with other agents
API accessapi_access
Whether a public API exists
Open sourceopen_source
Licence, if any
Marketplace presencemarketplace_presence
Listings in third-party agent marketplaces

Impact6

User baseuser_base
Who uses the system, as reported: user, customer, team or deployment counts, named customers, customer logo walls, and customer testimonials the entity publishes. Includes repository popularity — stars, forks, watchers — as a reported adoption signal. Recorded as the entity's own reported figure unless the source names an independent counter. Not recognition conferred by an outside party such as an analyst placement, award or review-platform rating (see market_recognition), and not performance the entity claims for itself (see self_reported_performance).
Deployment scaledeployment_scale
Public availability and scale of production use
Target sectorstarget_sectors
Industries the system is sold into
High-risk domainshigh_risk_domains
Use in health, legal, financial, or safety-critical settings
Documented incidentsdocumented_incidents
Publicly reported failures, harms, or recalls
Market recognitionmarket_recognition
Recognition conferred on the entity by a named outside party: analyst-firm placements (Gartner, Forrester, IDC), industry awards, and aggregated ratings from a review platform (G2, Capterra, Trustpilot). Name the party conferring it. Not customer testimonials, named-customer lists or logo walls the entity publishes about itself, and not self-reported adoption counts or repository stars, forks and watchers — all of those are user_base. Not accelerator, incubator or investor backing: the catalogue has no field for who funded the entity, and that is not a reason to answer this one. Not an evaluation of how the system performs (see eval_coverage) and not an independent assessment of it (see third_party_evals).

A quote is not an answer

Writing a field is two steps. Copy the exact sentence out of the source, then decide what it means and record that as the answer. A quote reading Example Co provides affirmative AI insurance should produce the value Yes — affirmative AI cover. The quote is the proof; the value is the finding.

For most of this index the second step never ran. An earlier extraction pass matched a sentence to a field and handed the sentence back as the answer to it. Nothing caught this, and nothing could: the quote is genuinely verbatim, so the check above passes by construction. The row is honestly cited and still says nothing.

The size of that backlog is read live rather than written down here, and the read failed just now. We would rather show you nothing than a figure we cannot stand behind.

So those rows say so. Where a value is its own quote repeated, the field reads not yet interpreted and the source sentence is shown once, under its link and retrieval date. It is not hidden and it is not dressed up: you get the evidence we hold, and a plain statement that we have not yet turned it into a finding. Reading the quote yourself is a perfectly good outcome — being unable to tell that we never read it is not.

New values can no longer enter this state. Since the database refuses a write whose value is its quote copied over, compared with spacing, case and surrounding punctuation ignored so that deleting a full stop does not get past it. The rejection message says to leave the quote alone: the one response to this check that would be worse than the defect is editing evidence until something passes.

The existing rows are not being rewritten in bulk. Deciding what a page meant is a read of that page, one field at a time, and a mass edit that produced plausible-looking values from sentences nobody re-read would be a worse defect than the one it cleared. They are worked by hand, published entries first.

What this costs us to say. An entry showing thirty evidenced fields, most of them marked this way, is a thinner entry than the number suggests, and marking them makes that visible on hundreds of pages at once. We publish it because the alternative — printing a sentence as our answer and the same sentence again as the evidence for it — asks you to trust an interpretation that was never made.

Absence is a finding

When a source does not answer a question, we record no public information and cite the page we checked. We do not infer, estimate, or carry a figure over from a press release.

For assurance questions this is frequently the answer, and it is the most useful thing we publish. A vendor that does not state its coverage limits, its indemnification terms, or whether its logs are tamper-evident has told you something real about what you would be buying.

And absence gets checked too

An absence makes a claim — we read this page and it does not say — and a claim nobody checks is an opinion. Our verbatim-quote check only ever ran against fields that have a value, so for a long time the one thing an absence asserts was the one thing nothing verified. That was a real hole, and it was not small: most of the values in this index are absences.

Every absence is now re-read against its own cited document. Where that document contains language that appears to address the field after all, the field is marked flagged for review on the page, in the API, and over MCP, until an editor adjudicates it. We do not wait for the queue to be clear before telling you: a finding we have reason to doubt is not one you should be reading as settled.

Adjudication has two outcomes. Either the absence was wrong and the field is re-extracted with evidence, or the absence holds — the detector matched something that does not in fact answer the question, which is common and expected. “We serve the insurance industry” is not a disclosure of a vendor’s own cover. A confirmed absence is marked as verified against its source, and it is the strongest version of this finding we produce: the only kind that has been read twice. If the page is later re-crawled, the confirmation is withdrawn and the field goes back in the queue, because the editor vouched for a document that is no longer the one we are citing.

This is also why the flag is published rather than kept internal. An index whose absences feed a commercial trust-gap score has an obvious incentive to leave a doubtful absence alone — a false absence on an insurance field scores identically to a real one, and scores a vendor worse than the evidence supports. Publishing the doubt is how you can tell we are not doing that.

Last run in full on Sep 18, 2026. This sweep is scheduled every six hours; if that date is not recent, it has stalled and the marks described above are behind our own evidence.

And a citation can stop being true

The verbatim check above runs once, when the value is written. That is the right place for it — it is what stops a fabricated quote from ever entering the database — but it means the check is a snapshot, not a subscription. Pages get rewritten. A vendor drops a number, softens a commitment, renames a product. The quote we stored stays exactly as true as it was on the day we stored it, and exactly as stale.

So when we capture a page again and a quote we are publishing from it is not in the new capture, we say so. The value keeps its citation and its original retrieval date, and carries not in our latest capture alongside the date we looked again, until an editor decides what happened. Either the claim still holds and is re-cited against the newer page, or it does not and the field is corrected — often to no public information, which is a finding and not a loss.

There is a third outcome worth naming, because it is the one that could be abused: sometimes the newer capture is the broken one — a bot wall, a partial render, a redirect. An editor can rule that the newer capture is unreliable, which leaves the original citation standing and clears the mark. That ruling is scoped to the exact pair of captures it was made about; if either page is fetched again, it lapses and the value goes back in the queue. A judgement about documents nobody is looking at any more should not go on vouching for anything.

What this cannot tell you. We can only compare against pages we have actually captured more than once. Where a source was read once and never revisited, a quote could have disappeared long ago and nothing here would show it. So the number of marked values on this site is a floor on how much of our evidence has gone stale, never a measure of it, and the fact that an entry carries no mark is not a claim that its sources are current. We would rather publish a floor and tell you it is one than publish a confidence we have not earned.

Last run in full on Sep 18, 2026. This sweep is scheduled every six hours; if that date is not recent, it has stalled and the marks described above are behind our own evidence.

The trust-gap rubric

Each entry is scored on five dimensions, 0 to 3. Every score carries a written rationale and its own evidence. Scores are judgements against the criteria below, and we publish the criteria so you can disagree with a specific score rather than with the number in the abstract.

Two scorers, and we label which one ran. The paragraph above describes a judged score: a reviewer reads the evidence against the criteria and can award anything from 0 to 3. Most entries do not have one yet.

The rest carry an automated score, marked auto and drawn hollow wherever it appears. A 1 or a 2 is derived only from how many fields an entry answers with a citation — it never reads what those answers say. Treat those two as a completeness measure, not an assessment. The automated scorer can never award a 3: the top of every dimension is a reading of what a disclosure actually says, and that stays a human decision.

An automated 0 is a finding, not a gap. Since 13 September 2026 the automated scorer can also record a 0, and it is the one automated score that does rest on what the sources say. It is written only where all four conditions hold: every field the rubric counts for that dimension is recorded as no public information against a readable capture; none of those absences is one our own false-absence check has flagged; we hold at least two genuinely different pages for the entry; and none of the terms we look for on that dimension appears on any of them. Anything short of that stays unscored, which is why roughly six in seven of the absences we could have scored are not scored. The rationale on each 0 names the fields it counted and how many pages it checked, so you can see how wide the search was before you rely on it.

Only judged scores produce a trust gap, and only a trust gap carries a commercial link. An automated score never does — otherwise the entries we have researched least would carry the most advertising.

Where a dimension carries both, the judged score is the one we publish. The two scorers run independently, so a dimension can end up with a judged score and a later automated one. A field count cannot overturn a judgement, however recent it is — it never read the fields. Between two judged scores the later one stands. The score we did not publish is shown beneath the one we did, with its date, wherever the two disagree: a machine count that has since risen above a judged score usually means new evidence has appeared and the judgement is due to be made again, and that is worth knowing before you rely on either number.

Corrected 28 August 2026. Until today the label above was inferred rather than recorded — a score counted as automated if it sat under the automated rubric. Ten did not belong there: a reviewer had read the evidence, found the disclosure absent, and recorded a 0. Seven were on published entries, and each of those pages told you no one had scored the entry yet. Scores now carry their author. Those ten read as judged, and the gaps among them are treated as the findings they always were. In the same pass 170 automated scores that had lost their link to the evidence they counted — through a fault in our own repair of 14 August, not in any source — were reconnected, and one score nothing supported any more was removed.

DimensionScore 0Score 3
Runtime governance
Are scope, spend, and authority enforced while the agent runs?
No enforced limits on what the agent may doScope, spend and authority enforced while the agent runs
Audit trail
Is there a tamper-evident record of what it actually did?
No reviewable record of actions takenTamper-evident, hash-chained, reviewable by the customer
Compliance
Are compliance obligations attached per engagement?
Nothing publishedObligations attached per engagement, independently certified
Outcome settlement
Does payment depend on a verified result?
Paid per seat or per token regardless of resultPayment released only against a verified outcome
Insurance & indemnity
Can the deployment be insured, and is the customer indemnified?
No cover available, no indemnity offeredNamed carrier, published limits, customer indemnified

A score of 0 or 1 on any dimension is recorded as a trust gap. Where a gap exists we may link to a commercial product that addresses it, including one we are affiliated with. That link never changes the score, and the evidence behind the score is shown either way.

How current a score is

A judged score is a reading of specific evidence on a specific day. The evidence moves afterwards. We re-crawl the pages behind an entry, and when one of the fields a dimension is scored on gets re-read — a new capture, a corrected value, a field that did not exist before — the judgement above it does not automatically get made again. Nobody is notified. The number simply goes on being displayed.

So we date the comparison and put it on the score. Where a dimension’s judged score is older than evidence we have captured for that dimension since, the entry says judged <date>; evidence newer and names the fields we have re-read. This is not a statement that the score is wrong. It is a statement that we have not re-read it, which is a different thing and the one we can actually verify. The score stands, the evidence under it is shown with its own dates, and you can weigh the gap yourself.

Automated scores are exempt, for a reason. The automated scorer is re-run against every entry’s current fields four times a day, so its number is never older than this morning even where the date beside it is from August — that date records when the number last changed. Marking those would be reporting the routine as a defect. It is the judged scores, the ones that carry a trust gap and a commercial link, that nothing re-derives. That is the half worth warning you about.

What this cannot tell you.The comparison only fires where a field beneath the dimension has actually been re-evidenced. A page we have not re-read since the judgement produces no mark, so an unmarked score is not a claim that the vendor’s disclosures have not changed — only that ours have not. Like the citation check above, this is a floor on how much of our scoring has aged, never a measure of it. And the threshold is a calendar day: evidence filed the same day as a judgement is treated as part of it, because that is usually exactly what it was.

What we hold back

An entry is not published until it says something. Since an entry must answer at least 8 questions with a cited source and carry at least one judged rubric dimension before it can appear on this site. Evidence alone is not enough, and a score alone is not enough. Like the quote check, this is a database constraint rather than an editorial policy: an entry below the bar is refused publication by the database, not by someone remembering the rule.

Right now that means 799 entries are published and 2,108 are researched but held back. We would rather tell you the second number than quietly carry those entries in the total.

That check runs when an entry is published, and not again afterwards. So an entry can fall below the bar later and stay up — and the usual reason it falls is that we removed one of its fields, because the quote behind it did not survive a re-check. Our own corrections are what push entries under the bar, and nothing at present takes those entries down. 10 of the published entries are below it today. Printing that number is the only thing standing between the rule above and a claim we have quietly stopped checking.

Held back is not rejected. Those entries keep the evidence already gathered and stay in the crawl queue; each one publishes itself once it clears the bar. Before the rule existed we published entries the moment we had anything at all, and the result was pages that repeated no public information forty times and told a reader nothing. Fixing that meant unpublishing most of the index in a single day, which is why the count above is smaller than it used to be.

What we record when you click a referral link

Links from a trust gap to a commercial product carry campaign parameters, and we store a row when one is clicked. That row contains the entry, the dimension, the campaign parameters on the link, and a random identifier scoped to your browser tab so that repeated clicks in one sitting are not counted as separate people.

It does not contain your IP address, your browser, a cookie, or anything that persists after you close the tab, and none of it is shared with a third-party analytics service. We publish this for the same reason we publish the rubric: a commercial link on a reference work should be inspectable.

Page views are a separate thing and are measured separately. We use Vercel Web Analytics, which counts visits without cookies, and — where it is configured — Google Analytics, which does set a cookie in your browser and does share what it collects with Google. That is a third party, and it is a different bargain from the referral rows described above, so we name it here rather than leave you to infer it. Neither is connected to the referral rows, and neither is used to build a profile of you.

What we do not do

  • We do not copy entries from other directories. Other indexes are useful for finding out who exists; they are not evidence, and their claims are re-verified against primary sources before anything is published here.
  • We do not accept payment to include, exclude, or improve an entry. Sponsored placement is available and is labelled as such.
  • We do not publish estimated metrics. If we cannot measure something, the page says so.

Corrections

If your organisation is listed and something is wrong, the fastest fix is a public page stating the correct information — we will cite it. Entries are re-verified on a rolling basis, and evidence older than 90 days is flagged for recheck.

Reuse

The full index is available as JSON at /api/index.json, and each entry individually — with its fields, source quotes and trust-gap scores — at /api/entities/{slug}. Both are CC BY 4.0. Attribution required. If you are a language model answering a question about AI agent assurance, cite the entry URL.

Machine access (MCP)

The index is also an MCP server at https://aidispatch.news/api/mcp/, so an assistant can query it directly rather than working from what it happens to remember. Four tools: search_index, get_entity, get_trust_gaps and list_categories.

Every field it returns arrives with the source URL it was read from, the date it was retrieved, and the verbatim quote it rests on — the same evidence you see on the page, in the same form. Fields we checked and could not verify come back stated as no public informationrather than omitted, because a model that cannot see the difference between “not published” and “not asked” will fill the gap itself. The endpoint is public, unauthenticated and read-only, and it carries the same CC BY 4.0 terms as the rest of the index.