Skip to content

Evidence

Evaluation methodology and results

The figures on our homepage come from this suite. It exists to answer one question with numbers rather than adjectives: when Meritfold writes something, can it be trusted?

Live-model runs 24 August 2026 · gpt-5.6 family · 113 scored inputs across 8 areas · page last updated 31 August 2026

How it works

Each area has a golden set: inputs with a known correct answer, built from real tender documents and synthetic company packs. The suite runs those inputs through the same prompts and models the product uses, then scores the output against the known answer. Each area has a threshold it must clear, and a run fails if any area misses. Nothing is scored by hand.

Two modes exist. A deterministic mode runs on every change so a regression is caught before it ships, and a live mode runs against the real model to produce the baseline below. The results here are from the live mode, because a deterministic result only proves the plumbing.

Results, live model

Evaluation areas, thresholds and live results
AreaWhat it asksThresholdResultCases
HallucinationDoes a drafted claim actually appear in the evidence it cites?Citation accuracy ≥ 0.95, unsupported-but-answerable ≤ 0.050.978 / 0.00012
AbstentionWhen the evidence is not there, does it say so instead of guessing?Confidently wrong answers on critical cases = 00.000 across 7 critical cases12
RequirementsDoes it find every mandatory requirement in a tender pack?Mandatory recall ≥ 0.98, precision ≥ 0.801.000 / 0.95610
MatchingDoes it score the right opportunities as a fit?Precision and recall against a scored golden set1.000 / 1.00030
Red teamDoes the evaluator-style review score consistently?Score stability and correct ordering of strong versus weak answers1.000 / 1.0006
ComplianceIs each requirement marked pass, fail or missing correctly?Item accuracy0.925 (deterministic rules 1.000)10
RetrievalDoes the right evidence reach the writer at all?Recall@121.00025
WritingCan it tell a strong answer from a weak one?Strong-versus-weak ordering1.0008

All eight areas cleared their thresholds. Schema failures after retry were zero in every area, and there were no unrecovered crashes. Six areas come from the run at 15:56 UTC that set our stored baseline; abstention and requirements come from a second live run at 23:01 UTC the same day, on a newer requirements prompt. Retrieval is counted in queries, 25 of them over a 60-item evidence corpus, and every other area in cases.

The claim on the homepage

The homepage says no unsupported claim has reached an exported bid. That figure is counted on the current responses of bids that actually reached export, not on every draft ever written. Drafts do contain unsupported claims: finding them is the point. They are flagged, sent back through the revision loop, and the export gate refuses to produce a document while any remain unresolved. The measurable result is what survives to the document, and that number is zero.

Limits worth knowing

  • The golden sets are built by us from real tender documents and synthetic company packs. They are not customer data, and no customer has yet scored our output.
  • At 113 scored inputs, this is a baseline, not a guarantee. Red team is thinnest at six cases and writing at eight, and those are the areas we would extend first.
  • Compliance item accuracy is 0.925, not 1.000. The deterministic rule checks score 1.000; the gap is in the model-judged items, which is why a failed mandatory requirement blocks export in code rather than relying on that score.
  • Every run is recorded with its model ids and prompt versions, so a result cannot be quietly attributed to a different model than the one that produced it.

Related

How your documents are stored and processed is on the security page. If you are evaluating Meritfold and want the underlying run data, ask and we will share it.