Evidence
Evaluation methodology and results
The figures on our homepage come from this suite. It exists to answer one question with numbers rather than adjectives: when Meritfold writes something, can it be trusted?
Live-model runs 24 August 2026 · gpt-5.6 family · 113 scored inputs across 8 areas · page last updated 31 August 2026
How it works
Each area has a golden set: inputs with a known correct answer, built from real tender documents and synthetic company packs. The suite runs those inputs through the same prompts and models the product uses, then scores the output against the known answer. Each area has a threshold it must clear, and a run fails if any area misses. Nothing is scored by hand.
Two modes exist. A deterministic mode runs on every change so a regression is caught before it ships, and a live mode runs against the real model to produce the baseline below. The results here are from the live mode, because a deterministic result only proves the plumbing.
Results, live model
| Area | What it asks | Threshold | Result | Cases |
|---|---|---|---|---|
| Hallucination | Does a drafted claim actually appear in the evidence it cites? | Citation accuracy ≥ 0.95, unsupported-but-answerable ≤ 0.05 | 0.978 / 0.000 | 12 |
| Abstention | When the evidence is not there, does it say so instead of guessing? | Confidently wrong answers on critical cases = 0 | 0.000 across 7 critical cases | 12 |
| Requirements | Does it find every mandatory requirement in a tender pack? | Mandatory recall ≥ 0.98, precision ≥ 0.80 | 1.000 / 0.956 | 10 |
| Matching | Does it score the right opportunities as a fit? | Precision and recall against a scored golden set | 1.000 / 1.000 | 30 |
| Red team | Does the evaluator-style review score consistently? | Score stability and correct ordering of strong versus weak answers | 1.000 / 1.000 | 6 |
| Compliance | Is each requirement marked pass, fail or missing correctly? | Item accuracy | 0.925 (deterministic rules 1.000) | 10 |
| Retrieval | Does the right evidence reach the writer at all? | Recall@12 | 1.000 | 25 |
| Writing | Can it tell a strong answer from a weak one? | Strong-versus-weak ordering | 1.000 | 8 |
All eight areas cleared their thresholds. Schema failures after retry were zero in every area, and there were no unrecovered crashes. Six areas come from the run at 15:56 UTC that set our stored baseline; abstention and requirements come from a second live run at 23:01 UTC the same day, on a newer requirements prompt. Retrieval is counted in queries, 25 of them over a 60-item evidence corpus, and every other area in cases.
The claim on the homepage
The homepage says no unsupported claim has reached an exported bid. That figure is counted on the current responses of bids that actually reached export, not on every draft ever written. Drafts do contain unsupported claims: finding them is the point. They are flagged, sent back through the revision loop, and the export gate refuses to produce a document while any remain unresolved. The measurable result is what survives to the document, and that number is zero.
Limits worth knowing
- The golden sets are built by us from real tender documents and synthetic company packs. They are not customer data, and no customer has yet scored our output.
- At 113 scored inputs, this is a baseline, not a guarantee. Red team is thinnest at six cases and writing at eight, and those are the areas we would extend first.
- Compliance item accuracy is 0.925, not 1.000. The deterministic rule checks score 1.000; the gap is in the model-judged items, which is why a failed mandatory requirement blocks export in code rather than relying on that score.
- Every run is recorded with its model ids and prompt versions, so a result cannot be quietly attributed to a different model than the one that produced it.
Related
How your documents are stored and processed is on the security page. If you are evaluating Meritfold and want the underlying run data, ask and we will share it.
