Published accuracy
Every AI coaching tool claims accuracy. Most give a round number and no method. Our published record shows the exact measurement, the exact sample and the exact language of the launch gate — and it moves only when we re-measure.
Published verbatim from the E2 launch gate.
Validated within one band on N of 20 real graded responses.
The set: 20 real UK construction tender responses — at least 12 social housing, at least 5 education — evaluated by the product and independently by the product owner. Bands align within one band on at least 18 of 20 for the product to launch.
Status: Pending — the 20-response graded golden set is being assembled. This page publishes the measured result the day the gate passes, and is updated only by re-measurement.
Once real assessment summaries are entered in the product, this section publishes live calibration from the Vault — the number of tenders with actual scores entered, and how close predictions have come at criterion level. Sample sizes are always shown.
Nothing here is fabricated. If a figure isn't on this page, we don't have it yet.
After launch, any change to the Charter, the band descriptors, the expert content or the underlying model is re-run against the 20-response golden set before release. A change that moves more than two responses by more than one band is rejected pending review.
Scores are an indicative guide and coaching — not a prediction of the award.