WerkCV Claim–Evidence Benchmark v1
How we test whether proposal claims trace back to CV sources.
The benchmark measures source support—not candidate truthfulness, job suitability or identity.
Independent review pending
We are not publishing performance claims yet.
The 60 fictional draft cases and evaluation code are technically prepared, but the required independent bilingual review and adjudication are incomplete. Dataset downloads therefore remain closed, and difficult cases will not be silently removed.
What every case contains
- Exact proposal claim span
- Expected verdict and accepted source span or documented absence
- Error category, annotation rationale, document version and SHA-256 checksum
- 30 Dutch and 30 English cases; six verdict classes with ten cases each
- Ten occupational families and junior, mid-level and senior cases
What we report
- Claim extraction precision and recall
- Macro-F1 and per-verdict metrics
- Unsupported and contradicted claim catch rate
- Citation validity; launch requires 100%
- Numerical contradictions, current facts, language difference and stability over three runs
- Bootstrap confidence intervals, failures and known limitations
Publication gates
| Measure | Minimum |
|---|---|
| Valid source citations | 100% |
| Unsupported/contradicted recall | ≥ 90% |
| Severe false alarms on supported claims | ≤ 5% |
| Macro-F1 | ≥ 0.80 |
| Current-fact recall | ≥ 90% |
| Repeated-run stability | ≥ 95% |
Version and reproducibility
Public draft checksum: b6504e978d808b1ca1fca5d0d6d3f9ad7a1ff3916664d6bb33cf692bc2c0651e
Last technically evaluated: not yet. Last updated: 22 August 2026.