2026-08-11 · 6 min read
Our regulated-advice detector scores 0.983. Its recall is 0.75.
Both numbers are real, they measure different slices of the same test set, and the one we would have put on a slide is the one that matters least.
The regulated_advice detector separates explaining a financial instrument from recommending one. It scores a mean per-language F1 of 0.995 across all 26 target languages. That is the number that would have gone on a slide, and it is close to useless on its own.
What the headline number is averaging
It is the mean of 26 per-language F1 values, computed over 622 positive examples in the test split. Twenty-five of those languages sit above 0.95. The weakest is hu at 0.957, which is Maltese, and Maltese is absent from the base model's pretraining, so that one is a property of XLM-RoBERTa rather than of our data.
Read that way it looks like a solved problem. The label breakdown says otherwise.
The same test set, sliced by label
| Label | Support | P | R | F1 |
|---|---|---|---|---|
| legal_advice | 102 | 0.804 | 0.726 | 0.763 |
| medical_advice | 520 | 0.998 | 0.821 | 0.901 |
Precision is excellent: 0.998 on medical advice, meaning when it fires it is right. Recall is 0.73 and 0.82. Between a fifth and a quarter of the advice in the test set walks straight past it.
A guardrail with high precision and mediocre recall is a specific kind of dangerous. It never annoys anyone, so nobody turns it off, and it produces a clean evidence record every time it misses.
The label that has no test data at all
One label has zero positives in the test split: financial_advice. That is the flagship case. The detector exists because a banking assistant should not tell someone which fund to buy, and the split we evaluated on contains no examples of it.
The 0.983 is not wrong. It is averaging over language, and language is not the axis on which this detector is weak. Nobody hid anything; the number was simply answering a question we were not asking.
Where it does hold up
The register breakdown is the part that survives scrutiny, and it is the part that was hardest to build. These are the positive registers:
| Register | Support | P | R | F1 |
|---|---|---|---|---|
| directive | 130 | 1.000 | 0.985 | 0.992 |
| personal_recommendation | 362 | 1.000 | 0.997 | 0.999 |
| suitability_claim | 130 | 1.000 | 0.977 | 0.988 |
And these are the hard negatives, written to sit as close to the class as possible without being in it. They contain no positives at all, so the only column that means anything is the false positive rate:
definitionat 0.0% false positivesgeneral_riskat 0.0% false positiveshistorical_factat 0.0% false positiveshypotheticalat 0.0% false positivesprocess_descriptionat 0.0% false positives
A definition of an ETF, a general statement about risk, a historical fact, a hypothetical, a description of a process. Those are the sentences a keyword filter destroys, and the detector leaves them alone almost perfectly. That is the thing worth reporting, and it is nowhere near the top of the page in the version of this we nearly published.
The corpus
11,436 examples across 26 languages, generated per language rather than translated, of which 5,718 are unlabelled negatives. Generated by claude-haiku-4-5 against prompt version advice_sys_v1, content hash 48de4bb155a8. Synthetic and in-distribution, which is what makes 26 languages affordable and is also why none of this is a claim about production traffic.
The threshold is 0.72 and it is the uncalibrated default. This detector predates the calibration sweep the others went through. Given that recall is the weak axis, a lower threshold is the obvious next experiment, and it has not been run.
What we changed
The benchmarks page leads with the caveats rather than the results, and every detector shows its positive count next to its F1 so a reader can see how much evidence is behind a number. The three things on the fix list for this detector are: get positives for the empty label into the test split, calibrate the threshold against recall rather than accept the default, and evaluate against real transcripts instead of generated ones.
Until then the per-language table is published in full, including Maltese at 0.957.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Bulgarianbg | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Croatianhr | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Dutchnl | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 24 | 1.00 | 0.92 | 0.00 | 0.957 |
| Irishga | 24 | 1.00 | 0.96 | 0.00 | 0.979 |
| Italianit | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 23 | 1.00 | 0.96 | 0.00 | 0.978 |
| Polishpl | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 23 | 1.00 | 1.00 | 0.00 | 1.000 |
| Romanianro | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Slovaksk | 24 | 1.00 | 0.96 | 0.00 | 0.979 |
| Sloveniansl | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 24 | 1.00 | 0.96 | 0.00 | 0.979 |
| Swedishsv | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
References
- [1]regulated_advice evaluation reportreports/artifacts/regulatedadvice-full/regulated_advice_eval.json, in the training repository. Every figure in this post is read from it at build time.
- [2]Corpus manifestreports/data/regulated_advice_manifest.json. 11,436 examples across 26 languages, content hash 48de4bb155a8.
- [3]How support, precision and recall are computedborder_train/metrics.py. Support is tp plus fn, so it counts positives only and the negatives are what the false positive rate measures.
- [4]The full per-language table
- [5]The use case this detector exists for
- [6]XLM-RoBERTaThe base model for every classifier in the set. Maltese is absent from its pretraining corpus, which is why Maltese is the weakest language.