Measured on 8 and 9 October 2026

Results and method

All our numbers, with their limits. They come from test sets the models did not see in training; where a test is less clean than it looks, we say so.

Speech recognition

Transcriber v4

Base: the w2v-BERT 2.0 family (MIT licence), trained by us for Baoulé. Measure: character error rate (CER) and word error rate (WER), after lower-casing and removing punctuation; tone marks are kept. Lower is better.

10.3%CER, speakers never heard

553 sentences read over the phone by 4 speakers who are not in the training data. WER 28.3%.

11.2%CER measured through the API

Same test, through our server, spaces not counted. 95% CI [10.5–11.8]. WER 28.3% [26.9–29.9].

Offline evaluation of transcriber v4 (9 October 2026)
Test setClipsCERWER
Phone, speakers never heard55310.3%28.3%
Field recordings, public test set4798.4%25.7%
Sentences read by volunteers, public test set ¹5628.0%20.1%
Read texts, public test set ²1924.1%11.6%
Bible reading, held-out books4204.1%9.8%

In this table, CER counts spaces as characters. ¹ Sentences differ from training, but all 15 speakers in this test also appear in our training data: an optimistic figure. ² The same kind of texts as the phone test; this set does not let us check whether its speakers also appear in its training part.

Measured through the API (9 October 2026, 95% confidence intervals)
Test setCER, no spacesCER, spaces countedWER
Phone, speakers never heard (553)11.2% [10.5–11.8]10.3% [9.7–10.9]28.3% [26.9–29.9]
Sentences read by volunteers (562) ¹8.7% [7.7–10.0]8.0% [7.0–9.2]19.9% [18.3–21.7]

Text overlaps

An automatic check compares the phone test with everything used in training. No shared recordings or speakers. But 16 test sentences appear word for word in our training texts (13 are one- to three-word expressions, 3 are full sentences found in earlier training data), and 3 more are near-duplicates.

Same kind of text

The phone test and one of the training sets added for v4 are readings of the same kind of text. Identical and near-identical sentences were removed from training, but part of the gain on this test may come from that similarity.

Two definitions of CER

Our internal evaluation counts spaces; the API benchmark does not. Measured the same way, the API and the internal evaluation agree (10.3%).

Translation

Baoulé ↔ French

“Everyday” test set: 316 sentence pairs never seen in training. Measure: chrF++ (how closely the characters and words match a human reference translation, 0 to 100). Higher is better. Baoulé is scored without tone marks.

“Everyday” test set, 316 sentences, chrF++
ModelBase and licenceBaoulé → FrenchFrench → Baoulé
Commercial (round 5, 9 October)MADLAD-400 family, Apache 2.0 licence: commercially usable34.734.1
Research (round 3, 8 October)NLLB-200 family, non-commercial licence (CC BY-NC 4.0)32.338.0
  • Confidence intervals. Measured through the API, the Research model scores 32.2 [30.3–34.3] and 37.5 [35.5–39.6] (the server decodes with slightly different settings). For the Commercial model, the gain over its previous round is +1.6 [0.6–2.6] into French and +2.2 [1.2–3.2] into Baoulé (bootstrap, 1,000 resamples).
  • Which one to use? The Research model is best into Baoulé; the Commercial model is best into French and the only one you can use in a paid product.
By part of the test (chrF++)
PartSentencesCommercial bci → fraCommercial fra → bciResearch bci → fraResearch fra → bci
Free conversation16631.629.728.232.2
Dictionary examples11237.438.837.347.0
Health questions and answers3843.846.543.453.1
Other registers (chrF++)
RegisterSentencesCommercial bci → fraCommercial fra → bciResearch bci → fraResearch fra → bci
Tales5423.125.720.628.4
Bible, held-out passages22737.835.047.439.5

A test close to training

The test sentences come from the same sources (same documents, same speakers) as part of the training data. The sentences themselves differ: any training pair identical or very close (90% similar or more) to a test sentence was removed. The numbers are therefore on the optimistic side.

Our weak spot

Our weakest part is free conversation, and tales are weaker still. That is exactly what the Voix Baoulé programme is meant to improve: real conversations, in real conditions.

Small tests, for now

316 sentences is not many: a gap of one or two points between two models can be chance. There is no standard public test set for Baoulé yet; we are building one.

Our method

Four steps, the same for speech and for text.

  1. 1. Gather

    Bilingual books scanned and read by our OCR, public recordings with their text, then our own recordings, consented and paid.

  2. 2. Unify the spelling

    Older texts are converted to the modern spelling (ɛ, ɔ…). Every rule is measured: on a 14,499-word trial corpus, the final conversion recognises 88.9% of words, against 70.8% with simple rules.

  3. 3. Verify the alignment

    Every audio clip is aligned with its text and scored, every sentence with its translation; doubtful pairs are removed before training, and samples are checked by hand.

  4. 4. Measure on held-out data

    Test sets are split off before training; an automatic check looks for shared recordings, speakers and sentences; every score comes with its confidence interval (1,000 bootstrap resamples).

What we don't publish

We do not publish our model weights, the detailed list of our data sources, or the details of our processing pipelines. A question about these measurements? Write to contact@abyshire.ci.

Last updated: 10 October 2026.