Most speech-to-text benchmarks test native speakers reading clean audio. This one measures what commercial STT does with non-native learner speech, English-accented French, German, Spanish, Italian, Portuguese, and Japanese, with mid-sentence code-switches into English (“je dois prendre un… appointment… demain”), filler words, false starts, and pauses. That’s the kind of speech a language learner actually produces, and the case no vendor publishes accuracy numbers for.
A synthetic (text-to-speech) tier of 1,230 clips across those six languages has already been streamed to nine STT configurations from seven vendors. The volunteer recordings collected here form the real-human tier that checks whether those synthetic results hold up on genuine speech.
Volunteer recording batteries for the real-human tier are below. Each language has one fixed battery of sentences; every recruited speaker of that language reads the same battery. Sample audio on each page illustrates the clip types.