CALLAII ASR v2 technical sheet
Speech recognition · A speech recognition model that separates call recordings into speaker channels and converts them into text. Analysis, coaching, and quality scoring operate on this transcription.
CALLAII ASR v2
Previous version
Released June 2026
Second version adapted to call center Turkish. The error rate on real recordings dropped to half of the first version.
What improved in this version
- Word error rate on 10 real recordings dropped from %14.25 to %6.51
- Character error rate dropped from %8.68 to %4.57
- Domain prompt added for call center expressions (announcements and institution names)
Limitations and responsible use
- Could enter a repetition loop at the end of long calls; fixed in v3
- Output was corrupted when the language label was incorrectly assigned in Arabic recordings; fixed in v3 with automatic language detection
Training and measurements
| Turkish validation | word error rate %7.02, character error rate %2.15 |
|---|
v1
v2
Lower is better. The same 10 real call recordings were fed to both versions.
Show values as a table
| v1 | v2 | |
|---|---|---|
| Word error rate (WER) | 14.25 | 6.51 |
| Character error rate (CER) | 8.68 | 4.57 |
Technical sheet
| Version | v2 |
|---|---|
| Release | June 2026 |
| Status | Replaced by v3 |
Customer data is never used to train the models. This sheet was generated on 27.09.2026. See callmenta.com for the current version. Versions and improvements