Speech recognition · A speech recognition model that separates call recordings into speaker channels and converts them into text. Analysis, coaching, and quality scoring operate on this transcription.
CALLAII ASR v4LiveReleased September 2026
Separates call recordings into speaker channels and converts them to text. This version was trained on ten languages and phone-line-simulated audio; error rates in call center conversations dropped significantly.
What improved in this version
Turkish word error rate in call center scenarios dropped from 25.70% to 6.32%
Error rates in Arabic, German, French, English, and Italian fell to one-quarter and one-fifth; per-language results shown in the chart
Improved on read Turkish speech as well: 4.72% → 4.47% (800 recordings)
Trained on phone-line, noisy-environment, and studio versions of the same text
Previous version's data included in training; new data added without degrading learned knowledge
What it does
Transcribes customer and agent separately in two-channel recordings
Detects each channel's language independently
Produces timestamped, speaker-labeled transcripts
Leaves silent sections blank; trained not to generate hallucinated sentences
Limitations and responsible use
Measurements were made on test sets; reference-based error rate on real phone calls has not yet been measured
Error increases with overlapping speech and very low-quality recordings
Each point is the average of 7 consecutive recordings.
Show values as a table
Training loss
175
10.8073
1,050
6.4649
1,925
5.6508
2,800
5.3350
3,675
5.0206
4,550
4.7431
5,425
4.6104
6,300
4.6143
7,175
4.5038
8,050
4.3882
8,925
4.4358
9,800
4.3362
10,675
4.2276
11,550
4.1843
12,425
4.2788
13,300
4.1626
14,175
4.1483
15,050
4.1727
15,925
4.0753
16,800
4.0031
17,675
4.0504
18,550
4.2040
19,350
4.0197
Validation loss
Measured on examples never seen in training; a falling curve shows the model learns rather than memorises.
Show values as a table
Validation loss
1,000
0.1911
2,000
0.1611
3,000
0.1446
4,000
0.1345
5,000
0.1259
6,000
0.1187
7,000
0.1135
8,000
0.1097
9,000
0.1068
10,000
0.1050
11,000
0.1025
12,000
0.1021
13,000
0.1010
14,000
0.0997
15,000
0.0994
16,000
0.0985
17,000
0.0985
18,000
0.0984
19,000
0.0980
19,369
0.0981
Error rate on validation (%)
Word error rate (WER)Character error rate (CER)
Lower is better. Measured on the validation set throughout training.
Show values as a table
Word error rate (WER)
1,000
15.03
2,000
13.22
3,000
12.06
4,000
10.78
5,000
10.51
6,000
10.01
7,000
9.60
8,000
9.72
9,000
9.52
10,000
9.43
11,000
9.36
12,000
9.35
13,000
9.38
14,000
9.52
15,000
9.23
16,000
9.36
17,000
9.35
18,000
9.36
19,000
9.24
19,369
9.41
Character error rate (CER)
1,000
4.24
2,000
3.85
3,000
3.47
4,000
3.13
5,000
3.01
6,000
2.94
7,000
2.87
8,000
2.89
9,000
2.89
10,000
2.85
11,000
2.88
12,000
2.81
13,000
2.92
14,000
2.93
15,000
2.80
16,000
2.83
17,000
2.77
18,000
2.80
19,000
2.72
19,369
2.82
Call center evaluation · Per-language word error rate (%)
v3v4
Lower is better. 300 samples per language.
Show values as a table
v3
v4
Arabic
67.59
13.86
German
54.58
15.37
English
59.92
12.58
French
62.36
15.57
Italian
38.88
8.42
Turkish
25.70
6.32
Call center evaluation · Per-language character error rate (%)
v3v4
Lower is better. 300 samples per language.
Show values as a table
v3
v4
Arabic
48.62
5.76
German
46.00
7.23
English
44.41
6.23
French
47.92
7.34
Italian
26.92
3.45
Turkish
10.47
1.96
Read Turkish speech · Per-language word error rate (%)
v3v4
Lower is better. 800 samples per language.
Show values as a table
v3
v4
Turkish
4.72
4.47
Read Turkish speech · Per-language character error rate (%)
v3v4
Lower is better. 800 samples per language.
Show values as a table
v3
v4
Turkish
1.11
0.95
Technical sheet
Version
v4
Release
8 September 2026
Input
Audio recording, single or dual channel
Output
Speaker-labeled, timestamped text
Language detection
Automatic per channel
Runs on
Servers managed by Callmenta
Customer data is never used to train the models.
This sheet was generated on 27.09.2026. See callmenta.com for the current version.
Versions and improvements