Speech recognition · A speech recognition model that separates call recordings into speaker channels and converts them into text. Analysis, coaching, and quality scoring operate on this transcription.
CALLAII ASR v3Previous versionReleased August 2026
Version with extensive multilingual training. Phone line simulation and silence detection were added in this version.
What improved in this version
Turkish word error rate dropped from 8.87% to 7.11% on the open benchmark set
Arabic dropped from 7.96% to 7.82%
Learned to leave silent sections blank; hallucinated sentences decreased
No degradation observed at the end of long calls (571-second real call)
Automatic language detection per channel has begun
Limitations and responsible use
The open benchmark set consists of read news sentences; it does not represent absolute success in phone calls
Error rates were high in languages other than Turkish in call center scenarios; resolved in v4
Training and measurements
11.540Step
1Epoch
0,1315Final validation loss
%10,49Final word error rate
%3,01Final character error rate
Training data
1,477,006 records, 2,190.9 hours of audio
Silence samples
25.7 hours
Audio types
phone band simulation (8 kHz), cross-mixing, dialogue segmentation
Steps
11,540 · 1 epoch
Training date
August 2026
Training loss
Each point is the average of two consecutive records.
Show values as a table
Training loss
100
2.6340
600
1.6663
1,100
1.6001
1,600
1.5598
2,100
1.5177
2,600
1.4338
3,100
1.4080
3,600
1.3531
4,100
1.3821
4,600
1.3530
5,100
1.3306
5,600
1.3333
6,100
1.3021
6,600
1.2923
7,100
1.2736
7,600
1.2613
8,100
1.2606
8,600
1.2679
9,100
1.2166
9,600
1.2256
10,100
1.2165
10,600
1.2068
11,100
1.1835
11,500
1.1665
Validation loss
Measured on examples never seen in training; a falling curve shows the model learns rather than memorises.
Show values as a table
Validation loss
2,000
0.1633
4,000
0.1528
6,000
0.1456
8,000
0.1383
10,000
0.1340
11,540
0.1316
Error rate on validation (%)
Word error rate (WER)Character error rate (CER)
Lower is better. Measured on the validation set throughout training.
Show values as a table
Word error rate (WER)
2,000
12.60
4,000
11.94
6,000
19.10
8,000
10.71
10,000
18.35
11,540
10.49
Character error rate (CER)
2,000
3.64
4,000
3.33
6,000
10.35
8,000
2.94
10,000
10.79
11,540
3.01
Open benchmark set (FLEURS) · Word error rate (%)
v2v3
Lower is better. 150 records per language, read news sentences. Valid for ranking; does not represent absolute success in phone calls.
Show values as a table
v2
v3
Turkish
8.87
7.11
Arabic
7.96
7.82
Open benchmark set (FLEURS) · Character error rate (%)
v2v3
Lower is better. 150 records per language.
Show values as a table
v2
v3
Turkish
2.83
1.67
Arabic
2.65
2.46
Internal test set · Language-based word error rate (%)
Lower is better. 200 samples per language.
Show values as a table
v3
Arabic
19.72
German
6.80
English
19.10
French
12.74
Italian
8.11
Turkish
10.49
Internal test set · Language-based character error rate (%)
Lower is better. 200 samples per language.
Show values as a table
v3
Arabic
7.41
German
1.95
English
11.89
French
3.75
Italian
2.22
Turkish
3.01
Technical sheet
Version
v3
Release
August 25, 2026
Status
Superseded by v4
Customer data is never used to train the models.
This sheet was generated on 27.09.2026. See callmenta.com for the current version.
Versions and improvements