Every call is read, not a sample of them
CALLAII is the speech analytics engine Callmenta builds in house. It transcribes the call with speech recognition trained for Turkish, reads the sentiment journey from the transcript, produces a 12 dimension quality score, writes coaching for the agent and escalates risky calls to managers. No sampling: every recorded call is processed.
What comes out of one call
Every field below is produced in a single analysis pass and sits on the call detail screen in the panel.
Turkish speech recognition and speaker separation
The recording is transcribed and speakers are split into time stamped segments. The engine was trained for Turkish, on real calls with regional accents, fast speech and background noise.
The sentiment journey, start to finish
How the customer entered the call and how they left it are read separately: calm, neutral, worried, disappointed, angry, happy, satisfied. Direction and intensity are written too. Every reading arrives with a short rationale from the transcript; it never just says "negative" and stops.
Critical moments, with the quote
Up to five moments that changed the course of the call, each with the customer's own sentence, the sentiment at that point and the trigger. You see what broke without listening to the recording.
Did de-escalation work
Whether a call that started tense actually softened is flagged as its own field. Whether the agent really calmed the customer becomes measurable.
12 dimension quality score
Warm welcome, active listening, isolating the reason for the call, understanding the issue, explaining the rationale, positioning outcomes, reassuring confidence, gaining agreement, positive language, staying composed and professional, going beyond expectations and the overall experience. Each dimension 0-100, with a weighted total.
Why the score is what it is
The strongest dimensions, the ones that lost the most points, and the transcript based reason for the loss. What should have been said for a full score is written out as an example sentence, and it is never left blank even on a high scoring call.
Operational risk
Churn risk, escalation need, risk category and priority. A risky call surfaces to the manager on its own rather than waiting for a report.
Coaching note
A strength, a critical development area and an action. The development area comes with the agent's own sentence, what went wrong in it and what could be said instead.
Vocal behaviour metrics
Next to the text, measurable behaviour is derived from the audio itself. This measurement is never asked of a language model; it is computed directly on the recording, so the margin of error comes from the signal rather than from interpretation.
The thresholds were not invented; they come from the distribution of real calls and are re-measured as the data grows. If we cannot map a voice to a person with confidence, the note is written without naming anyone. Telling the wrong person "you interrupted" is worse than saying nothing.
- Interruptions: starts that land on top of the other party still speaking
- Talk balance: each side's share of speaking time
- Dead air: time inside the call when nobody is speaking
- Overtalk: total time both sides speak at once
We do not infer emotion from voice
This is not a gap, it is a deliberate engineering and compliance decision. We write the reason out in the open, because the opposite is promised often in this market.
Emotion recognition on employees is prohibited
The EU AI Act prohibits systems that infer emotion from an employee's biometric data in the workplace. A product that produces an anger or stress score from an agent's tone of voice cannot lawfully be sold in Europe. We never opened that door.
Inference from a customer's voice is high-risk
Inferring emotion from a customer's voice sits in the high-risk annex of the same act. Landing in that class means conformity assessment, record keeping and audit obligations. We chose not to put the product there.
Our sentiment reading comes from text
The sentiment journey is read from the transcript. Text is not biometric data, which is why this reading falls outside the act's definition of emotion recognition. Behind every sentiment label you see in the panel there is a sentence.
Acoustic measurement never feeds the sentiment result
Interruptions, dead air and talk balance are not an input to the sentiment reading, they do not calibrate it, and they do not steer attention to "look at this minute" either. The two lines are kept apart on purpose and the boundary is protected by an automated check inside the code. Giving the same number an emotional name would put the product in the high-risk class.
Here is what we get in return: the measurement is raw behaviour, it is hard to argue with, and it can be deployed in Europe as well as inside a customer's own walls. Instead of telling an agent "you sounded angry", we can say "you interrupted seven times on this call and there are 42 seconds of dead air" — verifiable, and fixable.
How the line is set up
- Your existing phone system connects through a SIP trunk; your numbers and your carrier do not change.
- The system forwards a copy of each call through the SIPREC recording standard; CALLAII never sits in the call path.
- The recording is transcribed, the analysis is produced and the call lands on the customer record.
- Quality score, coaching and risk alerts appear in the panel without waiting for a report.
- If the data must not leave your organization, CALLAII Edge runs the analysis on your own servers.
How this differs from general purpose sentiment APIs
The text sentiment services of cloud providers return a label. What actually helps in a contact center is a score, a rationale, coaching and an action.
A report card, not a label
A positive or negative label on its own develops nobody. CALLAII produces the 12 dimension quality score, the reason points were lost and the example sentence together.
Built for Turkish
The engine was trained on Turkish calls. General purpose multilingual services miss the tone, the abbreviations and the floor jargon of Turkish contact center speech.
The panel and the workflow are included
The analysis is not a standalone API response; it arrives attached to the customer record, the agent scorecard, the manager alert and the shift plan.
Data residency is explicit
Data is held in the EU region and the process is KVKK and GDPR compliant. Where nothing may leave the building, there is the Edge deployment.
Frequently asked
Do you do sentiment analysis or not?
We do, but the source matters. The sentiment reading comes from the call transcript: how the customer started, how they ended, the direction, the intensity and the rationale. We do not infer emotion from tone of voice, pitch or any other biometric feature. That distinction is legally decisive: text is not biometric data.
Competitors say they detect anger from tone of voice. Why don't you?
Because emotion recognition on employees in the workplace is prohibited by the EU AI Act, and inference from a customer's voice sits in the high-risk class. A product built on that promise creates problems when you sell it into Europe or install it on premise. We went for measurable behaviour instead: interruptions, dead air, talk balance.
What percentage of calls gets analyzed?
All recorded calls. Quality assurance is not done by sampling; instead of listening to a handful of calls a month and judging the team on that sample, a score is produced for every one of them.
Do we have to change our phone system?
No. CALLAII connects to your existing system over a SIP trunk and listens to a copy of the call through the SIPREC standard. Your numbers, your carrier and your call flow stay the same.
Where are the recordings processed?
In the cloud deployment, data is held in the EU region. If the data must not leave your organization, CALLAII Edge runs the analysis on your own server, without an internet connection where required.
What happens when an analysis is wrong?
Scores are open to challenge: the manager sees the analysis, reads its rationale and corrects it where needed. If a speech recognition error is suspected, the transcript can be edited by hand on the correction screen. The rationale behind a result is always visible; it is not a closed box.
Test it on your own call
In the demo we run a real call recording from your own sector together. You see what the analysis produces on your data, not on a sample.