Not five calls a month,
all of them
Classic quality assurance rests on sampling: a reviewer listens to a handful of calls a month and the team is judged on that sample. CALLAII scores every recorded call and the human reviewer works on the same record. The argument about the sample ends and the conversation moves to what should be fixed.
Evaluation in two layers
The automatic score and human review do not replace each other, they stack.
Automatic: 12 dimensions
Warm welcome, active listening, isolating the reason for the call, understanding the issue, explaining the rationale, positioning outcomes, reassuring confidence, gaining agreement, positive language, staying composed, going beyond expectations and the overall experience. Each 0-100, with a weighted total.
The human form
Your own evaluation form stands separately; its criteria, weights and zero tolerance items work exactly as you define them.
CALLAII can fill the form too
For each criterion in the human form you can write what to ask the AI; CALLAII fills the form and the reviewer goes over it, approving or correcting.
Zero tolerance sits apart
Items like a forbidden phrase or a mandatory disclosure are flagged independently of the total, so they are not missed even on a high scoring call.
Scoring comes from one source
The same score used to be computed separately on three screens, and the screens disagreed. The calculation was moved to one place.
- The manager screen, the supervisor screen and the call record window show the same result.
- How a zero tolerance item affects the total is defined in exactly one place.
- The rationale is identical on every screen: which dimension lost how many points and why.
- Evaluation history stays on the record; if a score changes later, who changed it is visible.
The criteria scorecard
Whether a criterion was met is answered through two eyes: the rule and the model.
Rule based check
Mandatory sentences, forbidden phrases and disclosures are searched directly in the text rather than left to the model's judgement.
Turkish matching
The phrase may not appear word for word. Matching accounts for Turkish suffixes and differences in phrasing.
The model opinion is its own column
CALLAII's opinion does not replace the rule, it sits next to it. Where the two diverge, it is clear which call to look at.
Scorecard report
A breakdown per criterion by team and agent; which item is systematically missed shows up at a glance.
Performance management
A quality score alone is not performance. Criteria are assigned to match the actual work.
Criteria per industry
Performance criteria are assigned by the company's industry and team type; a company with a field team gets field criteria, and one without does not get pointless rows.
Recognition, with its source
Recognition received is broken down by source rather than shown as a single total. Recognition from a colleague does not carry the same weight as recognition from the Callmenta team.
Technical support apart
A technical support line is not judged by the same yardstick as sales or general support; resolution time and reopen rate are tracked separately.
It connects to coaching
A low score does not stay a line in a report; a coaching session opens and development is followed in the same system.
Frequently asked
Do we have to trust the score the AI gives?
You do not. Behind every score is which dimension lost what and why; a reviewer can go over it and correct it. Your own evaluation form stands separately and the final word belongs to a person.
Can we use our own quality form?
You can. Criteria, weights and zero tolerance items work to your definition. If you want, you also write what to ask the AI for each criterion, so the form arrives pre-filled and the reviewer only approves.
We have no QA staff. Can we still run this?
You can. The automatic layer scores every call and produces a rationale; human review steps in only for disputes, spot checks and zero tolerance items. You can start without building a quality team.
Can an agent challenge their score?
They can. The rationale and the quote it rests on are open to the agent. Where a speech recognition error is suspected the transcript is corrected by hand and the evaluation is rerun.
What percentage of calls is scored?
All recorded ones. There is no sampling; instead of judging a team from a few calls a month, every call carries its score and its rationale.
Let us set it up with your own form
In the demo we load your own quality form into the system and run it on a real call of yours. We look together at where the automatic score and the human score diverge.