Data stays inside the organisation
Call recordings, transcripts and evaluations never leave your organisation. The entire platform — panel, CRM, speech recognition and the CALLAII language model — runs locally, hybrid or fully offline. A single workstation suits SMBs, while a high-availability cluster fits the enterprise.
The screen above is a representative example, not real deployment data.
From SMB to enterprise one architecture
The same Callmenta platform runs in three different configurations, scaled to your size.
SMB Workstation
The full Callmenta platform running on a single machine, with CALLAII 12B.
Small businesses
Single-Node Server
A single server for mid-to-large operations, with CALLAII 26B.
Mid-to-large operations
HA Cluster
High availability and automatic failover.
Banking, insurance, public sector
16 GB on a single GPU is enough to start
For a small business, a single NVIDIA card with 16 GB of memory runs the entire platform with CALLAII 12B and speech recognition. As you grow, you move to a 48 GB card and a high-availability cluster.
Starter (SMB)
- 1× NVIDIA, 16 GB VRAM (for example RTX 4080 or RTX A4000)
- 32 GB DDR5 RAM
- 1 TB NVMe SSD
- Linux Ubuntu 22.04 / 24.04
CALLAII 12B · capacity measured before deployment based on your call volume
Enterprise HA
- 2+1 nodes, 48 GB GPU per node
- ECC RAM · NVMe RAID
- 10 GbE network
- 2+1 nodes, automatic failover, air-gapped option
CALLAII 26B · capacity measured before deployment based on your call volume
The model is selected to match the hardware during setup. On a card with 16 GB of memory, CALLAII 12B and speech recognition run together. CALLAII 26B requires a single 48 GB card or two 24 GB cards; 26B alone uses nearly the entire 24 GB card. Measurement (September 2026): on two 24 GB NVIDIA L4 cards (one running CALLAII 26B, the other running speech recognition), with real calls measured at average duration and two concurrent analyses, approximately 700 calls per day are analyzed under continuous load; speech recognition ran at approximately three times the speed of the audio.
Sized to your operation 3 packages
Hardware, licence and service level are set per package. Final pricing is built around your call volume and required precision.
Starter & Small Operation
Start with a single server. Ideal for SMBs and proof-of-concept deployments, powered by CALLAII 12B.
- 1× NVIDIA GPU, 16 GB
- Callmenta platform + CALLAII 12B
- Speech recognition on the same card
- Standard support (business days)
- Manually signed updates
Mid-Scale + HA
CALLAII 26B and fully redundant architecture for mid-size and large operations.
- 1× 48 GB or 2× 24 GB NVIDIA GPU
- Callmenta platform + CALLAII 26B
- Automatic failover, redundant storage
- Priority support (8×5)
- Edge + Management Cloud mode optional
Large Scale + Air-Gapped
Cluster architecture, fully offline operation - for critical infrastructure and regulated industries.
- 2+1 nodes, 48 GB GPU per node
- Automatic failover
- Fully air-gapped operation (offline license)
- 24/7 SLA + field engineer
- BDDK / KVKK / public-audit reporting
Your data never crosses the perimeter
Call recordings, transcripts and evaluations stay inside your organisation's boundary. No data ever leaves.
Data sovereignty
Recordings, transcripts and evaluations stay inside the organisation. Compliant with KVKK, GDPR and sector regulators (BDDK, SPK, health authorities).
No cloud round-trip
Analysis runs on your own server; audio and text never leave the premises. Risky calls are escalated to the manager from within the organisation.
License portal
Manage every node from one dashboard: licensing, versions, health metrics, alerts and audit logs.
No change to your phone system
CALLAII Edge connects to your existing phone system over a SIP trunk (SIPREC). The connection, recording and analysis all run on your own server. You change neither your carrier nor your numbers.
CALLAII Pro is also an Edge deployment
Pro is the training layer that sits on top of Edge. The model is trained on your GPU server with your own conversation data and runs there; the data does not leave your organisation.
The same server
If you already run Edge, Pro sits on top of it. No second stack is installed, the same nodes are used.
Trained on your data
Approved evaluations turn into training data in ALPACA, ChatML or JSONL format. The training run happens on your hardware.
If you have no GPU server
We can also run your model on our hardware. In that case an activation fee replaces the installation fee and the monthly amount covers hosting.
Frequently asked questions about CALLAII Edge
What is CALLAII Edge?
CALLAII Edge is an on-premise deployment that runs the entire Callmenta platform (panel, call center, CRM, industry modules, and CALLAII AI) on your own server. Which modules are enabled is determined by the license. Your data never leaves your organization.
Does my data leave the organization?
No. With an Edge deployment, data never leaves your own infrastructure, which suits the strictest data-privacy scenarios under KVKK and GDPR.
Who is it for?
It suits anyone who needs privacy and wants to run Callmenta on their own infrastructure, including data-sensitive sectors such as banking, insurance, healthcare, and the public sector.
Does it run fully offline?
Yes, if desired. In the standard setup, a limited connection to the internet is established solely for license verification; in a fully air-gapped setup, the license is verified offline. Call data is never sent out under any circumstances.
Does it offer the same capabilities as the cloud version?
Yes. The panel, call center, CRM, and industry modules also work on-premise; the license determines which modules are enabled. Connections that require the internet (WhatsApp, SMS, e-commerce) are opened only to the external addresses you authorize.
Let's plan an Edge architecture review together
Let us plan the hardware list, the licence model and the rollout together.