Most AI voice agent vendors want your audio on their servers and a bill per minute. If that doesn’t work for you, either because of where your data has to live or because per-minute pricing stops adding up above a certain volume, you’re left building it yourself. We build it instead, on the stack you already run.
Who this is for
You have a PBX in a rack or a VM, and it works. You’ve seen the demos and you can picture an agent handling the first thirty seconds of every call. Then you try to build one and hit the part nobody demos.
Getting a model to answer a question is the easy half. The hard half is everything around it. A caller interrupts and the agent talks straight over them. The first word gets clipped because synthesis started before the channel was ready. Someone says “okay, okay, okay” and the agent reads it as agreement instead of confusion. A transfer drops the call rather than handing it over. Latency drifts past a second and the whole thing starts to feel wrong in a way callers can’t name but do hang up on.
We’ve spent a long time on that half.
What we actually do
We deploy a self-hosted voice agent against your existing dialplan. Speech recognition, the language model and speech synthesis all run on hardware you control. Nothing leaves your network unless you decide it should.
The agent talks to Asterisk over AudioSocket, so it slots in as a routing target rather than replacing anything. You keep your trunks, your queues, your recording and your reporting. When the agent can’t handle a call, it transfers out like any other step in the dialplan.
Where it tends to earn its place: qualifying inbound calls before they reach a queue, taking the repetitive fraction that never needed a person, capturing details your agents currently type while the caller waits, and out-of-hours cover that does more than read an answering machine script.
Packages
| Package | What you get | From |
|---|---|---|
| Voice Agent Audit | We call your existing agent, record it, and measure what a caller actually hears. Latency at each stage, barge-in behaviour, where it breaks. You get the recordings and the numbers. | $1,500 |
| Pilot | Three weeks. One agent, one use case, running on your Asterisk against targets we agree before we start. Either it hits them or you walk away. | $5,000 |
| Implementation | The full build. CRM integration, dialplan work, personas, transfer logic, reporting, handover documentation and training for your team. | $20,000 |
| Managed | We keep it working. Monitoring, tuning as call patterns shift, model and voice updates, defined response times. | $2,000/mo |
| Partner enablement | For ITSPs and resellers who want to deliver this themselves. Training, architecture review and a licence to build on. | $15,000 |
Start with the audit if something is already running. Start with the pilot if nothing is. We’d rather you spend $5,000 finding out this doesn’t suit your call mix than $40,000 discovering it six months in.
Why us
We didn’t come to voice agents from the AI side. We’ve been building telephony software since 2005, and the agent we deploy was pulled out of ICTContact, our own contact center platform, where it runs the AI Voice Agent feature in production.
Then we open sourced the core of it. asterisk-ai-voice-agent is MIT licensed and on GitHub, alongside piper-tts-server for paced streaming synthesis and a benchmark tool that measures what a caller hears rather than what a dashboard claims.
You can read the code before you ever talk to us. Most vendors in this space can’t offer that. The rest of our open source work sits on the projects page, and our wider VoIP services cover the work that isn’t AI related.
How it starts
Tell us what your call flow looks like today and which calls you’d want an agent to take. We’ll tell you honestly whether it’s a good fit. Some call types aren’t, and we’d rather say so in the first conversation than the last one.
Questions
Does our audio ever leave our network?
Not in the self-hosted build. Recognition, the model and synthesis all run on your hardware. If you’d rather use a hosted model for quality reasons, that becomes a decision you make openly, not a default we bury in an architecture diagram.
What hardware do we need?
It depends on concurrency and which models you pick. A modest GPU covers more concurrent calls than most people expect. We size it during the pilot instead of guessing up front.
Can it transfer to a human?
Yes, and getting that right is most of the work. The agent hands back to your dialplan, so transfers run through your existing queues and routing rather than a parallel system.
We run FreeSWITCH, not Asterisk.
That’s fine. The audio path differs but the pipeline doesn’t. We build on both.
Do we have to buy your platform?
No. Everything above runs against whatever you already have. If ICTContact turns out to be a better fit than bolting an agent onto your current setup we’ll say so, but that’s a separate conversation and we won’t start with it.
How long does a build take?
The pilot is three weeks by design. A full implementation usually runs six to twelve weeks depending on how many systems the agent has to talk to. Integrations are what stretch the timeline, not the agent itself.
Talk to us about your call flow
Send us a short description of how calls arrive today and what you’d like to hand to an agent. We’ll come back with an honest read on fit and the package we think makes sense, or a reason not to bother.
