- Industry
- Inbound voice AI on PSTN / SIP
- Type
- Asterisk ExternalMedia + Python turn engine
- Role
- Architecture + media path + orchestration
- Turn path
- VAD → STT → FAQ / slots / Haiku → TTS
- Stack
- Asterisk · FastAPI · Groq · Anthropic · Inworld
- Proof
- One TTS path · scripted identity slots · barge-in
The problem
A hosted “AI agent” UI can look finished in a browser. The hard product is a real SIP session: a trunk answers, RTP has a peer, the greeting must not be heard as the caller, the model must not invent a name, and two TTS synthesizers must not talk over each other on the same turn.
Constraints
Delivery stayed inside the client's stack, tenancy, compliance and media-ownership boundaries. Where a choice was forced (suite vs owned plane, BYOC vs CPaaS, local vs cloud models), the architecture section below records the trade-off rather than a marketing rewrite.
- Scattered trees: duplicate
hooks/corepackages, two entrypoints, two playbacks per reply - Greeting played before the UDP peer exists — silent saluto
- LLM filling identity from bad STT
- Dead air while the model thinks, or fillers overlapping the first TTS syllable
- Echo treated as speech; barge-in cutting the assistant mid-sentence
The fix is not a longer prompt. It is one runtime, one import graph, and a turn machine that can skip the model.
Business requirements
One path: DID maps to a tenant profile, Asterisk parks the call in Stasis with ExternalMedia, Python binds RTP, greeting plays on the first inbound packet, then each utterance is FAQ, scripted slot, or Haiku — then a single paced TTS egress. Hangup is AMI/ARI cleanup, not the model.
This page does not publish invented containment rates or “minutes saved.” The engineering claim is a complete media and turn map.
The solution
Asterisk voice AI is a consolidation workspace. Telephony stays in Asterisk (SIP trunk such as Zadarma as advertised, dialplan, Stasis, bridges, ExternalMedia). Python consumes AMI and ARI. FastAPI plus an official python -m runtime.app entry is the process. Blocks are namespaced: shared audio helpers, orchestration, TTS. There is no production use of duplicate top-level hooks or core packages.
Four logical blocks:
- Telephony — SIP, Stasis, ExternalMedia on Asterisk. Python does not become the PBX.
- Ingest — RTP → G.711 decode → PCM 16 kHz, VAD, utterance assembly, Groq STT
- Orchestration — profiles, FAQ, intents, scripted name/phone, Anthropic Haiku, fillers, barge-in epoch, optional call artifacts
- TTS — Inworld streaming, G.711 encode, paced RTP callback, barge-in event
Contract that matters: orchestration returns text. Only the audio hook calls Block 4 synthesize. The model never fires a second speaker.
How the system flows
Asterisk ExternalMedia streams RTP into Python; VAD and Groq STT feed FAQ or Haiku; Inworld TTS returns paced audio — hangup stays on Asterisk.
flowchart LR SIP[Zadarma SIP] --> AST[Asterisk Stasis ExternalMedia] AST --> RTP[RTP G.711 to PCM] RTP --> VAD[VAD] VAD --> STT[Groq STT] STT --> FAQ[FAQ or Haiku] FAQ --> TTS[Inworld TTS] TTS --> RTP
Architecture
Asterisk owns the call. Python owns the turn. Inworld owns the voice asset. Haiku owns remaining language after FAQ and slots.
flowchart TB Trunk[SIP trunk] --> AST[Asterisk ARI AMI] AST --> Ingest[RTP ingest Python] Ingest --> VAD[VAD barge-in] Ingest --> STT[Groq STT] STT --> Logic[FAQ slots Haiku] Logic --> TTS[Inworld TTS G.711] TTS --> Ingest
Key components
Runtime
Official entry is the consolidation uvicorn app with a single root. Profiles, greetings, fillers, logs and optional call folders live under that root — not a second copy of the tree.
ARI + RTP
StasisStart resolves DID → JSON profile. Per-call audio hook, AI wrapper, Block 4 transport, UDP bind. AMI hangup aligns teardown.
Turn machine
FAQ cache, scripted name/phone read-back, else Haiku with slot context. Filler WAV loop only while Haiku runs.
TTS contract
Public surface is the TTS hook: register transport, synthesize, cleanup, barge-in. No duplicate synthesize from the AI hook.
Call path
- Establish — ARI StasisStart, profile by DID, ExternalMedia bridge, RTP socket. Greeting plays on the first inbound packet; VAD is paused so echo is not “the caller.”
- Ingress — PCMU/PCMA payload → PCM16 8 kHz → resample 16 kHz → audio hook chunks. Quiet frames can be boosted for VAD without rewriting the book.
- Utterance — Silero VAD; a max-utterance cap can force speech-end on long noise. Speech during an in-flight turn bumps an epoch so stale Haiku/TTS is dropped.
- STT — PCM wrapped as WAV to a Groq-compatible Whisper endpoint. Optional prompt when name or phone is the pending slot. Confidence on this path is a heuristic; Groq does not expose segment scores here.
- Decide — FAQ, scripted identity, or Haiku (see below).
- Speak — Block 4 streams µ-law through paced egress (~20 ms frames). Fillers use the same path and stop before TTS after a drain.
- Tear down — AMI Hangup / ARI StasisEnd: flush, end session, close RTP, TTS cleanup. Optional per-call transcript, STT JSON, inbound WAV, events log.
When the model is not allowed to invent
Reservation-class inbound is a slot machine with a language model on the remainder — not a chatbot that “books a table.”
- FAQ cache — instant text, no Haiku, no filler loop
- Scripted name — capture, read-back, confirm. Haiku skipped on that turn so a bad STT fragment does not become a person
- Scripted phone — same pattern, often spoken in blocks
- Else Haiku — system prompt carries collected vs pending slots and the next question template. Intent extract is lightweight (people, day, time). While Haiku runs, a filler manager can loop short bridge WAVs on RTP
Name and phone are never invented by Haiku when scripted capture is on. That is a product rule, not a hope.
TTS, fillers and barge-in
Inworld is the advertised neural TTS. Voice id is a tenant profile field (plus a default). Accent is the voice asset in the portal, not a temperature knob. Streaming frames are G.711 and paced so the caller hears a continuous line, not a dumped clip.
Fillers exist because Haiku has latency. They must drain before TTS or the first syllable clips. Barge-in sets an egress stop and signals the TTS hook; the next reply clears stop before playing. Greeting WAV is pre-recorded per tenant — not live TTS on the opening saluto.
Vendor names on this page (Zadarma, Groq, Anthropic, Inworld, Google Calendar, Twilio/WhatsApp) are described as advertised integrations in this build. WhatsApp and calendar are HTTP side channels when enabled — they are not on the RTP hot path.
Data (what this tree actually stores)
No central PostgreSQL in the consolidation repo. Operational data is JSON tenant profiles, optional per-client SQLite (memory, inventory), and external APIs. Call artifacts are optional files under the runtime root, not a warehouse.
Engineering challenges
- Two speakers per turn — AI hook used to synthesize as well. Consolidation: text out, one synthesize in.
- Greeting before peer — saluto waits for first RTP.
- Identity hallucination — scripted read-back, not Haiku, for name and phone.
- Filler vs TTS overlap — shared paced path, drain, then speak.
- Stale turns — utterance epoch drops in-flight results when the caller keeps talking.
My role
Architecture + media path + orchestration. ARI/RTP wiring, turn machine, TTS contract, tenant profiles. Live DIDs and transcripts stay off this page. The sibling production AI voice agent case is the broader SIP-agent story; Asterisk voice AI is this Asterisk ExternalMedia consolidation.
Technology stack
These technologies are the documented stack for this consolidation. This page does not add hosted agent brands that were not part of the work.
Production considerations
- Single RTP media path — no duplicate TTS streams on Asterisk.
- Latency budget across STT, LLM and TTS on live calls.
- Fail closed to human or hangup when RTP bridge drops.
Engineering outcomes
No invented answer rates or “calls handled.” What this system actually established:
- One runtime entry and one namespaced import graph
- A single TTS synthesize per reply
- FAQ and scripted identity turns that skip the model
- Fillers only on the Haiku wait path, drained before playback
- Greeting tied to RTP peer; barge-in that can stop paced egress
Screenshots and diagrams
Architecture on this page is the HTML diagram above. Related writing: How to build an AI voice agent with SIP and Vapi or Retell is the wrong question.
Architectural insight
Voice AI on a trunk is a turn machine with a speaker attached. If the LLM owns names, hangup and a second TTS, you will hear it on the first bad STT fragment. If Asterisk owns the call and Python owns one synthesize, you have a product you can debug.
FAQ — buyer & architecture questions
Ten common questions about this case study, fit, and engagement.
Building voice AI on Asterisk RTP?
Share the trunk, Stasis app, languages and which slots must not be hallucinated. The first reply is whether the gap is ExternalMedia, turn-taking or a second TTS.
Discuss Your Project