CONFIDENTIAL PROJECT

Asterisk voice AI — Asterisk voice AI on RTP

A consolidation workspace for Asterisk-mediated voice AI: ingest caller audio over RTP, detect speech, transcribe, decide the turn, speak back. One import graph. One TTS path. The model does not own hangup.

Asterisk ARI / AMIExternalMedia RTPZadarma (as advertised)Silero VADGroq WhisperAnthropic HaikuInworld TTSFastAPIJSON tenant profiles
Industry
Inbound voice AI on PSTN / SIP
Type
Asterisk ExternalMedia + Python turn engine
Role
Architecture + media path + orchestration
Turn path
VAD → STT → FAQ / slots / Haiku → TTS
Stack
Asterisk · FastAPI · Groq · Anthropic · Inworld
Proof
One TTS path · scripted identity slots · barge-in
Asterisk voice AI media path: Asterisk, VAD and STT, FAQ or LLM, Inworld TTS back to RTP

The problem

A hosted “AI agent” UI can look finished in a browser. The hard product is a real SIP session: a trunk answers, RTP has a peer, the greeting must not be heard as the caller, the model must not invent a name, and two TTS synthesizers must not talk over each other on the same turn.

Constraints

Delivery stayed inside the client's stack, tenancy, compliance and media-ownership boundaries. Where a choice was forced (suite vs owned plane, BYOC vs CPaaS, local vs cloud models), the architecture section below records the trade-off rather than a marketing rewrite.

  • Scattered trees: duplicate hooks / core packages, two entrypoints, two playbacks per reply
  • Greeting played before the UDP peer exists — silent saluto
  • LLM filling identity from bad STT
  • Dead air while the model thinks, or fillers overlapping the first TTS syllable
  • Echo treated as speech; barge-in cutting the assistant mid-sentence

The fix is not a longer prompt. It is one runtime, one import graph, and a turn machine that can skip the model.

Business requirements

One path: DID maps to a tenant profile, Asterisk parks the call in Stasis with ExternalMedia, Python binds RTP, greeting plays on the first inbound packet, then each utterance is FAQ, scripted slot, or Haiku — then a single paced TTS egress. Hangup is AMI/ARI cleanup, not the model.

This page does not publish invented containment rates or “minutes saved.” The engineering claim is a complete media and turn map.

The solution

Asterisk voice AI is a consolidation workspace. Telephony stays in Asterisk (SIP trunk such as Zadarma as advertised, dialplan, Stasis, bridges, ExternalMedia). Python consumes AMI and ARI. FastAPI plus an official python -m runtime.app entry is the process. Blocks are namespaced: shared audio helpers, orchestration, TTS. There is no production use of duplicate top-level hooks or core packages.

Four logical blocks:

  • Telephony — SIP, Stasis, ExternalMedia on Asterisk. Python does not become the PBX.
  • Ingest — RTP → G.711 decode → PCM 16 kHz, VAD, utterance assembly, Groq STT
  • Orchestration — profiles, FAQ, intents, scripted name/phone, Anthropic Haiku, fillers, barge-in epoch, optional call artifacts
  • TTS — Inworld streaming, G.711 encode, paced RTP callback, barge-in event

Contract that matters: orchestration returns text. Only the audio hook calls Block 4 synthesize. The model never fires a second speaker.

How the system flows

Asterisk ExternalMedia streams RTP into Python; VAD and Groq STT feed FAQ or Haiku; Inworld TTS returns paced audio — hangup stays on Asterisk.

flowchart LR
  SIP[Zadarma SIP] --> AST[Asterisk Stasis ExternalMedia]
  AST --> RTP[RTP G.711 to PCM]
  RTP --> VAD[VAD]
  VAD --> STT[Groq STT]
  STT --> FAQ[FAQ or Haiku]
  FAQ --> TTS[Inworld TTS]
  TTS --> RTP
Single-turn voice AI path on RTP.

Architecture

Asterisk owns the call. Python owns the turn. Inworld owns the voice asset. Haiku owns remaining language after FAQ and slots.

flowchart TB
  Trunk[SIP trunk] --> AST[Asterisk ARI AMI]
  AST --> Ingest[RTP ingest Python]
  Ingest --> VAD[VAD barge-in]
  Ingest --> STT[Groq STT]
  STT --> Logic[FAQ slots Haiku]
  Logic --> TTS[Inworld TTS G.711]
  TTS --> Ingest
Asterisk voice AI architecture.

Key components

Runtime

Official entry is the consolidation uvicorn app with a single root. Profiles, greetings, fillers, logs and optional call folders live under that root — not a second copy of the tree.

ARI + RTP

StasisStart resolves DID → JSON profile. Per-call audio hook, AI wrapper, Block 4 transport, UDP bind. AMI hangup aligns teardown.

Turn machine

FAQ cache, scripted name/phone read-back, else Haiku with slot context. Filler WAV loop only while Haiku runs.

TTS contract

Public surface is the TTS hook: register transport, synthesize, cleanup, barge-in. No duplicate synthesize from the AI hook.

Call path

  1. Establish — ARI StasisStart, profile by DID, ExternalMedia bridge, RTP socket. Greeting plays on the first inbound packet; VAD is paused so echo is not “the caller.”
  2. Ingress — PCMU/PCMA payload → PCM16 8 kHz → resample 16 kHz → audio hook chunks. Quiet frames can be boosted for VAD without rewriting the book.
  3. Utterance — Silero VAD; a max-utterance cap can force speech-end on long noise. Speech during an in-flight turn bumps an epoch so stale Haiku/TTS is dropped.
  4. STT — PCM wrapped as WAV to a Groq-compatible Whisper endpoint. Optional prompt when name or phone is the pending slot. Confidence on this path is a heuristic; Groq does not expose segment scores here.
  5. Decide — FAQ, scripted identity, or Haiku (see below).
  6. Speak — Block 4 streams µ-law through paced egress (~20 ms frames). Fillers use the same path and stop before TTS after a drain.
  7. Tear down — AMI Hangup / ARI StasisEnd: flush, end session, close RTP, TTS cleanup. Optional per-call transcript, STT JSON, inbound WAV, events log.

When the model is not allowed to invent

Reservation-class inbound is a slot machine with a language model on the remainder — not a chatbot that “books a table.”

  • FAQ cache — instant text, no Haiku, no filler loop
  • Scripted name — capture, read-back, confirm. Haiku skipped on that turn so a bad STT fragment does not become a person
  • Scripted phone — same pattern, often spoken in blocks
  • Else Haiku — system prompt carries collected vs pending slots and the next question template. Intent extract is lightweight (people, day, time). While Haiku runs, a filler manager can loop short bridge WAVs on RTP

Name and phone are never invented by Haiku when scripted capture is on. That is a product rule, not a hope.

TTS, fillers and barge-in

Inworld is the advertised neural TTS. Voice id is a tenant profile field (plus a default). Accent is the voice asset in the portal, not a temperature knob. Streaming frames are G.711 and paced so the caller hears a continuous line, not a dumped clip.

Fillers exist because Haiku has latency. They must drain before TTS or the first syllable clips. Barge-in sets an egress stop and signals the TTS hook; the next reply clears stop before playing. Greeting WAV is pre-recorded per tenant — not live TTS on the opening saluto.

Vendor names on this page (Zadarma, Groq, Anthropic, Inworld, Google Calendar, Twilio/WhatsApp) are described as advertised integrations in this build. WhatsApp and calendar are HTTP side channels when enabled — they are not on the RTP hot path.

Data (what this tree actually stores)

No central PostgreSQL in the consolidation repo. Operational data is JSON tenant profiles, optional per-client SQLite (memory, inventory), and external APIs. Call artifacts are optional files under the runtime root, not a warehouse.

Engineering challenges

  • Two speakers per turn — AI hook used to synthesize as well. Consolidation: text out, one synthesize in.
  • Greeting before peer — saluto waits for first RTP.
  • Identity hallucination — scripted read-back, not Haiku, for name and phone.
  • Filler vs TTS overlap — shared paced path, drain, then speak.
  • Stale turns — utterance epoch drops in-flight results when the caller keeps talking.

My role

Architecture + media path + orchestration. ARI/RTP wiring, turn machine, TTS contract, tenant profiles. Live DIDs and transcripts stay off this page. The sibling production AI voice agent case is the broader SIP-agent story; Asterisk voice AI is this Asterisk ExternalMedia consolidation.

Asterisk ARIRTP / G.711VADSTT pathSlot machineTTS contractBarge-inTenant profiles

Technology stack

AsteriskARI / AMIExternalMediaFastAPIPython 3Silero VADGroq WhisperAnthropic HaikuInworld TTSJSON profilesSQLite (optional)

These technologies are the documented stack for this consolidation. This page does not add hosted agent brands that were not part of the work.

Production considerations

  • Single RTP media path — no duplicate TTS streams on Asterisk.
  • Latency budget across STT, LLM and TTS on live calls.
  • Fail closed to human or hangup when RTP bridge drops.

Engineering outcomes

No invented answer rates or “calls handled.” What this system actually established:

  • One runtime entry and one namespaced import graph
  • A single TTS synthesize per reply
  • FAQ and scripted identity turns that skip the model
  • Fillers only on the Haiku wait path, drained before playback
  • Greeting tied to RTP peer; barge-in that can stop paced egress

Screenshots and diagrams

Architecture on this page is the HTML diagram above. Related writing: How to build an AI voice agent with SIP and Vapi or Retell is the wrong question.

Architectural insight

Voice AI on a trunk is a turn machine with a speaker attached. If the LLM owns names, hangup and a second TTS, you will hear it on the first bad STT fragment. If Asterisk owns the call and Python owns one synthesize, you have a product you can debug.

FAQ — buyer & architecture questions

Ten common questions about this case study, fit, and engagement.

Documented barge-in, multilingual fallback, tool use or telephony integration—not a demo webhook only.

No. Asterisk/FreeSWITCH handles media; AI orchestration sits beside call control.

Yes. Narrow IVR replacement or receptionist is a common first phase.

Retention and redaction policies are set per project; we implement access control and storage limits.

Architecture can route providers; latency and data residency drive the choice.

Call flows, CRM, languages, compliance rules and example calls that fail today.

Share the current stack (PBX, CRM, cloud), the failure or goal in one paragraph, peak call volume, carriers, and any deadline. Screenshots, a pcap, or a short Loom beat a 40-page RFP. Use the project brief or email hello@unifiedpbx.in.

Yes. Delivery is remote-first from Delhi with scheduled overlap for EU, UK, US and APAC stand-ups. Production changes use written runbooks, rollback steps and agreed maintenance windows.

Yes. Many engagements begin with a 1–2 week SIP trace review, tenant-isolation audit, or architecture assessment. If the fit is good, scope expands from evidence—not from a generic sales deck.

Founders, CTOs, telecom leads, MSPs and product teams building or fixing UCaaS, CCaaS, CRM+voice, AI voice, or vertical SaaS—not buyers who only need seats on a mass-market suite.

Building voice AI on Asterisk RTP?

Share the trunk, Stasis app, languages and which slots must not be hallucinated. The first reply is whether the gap is ExternalMedia, turn-taking or a second TTS.

Discuss Your Project