Vapi
API-first platform for inbound and outbound phone and web voice agents
A focused inbound agent is a credible weekend build with LiveKit Agents or Pipecat. Replacing Vapi as a product is much larger: it combines realtime orchestration, provider abstraction, phone-number and SIP workflows, outbound campaigns, tool calls, testing, observability, scaling, support, and compliance options. Self-hosting can remove Vapi's $0.05 per-minute platform fee, but model, carrier, number, infrastructure, failed-call, monitoring, and engineering costs remain.
Build verification: not recorded. How we judge buildability
What you give up
- managed realtime infrastructure, scaling, provider failover, and support
- assistant, squad, workflow, phone-number, and provider configuration APIs
- outbound campaigns, carrier integrations, spam mitigation, and production telephony debugging
- hosted simulations, evaluations, call logs, recordings, analytics, and retention controls
- enterprise SLAs, compliance add-ons, role controls, and data-residency options
Why people still pay
They pay to move from a working voice demo to dependable production calling without operating SIP, RTP, WebRTC, agent workers, provider fallbacks, observability, and capacity themselves.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- A user-owned server, LiveKit/SIP configuration, one provisioned inbound number/trunk, provider credentials, consent language and a human fallback destination. Confirm installed SDK APIs before implementation.
- Implementation components: Python and a documented supported LiveKit Agents release with one named worker. LiveKit Server, Redis and SIP services; explicit STT/LLM/TTS/VAD adapters and SQLite call outcome records.
- Scope boundary: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities.
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Optional external skill: modern-python — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: sharp-edges — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records
Project rule — preserve this invariant: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy.
Project rule — acceptance evidence: Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result.
Implementation plan
Phase 1
Scope and fixtures. Implement this bounded workflow: Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities.
Phase 2
Durable model. Model inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy.
Phase 3
Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Persist the call ID and tool receipts, cancel speech on interruption and bound tool latency. Treat uncertain tool outcomes as a handoff condition rather than repeating a side effect.
Phase 4
Permissions and integration failure. Authenticate SIP/agent connections and constrain callable tools and destinations. Caller speech is untrusted input and cannot grant access or override policy. The single lookup reads an allowlisted, non-sensitive support knowledge source only; no account-specific data or external writes are exposed. Authenticated account lookups and write actions require a separate scope and server-verified caller authorization. Keep provider keys server-side, bound call duration/spend and apply explicit recording/transcript consent and retention controls. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
Phase 5
Portable handoff. Retain only the chosen call records under a clear retention policy; export outcome/transcript data with access control. A worker restart must not resurrect a finished call. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
Phase 6
Acceptance scenarios. Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.
WORKING SLICE Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup. Build this scoped Vapi-inspired workflow with a documented data model and visible failure states. Architecture - Python and a documented supported LiveKit Agents release with one named worker. - LiveKit Server, Redis and SIP services; explicit STT/LLM/TTS/VAD adapters and SQLite call outcome records. Prerequisites and limits A user-owned server, LiveKit/SIP configuration, one provisioned inbound number/trunk, provider credentials, consent language and a human fallback destination. Confirm installed SDK APIs before implementation. Outside this release: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities. Data model and correctness inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records Invariant: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy. Persist the call ID and tool receipts, cancel speech on interruption and bound tool latency. Treat uncertain tool outcomes as a handoff condition rather than repeating a side effect. Security and privacy Authenticate SIP/agent connections and constrain callable tools and destinations. Caller speech is untrusted input and cannot grant access or override policy. The single lookup reads an allowlisted, non-sensitive support knowledge source only; no account-specific data or external writes are exposed. Authenticated account lookups and write actions require a separate scope and server-verified caller authorization. Keep provider keys server-side, bound call duration/spend and apply explicit recording/transcript consent and retention controls. Recovery and export Retain only the chosen call records under a clear retention policy; export outcome/transcript data with access control. A worker restart must not resurrect a finished call. Implementation order 1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities. 2. Phase 2 — Durable model. Model inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy. 3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Persist the call ID and tool receipts, cancel speech on interruption and bound tool latency. Treat uncertain tool outcomes as a handoff condition rather than repeating a side effect. 4. Phase 4 — Permissions and integration failure. Authenticate SIP/agent connections and constrain callable tools and destinations. Caller speech is untrusted input and cannot grant access or override policy. The single lookup reads an allowlisted, non-sensitive support knowledge source only; no account-specific data or external writes are exposed. Authenticated account lookups and write actions require a separate scope and server-verified caller authorization. Keep provider keys server-side, bound call duration/spend and apply explicit recording/transcript consent and retention controls. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results. 5. Phase 5 — Portable handoff. Retain only the chosen call records under a clear retention policy; export outcome/transcript data with access control. A worker restart must not resurrect a finished call. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README. 6. Phase 6 — Acceptance scenarios. Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder. Acceptance Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result. Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees. Optional agent guidance Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission. Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission. Project rule — data model: inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records Project rule — preserve this invariant: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy. Project rule — acceptance evidence: Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
Alternatives to building your own
no votes, no pay-to-list · just what's real
Vapi pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| build | $0 | — | At least 60 promotional call minutes at signup and 10 concurrent calls; ongoing calls are fully usage based.The $0 platform commitment is not a permanent free allowance. |
| scale | — | — | Custom annual usage commitment, concurrency, support and volume terms.Contact sales. |
free tierno permanent free tier; signup includes a promotional allowance of at least 60 call minutes
billingusage based with no base subscription for Build; Scale uses a custom annual commitment
hidden costsVoice calls add a $0.05/minute Vapi platform fee plus separate speech-to-text, model, text-to-speech and telephony charges; SMS/chat is $0.005/message; extra concurrency is $10/line/month; HIPAA is $2,000/month and zero-data-retention is $1,000/month.
pricing sources checked 2026-08-14 · pricing source ↗
Questions about Vapi
Can you build your own Vapi with AI?
Partly. A focused inbound agent is a credible weekend build with LiveKit Agents or Pipecat. Replacing Vapi as a product is much larger: it combines realtime orchestration, provider abstraction, phone-number and SIP workflows, outbound campaigns, tool calls, testing, observability, scaling, support, and compliance options. Self-hosting can remove Vapi's $0.05 per-minute platform fee, but model, carrier, number, infrastructure, failed-call, monitoring, and engineering costs remain.
What does the Vapi build prompt cover?
The prompt starts with this scope: Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup. Full-product capabilities excluded from the comparison include: managed realtime infrastructure, scaling, provider failover, and support; assistant, squad, workflow, phone-number, and provider configuration APIs; outbound campaigns, carrier integrations, spam mitigation, and production telephony debugging. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the Vapi prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this Vapi project take?
The catalogue estimate is weekend for one inbound agent; ongoing production operations for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing Vapi?
managed realtime infrastructure, scaling, provider failover, and support; assistant, squad, workflow, phone-number, and provider configuration APIs; outbound campaigns, carrier integrations, spam mitigation, and production telephony debugging; hosted simulations, evaluations, call logs, recordings, analytics, and retention controls; enterprise SLAs, compliance add-ons, role controls, and data-residency options. They pay to move from a working voice demo to dependable production calling without operating SIP, RTP, WebRTC, agent workers, provider fallbacks, observability, and capacity themselves.
What can I use instead of building Vapi?
Dograh: A Vapi-shaped workflow builder you can Docker-compose; phone and model meters remain. Check each option's license, hosting needs and feature limits.