Build voice agents for phone and web conversations with configurable speech providers, tools, Squads, transfers, testing, and call insights.
Vapi
Explore features, practical uses and pricing below.
Vapi is a developer platform for voice agents that make or receive phone calls and provide voice conversations inside web applications. It coordinates speech recognition, a language model, speech generation, telephony, and connections to external tools. The developer configures the assistant's behavior while Vapi supplies the voice application infrastructure. Typical documented uses include inbound support, appointment scheduling, information collection, and routing a caller to the appropriate next step.
The introduction describes the three core parts of a voice assistant: speech-to-text interprets the caller, the model decides what to say or do, and text-to-speech produces the reply. Vapi lets builders select providers and models for those components. A preset offers a starting combination, while more detailed configuration allows a team to tune the stack for its requirements.
A voice application is more than a chatbot with an audio button. It needs to handle interruption, uncertain recognition, periods of silence, a tool that responds slowly, and a caller who changes direction mid-sentence. Vapi is relevant when those interaction details belong inside a product or telephone workflow. Its documentation and APIs are particularly useful to teams that need control over the connected services and the call's actual outcome.
Vapi's current primary building block is an Assistant: a configured model, voice, transcriber, instructions, tools, and outputs. The dashboard offers templates, and developers can also create assistants programmatically. Instructions describe the role, the information to gather, and the conditions for using a tool or escalating. A first message establishes the opening of the conversation and should match the assistant's actual scope.
A useful assistant specification starts with a small set of outcomes. An appointment assistant might search availability, collect the user's preferred time, call the approved booking service, and confirm the result returned by that service. It should also know how to respond when no appointment is available or required information is missing. These are proposed workflow requirements, rather than automatic behavior implied by selecting a template.
Short, clear voice responses are especially important when the user cannot scan a long answer. Test dates, names, and email addresses through the actual voice path. A transcript that looks correct can still contain a spoken phrase the caller misunderstands. Review what the user hears and what the transcriber captures, then adjust the relevant component instead of treating every error as a prompt problem.
The phone quickstart covers creating an assistant, configuring a number, testing an inbound call, and optionally placing an outbound call. Current documentation explains that free Vapi numbers support inbound calling and that outbound calling requires an imported number from a supported provider. Check the current number availability, region, account requirements, and telephony integration for the intended deployment.
The web quickstart describes adding voice conversations to an application through Vapi's Web and Server SDKs. That route is useful when the user is already inside a web product and does not need to dial a telephone number. Microphone permission and the application's own interface become part of the experience, including how the user starts, ends, and understands a call.
Choose the channel based on the user journey. A business's published support number has different requirements from a voice feature inside an authenticated account. For the latter, decide which user context should be passed and which account actions need independent authorization. For a phone call, plan how the assistant identifies the relevant record without assuming the incoming number proves the caller's identity.
Vapi's tool documentation covers built-in tools, API Request tools, and Function tools. These allow an assistant to retrieve information or perform an operation through an external service. An order-status assistant can ask a backend for the real order state; a scheduling assistant can check a calendar. The model supplies arguments, while the connected service supplies the authoritative result.
Define tool inputs narrowly. A booking tool should specify the slot identifier, user details required by the service, and any confirmation condition. Avoid asking the model to construct arbitrary operations when a limited action will do. Review what happens if the service returns an error, an empty result, or a response in an unexpected format. The assistant should describe what happened rather than turning a failed request into a confident success message.
Tool latency also changes the conversation. The user needs to understand that a lookup is in progress without hearing a false claim that the action has completed. Test the waiting message and recovery path with a deliberately slow or unavailable test service. A practical implementation separates finding an option from committing it, so the caller can confirm the details before an action changes a record.
Squads combine specialized assistants that hand off during a conversation while retaining context. One assistant can gather information, another handle scheduling, and another manage a distinct support task. The documentation recommends Handoff tools to define destinations and the conditions for moving between them. This approach is useful when one very large prompt becomes difficult to maintain.
A handoff should correspond to a real responsibility. If every specialist repeats the same questions, the caller experiences the architecture as friction. Define which facts transfer and which must be confirmed again. For example, a scheduling specialist may need the chosen location and service type, while an account-support specialist may need a different authorization process. The team should inspect the complete conversation, including transitions.
The current docs include a migration guide for legacy Workflows. New implementations should use the current supported primitives and avoid assuming that an older visual workflow tutorial represents today's recommended path. For an existing application, follow the documented migration rather than replacing configuration by guesswork. Test the behavior of each routing condition after a change.
The Transfer Call tool documents transferring a caller to another telephone destination. This is distinct from a Squad handoff to another AI assistant. Human transfer is relevant when a request falls outside the automation's scope, the caller asks for staff, or the application requires a person to resolve an exception. Review the supported transfer mode and telephony requirements for the deployment.
Design escalation as a normal outcome. A caller should not need to repeat a long explanation solely because the automation reached its boundary. Decide what context can be passed, which destination is available, and what to say when staff cannot take the call. A transfer attempt is not the same as a successful conversation with a staff member, so inspect the actual call result.
For a first rollout, make it easy to reach the fallback path and measure how often it is used. Repeated escalations around one topic may suggest a missing knowledge source or an intentionally unsupported task. The right response can be to improve the assistant or keep that task with staff. Automation coverage should follow the process's requirements rather than a desire to keep every caller inside AI.
Imagine a repair company that wants an inbound assistant for service appointments. Start with the supported locations, service types, and booking rules. Configure an Assistant with a greeting, questions to collect the necessary information, an availability lookup, and a booking tool. Add an explicit rule that the assistant can only offer slots returned by the service. Decide which situations require a human, such as a service outside the covered region.
Connect a test booking system and call the assistant from a real phone. Ask for a normal appointment, change the date halfway through, interrupt a response, and provide an ambiguous address. Check that the selected slot corresponds to the user's final choice. The confirmation should follow the successful tool result and repeat the important details in a format a caller can understand.
Then test failure paths: no availability, a timeout, a slot taken by another booking, and a requested human transfer. Review the call transcript and tool results together. Add representative simulations and publish a reviewed assistant version. Launch with a limited scope and inspect exceptions before expanding. This example uses Vapi's assistant, telephony, tools, testing, and transfer functions as connected parts of one real workflow.
Structured outputs let teams extract defined information from conversations. Instead of reading every transcript to find a booking preference or disposition, an application can work with selected fields. Treat extracted values as data to validate, particularly when recognition was uncertain or the caller revised an earlier answer. The final result should reflect the conversation's actual outcome.
Vapi documents call logs, recordings, analysis, Evals, and Simulations. Simulations use an AI tester to exercise conversations. They complement real voice testing but do not reproduce every accent, noisy room, telephone network, or caller behavior. Build coverage around the workflow's outcomes and exceptions, then check important interactions through the intended channel.
Inspect different kinds of evidence together. The transcript may show that the user requested a booking; the tool response may show it failed; a structured disposition might incorrectly say booked. Resolve such disagreements before using extracted results for follow-up automation. The purpose of testing is to establish the application's behavior on meaningful cases, not merely to obtain a high score from an easy conversation.
The phone quickstart documents draft changes and published assistant versions. A draft does not affect calls until it is published. The GitOps guide describes keeping assistants, squads, tools, and tests as files, reviewing changes, and promoting configuration between environments. This route is useful for teams that need the voice application to follow their regular development process.
A change to the model, transcriber, voice, or tool schema can alter behavior even when the system prompt remains unchanged. Keep a small set of representative calls and simulations for checking those changes. Review the voice rendering of important details and the results of connected actions. That is a maintenance recommendation based on Vapi's configurable stack.
Assign owners for credentials, telephony settings, capacity, and exception review. A production assistant depends on several services, so a failure may come from a provider, a backend, or application configuration. Logs help identify which layer needs attention. Keep the customer-facing recovery route understandable while the technical team investigates.
The current pricing documentation distinguishes usage-based billing, optional Success Packages, and add-ons. Usage charges include a hosting fee plus the selected model, voice, transcriber, and telephony costs. A Success Package adds support, service commitments, and included limits while usage billing continues. The documentation also notes separate arrangements for customers on legacy pricing.
Use the official pricing calculator with the intended component choices and expected call minutes. Estimate both ordinary and exceptional conversations. A caller who needs a long explanation or several tool attempts can create a different cost profile from a short information request. Package access, concurrent call lines, retention, and support commitments should be reviewed for the actual subscription.
Vapi is best classified as a paid usage platform rather than unlimited free calling. A free number or initial experimentation allowance does not establish that all calls, providers, or production capacity are free. Confirm current account terms, any compliance add-ons required by the deployment, and the cost of the connected business services.
Voice quality and task success depend on the configured providers, instructions, tools, and environment. Vapi's vendor performance claims should be assessed against the team's own call conditions. A tool-enabled assistant can still mishear a name or misunderstand an intent. External services can fail, and a transfer destination may be unavailable. Test these cases as part of the application.
Yes. The platform documents inbound and outbound calling. Current free Vapi numbers are described as inbound-only; outbound use needs a supported imported number and the relevant configuration.
Yes. The web integration path uses SDKs to add a voice conversation to an application. Include microphone permissions and the start and end states in usability testing.
A Squad handoff moves the conversation between configured assistants. A call transfer routes to a telephone destination. Choose and test the appropriate mechanism for an AI specialist or human escalation.
No. Current documentation states that usage billing remains active. Review the package's support and capacity terms separately from the per-call component costs.