Skip to main content
← Journal·AI AutomationsOctober 7, 2026·12 min read

LiveKit Agents Guide 2026: Build Voice AI You Control

LiveKit Agents explained: how the open-source voice AI framework works, LiveKit Cloud pricing, phone calls over SIP, and when it beats Vapi on cost.

A

Founder of IImagined.ai

Quick answer

LiveKit Agents is an open-source Python and Node.js framework that runs your voice AI as a participant in a realtime room, with phone calls bridged in over SIP. Run it on LiveKit Cloud for a small per-minute fee, or host it yourself. It beats hosted platforms like Vapi when minute volumes are high and steady or when you need control over data and models.

LiveKit Agents is an open-source framework for building voice AI agents in Python or Node.js: your code joins a realtime LiveKit room as a participant and streams the caller's audio through speech-to-text, a language model and text-to-speech, or a realtime speech model. You can run it on LiveKit Cloud, which charges per agent session minute, or host the whole stack yourself under the Apache 2.0 license.

Facts and prices checked October 2026 against the LiveKit Agents documentation, LiveKit Cloud pricing, LiveKit billing docs, LiveKit telephony docs and Vapi pricing.

Hosted platforms such as Vapi, Retell and Bland get you a working phone agent fastest. LiveKit asks for more engineering in exchange for two things: a much smaller platform fee per minute and control over where audio runs and which models you use. This guide explains how the framework fits together, what LiveKit Cloud costs, and, with clearly labelled illustrative math, where the crossover sits against a hosted platform and against self-hosting. For the wider map of phone-agent options, start with our voice AI hub on building AI phone agents.

Who LiveKit Agents is for

  • Developers building a voice product that will run thousands of minutes a month and needs custom logic, tools and integrations in code.
  • Teams with data rules that want audio and transcripts on their own infrastructure, or a choice of model providers per region.
  • Agencies standardising on one stack for many client agents, where a per-minute markup across every client adds up.

It is not the best first step if you need a receptionist live this afternoon and nobody on the team writes Python or TypeScript. In that case, start with a hosted platform; our Bland AI vs Vapi vs Retell comparison covers the trade-offs, and you can move to LiveKit once volume justifies the work.

How LiveKit voice AI works end to end

LiveKit's documentation describes the agent as a stateful realtime bridge between AI models and users. Your agent code registers with a LiveKit server (Cloud or self-hosted) as an agent server. When a room is created, the server dispatches a job, a subprocess that joins the room. Users connect over WebRTC from a web or mobile app, or over the phone through LiveKit's SIP service. The agent talks to your backend over HTTP and WebSockets.

A LiveKit phone agent, end to end
  1. 01
    Caller dials your number

    Carrier or SIP trunk (Twilio, Telnyx) or a LiveKit Phone Number

  2. 02
    LiveKit SIP

    Creates a SIP participant; dispatch rules pick the room

  3. 03
    Room

    Realtime media between caller and agent

  4. 04
    Agent job

    Your Python or Node.js AgentSession: VAD, turn detection, interruptions

  5. 05
    Models

    STT, LLM, TTS, or one realtime speech model

  6. 06
    Your backend

    Tools call your CRM, calendar or database over HTTP

Inside the job, an AgentSession orchestrates the conversation. The framework handles the hard realtime parts: streaming audio through the pipeline, deciding when the caller has finished speaking with its turn detection model, stopping playback when the caller interrupts, and calling tools you define. Multi-agent handoffs let you split a flow into stages, for example a triage agent that hands off to a booking agent with the context preserved. Plugins cover most major model providers, and LiveKit Inference on Cloud lets you call models without managing separate API keys.

Build your first LiveKit agent: the steps

The official Voice AI quickstart gets a basic assistant running in Python or Node.js. The shape of the work, from blank folder to a phone number that answers, looks like this:

From zero to a phone agent
  1. 1
    Create a LiveKit Cloud project

    Free Build plan. Note the project URL, API key and secret.

  2. 2
    Scaffold an agent from the quickstart

    Python or Node.js. Expected result: a local agent server that registers with your project.

  3. 3
    Pick the pipeline

    STT, LLM and TTS plugins (or LiveKit Inference), or a single realtime speech model.

  4. 4
    Write instructions and tools

    System prompt, plus functions for booking, lookup or transfer.

  5. 5
    Test in the browser first

    Talk to the agent over WebRTC from a web frontend before touching phones.

  6. 6
    Add telephony

    Buy a LiveKit number or connect a SIP trunk; create an inbound trunk and a dispatch rule.

  7. 7
    Deploy

    Deploy to LiveKit Cloud, or run the agent server in your own containers.

Keep the first version narrow: one job, such as answering opening-hours questions and booking a slot, with a transfer to a human for everything else. Turn detection and latency matter more to callers than clever prompts, so test with real phone audio early. The STT and TTS choice also changes how the agent sounds and how quickly it replies.

Tools are where most of the value sits. A tool is a function your agent can call mid-conversation; the patterns are the same as LLM function calling elsewhere, which our guide to Claude API function calling walks through.

Phone calls: SIP, trunks and dispatch rules

LiveKit's telephony docs add two objects to rooms and participants. Trunks connect your SIP provider to LiveKit: inbound trunks accept calls and can be restricted by IP or number, and outbound trunks place calls. Dispatch rules decide which room each inbound caller lands in. Outbound calls are made by creating a SIP participant through the API. Supported features include DTMF, cold transfer via SIP REFER, warm transfer and caller ID. Connectors also bridge Twilio calls and WhatsApp Business calls into rooms without a SIP trunk.

If you place outbound sales calls with an AI voice, consent rules apply in the US: the FCC has confirmed that AI-generated voices count as artificial voices under the Telephone Consumer Protection Act. The outbound section of our platform comparison links the official FCC and FTC pages.

LiveKit Cloud pricing

From LiveKit's pricing page, checked October 2026. The billing docs add that agent session minutes are metered in one-minute increments from when the agent joins the room, and that on the free Build plan the allowance is a hard cap rather than an overage.

BuildShipScale
Monthly price$0From $50From $500
Agent session minutes1,000 included5,000 included, then $0.01/min50,000 included, then $0.01/min
Concurrent agent sessions520Starts at 50, up to 600
Third-party SIP minutes1,000 included5,000 included, then $0.004/min50,000 included, then $0.003/min
US local numbers1 free1 free, then $1.00/month1 free, then $1.00/month
Inference credits$2.50$5$50

Model costs sit on top, either through LiveKit Inference or your own provider keys. The pricing page lists per-minute equivalents for Inference models; for example, Deepgram Nova-3 at $0.0048 a minute for speech-to-text on Build and Ship, GPT-5.4 mini at $0.0030 a minute and Inworld Realtime TTS 2.0 Flash at $0.0090 a minute. Carrier charges from your SIP provider are separate too. Session recording and observability have their own allowances per plan.

LiveKit vs Vapi: platform fee and control

Vapi's pricing page lists $0.05 per minute for hosting, with models passed through at provider cost and telephony billed by the carrier (checked October 2026). LiveKit Cloud's comparable platform charges are $0.01 per agent session minute plus $0.003 to $0.004 per third-party SIP minute after the plan allowance. Both let you pick models, so the model bill can be held equal; the difference is the platform layer and how much you build yourself.

Vapi if
  • You want a dashboard-first build
  • Outbound campaigns from a CSV out of the box
  • Volume is modest or unpredictable
  • Nobody wants to own realtime infrastructure
LiveKit if
  • Engineers will own the agent in code
  • Minutes are high and steady
  • You may self-host for data control later
  • You also need web, mobile or video agents

The cost crossover: illustrative math

This section is an illustrative calculation from published rates, not a measurement. It compares only the platform layer, because model and carrier costs are similar whichever platform you choose. The self-hosted figure is an assumed input for servers, monitoring and the engineering time to run them; replace it with your own.

Minutes per monthVapi hostingLiveKit ShipLiveKit ScaleSelf-hosted (assumed)
10,000$500$120$500$4,000
100,000$5,000$1,380$1,150$4,000
500,000$25,000$6,980$6,350$4,000

What the table says, on these inputs. LiveKit Cloud's platform layer is cheaper than Vapi's hosting fee from low volumes, but at 10,000 minutes the $380 difference will not pay for much engineering, so convenience can rightly win. By 100,000 minutes a month the gap is several thousand dollars and the build effort starts to pay. Against an assumed $4,000 self-hosting budget, Vapi crosses over at 80,000 minutes (80,000 x $0.05), while LiveKit Scale only crosses at about 319,000 minutes ($500 plus 269,000 x $0.013). Ship's 20 concurrent sessions would also cap a busy line well before the higher volumes, which pushes you to Scale anyway.

Two caveats. First, the self-hosted line is the weakest number here: one engineer's time on call can exceed it, and a second region doubles parts of it. Second, data control is not a cost line. If a contract requires audio to stay inside your network, self-hosting can be the answer at any volume. For the open-source side of the stack in general, our best voice AI platforms roundup covers how LiveKit compares with the other developer options.

What self-hosting actually involves

LiveKit's pricing FAQ states that the Agents framework and the media server are open source and can run on your own infrastructure. The self-hosted SIP server docs show the moving parts: a LiveKit server, a SIP server and Redis, with the SIP signalling port 5060 and media ports 10000 to 20000 reachable from the internet. On top of that you run the agent servers, scale them for long-lived sessions, drain them during deploys and keep logs and transcripts somewhere.

Before you self-host LiveKit
  • Steady volume high enough that the crossover math holds with your own numbers
  • Someone on call for realtime infrastructure, not just the app
  • Firewall rules for SIP signalling and the media port range
  • Autoscaling that respects long sessions and drains before deploys
  • Logging, transcripts and recordings stored to your retention policy
  • A tested fallback route to a human when the agent pool is down

A common middle path is to build on LiveKit Cloud from day one, since the same agent code runs on either, and move to self-hosting only once the numbers and the team are there. If you already self-host other automation tools, our guide to running n8n with Docker Compose covers similar container and networking basics.

Mistakes that cost the most time

  • Testing only in the browser. Phone audio is narrower and noisier than WebRTC. Test over a real trunk before judging the STT and turn detection.
  • Leaving agents connected. Billing runs until the room ends or the agent disconnects; LiveKit's billing docs recommend calling ctx.shutdown() to end a session explicitly.
  • Choosing the biggest model by default. A larger model adds latency and cost per minute. Start small, then upgrade the parts of the flow that need it.
  • No human fallback. Every production line needs a transfer path and a plan for when the agent cannot help.

If you are turning voice agents into a product for clients, our AI SaaS Builder program covers scoping, pricing and shipping automation products.

LiveKit Agents: FAQ

What is LiveKit Agents?

LiveKit Agents is an open-source framework, released under the Apache 2.0 license, for building realtime voice, video and text AI agents in Python or Node.js. Your agent joins a LiveKit room as a participant and streams audio through a speech-to-text, language model and text-to-speech pipeline, or through a realtime speech model, with turn detection, interruptions and tool calls handled by the framework.

Is LiveKit free to use?

The framework and the LiveKit media server are open source, so you can run both on your own servers without a licence fee; you still pay for compute, carriers and model providers. LiveKit Cloud has a free Build plan with 1,000 agent session minutes a month as a hard cap, then paid Ship and Scale plans starting at $50 and $500 a month (checked October 2026).

LiveKit vs Vapi: which is cheaper?

On published rates, LiveKit Cloud's platform layer is cheaper per minute: $0.01 per agent session minute plus $0.003 to $0.004 per third-party SIP minute after plan allowances, against Vapi's $0.05 per minute hosting fee (checked October 2026). Model and carrier costs come on top of both. Vapi is quicker to configure, so the saving has to cover the extra engineering time.

Can LiveKit agents make and receive phone calls?

Yes. LiveKit telephony bridges phone networks into LiveKit rooms over SIP. You connect a SIP trunk from a provider such as Twilio or Telnyx, or buy US local and toll-free numbers through LiveKit Phone Numbers on LiveKit Cloud. Inbound calls are routed to rooms by dispatch rules; outbound calls are placed by creating a SIP participant through the API. Warm and cold transfers and DTMF are supported.

Should I self-host LiveKit or use LiveKit Cloud?

Use LiveKit Cloud until your volume or data rules force a change. Cloud handles agent scaling, draining during deploys, observability and global routing for a per-minute fee. Self-hosting removes that fee but adds servers, Redis, a SIP server with open ports, monitoring and on-call work. It pays off at high, steady minute volumes or when audio must stay inside your own network.

What languages can I write LiveKit agents in?

LiveKit's documentation covers Python and Node.js SDKs for the Agents framework, and LiveKit Agent Builder lets you prototype and deploy a voice agent in the browser without code. Python has the deepest set of examples and plugins. Your frontend can use LiveKit's client SDKs for web and mobile, or callers can reach the agent through a phone number with no app at all.

All Access · all four programs · $99/mo

Building on open infrastructure? Learn the whole automation stack.

All Access includes the AI SaaS Builder program alongside the other three programs, live coaching and the private community, so you can plan, price and ship automation products in one place.

Start All Access — $99/mo →30-day money-back guarantee
Free · no signup

Talk shop with other builders

Join the free Discord to compare voice AI stacks, LiveKit setups and what is working right now.

About the author

Written by Anyro, Founder of IImagined.ai. IImagined.ai is a founder-led education platform teaching Instagram growth, AI influencers, digital products, and AI automation.

Results vary; no income is guaranteed.

All-Access subscription

Every program. Member benefits.
One subscription.

Use all four premium programs with weekly live coaching, a private community, and the resource vault.

Confirm current lessons, downloadable resources and member-benefit arrangements before purchasing.

  • All 4 premium programs plus free Futures Trading
  • Weekly live coaching calls
  • Private community access
  • Resource vault and templates
  • 30-day money-back guarantee, cancel anytime
$99/ month
$99 for the first month · $702 to buy all four standalone
Start All-AccessOr browse standalone programs
30-day money-back guarantee · $99/month · cancel anytime