Introduction
What VoiceGateway measures, the two calls it adds to your agent, and when something else is the better fit.
A voice agent runs three meters at once. STT bills by audio seconds, the LLM by tokens, TTS by characters. No provider dashboard shows what one conversation cost across all three.
VoiceGateway does. Add one line to a LiveKit or Pipecat agent and every STT, LLM and TTS call is priced, timed and stored, per call and per project. It runs on your machine, with your keys. Nothing leaves your infrastructure.
from livekit.agents import AgentSession
from livekit.plugins import cartesia, deepgram, openai
import voicegateway
session = AgentSession(
stt=deepgram.STT(model="nova-3"),
llm=openai.LLM(model="gpt-4o-mini"),
tts=cartesia.TTS(model="sonic-3"),
)
voicegateway.attach(session)Two calls, two jobs
| Call | Job | What it does |
|---|---|---|
attach(session) | observe | Meters every provider call on the session: cost, latency, units. Never touches a request. |
guard(provider) | control | Wraps one provider with fallback, a rate limit and a spend cap. Writes no metrics. |
attach() is the only meter, so using both never double-counts. Most agents start with attach() alone.
Where the numbers come from
Prices come from voice-prices, an open catalog of STT, LLM and TTS rates. Every recorded row names the catalog entry that priced it. When a number has to be exact, voicegw reconcile checks it against the provider's own usage export.
When something else fits better
Building a text-only LLM app with no voice? LiteLLM is the better fit. VoiceGateway is for agents that listen and speak.
VoiceGateway is alpha software. Public APIs can change between minor releases, and the changelog says when they do.