VoiceGateway

Costs and reconciliation

How a recorded cost is computed, what a zero means, and how to check the numbers against a provider's invoice.

VoiceGateway records an estimate: usage units times a catalog rate, at the moment of the call. Your provider bills from its own meter, with your plan, discounts and credits. The two will not match to the cent, and this page says by how much they should differ.

Every cost names its source

Rates come from voice-prices, in US dollars. Each row stores the units it was priced from and a pricing source, so a number is never anonymous:

Pricing sourceMeaning
voice-prices@<version>The catalog priced every unit on the row
voicegateway-localA self-hosted model (local/*, ollama/*). Zero, because it runs on your hardware
voice-prices-unratedThe catalog knows the model but has no rate for these units. The zero is not a price
emptyThe catalog does not know the model

A zero is only a price when the source says so. The other zeros are gaps: voicegw prices gaps lists the unpriced models, heaviest traffic first, and voicegw prices set declares a rate for one.

Rates move. Upgrading the catalog is a package upgrade:

pip install --upgrade voice-prices

LiveKit session totals

At session close, VoiceGateway compares cumulative usage with the individual metric rows and records only missing usage. Provider display names and the standard OpenAI API hostname resolve to the same provider IDs as those rows. Custom endpoint hostnames stay distinct. A different spelling alone does not create a second charge.

Check against the invoice

voicegw reconcile compares recorded usage with a provider's usage export, per model, on units and on cost. It supports three providers today, one per modality: openai (LLM), deepgram (STT) and cartesia (TTS).

Download the provider's usage export for the window and reshape it into a CSV or JSON file with one row per model:

ProviderColumns
openaimodel, input_tokens, output_tokens, cost_usd, n_requests
deepgrammodel, audio_seconds, cost_usd, n_requests
cartesiamodel, characters, cost_usd, n_requests

Then run it for the same window:

voicegw reconcile --provider openai \
  --start 2026-05-01 --end 2026-05-31 \
  --provider-usage-file openai-may.csv

A row is flagged when its cost differs by more than --threshold percent (default 5). Add --format csv or --format json for machine-readable output.

Reading the difference

ModalityExpected driftLook closer at
LLMwithin 5%more than 5% on cost, or any drift on units
STTwithin 1%more than 2% on cost, or any drift on units
TTSwithin 2%more than 3% on cost, or any drift on units
  • Units agree, cost differs. The rate moved, or your account has a price the public catalog does not. Upgrade voice-prices, or declare your rate with voicegw prices set.
  • Units differ. A request ended before its usage was reported, or the provider bills a unit the catalog splits differently.
  • A model on one side only. Provider exports often lag 24 to 72 hours, so re-pull later. If VoiceGateway has no rows for it, another client may share the API key.

Reconcile once in the first month, after a provider changes its rates, and before you pass costs on to anyone. It is a spot check, not a monthly chore.

On this page