Costs and reconciliation
How a recorded cost is computed, what a zero means, and how to check the numbers against a provider's invoice.
VoiceGateway records an estimate: usage units times a catalog rate, at the moment of the call. Your provider bills from its own meter, with your plan, discounts and credits. The two will not match to the cent, and this page says by how much they should differ.
Every cost names its source
Rates come from voice-prices, in US dollars. Each row stores the units it was priced from and a pricing source, so a number is never anonymous:
| Pricing source | Meaning |
|---|---|
voice-prices@<version> | The catalog priced every unit on the row |
voicegateway-local | A self-hosted model (local/*, ollama/*). Zero, because it runs on your hardware |
voice-prices-unrated | The catalog knows the model but has no rate for these units. The zero is not a price |
| empty | The catalog does not know the model |
A zero is only a price when the source says so. The other zeros are gaps: voicegw prices gaps lists the unpriced models, heaviest traffic first, and voicegw prices set declares a rate for one.
Rates move. Upgrading the catalog is a package upgrade:
pip install --upgrade voice-pricesLiveKit session totals
At session close, VoiceGateway compares cumulative usage with the individual metric rows and records only missing usage. Provider display names and the standard OpenAI API hostname resolve to the same provider IDs as those rows. Custom endpoint hostnames stay distinct. A different spelling alone does not create a second charge.
Check against the invoice
voicegw reconcile compares recorded usage with a provider's usage export, per model, on units and on cost. It supports three providers today, one per modality: openai (LLM), deepgram (STT) and cartesia (TTS).
Download the provider's usage export for the window and reshape it into a CSV or JSON file with one row per model:
| Provider | Columns |
|---|---|
openai | model, input_tokens, output_tokens, cost_usd, n_requests |
deepgram | model, audio_seconds, cost_usd, n_requests |
cartesia | model, characters, cost_usd, n_requests |
Then run it for the same window:
voicegw reconcile --provider openai \
--start 2026-05-01 --end 2026-05-31 \
--provider-usage-file openai-may.csvA row is flagged when its cost differs by more than --threshold percent (default 5). Add --format csv or --format json for machine-readable output.
Reading the difference
| Modality | Expected drift | Look closer at |
|---|---|---|
| LLM | within 5% | more than 5% on cost, or any drift on units |
| STT | within 1% | more than 2% on cost, or any drift on units |
| TTS | within 2% | more than 3% on cost, or any drift on units |
- Units agree, cost differs. The rate moved, or your account has a price the public catalog does not. Upgrade
voice-prices, or declare your rate withvoicegw prices set. - Units differ. A request ended before its usage was reported, or the provider bills a unit the catalog splits differently.
- A model on one side only. Provider exports often lag 24 to 72 hours, so re-pull later. If VoiceGateway has no rows for it, another client may share the API key.
Reconcile once in the first month, after a provider changes its rates, and before you pass costs on to anyone. It is a spot check, not a monthly chore.