VoiceGateway

attach()

The one meter. Binds to a LiveKit AgentSession or a Pipecat PipelineTask and records cost, latency and usage for every STT, LLM and TTS call.

attach() watches. It never reroutes, throttles or blocks a call; control lives in guard(). Call it once per session, before the session starts.

voicegateway.attach(
    session,                        # LiveKit AgentSession or Pipecat PipelineTask
    *,
    project: str | None = None,     # else VOICEGW_PROJECT, else "default"
    agent_id: str | None = None,    # else VOICEGW_AGENT_ID, else the hostname
    tenant_id: str | None = None,   # per-customer attribution, optional
    revision: str | None = None,    # which build of this agent; else VOICEGW_AGENT_REVISION
    channel: str | None = None,     # "telephony" or "web"; detected from the transport
    policy: str | None = None,      # what to capture beyond cost; see below
    collector_url: str | None = None,  # push to a shared collector; else VOICEGW_COLLECTOR_URL
    api_key: str | None = None,        # the collector's key; else VOICEGW_API_KEY
) -> str                            # the session id stamped on every row

It returns immediately. There is nothing to await.

One row per request

FieldWhat it holds
modality, provider, modelstt, llm or tts, and the model id your plugin reported
unitsSTT audio seconds, LLM prompt, completion and cached tokens, TTS characters
costUSD, priced from those units through voice-prices
latencytime to first byte and total
correlationsession id, project, agent, tenant, revision, channel

LLM and TTS units come from the framework's own usage metrics. STT is derived from audio duration. On Pipecat, that means counting the audio frames each STT service receives.

attach() picks its framework from the type of the object you pass. Anything that is not a real AgentSession or PipelineTask records nothing, silently.

Realtime and GPT-Live

LiveKit realtime models emit metrics on AgentSession. VoiceGateway preserves the provider and model on each metric, so a GPT-Live voice session and its Responses backend appear as separate services.

GPT-Live reports incremental session seconds, including a final usage update at close. These become audio_seconds in accounting quantities. VoiceGateway reconciles against cumulative session.usage and deduplicates repeated metric IDs. Keep the normal session shutdown path so the final usage can arrive.

For the exact openai/gpt-live-1 model, a dated fallback estimates voice cost at $0.05 per minute, billed per second, from OpenAI's model pricing (checked October 5, 2026). Backend model tokens are additional. The exact openai/gpt-5.6-luna backend also has a dated official fallback: $0.20 per million input tokens, $0.02 per million cache reads, $0.25 per million cache writes, and $1.20 per million output tokens. Prompts above 272,000 input tokens apply the published 2x input and 1.5x output multipliers. Native cache-write counts are preserved; reasoning tokens are already included in output tokens. This estimate excludes transport, hosting, taxes, and account-specific discounts. Future model names do not inherit this rate.

Token-based realtime models keep text, audio, and cached usage separate. Missing quantities or rates set metadata.pricing_complete to false; a legacy zero cost column then means unpriced usage, not free usage. InsideCall.snapshot() exposes allowlisted measurements and pricing_complete for visitor interfaces.

Unavailable native first-token times, including -1, remain null. They are not zero-latency observations or end-to-end response measurements. Legacy fixed-token sell-price cards cannot represent duration or mixed audio pricing: those rows are marked unsupported-realtime-fixed:pass-through. Use the dimension-specific accounting ledger for those contracts. Cost-plus cards still apply.

Revision separates builds

revision names the build of this agent's configuration: the prompt, the models, the voice. Pass a git sha, a content hash or a semver string. Two revisions live in the same window stay separate instead of blending into one p95. Leave it unset rather than inventing one; a made-up revision splits numbers that belong together.

Policy names what to capture

Cost and latency are always recorded. Everything else is a choice, and policy names the set:

PolicyTranscriptTurn timingDead airSnapshots
standard (default)onononoff
timing_onlyoffononoff
leanoffonoffoff
debugonononon
offoffoffoffoff

A transcript is what the caller said. A snapshot is your own system prompt and every tool payload, which is why only debug carries one. An unknown policy name raises instead of falling back to standard: a typo must not turn "nothing the caller said" into transcript capture.

Environment kill-switches beat every policy: VOICEGW_TRANSCRIPTS=0, VOICEGW_TURNS=0, VOICEGW_DEAD_AIR=0, VOICEGW_SNAPSHOTS=0. Transcripts, turns, dead air and snapshots are captured on LiveKit today; Pipecat accepts the flags and records cost and latency only.

Where rows go

Local SQLite by default, at ~/.config/voicegateway/voicegw.db (override with VOICEGW_DB_PATH). Set VOICEGW_COLLECTOR_URL and VOICEGW_API_KEY and rows are batched to a shared collector instead. attach() reads only these environment variables, never voicegw.yaml.

On this page