attach()
The one meter. Binds to a LiveKit AgentSession or a Pipecat PipelineTask and records cost, latency and usage for every STT, LLM and TTS call.
attach() watches. It never reroutes, throttles or blocks a call; control lives in guard(). Call it once per session, before the session starts.
voicegateway.attach(
session, # LiveKit AgentSession or Pipecat PipelineTask
*,
project: str | None = None, # else VOICEGW_PROJECT, else "default"
agent_id: str | None = None, # else VOICEGW_AGENT_ID, else the hostname
tenant_id: str | None = None, # per-customer attribution, optional
revision: str | None = None, # which build of this agent; else VOICEGW_AGENT_REVISION
channel: str | None = None, # "telephony" or "web"; detected from the transport
policy: str | None = None, # what to capture beyond cost; see below
collector_url: str | None = None, # push to a shared collector; else VOICEGW_COLLECTOR_URL
api_key: str | None = None, # the collector's key; else VOICEGW_API_KEY
) -> str # the session id stamped on every rowIt returns immediately. There is nothing to await.
One row per request
| Field | What it holds |
|---|---|
| modality, provider, model | stt, llm or tts, and the model id your plugin reported |
| units | STT audio seconds, LLM prompt, completion and cached tokens, TTS characters |
| cost | USD, priced from those units through voice-prices |
| latency | time to first byte and total |
| correlation | session id, project, agent, tenant, revision, channel |
LLM and TTS units come from the framework's own usage metrics. STT is derived from audio duration. On Pipecat, that means counting the audio frames each STT service receives.
attach() picks its framework from the type of the object you pass. Anything that is not a real AgentSession or PipelineTask records nothing, silently.
Realtime and GPT-Live
LiveKit realtime models emit metrics on AgentSession. VoiceGateway preserves the
provider and model on each metric, so a GPT-Live voice session and its Responses
backend appear as separate services.
GPT-Live reports incremental session seconds, including a final usage update at
close. These become audio_seconds in accounting quantities. VoiceGateway
reconciles against cumulative session.usage and deduplicates repeated metric IDs.
Keep the normal session shutdown path so the final usage can arrive.
For the exact openai/gpt-live-1 model, a dated fallback estimates voice cost at
$0.05 per minute, billed per second, from
OpenAI's model pricing
(checked October 5, 2026). Backend model tokens are additional. The exact openai/gpt-5.6-luna
backend also has a dated official fallback: $0.20 per million input tokens,
$0.02 per million cache reads, $0.25 per million cache writes, and $1.20 per
million output tokens. Prompts above 272,000 input tokens apply the published
2x input and 1.5x output multipliers. Native cache-write counts are preserved;
reasoning tokens are already included in output tokens. This estimate
excludes transport, hosting, taxes, and account-specific discounts. Future model
names do not inherit this rate.
Token-based realtime models keep text, audio, and cached usage separate. Missing
quantities or rates set metadata.pricing_complete to false; a legacy zero cost
column then means unpriced usage, not free usage. InsideCall.snapshot() exposes
allowlisted measurements and pricing_complete for visitor interfaces.
Unavailable native first-token times, including -1, remain null. They are not
zero-latency observations or end-to-end response measurements. Legacy fixed-token
sell-price cards cannot represent duration or mixed audio pricing: those rows are
marked unsupported-realtime-fixed:pass-through. Use the dimension-specific
accounting ledger for those contracts. Cost-plus cards still apply.
Revision separates builds
revision names the build of this agent's configuration: the prompt, the models, the voice. Pass a git sha, a content hash or a semver string. Two revisions live in the same window stay separate instead of blending into one p95. Leave it unset rather than inventing one; a made-up revision splits numbers that belong together.
Policy names what to capture
Cost and latency are always recorded. Everything else is a choice, and policy names the set:
| Policy | Transcript | Turn timing | Dead air | Snapshots |
|---|---|---|---|---|
standard (default) | on | on | on | off |
timing_only | off | on | on | off |
lean | off | on | off | off |
debug | on | on | on | on |
off | off | off | off | off |
A transcript is what the caller said. A snapshot is your own system prompt and every tool payload, which is why only debug carries one. An unknown policy name raises instead of falling back to standard: a typo must not turn "nothing the caller said" into transcript capture.
Environment kill-switches beat every policy: VOICEGW_TRANSCRIPTS=0, VOICEGW_TURNS=0, VOICEGW_DEAD_AIR=0, VOICEGW_SNAPSHOTS=0. Transcripts, turns, dead air and snapshots are captured on LiveKit today; Pipecat accepts the flags and records cost and latency only.
Where rows go
Local SQLite by default, at ~/.config/voicegateway/voicegw.db (override with VOICEGW_DB_PATH). Set VOICEGW_COLLECTOR_URL and VOICEGW_API_KEY and rows are batched to a shared collector instead. attach() reads only these environment variables, never voicegw.yaml.