Member Junction
    Preparing search index...

    Google Gemini implementation of BaseRealtimeClient: a browser-direct Gemini Live websocket session authenticated with the server-minted ephemeral auth token (v1alpha auth_tokens mechanism — the token is passed as the SDK apiKey).

    Registered with the ClassFactory under the key 'gemini' — the Provider string the server's GeminiRealtime driver stamps on its ClientRealtimeSessionConfig — so hosts resolve it without referencing this class directly.

    Owns ALL Gemini wire concerns (the behavioral twin of OpenAIRealtimeClient, adapted to Gemini's client-owned audio plane — there is no WebRTC here):

    • Audio in: mic PCM16 @ 16 kHz captured via an inline-Blob AudioWorklet (createMicCapture seam) and streamed with sendRealtimeInput.
    • Audio out: model PCM16 @ 24 kHz chunks scheduled through GeminiPcmPlayback (createPlayback seam); IsAudioPlaying is computed from the playout clock.
    • Event translation (Gemini → contract): inputTranscription → User deltas/finals, outputTranscription → Assistant deltas/finals (accumulated like the OpenAI driver), toolCall.functionCallsOnToolCall (callID→name cached for SendToolResult), interrupted → playback flush + 'listening', turnComplete → busy cleared + queued sends flushed, usageMetadataOnUsage (per-turn prompt/response token deltas — see handleUsageMetadata).
    • Busy mapping: Gemini has no response.created frame, so IsBusy is set EAGERLY when this client triggers a response (text / narration / tool result) and on the first model output of a turn (audio part or output-transcription delta); cleared on turnComplete, and on a toolCall frame (the model has yielded the floor pending the tool result — so a slow turnComplete can never deadlock the queued result).
    • Collision safety: ANY sendClientContent interrupts in-flight Gemini generation (per the Live API contract), so text / narration / context-note / tool-result sends issued while a turn is in flight are queued and flushed in order on turnComplete (the flush stops at the first send that starts a new response).
    • Open-turn commit: context notes ride as turnComplete: false client content, which tells the Live API MORE INPUT IS COMING — the server holds ALL generation (including the normally-automatic continuation after a tool response) until a turnComplete: true commit. SendToolResult therefore follows sendToolResponse with an empty-turn commit whenever a note left the turn open, so the model speaks the result immediately (the behavioral equivalent of the OpenAI driver's explicit response.create).
    • Triggering turns ride realtime text: typed text (SendText) and narration triggers (RequestSpokenUpdate) are sent via sendRealtimeInput({ text }) — the Live API's documented in-conversation text path. Native-audio Live models treat sendClientContent as initial-history seeding only: a mid-call turnComplete: true client turn appends to history WITHOUT starting generation (the model stays silent until the user's next spoken turn), while realtime text triggers an immediate response on every model generation. Context notes stay on sendClientContent (turnComplete: false) — the silent history-append is exactly the contract they want.
    • Narration tagging: RequestSpokenUpdate has no per-response-instructions equivalent on Gemini, so it is emulated as a realtime-text user turn carrying the instructions; the response kind is stamped 'narration' at send time (sends ARE the turn triggers on Gemini, unlike OpenAI where response.created confirms) and reset to 'normal' when the turn completes.

    Hierarchy (View Summary)

    Index

    Constructors

    Accessors

    • get IsAudioPlaying(): boolean

      Returns boolean

      directly from the playout engine's playhead clock — this client OWNS the output buffer (no WebRTC playback events exist on Gemini), so "audibly playing" is precisely "scheduled audio extends beyond the audio context's current time".

    • get IsBusy(): boolean

      true while a model response is in flight (generation started and not yet done). Distinct from IsAudioPlaying: generation runs ahead of playback. Hosts use this to gate interim narration so it never interrupts a reply.

      Returns boolean

    Methods

    • Returns void

      has no explicit cancel frame — the client OWNS the audio plane, so cancelling means: flush the local playout queue (GeminiPcmPlayback) so speech stops immediately, mark the in-flight turn inactive, and drain queued sends (a queued tool result or context note takes the floor next — tool-result delivery is never dropped by a cancel). Server-side, the next client content sent naturally interrupts any residual generation per the Live API contract. The interrupted turn's accumulated transcript is kept — the provider's trailing frames finalize what WAS spoken. No-op when nothing is active.

    • Opens the client-direct Gemini Live session: creates the playout engine, connects with the ephemeral token + the server-built SessionConfig ({ model, config } — the same values the server LOCKED into the token, so tampering is ignored by the API), then wires the mic-capture worklet. Reports 'listening' once audio is flowing.

      Parameters

      Returns Promise<void>

    • Creation seam for the mic-capture pipeline. Production delegates to the shared createPcmMicCapture (a 16 kHz AudioContext, inline-Blob capture worklet, and a zero-gain tail; each worklet block is PCM16-encoded and handed to onPcmChunk as base64). Unit tests override this with a no-op fake (and may capture onPcmChunk to simulate mic frames).

      Parameters

      • micStream: MediaStream
      • onPcmChunk: (base64Pcm16: string) => void

      Returns Promise<IPcmMicCapture>

    • Tears down the session, mic capture, mic tracks, and playout engine, resets the response state machine, and emits a final 'closed' (unless already 'error'). Safe to call more than once.

      Returns Promise<void>

    • Returns the AGENT's remote-audio MediaStream when this driver owns a tappable remote-audio plane (e.g. a WebRTC peer-connection driver routes the model's audio track here), or null otherwise. Hosts use it to MIX the agent's voice into a browser-side recording alongside the mic; a null return (the default, and what every non-WebRTC driver gives) degrades gracefully to mic-only capture.

      Optional capability: the method itself is optional — call sites must use client.GetRemoteMediaStream?.() ?? null. Drivers that can't expose a remote stream simply don't implement it (or return null).

      Returns MediaStream

    • Registers the (single) interruption handler.

      True barge-in only: fires ONLY when USER INPUT CUT OFF ACTIVE MODEL OUTPUT — a model response in flight or audio audibly playing when the user took the floor. A user simply taking their normal turn while the model is idle is NOT an interruption, and drivers must not report it as one (e.g. a raw "speech started" frame must be gated on whether a response is actually active or audio is playing).

      On interruption, drivers must also flush locally-owned playback and report IsAudioPlaying === false promptly (driver obligation #3). Hosts typically use this hook to abort in-flight delegated work so a stale result is never narrated into a conversation that has moved on.

      Parameters

      • handler: () => void

      Returns void

    • Optional capability: registers a handler invoked when the agent's remote-audio stream becomes available — immediately if it has already landed, otherwise when the WebRTC track arrives (typically AFTER Connect resolves). Lets a host attach the agent voice to a recording that began before the track landed. Call as client.OnRemoteMediaStream?.(cb); drivers without a remote stream simply don't implement it.

      Parameters

      • handler: (stream: MediaStream) => void

      Returns void

    • Registers the (single) remote-VIDEO handler — the model/avatar's video track for a VIDEO session (a talking-head the host renders, e.g. as the agent's tile). Invoked once the provider publishes its video track.

      Optional capability: audio-only drivers (the default) never emit — registering is always safe, but hosts must not assume a video track arrives. Video-capable drivers (BaseRealtimeModel.SupportsVideo) call emitRemoteVideo when the track is live.

      Parameters

      • handler: (stream: MediaStream) => void

        Invoked with the remote video MediaStream when it becomes available.

      Returns void

    • Registers the (single) usage handler.

      Emissions carry token deltas for the response/turn that just completed (see RealtimeClientUsage — deltas preferred; cumulative-only providers must convert in the driver). Hosts accumulate and relay/persist on their own cadence (e.g. the voice session service debounces a RelayRealtimeUsage mutation onto the co-agent prompt run).

      Optional capability: drivers whose provider exposes no usage telemetry simply never emit — registering a handler is always safe, but hosts must not assume emissions arrive. See RealtimeClientUsage for per-driver availability.

      Parameters

      Returns void

    • Triggers ONE short spoken update. Gemini has no per-response instructions (OpenAI's response.create.instructions), so this is EMULATED: the instructions ride as a realtime-text user turn (sendRealtimeInput({ text }) — the path that triggers generation on native-audio models, where mid-call sendClientContent is inert), and the resulting turn is stamped Kind: 'narration' at send time (reset on turnComplete) — mirroring the OpenAI driver's narration semantics. Queued behind any in-flight turn so it can never interrupt a pending reply.

      Parameters

      • instructions: string

      Returns void

    • Injects background context as a user turn with turnComplete: false — appended to the conversation WITHOUT starting generation, so the model draws on it the next time it speaks. Gemini Live turns have no system role, so the user role carries it (the host owns prefixing policy). Queued while a turn is in flight because ANY client content interrupts in-flight generation on Gemini (a divergence from the OpenAI driver, which can inject items mid-response safely).

      SIDE EFFECT: turnComplete: false leaves the client content turn OPEN — the server holds generation until a turnComplete: true commit arrives. openClientTurn tracks this so SendToolResult can commit the turn and unblock the spoken reply.

      Parameters

      • text: string

      Returns void

    • Injects typed text as a realtime-text user turn (sendRealtimeInput({ text }) — Gemini's "respond now" trigger on every Live model generation, including native-audio models that ignore mid-call sendClientContent). No-op when the session is not open.

      SendText implies barge-in (base-contract rule): an active spoken response is cancelled via CancelActiveResponse first — playback flushed, turn marked inactive, queued sends drained — so the typed turn takes the floor immediately. If a drained queued send (e.g. a pending tool result) starts a new turn, the text queues behind it, preserving the tool-result delivery invariant. On the wire, sending the user turn itself interrupts any residual server-side generation (Gemini Live's any-client-content-interrupts contract), so no explicit cancel frame exists or is needed.

      Parameters

      • text: string

      Returns void

    • Feeds an executed tool's result back via sendToolResponse, supplying the function name cached from the originating tool call (Gemini requires it; the contract only carries the callID). Sent immediately when idle — Gemini speaks the result as its next turn — otherwise queued until the in-flight turn (e.g. a progress narration) completes.

      If context notes left a client content turn OPEN (turnComplete: false), the tool response is followed by an empty-turn commit (sendClientContent({ turnComplete: true })) — without it the server keeps waiting for more client input and NEVER starts the spoken reply (observed live as the model staying silent after a delegated agent's result).

      Parameters

      • callID: string
      • outputJson: string

      Returns void

    • Mutes / unmutes by toggling the mic tracks' enabled flag: the capture pipeline stays up and streams SILENCE while muted (chosen over gating the worklet send so the provider's VAD sees a continuous stream and the un-mute is glitch-free — same policy as the OpenAI client driver).

      Parameters

      • muted: boolean

      Returns void