Member Junction
    Preparing search index...

    OpenAI implementation of BaseRealtimeClient: a browser-direct WebRTC connection to OpenAI's Realtime API, authenticated with the server-minted ephemeral client secret.

    Registered with the ClassFactory under the key 'openai' — the Provider string the server's OpenAIRealtime driver stamps on its ClientRealtimeSessionConfig — so hosts resolve it without referencing this class directly.

    The OpenAI-protocol brain (event translation, response state machine, narration-kind tagging, tool-result queueing, the outbound actions) lives in the shared OpenAIProtocolRealtimeClient; this class owns only the WebRTC transport:

    • Connect: mic tracks → peer connection, remote audio → hidden <audio> sink, the 'oai-events' data channel, and the GA SDP handshake (see performSdpHandshake).
    • Audible-playback tracking (IsAudioPlaying) from the provider-managed output_audio_buffer.* events (this transport does NOT own playback — the peer connection plays the remote track).
    • The remote-stream APIs a host uses to mix the agent's voice into a recording.

    Hierarchy (View Summary)

    Index

    Constructors

    Properties

    activeResponseKind: "normal" | "narration" = 'normal'

    The kind of the response currently in flight. Event ordering (confirmed against the live OpenAI API): response.created → transcript deltas → *_audio_transcript.doneresponse.done. The transcript-done frame therefore arrives while the kind is still set, letting onAssistantDone classify the turn; response.done then resets it.

    confirmedResponseActive: boolean = false

    True while a response the provider has CONFIRMED (response.created seen, response.done not yet) is in flight. Distinct from responseActive, which is set EAGERLY before a local response.create is even sent: if the provider REJECTS that create (an error frame with no response.created), responseActive would otherwise stay stuck true forever. This flag lets onErrorFrame tell "the eager flag is a phantom for a rejected create" (no confirmed response) from "a real response is genuinely active" (e.g. a concurrent VAD turn), so it clears the phantom without disturbing a live turn. Set on response.created, cleared on response.done.

    currentState: RealtimeClientState = 'closed'

    The client's own view of the session state — mirrors what emitStateChange last reported, EXCEPT after a tool call is emitted: the host typically shows its own busy state then, so the client silently leaves 'speaking' (no emission) to preserve the host's indicator until the result reply starts (see onToolCallFrame).

    dataChannel: IRealtimeDataChannel = null

    Protected so test subclasses can inspect/inject; production code treats it as private.

    micStream: MediaStream = null

    The mic stream owned by the current connection (used by the shared SetMuted).

    pendingAssistantText: string = ''

    Accumulates the in-flight assistant transcript across delta frames.

    pendingLocalResponseCreates: number = 0

    Count of locally-initiated response.creates whose response.created echo has not arrived yet. While non-zero, a response.done belongs to an EARLIER (typically just-cancelled) response — it must not clear the busy flag, flush the queued trigger, or flip the state for the response we just started (usage is still emitted). Consumed by response.created.

    pendingNarrationKind: boolean = false

    Set by RequestSpokenUpdate just before it sends its response.create, and CONSUMED by the very next response.created frame, which stamps activeResponseKind for that turn. Narration is only requested while the model is idle (hosts gate on IsBusy), so under normal ordering the next response.created is ours.

    pendingResultResponse: boolean = false

    Set when a tool result is ready while a response is active; sent on the next response.done.

    responseActive: boolean = false

    True while the model has a response in flight; gates narration + queues the tool result.

    sessionConfig: JSONObject = null

    The server-built session config applied verbatim via session.update when the data channel opens. Protected so test subclasses can seed it without a full Connect.

    Accessors

    • get IsAudioPlaying(): boolean

      true while model audio is AUDIBLY playing in the browser. Distinct from IsBusy — audio plays at realtime while generation finishes early, so the model can be "idle" while speech is still coming out of the speaker. Hosts must gate narration on BOTH, or queued utterances come out late and stale.

      Returns boolean

    • get IsBusy(): boolean

      true while a model response is in flight (generation started and not yet done). Distinct from IsAudioPlaying: generation runs ahead of playback. Hosts use this to gate interim narration so it never interrupts a reply.

      Returns boolean

    • get providerDebugLabel(): string

      Debug label used in the shared console diagnostics.

      Returns string

    • get releasesBusyFlagOnToolCall(): boolean

      Whether a completed tool call CLEARS responseActive. WebRTC keeps the flag (the provider reliably follows with response.done); websocket transports clear it as a deadlock guard so a queued SendToolResult can never wedge if the endpoint skips the trailing frame.

      Returns boolean

    Methods

    • Adopts + wires the events data channel: applies the session config and reports 'listening' on open; translates inbound frames through the shared protocol brain; reports transport errors / closure. Protected so test subclasses can inject a fake channel directly.

      Parameters

      Returns void

    • Returns void

      response.cancel (only when a response is actually in flight) and silences already-generated speech via the transport's stopAudioOutput. Resets the local response state machine (active flag, narration kind, accumulated transcript) but PRESERVES any queued tool-result trigger: delegated work is never affected by a floor-control cancel, and the queued trigger still fires on the cancelled response's trailing response.done. No-op when idle or when the transport is not open.

    • Handles one base64 PCM16 audio delta. Default: no-op — on the WebRTC transport the agent's audio rides the peer connection's remote track, not data-channel frames. Websocket transports override to enqueue into their local playout engine.

      Parameters

      • _deltaBase64: string

      Returns void

    • Registers the (single) interruption handler.

      True barge-in only: fires ONLY when USER INPUT CUT OFF ACTIVE MODEL OUTPUT — a model response in flight or audio audibly playing when the user took the floor. A user simply taking their normal turn while the model is idle is NOT an interruption, and drivers must not report it as one (e.g. a raw "speech started" frame must be gated on whether a response is actually active or audio is playing).

      On interruption, drivers must also flush locally-owned playback and report IsAudioPlaying === false promptly (driver obligation #3). Hosts typically use this hook to abort in-flight delegated work so a stale result is never narrated into a conversation that has moved on.

      Parameters

      • handler: () => void

      Returns void

    • Registers a handler invoked when the agent's remote-audio stream becomes available — either later via the WebRTC ontrack, or IMMEDIATELY if the track has already landed. Lets a host attach the agent voice to a recording that started (mic-only) before the track arrived.

      Parameters

      • handler: (stream: MediaStream) => void

      Returns void

    • Registers the (single) remote-VIDEO handler — the model/avatar's video track for a VIDEO session (a talking-head the host renders, e.g. as the agent's tile). Invoked once the provider publishes its video track.

      Optional capability: audio-only drivers (the default) never emit — registering is always safe, but hosts must not assume a video track arrives. Video-capable drivers (BaseRealtimeModel.SupportsVideo) call emitRemoteVideo when the track is live.

      Parameters

      • handler: (stream: MediaStream) => void

        Invoked with the remote video MediaStream when it becomes available.

      Returns void

    • The user started speaking. TRUE barge-in only when it cut off active model output (a response in flight or audio audibly playing) — a normal turn while the model is idle is NOT an interruption, so the emission is gated (base-contract rule). Transports that own their playback override to also flush the local playout queue.

      Returns void

    • Registers the (single) usage handler.

      Emissions carry token deltas for the response/turn that just completed (see RealtimeClientUsage — deltas preferred; cumulative-only providers must convert in the driver). Hosts accumulate and relay/persist on their own cadence (e.g. the voice session service debounces a RelayRealtimeUsage mutation onto the co-agent prompt run).

      Optional capability: drivers whose provider exposes no usage telemetry simply never emit — registering a handler is always safe, but hosts must not assume emissions arrive. See RealtimeClientUsage for per-driver availability.

      Parameters

      Returns void

    • Emits the user's spoken-input transcription. Default (OpenAI/HuggingFace): each .completed frame is one final caption. Providers that STREAM the completed event (Grok re-sends the full growing text each time) override to collapse the stream into a single in-place-updating bubble.

      Parameters

      • transcript: string

      Returns void

    • POSTs the raw SDP offer to OpenAI's Realtime WebRTC endpoint and returns the answer SDP.

      GA browser flow (confirmed against the OpenAI Realtime WebRTC guide): POST to https://api.openai.com/v1/realtime/calls with no query params and no OpenAI-Beta header. The ephemeral client secret already encodes the model + session config (set server-side at mint), so the browser must not specify the model — passing ?model= returns an empty 400. The answer comes back as raw application/sdp.

      Parameters

      • offerSdp: string

        The local SDP offer.

      • ephemeralToken: string

        The server-minted ephemeral client secret.

      Returns Promise<string>

      The answer SDP.

    • Asks the model to speak (a tool result or typed-text reply) — immediately if it's idle, otherwise queued until the current response finishes. An immediate trigger also CONSUMES any queued trigger debt: every payload item is already in the conversation, so one response.create voices everything (e.g. typed text barging in over a narration that had tool results queued behind it).

      Returns void

    • Triggers ONE short spoken update with the given instructions. Marks the upcoming response as 'narration' (flag consumed by the next response.created) so its transcripts are emitted with Kind: 'narration' — ephemeral by contract. Sets responseActive eagerly so a tool result landing mid-narration queues instead of colliding.

      Skips when busy (base-contract collision rule — drivers MUST queue or skip): a response.create sent while a response is in flight would be rejected/garbled by the provider, and narration is disposable by contract, so the update is dropped with a debug log rather than queued to come out late and stale. Hosts SHOULD still gate on IsBusy / IsAudioPlaying for timing quality.

      Parameters

      • instructions: string

      Returns void

    • Injects a system-role context item the model can draw on the next time it speaks, WITHOUT forcing a reply. Item creation is always safe mid-response.

      NOTE: role must be 'system' — gpt-realtime (and the compatible endpoints) reject 'developer' items ("Developer messages are only supported for quicksilver sessions").

      Parameters

      • text: string

      Returns void

    • Injects typed text as a user-role message conversation item, then triggers a reply through the SAME collision-safe path tool results use (requestResultResponse). No-op when the transport isn't open.

      SendText implies barge-in (base-contract rule): an active spoken response is cancelled via CancelActiveResponse before the text is injected, so the typed turn takes the floor immediately instead of waiting behind stale speech. When nothing is active the cancel is a no-op and the reply triggers immediately.

      Parameters

      • text: string

      Returns void

    • Sends the tool result back as a function_call_output conversation item, then triggers a reply — immediately if the model is idle, otherwise queued until the current response (e.g. a progress narration) finishes. Without the queueing the result's response.create would collide with an in-flight narration and be dropped, leaving the model silent when delegated work comes back.

      Parameters

      • callID: string
      • outputJson: string

      Returns void

    • Mutes / unmutes by toggling the mic tracks' enabled flag: the transport stays up and streams SILENCE while muted (the provider's VAD sees a continuous stream and the un-mute is glitch-free — the same policy across the client driver family).

      Parameters

      • muted: boolean

      Returns void

    • Returns void

      transport does NOT own playback (the peer connection plays the remote track), so a cancel asks the provider to stop + clear its managed output buffer via the WebRTC-only output_audio_buffer.clear client event — sent only when audio is audibly playing.