Member Junction
    Preparing search index...

    Scriptable, call-recording LLM driver extending the real BaseLLM.

    Scripted outcomes are consumed one per ChatCompletion call, in order. When the script is exhausted the default outcome is used (or, with RepeatLastOutcome, the final scripted outcome repeats). Every call's real ChatParams is recorded on Calls.

    Hierarchy (View Summary)

    Index

    Constructors

    Properties

    _additionalSettings: Record<string, any>

    Protected property to store additional provider-specific settings

    RepeatLastOutcome: boolean = false

    When true and the script is down to its final outcome, that outcome keeps repeating for every subsequent call instead of falling back to the default outcome ("advance through the script, then repeat the last entry").

    thinkingStreamState: ThinkingStreamState | null

    State tracking for streaming thinking extraction Providers should initialize this if they support thinking models

    Accessors

    • get AdditionalSettings(): Record<string, any>

      Get the current additional settings

      Returns Record<string, any>

    • get apiKey(): string

      Only sub-classes can access the API key

      Returns string

    • get SupportsPrefill(): boolean

      Whether this LLM provider supports assistant prefill (pre-seeding the start of the model's response). Providers that support prefill should override this to return true. This is used as a code-level default when database metadata (AIModelType/AIModel/AIModelVendor.SupportsPrefill) is null. Database values of true/false override this getter.

      Returns boolean

    • get SupportsStreaming(): boolean

      Check if this provider supports streaming

      Returns boolean

      true if streaming is supported, false otherwise

    • get SupportsTools(): boolean

      Whether this driver implements native tool/function calling — i.e. whether it maps ChatParams.tools onto its SDK and normalizes tool calls back into ChatCompletionMessage.toolCalls.

      This is a CODE-level capability ("has the mapping been written for this driver?"), distinct from the METADATA-level capability ModelConfiguration.LLM.SupportsNativeToolCalling ("does this model, on this vendor, support tools at all?"). Both must hold for native mode.

      A driver returning false ignores any tools passed to it and records the fact in the result's modelSpecificResponseDetails — the prompt runner's gate should keep that from happening, but the layer is safe standalone.

      Returns boolean

    Methods

    • Process multiple chat completion requests in parallel. This is useful for:

      • Generating multiple variations with different parameters (temperature, etc.)
      • Getting multiple responses to compare or select from
      • Improving reliability by sending the same request multiple times

      Parameters

      Returns Promise<ChatResult[]>

      Promise resolving to an array of ChatResults in the same order as the input params

    • Clear all additional settings This is useful for resetting the state of the provider or when switching between different configurations.

      Returns void

    • Extract thinking content from non-streaming content This method handles case-insensitive extraction of thinking blocks

      Parameters

      • content: string

      Returns { content: string; thinking?: string }

    • Create the final response object from streaming results

      Parameters

      • accumulatedContent: string | null | undefined

        The complete content accumulated from all chunks

      • _lastChunk: string | null | undefined

        The last chunk received from the stream

      • _usage: ModelUsage | null | undefined

        The usage information (tokens, etc.)

      Returns ChatResult

      A complete ChatResult object

    • Returns (and clears) any user-visible content the thinking-tag stripper is still holding back at the end of a stream. Mid-stream, processStreamChunkWithThinking holds back a trailing fragment that could be the start of a <think>/</think> tag so a split tag never leaks as partial text; once the stream ends, such a fragment is real content and must be emitted (bug A5).

      Only flushes when NOT inside a thinking block: an unterminated <think> block's buffered text is reasoning, not answer, and is left held back (never surfaced as visible content). Returns '' when thinking extraction isn't active (no state) or there's nothing to flush.

      Returns string

    • Get the thinking tag format for this provider Providers can override this to customize the thinking tag format

      Returns { close: string; open: string }

    • Template method for handling streaming chat completion This implements the common pattern across providers while delegating provider-specific logic to abstract methods.

      Parameters

      Returns Promise<ChatResult>

    • Initialize thinking stream state for streaming extraction

      Returns void

    • Process streaming chunk with thinking extraction This method handles case-insensitive extraction across chunk boundaries

      Parameters

      • rawContent: string

      Returns string

    • Hook invoked at the start AND end (in finally) of every streaming chat completion to reset per-request streaming state. Default is a no-op; providers that maintain instance-level streaming state (e.g., Anthropic / OpenAI thinking-block accumulators) MUST override this. See audit R2-C5 for context — without this, state from a prior request bleeds into the next one and accumulated buffers grow unbounded across the singleton's lifetime.

      Returns void

    • Reset thinking stream state

      Returns void

    • Set additional provider-specific settings Subclasses should override this method to validate required settings

      Parameters

      • settings: Record<string, any>

        Provider-specific settings

      Returns void

    • Sets the code-level prefill default this driver reports.

      Parameters

      • value: boolean

      Returns this

    • Enables/disables the real BaseLLM streaming path for this driver.

      Parameters

      • value: boolean

      Returns this

    • Check if the provider supports thinking models Providers should override this to return true if they support thinking extraction

      Returns boolean