Member Junction
    Preparing search index...

    Module @memberjunction/ai-elevenlabs - v6.1.0

    @memberjunction/ai-elevenlabs

    MemberJunction AI provider for ElevenLabs. The package ships two drivers:

    • ElevenLabsAudioGeneratorBaseAudioGenerator implementation for text-to-speech, voice management, and pronunciation dictionaries (documented below).
    • ElevenLabsRealtimeBaseRealtimeModel driver for the ElevenLabs Agents Platform, powering realtime full-duplex voice sessions (see the next section).

    ElevenLabsRealtime (src/elevenLabsRealtime.ts, registered via @RegisterClass(BaseRealtimeModel, 'ElevenLabsRealtime')) exposes ElevenLabs' orchestrated STT→LLM→TTS Agents stack as a standard MJ realtime model, resolved for MJ: AI Models rows typed Realtime. ElevenLabsRealtimeSession is the IRealtimeSession backing the server-bridged topology; the matching browser-direct client driver (ElevenLabsRealtimeClient, ClassFactory key 'elevenlabs') ships in @memberjunction/ai-realtime-client.

    Concern How this provider does it
    What you connect to A pre-configured server-side agent — there is no bare-model realtime socket
    Managed-agent strategy RealtimeSessionParams.Model starting with agent_ is used verbatim (a deployment-managed agent); any other value is the name of the driver-managed agent: find-by-name → create-if-missing (with the session's client-tool set + per-session prompt-override enablement) → PATCH on order-insensitive tool-fingerprint drift; cached per name + fingerprint. The seeded metadata row's APIName is MJ Realtime Co-Agent
    Prompt authority Per-session: the managed agent stores only a placeholder prompt and enables the conversation_config_override.agent.prompt.prompt override; every session sends the real system prompt in its conversation_initiation_client_data frame
    Client-direct Supported — the signed websocket URL minted via GET /v1/convai/conversation/get-signed-url is the ephemeral credential (~15-minute open window; no API key leaves the server)
    Session-ready gate StartSession/Connect resolve only after conversation_initiation_metadata confirms the config is applied
    Transcripts Finals only (whole-utterance user_transcript / agent_response); agent_response_correction re-finalizes a barged-in agent turn with the text actually spoken
    Tool calling Inline client tools on the agent config (expects_response: true, 120 s timeout); client_tool_result carries structured JSON when parseable
    Tools mid-session Not re-declarable on an open conversation — an identical RegisterTools set no-ops; a different set warns and is ignored (the next session's ensure flow picks it up)
    Context notes (SendContextNote) Nativecontextual_update, the platform's purpose-built non-interrupting channel
    Narration (RequestSpokenUpdate) Emulated as a user_message turn, queued behind in-flight responses (fidelity caveat: the model may reference the instruction as a user message)
    Audio Base64 PCM16; rates negotiated from the initiation metadata (pcm_<rate>, platform default 16 kHz)
    Usage events None — the Agents websocket reports no token usage; OnUsage never fires (accounting lives in the platform dashboard)
    • API key env alias: AI_VENDOR_API_KEY__ElevenLabsRealtime
    • RealtimeSessionParams.Config.llm (string) selects the managed agent's underlying LLM; for full provider-side control, use a verbatim agent_… id
    • RealtimeSessionParams.Config.voice (string) sets the per-session ElevenLabs voice id, sent as the tts.voice_id conversation-config override on both topologies. voice is the driver-neutral key AssemblyAI and Inworld also read, and it reaches the driver from the effective config as realtime.voice.providers.elevenlabs.voice. Omit it and the agent's own configured voice is used, exactly as before. Blank or non-string values are ignored rather than sent
      • The managed agent enables this override automatically; an agent provisioned by an earlier MJ version is re-PATCHed on next use to enable it
    • RealtimeSessionParams.Config.firstMessage (string) makes the agent speak first, sent as the agent.first_message conversation-config override on both topologies. Spoken VERBATIM — it is the literal opening utterance, not an instruction about how to open. Authorable on the persona as realtime.voice.default.firstMessage, from where it reaches the driver on the neutral firstMessage key. Omit it (or leave it blank) and the agent waits for the user to speak first, exactly as before
      • Same enablement rule as the voice override: written on every managed agent this driver provisions, and an agent provisioned by an earlier MJ version is re-PATCHed on next use rather than silently dropping it
      • Conversation-start behavior is not prompt-addressable on this provider: with no first_message, the agent produces no audio until it receives user audio, whatever the persona prompt says
    • InitialContext is injected as a contextual_update once the session is confirmed (the protocol has no history-seeding channel)

    For the full architecture (topologies, co-agent model, capability matrix across all four realtime providers), see guides/REALTIME_CO_AGENTS_GUIDE.md.


    This driver implements the BaseAudioGenerator interface to provide high-quality voice synthesis, voice management, and pronunciation dictionary support.

    graph TD
        A["ElevenLabsAudioGenerator
    (Provider)"] -->|extends| B["BaseAudioGenerator
    (@memberjunction/ai)"] A -->|wraps| C["ElevenLabsClient
    (elevenlabs-js SDK)"] C -->|provides| D["Text-to-Speech"] C -->|provides| E["Voice Management"] C -->|provides| F["Model Listing"] C -->|provides| G["Pronunciation
    Dictionaries"] B -->|registered via| H["@RegisterClass"] style A fill:#7c5295,stroke:#563a6b,color:#fff style B fill:#2d6a9f,stroke:#1a4971,color:#fff style C fill:#2d8659,stroke:#1a5c3a,color:#fff style D fill:#b8762f,stroke:#8a5722,color:#fff style E fill:#b8762f,stroke:#8a5722,color:#fff style F fill:#b8762f,stroke:#8a5722,color:#fff style G fill:#b8762f,stroke:#8a5722,color:#fff style H fill:#b8762f,stroke:#8a5722,color:#fff
    • Text-to-Speech: High-quality voice synthesis with customizable voice settings
    • Voice Management: List and browse available voices with labels and preview URLs
    • Model Discovery: Query available audio models with capability metadata
    • Pronunciation Dictionaries: Manage custom pronunciation dictionaries with paginated retrieval
    • Streaming Audio: Audio output returned as base64-encoded buffers
    • Text Normalization: Optional text normalization for improved speech output
    npm install @memberjunction/ai-elevenlabs
    
    import { ElevenLabsAudioGenerator } from '@memberjunction/ai-elevenlabs';

    const tts = new ElevenLabsAudioGenerator('your-elevenlabs-api-key');

    const result = await tts.CreateSpeech({
    text: 'Hello, welcome to MemberJunction!',
    voice: 'voice-id-here',
    model_id: 'eleven_turbo_v2'
    });

    if (result.success) {
    // result.content contains base64-encoded audio
    // result.data contains raw Buffer
    console.log('Audio generated successfully');
    }
    const voices = await tts.GetVoices();
    for (const voice of voices) {
    console.log(`${voice.name} (${voice.id}): ${voice.category}`);
    }
    const models = await tts.GetModels();
    for (const model of models) {
    console.log(`${model.name}: TTS=${model.supportsTextToSpeech}`);
    }
    Method Description
    CreateSpeech Convert text to speech audio
    GetVoices List available voices
    GetModels List available audio models
    GetPronounciationDictionaries List pronunciation dictionaries
    • SpeechToText is not yet implemented
    • ElevenLabsAudioGenerator via @RegisterClass(BaseAudioGenerator, 'ElevenLabsAudioGenerator')
    • ElevenLabsRealtime via @RegisterClass(BaseRealtimeModel, 'ElevenLabsRealtime')
    • @memberjunction/ai - Core AI abstractions (BaseAudioGenerator, BaseRealtimeModel)
    • @memberjunction/global - Class registration
    • @elevenlabs/elevenlabs-js - Official ElevenLabs SDK (REST agent management + TTS; the realtime conversation websocket is spoken raw — the SDK's high-level wrapper owns audio devices, which a server bridge must not)

    Classes

    ElevenLabsAudioGenerator
    ElevenLabsRealtime
    ElevenLabsRealtimeSession

    Interfaces

    ElevenLabsConnectArgs
    ElevenLabsRealtimeSocket
    ElevenLabsServerEvent

    Variables

    ELEVENLABS_SUPPORTED_TURN_SETTINGS

    Functions

    MapTurnEagerness
    SanitizeToolParametersForElevenLabs