Skip to main content
@komaa/standin-sdk/elevenlabs puts ElevenLabs on a real Microsoft Teams call. StandIn answers the call and dials your worker; this plugin answers that dial, opens one session per call, and relays the audio both ways.

Install and run

There is nothing to install beyond the SDK. ElevenLabs is reached over an ordinary WebSocket, so this plugin adds no dependency.
Expose port 9442 at the /msteams/calling path and register the public wss:// URL as your StandIn identity’s agent voice URL. The full walkthrough is in the Quickstart, and there is a runnable example at examples/elevenlabs-msteams-connector.

The audio format

In the ElevenLabs dashboard, set the agent’s input and output audio format to pcm_16000. An agent set to anything else is refused at the first frame rather than producing a whole call of garbled audio.

What the agent can do about the call

Declare these as client tools on the agent. Nothing to implement: the plugin answers them. clientTools() returns the declarations, so there is nothing to retype into the dashboard:
show_image is widened here, because ElevenLabs is the one provider that can hand you the bytes inline: either form will do, so neither argument is required. A URL is chosen by the model, which the caller steers, so it is fetched through the SDK’s guard: public hosts only, no private or link-local addresses, and the address is re-checked at connect time to close the DNS rebind window.
The fullscreen or overlay choice is not selectable on this plugin. The other provider plugins route show_image through CallTools.dispatch, which honours a display parameter. This one has its own dispatch, and it reads mode. clientTools() still emits display with its fullscreen / overlay enum, because that comes from the shared tool spec, so the declaration you paste will offer the model a choice that is then dropped in silence. Deleting the display property from the pasted declaration changes nothing except that the model stops being offered it, which is the honest version. Either way the picture goes up with the default placement. Fullscreen or beside your face covers what the two placements are for.
look and look_back upload the frame into the ElevenLabs conversation, which stores the caller’s screen with a third party. They therefore work only while the Microsoft Teams call is being recorded. When it is not, the agent is told why rather than being left with silence.

Inside your own worker

The plugin is a CallHandler like any other:
This handler takes its configuration as a positional argument, where the other provider plugins take an options object. Read it once, then close over it: handlerFactory runs once per call, so new ElevenLabsHandler() with nothing passed reads the environment again on every call, and a key removed after startup would fail the next caller rather than failing you. serve(), which is what npx standin-elevenlabs runs, does exactly this, then waits for SIGINT or SIGTERM before calling server.aclose(). server.start() returns as soon as the listener is bound, so a script that ends there exits before a single call arrives, and aclose() is what drains live calls and releases the port.

Configuration

The configuration is read once when the worker starts, not per call, so a missing key stops the worker at startup rather than surprising the first caller. ELEVENLABS_HOST is checked against the elevenlabs.io suffix, because your API key and agent id travel to it: a wrong host would be credential leakage rather than a failed call. The reasoning behind that shape is on Configuration.

Next

Realtime providers

The startup buffer, the echo guard and barge-in this plugin sits on.

Vision and the avatar

What look, show_image and express reach.