TL;DR: ElevenLabs expanded its MCP server beyond speech to cover music, sound effects, images, and video, letting AI assistants generate full multimedia from a single integration.
Summary: ElevenLabs announced that its Model Context Protocol server now supports voice, music, image, and video generation in addition to its existing audio capabilities. The MCP lets assistants generate speech, transcripts, dubs, music, sound effects, images, and video directly from the chat client a developer already uses. This turns the ElevenLabs MCP into a broader multimodal generation endpoint rather than a speech-only tool.
Why it matters: Builders using MCP-compatible assistants can now wire one server for multimodal asset generation instead of stitching together separate image, music, and video APIs. Worth testing whether the new modalities fit existing agent workflows and how latency and cost compare to dedicated providers.
Source: x_com