The chain is Kling → ElevenLabs (+ optionally Freepik and Descript), about 30 minutes: Kling gives your photo motion, ElevenLabs reads your line like a person, captions seal it. Keep the motion small and direct the voice — those two choices decide the result.
A reel that moves and talks normally takes three crafts: animating a still, recording a voice, and editing the two together. Each MCP connector only knows its own tool — Kling makes video, ElevenLabs makes voice — and neither knows the other is in the chain. The repos document one tool each; an app-store listing just hands you the install. Nobody covers the seam.
This recipe is the seam. You bring one photo (or generate one in Freepik), Kling gives it motion, ElevenLabs reads your line in a real-sounding voice, and you lay the two together into a vertical clip. The gotchas above — keep the motion small so faces don’t warp, ask ElevenLabs for a conversational delivery, mind two separate credit meters — are the parts no single tool’s docs mention, because they only show up when you chain the tools.
Checked 2026-06-14: Kling and ElevenLabs connectors both reachable. Kling is community (needs your Kling API keys); ElevenLabs is the official connector.