feat: add voice and audio generation tools to generative reference
Covers ElevenLabs (voice cloning, best quality), OpenAI TTS (cheap at scale), Cartesia Sonic (40ms latency), PlayHT, Resemble AI, WellSaid Labs, Fish Audio, and cloud providers. Includes comparison table, decision tree, and voice+video layering workflow with ffmpeg. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -122,7 +122,8 @@ For detailed specs and format variations, see [references/platform-specs.md](ref
|
||||
For image and video ad creative, use generative AI tools and code-based video rendering. See [references/generative-tools.md](references/generative-tools.md) for the complete guide covering:
|
||||
|
||||
- **Image generation** — Nano Banana Pro (Gemini), Flux, Ideogram for static ad images
|
||||
- **Video generation** — Veo, Kling, Runway, Sora, Higgsfield for video ads
|
||||
- **Video generation** — Veo, Kling, Runway, Sora, Seedance, Higgsfield for video ads
|
||||
- **Voice & audio** — ElevenLabs, OpenAI TTS, Cartesia for voiceovers, cloning, multilingual
|
||||
- **Code-based video** — Remotion for templated, data-driven video at scale
|
||||
- **Platform image specs** — Correct dimensions for every ad placement
|
||||
- **Cost comparison** — Pricing for 100+ ad variations across tools
|
||||
|
||||
Reference in New Issue
Block a user