Paste a script. Gemini works out who speaks each line and casts a voice, reads it aloud, and sketches the visual beats — then stitches the voices into one audio track and makes your video. Runs entirely in your browser.
Settings
Saved projects
Saved on this device, in this browser — your script, voices and images are kept so you can close the tab and come back. Clearing browser data (or using another device) removes them.
Your Gemini key
Analyzing a script is already covered by a shared key, so you can try it without your own. To make voices or images with Gemini, add your own key — or use the free providers below (Kokoro, Pollinations, Puter), which need no key at all. Get a free key. Stored in this browser only.
Your key stays in this browser only (saved to localStorage on this device). Never commit it to the repo — GitHub Pages is public. For safety, restrict the key to your Pages URL in Google Cloud Console → Credentials → HTTP referrers.
Use my own key for:
Unticking a box means that part won't use your key: the Analyzer falls back to the shared key, and Audio/Images ask you to pick a free provider instead. The Analyzer also uses the shared key automatically if you haven't entered a key.
Audio generation
Point this at any OpenAI-compatible TTS server. Examples: Pollinations — URL https://gen.pollinations.ai/v1/audio/speech, get a free key at enter.pollinations.ai; a self-hosted Edge-TTS (travisvn/openai-edge-tts) at your own URL; or OpenAI's own https://api.openai.com/v1/audio/speech with your key (note: OpenAI may block direct browser calls). One voice for all characters.
Voices are auto-assigned per character from ElevenLabs' default set. Enter your key in the ElevenLabs API key box below.
Kokoro is a high-quality open-source voice model that runs entirely in your browser — no key, no account, no quota, and it works offline once loaded. The first time you generate or preview a Kokoro voice it downloads the model (~80 MB, one time only — your browser caches it afterward). It runs on your CPU, so it's slower per line than the cloud options, and that first download is hefty on mobile data. Voices are auto-assigned per character (28 English voices, American & British); the Voice-model dropdown above doesn't apply.
Experimental. Uses your device's built-in voices — no key, no quota, no account. To actually save it into downloads/video, the browser has to record this tab's audio: when you click Generate voices, a share box appears — choose This Tab and tick Share tab audio. It plays out loud and records in real time, so it's slow, and it works best in Chrome (other browsers may not allow tab-audio capture). If no sound gets captured, you'll get a clear error.
Google Cloud TTS uses the same Google key as Gemini but is a separate service with its own (large, free-tier) quota — the best way around the Gemini 100/day TTS cap. One-time setup on your own Google account: enable the Cloud Text-to-Speech API on the same project as your key (free to enable; ~1M characters/month free, billing must be on the project). It auto-assigns a Neural2 voice per character.
ElevenLabs needs its own free key (separate from Google) and is the highest quality. All the Puter options are free and need no key (a quick Puter sign-in may pop up on first use) and don't touch your Gemini quota. "Puter · Gemini" keeps your exact cast voices; the others use their own voices, auto-assigned per character. The Google API Voice model dropdown only shows for the Gemini-voice options.
Only needed if you choose ElevenLabs as the voice provider above. Get a free key (free tier ~10k characters/month). Stored only in this browser, never sent anywhere except ElevenLabs.
Script Analyzer
The model that reads your script on Analyze — it works out who speaks each line, casts a voice per character, and writes the visual prompts. It runs once (not per line), so it barely touches your quota. Bump it to Pro if a messy script gets parsed or cast wrong.
Animation generation
Needs a Pollinations key & Pollen credits (not no-key). Make a free account at enter.pollinations.ai — new accounts get some free Pollen to start, then it's pay-as-you-go ($1 ≈ 1 Pollen) from your own balance. Renders one clip at a time. Your key is stored only in this browser.
Video is in alpha on Pollinations. If a model id is rejected, check the model list. Clips are per-shot — make them with the 🎬 Animate button under each visual; they aren't auto-stitched, so download what you want. Length is set by the model — there's no separate length control (video is alpha).
No key needed. A quick Puter sign-in popup appears on first use, then you get free credits — Puter's "User-Pays" model means it's free to start, and after the free quota your own Puter account covers usage (cheapest is Wan). Want Kling, PixVerse, Seedance, Hailuo and others? Pick Custom and paste the id from Puter's video docs. Clips render one at a time (~10s–2 min each), so animate per shot with the 🎬 Animate button under each visual. Clips aren't auto-stitched into the in-app video yet — download the ones you want. Clip length is set by the model (e.g. Wan ≈5s, Veo ≈8s, Sora longer) — Puter has no separate length control, so pick a model for the length you want.
This costs real money — billed to your own fal.ai account, about $0.04/sec at 1080p, so one 6-second clip ≈ $0.24 (a 20-shot script ≈ $5). Clips render one at a time and take ~10–60s each, so animation is per-shot, on demand — use the 🎬 Animate button under each visual, not a bulk button. Your key is stored only in this browser. Clip length: type any value 6–20s (LTX-2 uses even numbers, so odd/out-of-range values snap to the nearest); over 10s runs at 1080p only, and cost scales with length (a 20s clip ≈ $0.80 at 1080p Fast).
LTX-2 (Lightricks) is cheapest and uses the length/resolution above. Custom lets you run any fal video model — paste an id from the fal model gallery (Kling, Veo, Wan, Hailuo, PixVerse, Seedance…); it sends just the prompt and uses that model's defaults, and cost depends on the model. Get a fal.ai key (needs billing). Clips come back as .mp4 links per shot — not auto-stitched into the in-app video yet, so download what you want (fal links are temporary).
Image generation / Video output
Get a key at api.together.xyz. The FLUX.1 schnell (Free) model is free (rate-limited); the others bill your Together account. Images come back as base64 and are stored as a data URL. Your key stays in this browser. If you get a CORS/network error, tell me and I'll route it through your proxy like Craiyon.
Needs a Prodia account/token from app.prodia.com/api (the v2 API may require a paid Pro plan — their free tier is unclear). The image comes back directly and is stored as a data URL. Your token stays in this browser. If you get a CORS/network error, tell me and I'll route it through your proxy like Craiyon.
Routed through your server proxy (craiyon.php next to analyze.php) because Craiyon has no browser API. It's free, no key, but the quality is dated (DALL·E-mini era) and a generation can take 30–60s. The proxy returns one image per shot as a data URL.
Pollinations is free with no key. If images start failing from rate limits, get a free publishable key (starts pk_) at enter.pollinations.ai — paste it here. Stored only in this browser.
Any OpenAI-compatible image server: a local Stable Diffusion bridge, LocalAI, or a hosted endpoint. The app POSTs to /v1/images/generations and reads b64_json (or a returned URL).
The image provider draws the visuals used in the storyboard and video. Pollinations and Puter need no key (Puter may ask you to sign in to your free Puter account, which pays for its own usage). Motion animates those stills while the video records — a slow zoom/pan (Ken Burns) and a fade between shots, done in your browser with no API and no credits. Lower resolution makes the file smaller.
STEP 1
Your script
Paste a script in any format — a screenplay, labelled dialogue, or plain prose. Gemini works out who's speaking, the on-screen text, and the shots.
What can I paste? formats & examples
Plain prose / a story — read as narration; dialogue in quotes is split out to characters. Who speaks — start a line with a name and a colon; each name gets its own voice: NARRATOR:REES:DEV: On-screen / system text — label it for the flat robotic readout: ON SCREEN:SYSTEM: Delivery tone — add a note in brackets: REES: (whispering) You notice more when you write it down. Screenplay format works too — scene headings like INT. KITCHEN – NIGHT, action lines (turned into shots), and (V.O.) or (O.S.) after a name.
Visuals are added for you (about one every 3–5 lines). If something's parsed wrong, set the Analyzer to Pro in Settings ⚙, or fix any line with ✎ Edit after Analyze.
STEP 2
Breakdown & casting
Each character is cast to a voice. Press ▶ to hear the voice this character will use with your current provider, change any of them, then generate.
Long script? The free tier caps requests per minute, so some lines/images may fail at first — that's normal. Wait a few seconds and click again; it only fills what's missing. Repeat until done.
Your narration always plays at full volume — the music sits behind it at the level above, looped to fit and faded out at the end. For no Content-ID claims, use your YouTube Audio Library (in YouTube Studio) or Uppbeat — not audio ripped from YouTube videos.