Soku AI
All tools

Turn a Script Into a Video One Scene at a Time

Paste the scene as it is written — the action, then the line — and get it back as a clip with the line actually spoken.

Script to VideoSpoken Dialogue
  • 4–15 seconds per scene
  • Dialogue spoken in-shot
  • 9:16, 1:1 and 16:9

Three deliberate steps

01

Write the scene, not the whole script

One beat per run: where we are, who is on screen, what happens, and the line to be said. Keep it to what could plausibly play in under fifteen seconds.

02

Put the line in quotes

Anything you write as spoken dialogue is performed in the clip. Describe the delivery if it matters — flat, urgent, amused — the same way you would in a shooting script.

03

Pick the shape and generate

Choose the aspect ratio for the placement you are cutting to and the length of the beat. Send the result into Creatives to assemble the scenes into a finished ad.

Use cases

Who reaches for this, and when

Copywriters

See the line before you sell it

A script reads fine on the page and dies out loud. Generating the scene is the cheapest way to hear whether the hook lands in the first three seconds, before anyone books a shoot around it.

Performance teams

Test the script, not just the cut

Run the same scene with three different lines and let the account decide which message works. Varying copy is usually the cheapest variable to test and the hardest one to produce footage for.

Agencies

Pitch with a moving board

Turn the scripted scenes of a treatment into clips so a client watches the idea instead of reading it. It is closer to the finished thing than an animatic and takes minutes rather than days.

Questions before you run it

Can I paste my entire script?
No — and this is the main thing to know before you start. A single run produces between 4 and 15 seconds, so it renders one scene. A full thirty- or sixty-second script is several runs, one per beat, assembled afterwards. Pasting the whole script into one prompt gives you a muddled clip of the first idea rather than the whole ad.
Does it actually speak the lines?
Yes. Audio is generated together with the picture, so dialogue written into the prompt is performed by the character on screen and is lip-timed to the shot. You do not add a voiceover afterwards.
Can I choose the voice?
Not from a preset list on this tool. The voice comes from how you describe the speaker — their age, manner and delivery — so casting happens in the writing. If you need one specific, consistent voice across scenes, record or clone it and use a lip-sync tool instead.
Is script to video free?
You start with free credits when you sign up on this page. After that a run is priced per second of output, so a 4-second scene costs roughly a third of a 12-second one — which is why testing at the short end is worth doing before you commit.
Will the same character look the same in the next scene?
Not reliably. Each run is generated independently, so continuity of face and wardrobe across scenes is not guaranteed. Ads that cut between locations or use a voice over changing footage handle this well; a two-hander that stays on one actor for thirty seconds does not.