Soku AI
All tools

Remove Text From a Video Without Cropping It Away

Burned-in text is the reason most reusable footage gets thrown away: a stale price, a hard-coded subtitle, a stock watermark, a legal super for a market you are no longer running in.

Video InpaintingWatermark Removal
  • Up to 15 seconds per clip
  • No mask drawing or keyframing
  • Audio is left untouched

Three deliberate steps

01

Upload the clip

Up to 15 seconds, any common video format. The whole clip is processed — there is no in/out range to set.

02

Describe the lettering

Say what the text is and roughly where it sits: "the white subtitles at the bottom", "the watermark in the top-right corner". Position helps the model more than quoting the words does.

03

Download the clean plate

The text is tracked and painted out frame by frame, with the background reconstructed underneath. The result is a full-frame clip you can re-caption or re-price.

Use cases

Who reaches for this, and when

Performance teams

Re-run last quarter’s winner at this quarter’s price

A creative that beat the account average is usually retired for one reason: the offer is stamped into the pixels. Erase the old number and the winning edit, hook and pacing all survive into the new promotion.

Localization

Strip hard subtitles before re-voicing

Footage that arrives with English subtitles baked in cannot be dubbed for another market — the old line stays on screen under the new audio. Removing the lettering first is what makes the clip re-usable in every other language.

Social and UGC

Clear another platform’s watermark

Creator footage delivered with an app watermark is not a clean asset. Take the mark out rather than cropping in, which on a 9:16 clip costs you the top or bottom of the composition where the hook usually lives.

Questions before you run it

Is removing text from a video free?
You can start free — sign up on this page and you get free credits to run it. After that each run costs a small number of credits, charged per clip rather than per second.
Do I need to type the exact words that are on screen?
No, and it is usually better not to. The model is segmenting a region, not reading the text, so describing what the lettering is and where it sits ("the white subtitle line at the bottom") works better than quoting it exactly.
Is this a crop or a blur?
Neither. It is generative inpainting: the lettering is masked out and the area behind it is reconstructed, so the frame stays the same size and there is no blurred patch left where the text was. That also means the area behind the text is synthesised rather than recovered — it is a plausible background, not the one the camera actually saw.
What kind of text is hardest to remove?
Text over a moving, detailed background — a busy street, water, a crowd — because the fill has to invent motion as well as texture. Text sitting on a flat or slow-moving area comes out cleanest. Very large lettering covering most of the frame is also difficult, because there is little real background left to infer from.
Can I remove several pieces of text at once?
Yes. Name them together in one prompt — "the watermark in the corner and the subtitles at the bottom" — or run the tool a second time on the result if one of them survives.