Record your screen once, talking as you go. VideoPolisher removes the dead time, rewrites and re-voices the narration, keeps the voice in step with what is on screen, and hands back a YouTube-ready video plus vertical shorts, captions, title and tags.
You approve every word of the script before anything is rendered. No re-recording, no timeline editing.
Pick a card. Every one says what you get back.
Long pauses while the page loads, the sentence you started twice, the click you had to redo. Cutting it by hand takes longer than recording it did.
A better script means a second take, a third take, and a room that is quiet enough. The screen part was fine. Only the voice needed work.
YouTube wants 16:9. Shorts, TikTok and Reels want 9:16 with captions. LinkedIn wants a one-minute highlight. That is three more edits of the same thing.
Subtitles, a title that people search for, a description, tags, a thumbnail, an intro and an outro. Each one small, together an afternoon.
Pick what you want, see what will run, approve the script, download. No timeline, no keyframes.
Choose a card, answer a question or two, and add your recording. Tick the outputs you want; mark cuts, an intro or outro to keep, or a region to crop if you like. We show exactly what will run and what it costs before you commit.
We transcribe what you said and rewrite it into clean narration that follows the order of events on screen. You read it, edit any line, and approve. Nothing is voiced or rendered until you do. Want to keep your own words? Approve as is.
A natural voice reads the approved script. The video is re-timed so the screen keeps pace with the narration, dead time disappears, and every format you picked renders. Download, or publish straight to YouTube.
For when you want to know how it works before you try it. Every step is a separate piece of work: change one setting and only the steps that depend on it run again.
Normalises the frame rate, scales or letterboxes to a standard size, crops to a region you draw, and removes dead time. Cuts are detected by AI, or you mark them on a timeline and they are used exactly as drawn.
Transcribes your voice with word timings and picks the key frames where the screen changed. Both feed the script and the sync.
Rewrites the transcript into narration that follows the on-screen order, keeps your technical terms, and marks each sentence to the moment it belongs to. Keeps the spoken language or translates to another. You edit and approve — or take just the script and record it yourself.
Natural voices from OpenAI or ElevenLabs, with pronunciation overrides for acronyms. The video is re-timed scene by scene so the narration lands on the action, speeding through dead stretches and never cutting a scene in half. Prefer the original length? Pauses are inserted instead.
The polished 16:9 video, a transcript aligned to the final cut, SRT captions, YouTube metadata and a thumbnail. GPU encoding when available.
Reframes to 1080×1920 with a camera that follows the cursor or interprets the scene, stays inside the content, avoids black bars, and respects the safe zone under Shorts and TikTok overlays.
An AI-driven camera on the 16:9 output: zooms toward what the narration talks about, pans smoothly, snaps to edges, and holds still when nothing moves.
Karaoke-style captions on the phrases that matter, in a font that fits the language. Then 30–60 second highlights with their own titles, for Shorts and for LinkedIn.
Mark an intro or outro to keep and it is passed through untouched and re-attached in every format. Or generate a branded intro and recap outro from the script. Splice other clips in at chosen times, loudness-matched.
Detects and tracks the webcam bubble in a recording, covers the original with a card or logo so the camera can move freely, and places a clean still or your own headshot where you want it.
Start from a script, an article, or a link to one. It is rewritten for listening with chapters and show notes, fact-checked in a second pass, and narrated: an MP3 episode, plus video in 9:16, 1:1 and 16:9 with slides or images, captions and clips.
Give it your site. It finds the core product pages and the demo clips and images already on them, writes one scene per feature, narrates it, and lays it out in 16:9 and 9:16 from a single script. Edit any scene's words and re-render only that scene.
Connect YouTube once and publish any output with its metadata. Titles, tags and descriptions are checked against YouTube's limits before upload.
Every output remembers what it was made from. Change a voice, a caption style or one line of script, and only the steps downstream of that change run again. Your approved script is never regenerated behind your back.
Credits per minute of output. Start on the Free plan: polish recordings up to three minutes, with a watermark, no card needed. Plans add longer videos, every format and no watermark.
No. The screen recording is the source of truth. Only the narration is rewritten and re-voiced, and the video is re-timed to match it. If you would rather use your own voice, take the script and record it yourself.
It rewrites for clarity and keeps the order of what happens on screen. It does not invent steps. You read the whole script and approve or edit every line before it is voiced.
Transcription detects the spoken language. Narration keeps that language by default or translates to another, with captions in a font that fits the script.
Screen recordings in common formats such as MP4 and MOV, with your voice on the track. For podcasts, an audio file, a script, or an article URL. For a website tour, the site's address.
Text-to-speech voices from OpenAI and ElevenLabs, chosen per project. Pronunciations of product names and acronyms can be pinned.
Uploads are stored in our own cloud storage in the US and processed by AI providers to transcribe, write and narrate. Your files are yours, exportable and deletable on request. The privacy policy lists every provider we use.