Shorts Maker: Turn Landscape Video into Vertical 9:16 with Auto Captions
Turn a video you shot in landscape into a vertical clip that's ready for YouTube Shorts, Instagram Reels, and TikTok. Choose an automatic face-following crop, a blurred-background mode that keeps the entire original frame visible, optional title text, and AI-generated captions that you can edit before exporting. Everything happens inside this browser, on your own device - nothing is uploaded to a server.
Unlike tools that require installing a desktop app or uploading your footage to someone else's server, this one keeps every step local: the face detector, the speech-recognition model, and the final video encoding all run using your browser's own machine-learning and recording APIs. The page can even keep working offline once it and its small AI models are cached, and your original footage is never transmitted anywhere.
Loading the tool...
How to use
- Drop in the video you want to convert, or click to choose a file. (MP4, MOV, WEBM · up to 3 minutes)
- Pick an output ratio (9:16, 4:5, or 1:1) and a crop mode: Follow face, Center, or Blurred background + full frame.
- Optionally enter a trim start and end time to use only part of the clip.
- If you picked "Follow face", press "Analyze face position" to scan the clip and find where faces are before rendering.
- Optionally add a title, then turn on "Auto-generate subtitles" and press "Generate subtitles" to let the AI transcribe the audio. Review the result in the editable list below - you can fix any line, or add and remove lines.
- Use the preview slider to check how the crop and captions look at different points in the clip, then press "Convert & save". Rendering happens roughly in real time relative to the clip length, so keep the tab open until it finishes.
- Download the finished vertical video, and optionally export the captions separately as a .srt file.
How it works
This tool never sends your video to a server - everything runs locally in the browser. To find where to crop, an on-device face-detection AI samples the video roughly every half second, looking for the horizontal center of any face in frame. That raw position is smoothed over time (small dead-zone plus gentle exponential smoothing) so the final crop glides rather than jitters when someone moves, and the smoothed centers are interpolated between samples while rendering so the motion stays continuous frame to frame. In Blurred background mode, the source frame is scaled up and blurred to fill the entire vertical canvas, and then the full, unmodified source frame is drawn on top at a smaller size so nothing is cropped out. Captions are produced by an on-device speech-recognition model that listens to the audio track and writes out timestamped text, which is then burned directly onto the video frames during export. Rendering takes roughly as long as the clip itself, since the browser has to decode, redraw, and re-encode every frame in real time.
Tips for Shorts, Reels, and TikTok specs
Most Shorts, Reels, and TikTok videos default to a 1080×1920 vertical 9:16 frame. These apps typically overlay their own interface - like buttons, captions, and account info - near the top and bottom edges of the screen, so it's generally safer to keep any text you add, such as a title or caption, inside a central "safe area" rather than right at the very top or bottom. This tool places captions at roughly 72% of the frame height and titles near the top 10%, which tends to stay clear of common UI overlays, though you should always preview the result before posting since every app's layout can differ. If you're posting to the regular Instagram feed instead of Reels, a 4:5 ratio shows more of the original frame than a tight 9:16 crop; a 1:1 square works well for feed posts or square thumbnails. Exact maximum video lengths, file size caps, and other platform policies change fairly often and differ by app and by content type, so it's worth checking each platform's current help documentation before you publish, rather than relying on any single fixed number.
Choosing a crop mode
"Follow face" works well for interviews, talking-head videos, and vlogs where a person moves left and right across the frame - the crop will glide to keep them roughly centered. "Center" is a good default when the camera barely moves or when the important content always stays in the middle of the frame, and it skips the face-analysis step entirely. "Blurred background + full frame" is the right choice when you don't want to crop anything out at all - for example clips with on-screen text, tables, graphs, or multiple people who all need to stay visible - since it keeps the entire original frame inside the vertical canvas and fills the empty space above and below with a softly blurred, zoomed-in copy of the same footage.
Why automatic subtitles help
A large share of viewers watch short-form video with the sound off, especially while scrolling in public or at work, so burned-in captions can meaningfully increase how much of a video people actually watch and understand. Manually typing out every line of dialogue and timing it to match the audio is slow, so this tool automates the first draft: the speech-recognition model listens to the trimmed audio in overlapping windows and produces a list of timestamped lines, offsetting and stitching them together so the result covers the whole clip without obvious gaps or duplicated phrases. You stay in control of the final result - every line is editable, you can delete lines that aren't needed, add new ones by hand, and adjust the start and end time of any line before it gets burned into the final video or exported as a standalone .srt file.
FAQ
Is there a length or resolution limit?
The trimmed section can be up to 3 minutes long. Anything larger than 1080p is automatically scaled down to 1080p while processing, and if the source is 720p or smaller, the output uses a smaller 720p-based vertical resolution instead.
How does "Follow face" work?
The video is sampled roughly every 0.5 seconds to find the horizontal position of any face on screen, the result is smoothed so the crop does not jump around, and the vertical frame is cropped around that smoothed center. When no face is found, it falls back to the last known position or the center.
How accurate are the AI captions?
The speech-recognition AI generates captions automatically, but accuracy can vary with accents, background noise, and specialized vocabulary. You can edit any line, or add and delete lines by hand, so always review the captions before sharing.
Which ratio should I pick: 9:16, 4:5, or 1:1?
YouTube Shorts, Instagram Reels, and TikTok mostly default to vertical 9:16. 4:5 shows a bit more of the frame and works well for the Instagram feed, while 1:1 suits square thumbnails or feed posts.
Is my video uploaded to a server?
No. Face position analysis, caption generation, and the video conversion itself all run entirely inside this browser, on your own device. The video file is never sent anywhere.
Open source used
Face position analysis runs on-device using Google's MediaPipe Face Detector (Apache-2.0 license). Automatic captions are generated with a Whisper speech-recognition model (MIT license), run directly in the browser through Transformers.js (Apache-2.0 license) and ONNX Runtime Web (MIT license). The model files and runtime are all served directly from Dagochim's own servers, and no video, audio, or face data is ever transmitted anywhere else.
⚠️ Speech recognition can make mistakes, so always check the captions before sharing. Processing takes roughly as long as the video itself and runs in real time, so please keep the tab open until it finishes. Your file never leaves your device, but please only use videos you have the rights to.