Audio Description Workflow

Most videos are made with the assumption that you can see them, which leaves a real gap when the visual action carries the story — a silent gesture, a facial expression, an on-screen caption, an establishing shot with no dialogue at all. The Audio Description Workflow closes that gap by generating a full narration track describing what's happening visually, timed to match the video itself.

This is one of the few tools in the app that requires you to bring your own AI access: if you haven't yet added a personal Google AI API key under AI Settings, the workflow stops you with a clear "API Key Required" message and a direct "Open AI Settings" link, rather than letting you proceed and fail partway through. Once that's set up, though, the flow is straightforward. Tap "Select Video File" to choose a video from your device, and it loads into a small preview player with play and pause controls so you can confirm you've picked the right one.

Three settings then shape how your description track turns out. Language lets you choose the spoken language of the narration itself from a genuinely wide list — seventeen options in total, including English, Hindi, Spanish, French, German, Arabic, Portuguese, Japanese, Chinese, Bengali, Tamil, Telugu, Kannada, Malayalam, Odia, Punjabi, and Marathi — so the description doesn't have to match the video's original spoken language at all. Interval controls how frequently the tool pauses to describe what's visually happening, from a tight 2 seconds up through 3, 5, 8, 10, 15, 20, or a more relaxed 30 seconds, depending on how visually dense the content is. Detail Level determines how much is said at each of those intervals: "Short" gives a single, efficient sentence; "Medium" expands to two or three sentences; "Explanatory" adds context — setting, emotion, the texture of a scene — across three to five sentences; and "Very Detailed" aims for a comprehensive, almost play-by-play account of visual elements, expressions, and motion.

Once your settings are chosen, "Generate Audio Descriptions" uploads the video for analysis and produces a timed set of narration segments — each with a start time, an end time, and its own description — saved to your device under a dedicated audio description folder in Downloads. From there, a results screen gives you real editorial control: "View & Edit Segments" lets you open the full list, edit the wording of any individual segment, remove one that doesn't add anything, or add a new one of your own if something was missed. "Watch with AD" opens the video with the narration synced in as you play it, giving you the full experience of watching alongside the description. "Share Track" lets you export just the description data to share with someone else working on the same video, and "New Video" resets the workflow to start again from scratch. It's a genuinely powerful tool for making your own video library — home movies, downloaded lectures, shared clips from friends — accessible on your own terms, rather than waiting for someone else to add description that may never come.