Edit footage, style captions, generate presenter videos and create short clips with Captions’ AI Edit, avatars and chat-based revisions.
Captions
Explore features, practical uses and pricing below.
Captions is an AI video creation and editing product for turning footage, scripts and prompts into videos. Users can edit their own recordings, add and style captions, generate presenter footage with avatars, or produce shorter clips from a longer video. It brings these jobs into an editing workflow that includes both direct controls and changes requested through chat.
Its name describes one important function, but the current product extends well beyond subtitles. AI Edit assembles an edited treatment of supplied footage, while avatar and prompt-based workflows generate new material. Those are different starting points: one changes a recording you already have, and the other creates a performance around an approved script.
Captions suits creators and small teams that repeatedly produce explanatory or promotional videos. The useful question is which parts of production need assistance: selecting a moment, correcting words, assembling supporting visuals or presenting a script. Choosing the starting workflow deliberately helps avoid spending credits on a generated video when a straightforward edit would serve the audience.
The current AI Edit documentation describes automatic captions, transitions, B-roll, music and motion graphics around uploaded footage. Users choose an editing style and can adjust the resulting project. The web workflow also accepts a direction for the edit and optional media or reference links to provide context.
Before generating, the web workflow presents a plan with recommended cuts and settings such as style, edit intensity and color. Users can restore cuts they want to retain and decide whether to supply custom media or generate supporting media. This is useful when the original timing is important, or when a demonstration needs a particular image rather than an approximate illustration.
The documentation recommends short, single-speaker, previously unedited footage and accepts vertical or horizontal orientation. Check the current duration and import requirements for the route being used. A video with several speakers, extensive existing effects or a complicated sequence of demonstrations deserves a different evaluation from a simple presenter clip.
Review the finished treatment for the intended message. A cut that removes a hesitation can improve pacing, but a cut that removes a qualifier can alter the meaning. Supporting media should clarify the explanation. If a presenter describes an actual product screen, an invented interface can confuse viewers even when it looks visually polished.
Use edit intensity as an editorial choice. A concise tutorial may need relatively restrained graphics, while a promotional introduction may benefit from a more energetic style. The best choice is the one that helps the viewer follow the content. A consistent visual treatment matters less than whether the video accurately explains what the viewer came to learn.
Chat-based editing lets users request changes in ordinary language. This complements the project's timeline and visual controls rather than requiring every adjustment to begin with a new generation. A specific request about caption size or an unwanted shot is easier to assess than a broad request to improve the entire video.
Work through changes in an order that reflects the purpose of the video. First check the spoken message and sequence. Then revise supporting media and timing. Finally, adjust caption placement, color and other presentation details. Otherwise, a team can spend time polishing a section that will later be removed because it does not support the explanation.
Make each request concrete: identify the element, the desired change and the relevant moment. Afterward, preview the affected passage and the transition on either side. Some changes have consequences beyond the element named in the request, such as a longer caption obscuring an image or a new visual changing the pacing.
Keep a distinction between correction and creative variation. Fixing a misspelled name should produce one accurate result. Trying several styles is an experiment with alternatives. Both can be useful, but treating every revision as an open-ended experiment can make review difficult and increase AI usage without improving the message.
Captions generates text from spoken audio and supports editing the words and their timing. Its caption workflow guide recommends reviewing names and specialized terms, correcting errors and checking the result on a phone. The style controls cover fonts, colors and effects for presenting that text.
Check captions as information, not only decoration. They should retain the speaker's intended meaning, appear at a useful pace and remain readable against the background. A bright style can still be difficult to read over a busy demonstration. Avoid placing words over the object or control the viewer needs to see.
Translated subtitles and dubbing serve different audience needs. Subtitles retain the original audio while displaying translated text. Dubbing replaces the speech with a translated voice track. Captions' dubbing overview describes this distinction, and Lipdub additionally adjusts lip and facial movement to the translated audio.
Choose a target language supported by the specific function, rather than assuming every caption language is available for every voice or translation workflow. Review terminology, tone and timing with someone who understands the target audience. A literal translation can be grammatically plausible while changing a product promise or making a practical instruction hard to follow.
Retain the original-language export alongside each localized version. Identify the approved script and product version used for all of them. When a feature name or instruction changes, this record helps a team update the correct videos instead of overlooking a translation that still teaches the previous behavior.
The current avatar documentation describes a library of ready-made presenters, custom avatars created from text prompts and AI Twins based on a user's own recorded material. Saved private avatars appear in the account's library, and users can create additional looks. Captions uses Mirage generation technology within these workflows.
An AI Twin can help present an approved script when the same person is not available to record each version. Creating it requires suitable reference material and review of the generated appearance and voice. Follow the current setup instructions for the device and account; different workflows may request different calibration inputs.
Write the script before evaluating the generated performance. For an educational video, establish the learning objective and the steps the presenter must explain. For a product announcement, verify what is available and how the audience can access it. An engaging delivery cannot correct an inaccurate script.
Review facial appearance, pronunciation and the relationship between the voice and visuals. If a product name is spoken incorrectly, revise the relevant direction or script and assess the next output. If a generated look implies a real location or event, decide whether that representation fits the communication. Generated footage should serve the message rather than supply unsupported evidence of an experience.
Use your own likeness or material you have permission to use, and make the generated nature clear when it matters to the audience. A synthetic presenter delivering a factual tutorial is different from an actual customer describing personal experience. Do not let avatar delivery turn drafted promotional copy into an apparent testimonial.
The current Clips documentation describes uploading a video or using a supported public link, optionally directing the selection with a prompt, reviewing the video plan and generating a batch of clips. Users can adjust aspect ratio and titles, then edit the resulting clips. The workflow includes reframing for different kinds of source material.
A prompt is useful when a specific topic matters more than a general highlight. For a webinar, request the section that answers a recurring customer question. For a gaming session, describe the type of moment wanted. Review the proposed selections with the full context of the source rather than accepting every clip in the batch.
Captions ranks clips using its own scoring. Treat that as a selection aid, not a prediction of views or a guarantee of performance. A high-ranked excerpt can still need a clearer opening, a corrected caption or more context. Keep clips that have a complete point and a natural ending for the intended audience.
Reframing also deserves review. A presenter-centered crop may miss a detail on a shared screen, while a gameplay clip may need both the action and the speaker's reaction. Inspect the actual export at the intended display size. The best section of the recording is useful only if the audience can see and understand it.
Imagine a small booking-software company preparing a video about rescheduling an appointment. Its support specialist records a short explanation, and the team has an approved screenshot of the relevant control. Begin by checking that the script describes the current interface and does not promise a behavior the software cannot perform.
Import the recording and choose an AI Edit style that leaves room for the demonstration. Supply the actual screenshot as supporting media and give a clear instruction about where it belongs. Review the proposed cuts before generating, retaining any qualification about who can reschedule or when a change is allowed.
Watch the draft from beginning to end. Correct the appointment terminology in the captions, ensure the screenshot remains legible and remove a visual that suggests the wrong interface. Use chat-based editing for specific adjustments, such as reducing caption size during the screenshot or changing an unnecessarily busy transition.
Create a shorter promotional excerpt only if it still explains the benefit accurately. If the team has a longer product webinar, use Clips to locate the appointment section and review its boundaries. Do not assume that an excerpt from the middle of a conversation will make sense to someone who has never seen the full session.
For another language, decide whether subtitles or a dubbed version better serve customers. Have a suitable reviewer check the translated instruction and the name of the rescheduling control. If Lipdub is used, inspect the face and timing as well as the translation. Save separate approved exports for each audience.
If the team later chooses an AI Twin for routine update videos, create and review it separately from this feature explanation. Reuse the approved script structure, while checking every new product detail. The workflow then has a repeatable editorial process without assuming that generated footage can be published without review.
Captions fits social creators, educators, entrepreneurs and marketing teams producing short explanatory videos, promotional variations or localized presenter content. It is also relevant to teams repurposing recorded expertise. A producer working mainly on complex films or detailed multitrack post-production should compare the exact editing controls with that project's requirements.
The pricing page lists a limited Free option alongside paid Max, Frontier and Enterprise access. Its public figures explicitly describe iOS plans. The subscription documentation provides additional plan detail. Confirm the offer and feature availability on the device and checkout route you intend to use.
Free access provides limited editing, while broader automatic editing and generation depend on paid access. The credit guide distinguishes the free account's one-time allocation from renewing paid allocations. AI generation consumes credits even when the resulting project is not exported or saved.
Usage depends on the selected AI functions, video duration and models. Chat-based requests and generated media can add costs beyond the initial project. Evaluate one representative video and its revisions before estimating a full month of production. The number of videos a team can make is not determined by subscription name alone.
The guide also describes rollover and account-specific top-up eligibility. Confirm those rules before assuming unused credits remain indefinitely or that any account can buy more. Keep a modest revision budget when planning a campaign, since a script change or a replacement generated shot may require further processing.
Different Captions functions have different input and output requirements. AI Edit for a short presenter recording, Clips for a longer source and Lipdub for localization should be evaluated separately. Check the supported language, format, duration and device for the actual workflow, then use the export documentation to confirm the needed deliverable.
Review generated words, visuals and speech together. Correct captions can accompany the wrong image; an attractive avatar can pronounce a term incorrectly; a fluent translation can change the instruction. Approval should assess the complete video, including what the audience is likely to infer from it.
No. Its current offering includes AI editing, generated avatars and presenter footage, clips, chat-based revisions and language workflows alongside captions.
Yes. AI Edit works with supplied footage, and users can revise the resulting project. Match the source material to the documented requirements.
They change different parts of the video: displayed text, spoken audio and lip or facial movement. Choose the combination your audience needs.
No. Review the excerpt's message and presentation, then evaluate actual audience response after publication.