Prepare assets
Upload the reference video and product image. Product name, selling point, and audience are optional and can be inferred by AI.
Upload a reference video and a product image. The agent recovers shot, action and voice beats; perfect every shot’s first/last frames, or let a 3×3 storyboard grid produce a whole segment at once — long videos auto-split and chain on the previous tail frame. Product truth stays locked; every paid task asks first.
STEP 01
Upload the reference video and product image. Product name, selling point, and audience are optional and can be inferred by AI.
01 / SOURCE
02 / PRODUCT
Analysis uses Kimi K3 multimodal understanding by default. Empty responses, timeouts or invalid JSON trigger an automatic switch to MiniMax M3 or degraded retries (fewer evidence frames, longer waits).
Only the reference video and product image are required. AI can infer the product name, benefit, and audience from the assets.
Upload the reference video and product image. All text fields are optional.
The real critical path
Every step produces something you can inspect: evidence frames, product truth, storyboard and voiceover. Unhappy with one step? Redo only that step.
Upload the reference video and product image. Product name, selling point, and audience are optional and can be inferred by AI.
Review the shots, actions, voiceover, and marketing beats reconstructed from the reference video.
Review the new-product storyboard, per-shot voiceover and immutable product rules, then choose per-shot precision mode or nine-grid storyboard mode.
Play shot-by-shot results and review motion, product state, and voice synchronization.
Not a one-click black box
Every recovered beat traces back to evidence frames from your reference video, so the breakdown can be checked and corrected.
Product name, selling points and non-deformable structure are locked as hard constraints no generation may violate.
Every paid generation asks for explicit confirmation — failed paid tasks are never auto-retried.
Precision mode crafts each shot’s first/last frame; grid mode turns a 3×3 storyboard image into a whole segment, auto-splitting long videos and chaining on the previous tail frame. Voiceover and AI sound optional.
Your next product video