Specific beats clever
Describe visible subject, action, setting, camera, and light instead of abstract hype words.
Answer-first beginner guide
Start with one audience, one outcome, and one short visible action. Choose text or image input, describe the shot like a director, generate a short clip, review defects, then add sound, captions, and editing outside the model.
Researched and updated August 20, 2026

Specific beats clever
Turn a vague idea into one shootable prompt. Keep the task small enough to review before spending more generations.
Describe visible subject, action, setting, camera, and light instead of abstract hype words.
If the prompt asks for walking, turning, waving, and sitting, motion and anatomy usually break.
Check identity, hands, physics, text, products, claims, and continuity before making variants.
Write who the video is for, what they should understand, and where it will publish.
Use text for invention; use an image when composition, product, or character appearance must stay recognizable.
Specify subject, one action, setting, camera, lighting, style, duration, and what must not change.
Start with a short clip and a cost-appropriate model. Do not make duration longer to rescue a broken prompt.
Change one main variable, generate a small set, and keep the cleanest take with notes.
Add sound, captions, cut points, accessibility, rights review, disclosure, and platform-specific export.
| Problem | Likely cause | Next test |
|---|---|---|
| Melting or duplicate objects | Too many actions or subjects | Reduce to one subject and one action |
| Product or face drifts | Text-only prompt has weak visual constraints | Use a clean image or references |
| Camera feels random | Movement and end frame are unspecified | Name one camera move and final composition |
| Clip feels generic | Prompt lacks concrete setting and light | Add specific physical details, not more adjectives |
| Long clip gets worse | The model cannot sustain the motion | Split it into shorter shots |
A repeatable workflow reduces waste, but the same prompt can still produce different results.
You need a clear goal, a prompt or image, a suitable generation mode, an aspect ratio, and time to review and edit the output.
Start with the shortest useful duration, often five to ten seconds. Longer clips magnify motion and continuity problems.
Describe one subject, one visible action, setting, camera movement, light, style, and the details that must remain stable.
Editing adds timing, sound, captions, branding, factual review, accessibility, and a publishable structure that raw generation does not provide.