WAN Video GeneratorWAN Video Generator

Wan AI Image-to-Video Prompt Guide: Complete Tutorial with Prompt Examples

Jacky Wangon 6 hours ago

Introduction

I've been working with AI video tools for about a year now, and one thing keeps surprising me: the difference between a mediocre AI video and a great one often comes down to the prompt. This is especially true for image-to-video generation, where you're not starting from scratch — you have a reference image that sets the scene, character, or style, and your prompt tells the AI how to breathe life into it.

Wan AI's image-to-video capability is one of the most accessible options right now. You upload a photo — a product shot, a character portrait, a landscape — and describe how it should move. The result is a short video clip that preserves the original image's identity while adding motion, atmosphere, and narrative.

But getting the prompt right takes practice. Too vague, and the movement is random. Too specific, and the AI fights against itself. In this guide, I'll share the prompt structures I've tested extensively, with real examples you can adapt for your own projects.

TL;DR

  • Wan AI image-to-video uses your photo as the first frame — the prompt controls motion, atmosphere, and transitions
  • Best prompt structure: [subject] + [action] + [camera movement] + [atmosphere]
  • Keep prompts under 120 words — shorter focused prompts outperform long lists of instructions
  • Include motion verbs — "flowing," "drifting," "panning," "emerging" — to guide the AI's movement generation
  • Reference images with clean composition produce the best results — busy backgrounds confuse the AI

How Wan AI Image-to-Video Works

Wan AI takes your uploaded image and uses it as the first frame of the generated video. The model then:

  1. Analyzes the image composition — identifies the subject, background, depth, and lighting
  2. Interprets your prompt — maps text descriptions to motion patterns, camera movements, and atmospheric changes
  3. Generates subsequent frames — creates a sequence that starts from your image and evolves according to your prompt
  4. Maintains consistency — preserves the subject's identity, colors, and style across frames

The key insight is that your image anchors the AI. Unlike text-to-video where the AI designs everything from scratch, image-to-video starts with a concrete reference. This means your prompt doesn't need to describe what things look like — it only needs to describe how they move and change.

The Anatomy of a Great Image-to-Video Prompt

After testing dozens of prompt variations, I've settled on this structure:

[Subject reference] + [Motion/Action] + [Camera movement] + [Atmosphere/Lighting] + [Duration context]

1. Subject Reference

Keep this minimal — your image already defines the subject. Just name it for context:

  • "A woman in a red dress"
  • "A ceramic coffee cup on a wooden table"
  • "Mountain landscape at sunrise"

2. Motion/Action

This is the most important element. Be specific about movement:

  • Natural motion: "Hair flowing gently in the breeze," "Leaves rustling"
  • Directional: "The river flows from left to right"
  • Transformation: "The flower petals slowly opening," "The ice melting into water"
  • Subtle motion: "Steam rising from the coffee," "Gentle waves lapping the shore"

3. Camera Movement

If you want the camera itself to move:

  • Pan: "Camera slowly pans across the scene from left to right"
  • Zoom: "Gentle zoom in on the subject's face"
  • Dolly: "Camera moves forward into the landscape"
  • Static: Default — the camera stays still while elements move

4. Atmosphere/Lighting

Add mood through lighting and atmosphere:

  • "Soft golden hour lighting"
  • "Misty morning fog"
  • "Dramatic storm clouds rolling in"
  • "Neon-lit night scene"

5. Duration Context (Optional)

For longer clips, describe how the scene evolves:

  • "The scene gradually transitions from sunset to twilight"
  • "Over 10 seconds, the crowd slowly disperses"

Prompt Examples by Use Case

Product Photography

Reference image: Product photo on a clean background

Prompt: "The product rotates slowly on its axis, revealing all angles. Soft studio lighting catches the surface texture. Gentle shadow moves beneath it as it turns. Clean white background remains consistent."

Why it works: The prompt focuses on a simple, predictable motion (rotation) that the AI can generate reliably. It specifies lighting behavior and explicitly asks the background to stay consistent — preventing unwanted scene changes.

Portrait to Cinematic Clip

Reference image: Portrait photo with blurred background

Prompt: "The subject breathes naturally, eyes blinking softly. Hair moves slightly in a gentle breeze. Background bokeh shifts subtly. Warm golden light plays across the face. Camera slowly pushes in for an intimate feel."

Why it works: Combines realistic human movement (breathing, blinking) with cinematic lighting direction and camera movement. The result feels like a film clip rather than an animated photo.

Landscape to Atmospheric Scene

Reference image: Landscape photo of mountains and lake

Prompt: "Clouds drift slowly across the mountain peaks. Water ripples gently with soft wind. Sunlight breaks through clouds creating moving light patches on the valley. Autumn leaves drift down. Camera pans right to reveal more of the panorama."

Why it works: Multiple layers of motion at different depths (clouds in background, water in foreground) create a rich, immersive scene. The camera movement adds narrative scope.

Architecture to Living Scene

Reference image: Building or interior photo

Prompt: "Sunlight streams through windows, shifting across the floor over time. Curtains sway gently in the breeze. Shadows lengthen gradually as if hours are passing. The space feels alive and lived in."

Why it works: Time-based cues ("shadows lengthen gradually") give the video a temporal narrative. The result transforms a static architectural photo into a living space.

Creative and Artistic

Reference image: Artwork or illustration

Prompt: "The painted scene comes alive. Brushstrokes blend and shift like liquid color. Elements of the composition drift apart and reform. Dreamlike, surreal atmosphere with colors pulsing gently. No realistic elements — keep the painterly style."

Why it works: This prompt leans into the artificial nature of the medium rather than fighting it. By requesting "painterly" and "surreal" movement, it turns the AI's limitations into artistic choices.

Common Mistakes and How to Fix Them

Mistake 1: Overcrowding the Prompt

❌ "A man in a blue shirt walks across a busy street while cars honk and birds fly overhead and the sun sets and a dog runs past..."

✅ "A man in a blue shirt walks confidently across the street. Traffic moves slowly around him. Warm sunset lighting."

Less is more. Pick 2-3 motion elements and execute them well.

Mistake 2: Ignoring the Reference Image's Limitations

If your reference image has a solid white background, don't prompt "a busy marketplace unfolds around the subject." The AI has no marketplace content to extend from.

Match your prompt ambition to what the image actually contains. Images with simple backgrounds work best with subtle motion prompts.

Mistake 3: Contradictory Motion Instructions

❌ "The camera zooms in while also pulling back"

The AI can't do two opposite things at once. Choose one camera movement per clip.

Mistake 4: Vague Action Verbs

❌ "The scene moves"

✅ "Leaves drift down from the trees. Water ripples outward from a falling leaf."

Specificity gives the AI a clear pattern to generate. For a closer look at how it stacks up against other models, see Kling 2.6 Motion Control vs Wan 2.2 Animate.

Image Preparation Tips for Better Results

Use Clean Compositions

Images with one clear subject and a simple background produce the most coherent videos. If your image is cluttered, consider using an AI background remover to isolate the subject first.

Optimize Resolution

Higher resolution images give the AI more pixel data to work with. If your reference is small, upscale it. Most AI video tools work best with images at least 1024×1024 pixels.

Avoid Extreme Crops

If the subject is cropped awkwardly (e.g., head cut off at the forehead), the AI will struggle to animate it naturally. Full-composition images produce better motion.

Test with Subtle Motion First

For your first attempt at any image, use a conservative prompt like "gentle motion, subtle atmosphere." Once you see how the AI interprets the reference, you can push for more dramatic movement in subsequent attempts. If you want to test it without installing anything, this free Wan 2.2 video generator works in the browser. If you want to test it without installing anything, the Wan 2.2 Animate tool works in the browser.

Quick Reference: Prompt Templates

Image Type Prompt Template
Product still [Product] [rotates/reveals] with [lighting description]. Background [stays consistent/behaves].
Portrait [Subject] [breathes/moves subtly]. [Atmosphere/lighting]. [Camera movement].
Landscape [Elements] drift/move. [Weather/lighting changes]. Camera [pans/zooms].
Architecture [Light effects]. [natural elements like curtains/trees] move. Time feels [passing/compressed].
Artwork [Artistic motion description]. Keep [style] consistent. [Surreal/dreamlike atmosphere].

The Bottom Line

Image-to-video prompting with Wan AI is a skill that rewards practice. The more you experiment with different motion verbs, camera directions, and atmospheric cues, the more you'll develop an intuition for what the model responds to.

Start simple. Take one reference image and try three different prompts — one minimalist ("gentle motion"), one detailed (with camera movement and atmosphere), and one creative (experimental movement). Compare the results and you'll quickly learn which direction to push your prompts for your specific use case.

And if you want a library of ready-to-use prompt templates, check out the pre-built prompt examples to kickstart your next project.

Related guides

FAQ

What is Wan AI image-to-video?

Wan AI image-to-video is a feature that turns static images into short video clips. You upload a photo and provide a text prompt describing how the image should animate — the AI generates video frames starting from your image.

How long are Wan AI image-to-video clips?

Clip length depends on the platform and settings. Most tools generate 4-10 second clips. Longer durations may reduce motion quality.

What makes a good reference image for AI video generation?

Images with a clear subject, simple background, good lighting, and clean composition produce the best results. Busy, cluttered images or images with extreme cropping are harder for the AI to animate naturally.

Can I use Wan AI image-to-video for commercial projects?

Usage terms vary by platform. Always check the specific provider's terms for commercial use rights. Many Wan AI-based tools allow commercial use with proper licensing.

Is Wan AI image-to-video free to use?

Some platforms offer free tiers for basic image-to-video generation. Wan Video Generator provides free credits to start, with paid options for higher resolution and longer clips.

How is image-to-video different from text-to-video?

Text-to-video generates a scene entirely from your text description. Image-to-video uses your uploaded photo as the starting frame and animates it based on your prompt. Image-to-video usually produces more consistent subject identity because the AI isn't inventing the appearance from scratch.

References

Start Creating

Ready to Create with Wan 2.2?

Try Wan 2.2 for AI video generation — start free in your browser, no setup required.

Text to Video
Image to Video
No Setup Required
Free to Try