How to Make a “Day in Seoul” AI Video from One Photo: Seedance Prompts and Real Examples
Learn character consistency, early-2000s DV texture, 30-second story structure, and ambient audio from real Seoul videos in the Pixocto Gallery, with reusable Seedance prompts.
How to Make a “Day in Seoul” AI Video from One Photo: Seedance Prompts and Real Examples
A convincing “Day in Seoul” AI video is not a person placed in front of N Seoul Tower or Myeongdong with a retro filter. What makes it feel like an old recording is more specific: a quiet residential lane, an ordinary action, a camera held casually by a friend, and autofocus and exposure that occasionally get things wrong.
This guide draws on multiple Seedance videos from the Pixocto Gallery. Upload one character photo, copy a prompt below, then replace the person, season, and events to create your own early-2000s Seoul home video.
Quick answer: what is this video made of?
The basic formula is:
- A consistent main character
- A quiet, everyday Seoul setting
- Four to six small events
- Early consumer DV camcorder imperfections
- Natural ambient sound
- A clear timeline
- Consistent character, clothing, and objects
- An abrupt cut to black
It may look like an entire day, but the number of events is not what matters. The feeling you want is that someone happened to pick up a camcorder and record an insignificant piece of everyday life.
Upload a photo and start making your Day in Seoul →
Start with the closest Gallery example
Watch the full video and view its original prompt
There are no palaces, shopping districts, or sweeping city views here. A young woman hears a bicycle bell, notices a fallen leaf, tries to balance it on a bicycle seat, and eventually walks home laughing.
The story works because every detail helps it read as a real recording:
- Concrete lanes, older apartments, plants, bicycles, and utility poles create a lived-in setting.
- The camera shakes, loses focus, shifts exposure, and zooms awkwardly.
- Reactions are small; the subject is not performing for the camera.
- The soundtrack is only footsteps, insects, wind, leaves, and distant traffic.
- The recording cuts off abruptly instead of fading out.
Step 1: choose a suitable reference photo
To make yourself the subject, use one clear, natural single-person photo without heavy filters.
- Keep the face unobstructed; front-facing or a slight angle both work.
- A half-body or full-body image helps the model understand clothing and proportions.
- Everyday clothes fit this format better than formal or stage outfits.
- The original background does not need to be Seoul; the prompt will rebuild it.
- Avoid group photos, dark images, or already distorted AI portraits.
At the beginning of the prompt, tell the model to treat the upload as the character source of truth. Explicitly lock the face, hair, skin tone, body proportions, outfit, and accessories.
Use the uploaded photo as the exact character reference. Preserve the same face, hairstyle, skin tone, body proportions, outfit, accessories, and identity throughout the entire video. No character redesign, wardrobe change, facial change, or identity drift.
Test character consistency in the Seedance workspace →
Step 2: choose everyday Seoul, not sightseeing Seoul
Watch the full video and view its original prompt
This popular Gallery example stays in an older Seoul neighborhood: low-rise homes, rooftop terraces, outside stairs, laundry lines, plants, bicycles, overhead wires, and concrete walls. The subject drinks on a rooftop, walks through an alley, misses a basketball shot, and runs for cover in a sudden summer shower.
Locations like these create a memory more easily than landmarks. If N Seoul Tower, Myeongdong, Gyeongbokgung, street food, and neon shopping streets all appear in the same clip, the result starts to feel like a tourism commercial.
This cinematic Seoul travel example shows that alternative direction: vertical framing, HDR, smooth tracking shots, landmarks, and polished grading. It works as a travel reel, but not as the forgotten home video we are building here.
Use this for the neighborhood:
Location: A quiet older residential neighborhood in Seoul during a warm summer afternoon. Narrow concrete alleys, low-rise apartments, rooftop terraces, external staircases, potted plants, laundry lines, parked bicycles, utility poles, overhead wires, mature trees, old walls, and a tiny neighborhood shop. Lived-in, peaceful, and authentic. No tourist landmarks, crowds, advertisements, recognizable brands, or polished commercial activity.
Step 3: do not make it too beautiful
A retro color filter is not enough. The home-video look also comes from how the camera behaves.
| Element | What to request | What to avoid |
|---|---|---|
| Movement | Friend-held camera, slight shake, occasionally imperfect framing | Gimbal, drone, smooth orbit |
| Focus | Autofocus briefly hunts, then finds the face | Perfect sharpness throughout |
| Exposure | Small shifts between sun and shade | HDR and studio lighting |
| Color | Faded, low-contrast, imperfect white balance | Saturated commercial grading |
| Image | Soft detail, mild noise and compression | Oversharpened modern digital image |
| Operator | Reacts slightly late or clips the subject | Perfectly anticipates every action |
Use this camera block directly:
Camera style: Raw footage recorded by a friend using a cheap early-2000s consumer DV camcorder. Natural handheld shake, imperfect framing, autofocus hunting, slight exposure shifts between sunlight and shade, faded colors, soft contrast, mild digital noise, natural motion blur, imperfect white balance, accidental zooms, and subtle camera-handling noise. The camera operator occasionally reacts late or briefly loses focus. No stabilization, gimbal, drone, slow motion, cinematic choreography, dramatic lighting, HDR, or modern commercial color grading.
Step 4: use tiny events instead of a dramatic plot
The most common mistake is packing too many places and conflicts into 30 seconds. The model must maintain the character, objects, environment, and motion at the same time; every complicated event increases the chance of visual errors.
Reliable events contain one action, one small reaction, and one audible sound:
- A cold bottle touches the subject’s cheek and they laugh.
- A leaf lands in their hair; a friend points it out.
- A basketball misses and the subject quietly shakes their head.
- Wind rings a bicycle bell and the subject looks for the sound.
- A jar refuses to open, then finally pops loose on the third try.
Watch this 15-second indoor example
The whole video is about opening one stubborn jar. Attempt, failure, another attempt, and success already make a complete short story. The smaller the scene, the easier it is for subtle reactions and room sound to feel real.
15-second starter prompt
Do not begin with a whole day. First test character consistency, camera texture, and sound with one location and one small event.
Create a 15-second, 16:9 ultra-realistic early-2000s DV home video using the uploaded photo as the exact character reference.
SUBJECT
Preserve the same face, hairstyle, skin tone, body proportions, casual outfit, and identity throughout. Natural skin texture, relaxed expression, no beauty filter, no identity drift or wardrobe change.
LOCATION
A quiet older Seoul residential alley on a warm late-summer afternoon: narrow concrete lane, low-rise apartments, potted plants, bicycles, utility poles, overhead wires, leafy trees, and old walls. No landmarks, crowds, advertisements, or recognizable brands.
CAMERA
Friend-recorded footage from a cheap early-2000s consumer DV camcorder. Handheld shake, imperfect framing, autofocus hunting, faded colors, soft contrast, slight exposure shifts, mild digital noise, natural motion blur, imperfect white balance, and one awkward accidental zoom. No stabilization, gimbal, drone, slow motion, cinematic lighting, HDR, or modern commercial color grading.
00:00–00:04 — The subject walks slowly down the lane holding a cold bottled drink. A breeze moves their hair and clothes. The camera briefly struggles to focus.
00:04–00:08 — A dry leaf lands on the subject's shoulder. They do not notice at first. The camera operator quietly laughs.
00:08–00:12 — The subject notices the leaf, removes it, looks toward the camera with a mildly embarrassed smile, and laughs naturally.
00:12–00:15 — The subject continues walking, turns back for a tiny wave, then looks away. The recording abruptly cuts to black mid-motion. No fade-out.
AUDIO
Natural location sound only: footsteps, distant traffic, summer insects, leaves, wind, bicycle bell, neighborhood voices, and subtle camera-handling noise. No music, narration, subtitles, captions, logos, watermarks, or artificial sound effects.
REALISM
Believable walking, hand movement, hair, clothing, leaf physics, and foot contact with the ground. No distorted hands, extra fingers, duplicated people, teleportation, floating objects, CGI appearance, exaggerated acting, or sudden transformations.
Copy the starter prompt and generate a 15-second video →
Complete 30-second prompt
Once the character and look are stable, expand to 30 seconds. This version uses six connected small moments and is suitable for Seedance 2.5.
Create a 30-second, 1080p, 16:9 ultra-realistic early-2000s DV home video showing an ordinary summer afternoon in the life of the person in the uploaded reference photo. The footage should feel spontaneous, intimate, imperfect, and genuinely observed rather than performed.
MAIN SUBJECT
Use the uploaded photo as the exact character reference. Preserve the same face, hairstyle, skin tone, body proportions, casual outfit, shoes, bag, accessories, and identity from beginning to end. Natural skin texture, minimal styling, relaxed expressions. No character redesign, wardrobe change, facial change, or identity drift.
LOCATION
A quiet older residential neighborhood in Seoul: narrow concrete alleys, low-rise homes, rooftop terraces, external staircases, potted plants, laundry lines, parked bicycles, utility poles, overhead wires, mature trees, old walls, and a tiny neighborhood shop. Authentic and lived-in. No crowds, tourist landmarks, advertisements, recognizable brands, or polished commercial activity.
CAMERA STYLE
Raw personal footage recorded by a friend using a cheap early-2000s consumer DV camcorder. Natural handheld shake, imperfect framing, autofocus hunting, exposure shifts between sunlight and shade, faded colors, soft contrast, mild digital noise, natural motion blur, imperfect white balance, accidental zooms, and subtle camera-handling noise. The camera operator sometimes reacts late or briefly loses focus. No stabilization, gimbal, drone, slow motion, cinematic choreography, dramatic lighting, HDR, or modern commercial color grading.
00:00–00:05 — LEAVING HOME
The subject steps out of an old apartment, checks their keys, adjusts their bag, notices the camera, gives a small smile, and starts walking down the lane. The camera takes a moment to find focus.
00:05–00:10 — THE COLD DRINK
At a tiny neighborhood shop, the subject buys a cold bottled drink. Outside, they press the bottle briefly against their cheek because of the summer heat and laugh at the camera operator's awkward zoom.
00:10–00:15 — A SMALL INTERRUPTION
While walking through a quieter alley, a bicycle passes. The subject steps aside naturally, waits, and continues. A breeze moves their hair and clothes. They glance back but do not pose.
00:15–00:20 — THE LEAF
The subject sits on a low wall beside a small apartment garden. A dry leaf lands in their hair. The camera operator points it out. The subject removes it and gives the camera an embarrassed, amused look.
00:20–00:25 — SUMMER RAIN
A light summer shower suddenly begins. The subject looks toward the sky, laughs, protects the drink with one hand, and walks quickly toward a covered entrance. Show believable rain, wet concrete, soft reflections, damp hair, and realistic foot contact.
00:25–00:30 — GOODBYE
Under the apartment entrance, the subject catches their breath, wipes rain from their forehead, gives a small genuine smile, and says, “See you tomorrow.” They walk inside. The camera remains pointed at the closing door for one second, then abruptly cuts to black mid-motion. No fade-out.
AUDIO
Natural environmental audio only: footsteps on concrete, distant traffic, birds, summer insects, leaves, bicycle wheels, shop refrigerator hum, bottle opening, light rain, water dripping from rooftops, quiet neighborhood voices, and subtle camera-handling noise. Dialogue should sound casual and local. No music, narration, soundtrack, subtitles, captions, logos, watermarks, or artificial sound effects.
CONTINUITY AND PHYSICAL REALISM
Maintain believable anatomy, movement, object permanence, clothing, hair, rain interaction, reflections, and environmental continuity. Hands and fingers remain natural. The drink stays a separate physical object and never intersects with the face or hand. No duplicated people, distorted anatomy, teleportation, disappearing objects, floating objects, sudden transformations, or CGI appearance.
FINAL FEEL
Like a forgotten personal recording of an ordinary summer day. Not a commercial, fashion film, music video, or travel advertisement. Quiet, warm, nostalgic, spontaneous, and deeply human. The charm comes from small moments and the feeling that the camera happened to be there.
Use Seedance 2.5 to make the full 30-second version →
How to keep interaction natural with a friend or partner
Watch the full video and view its original prompt
For two people, do not merely ask for “happy interaction.” Give each exchange one executable action: pass a drink, point out a leaf, step aside for a bicycle, or wave goodbye.
This Korean couple vlog example packs a photo booth, convenience store, arcade, and hand-holding into 30 seconds. It is richer but harder to keep consistent. For a first two-person attempt, stay within one or two locations and repeat that both people’s appearances, outfits, and relationship remain unchanged.
Keep the same two people throughout. Preserve both faces, hairstyles, body proportions, outfits, accessories, and relationship continuity. Their chemistry should feel familiar and unstaged. Use small physical interactions—passing a drink, briefly blocking the camera, pointing out a leaf, sharing a laugh—rather than exaggerated romantic poses. No face swapping, duplicated people, wardrobe changes, or sudden changes in height and age.
Reusable template: change Seoul to any city
Do not change only the city name. Regional character comes from housing, street objects, weather, sound, and ordinary local actions.
Create a [15/30]-second [ASPECT RATIO] ultra-realistic early-2000s DV home video using the uploaded photo as the exact character reference.
Subject: [AGE / APPEARANCE / OUTFIT]. Preserve the same identity, face, hairstyle, proportions, outfit, and accessories throughout.
Location: A quiet older residential neighborhood in [CITY] during [SEASON / TIME / WEATHER]. Include [LOCAL HOUSING], [STREET MATERIAL], [PLANTS], [UTILITY DETAILS], [SMALL LOCAL SHOP OR DOMESTIC SPACE], and [DISTANT AMBIENT ACTIVITY]. No tourist landmarks, crowds, recognizable brands, or advertisements.
Camera: Raw friend-recorded early-2000s consumer DV footage—handheld shake, imperfect framing, autofocus hunting, exposure shifts, faded colors, soft contrast, mild digital noise, imperfect white balance, accidental zooms, and natural camera mistakes. No stabilization, gimbal, drone, HDR, or polished cinematic grading.
Timeline:
00:00–[TIME] — [ONE SIMPLE ACTION].
[TIME]–[TIME] — [SMALL INTERRUPTION OR DISCOVERY].
[TIME]–[TIME] — [NATURAL REACTION].
[TIME]–[END] — [QUIET GOODBYE OR WALKING AWAY], followed by an abrupt cut to black.
Audio: Natural location sound only—[CITY-SPECIFIC AMBIENCE], footsteps, distant traffic, wind, birds or insects, room tone, and subtle camera-handling noise. No music, narration, subtitles, logos, or artificial sound effects.
Realism: Consistent identity and environment, natural anatomy, believable object physics, subtle acting, and no teleportation, duplication, distortion, or sudden transformations.
Replace the city and events, then generate your own version →
For Taipei, add arcades, older apartments, scooters, drink-shop refrigerators, and a humid summer evening. For Tokyo, use low-rise homes, narrow lanes, bicycle parking, distant trains, and a convenience-store door chime.
Common failure modes
1. Asking for an old tape and a cinematic blockbuster at once
cinematic lighting, HDR, perfect composition, and smooth gimbal movement contradict the home-video direction. Explicitly exclude them.
2. Packing too many locations into 30 seconds
An apartment, subway, Myeongdong, palace, market, and the Han River will become a sightseeing montage and increase scene jumps and identity drift. Keep the story in one neighborhood.
3. Describing emotion instead of action
“She has a happy day” cannot direct a shot. Replace it with something filmable: “She tastes a sour drink, winces, then laughs toward the camera.”
4. Removing every imperfection
Missed focus, late reactions, off-center framing, and exposure shifts are authenticity signals. The cleaner the image, the more it resembles a modern advertisement.
5. Ignoring object continuity
If the subject carries a bottle, bag, or umbrella throughout, lock it in the continuity section. Otherwise it may switch hands, deform, or disappear.
6. Covering natural sound with music
Insects, footsteps, refrigerator hum, bottle caps, and rain on concrete establish space better than a soundtrack. Generate clean ambient audio first; add music later only if the publishing platform needs it.
Generate it in Pixocto
- Open the video workspace and choose Seedance.
- Upload one clear single-person reference photo.
- Start with the 15-second prompt to test the character and visual direction.
- Choose 16:9 for a home-video feel, or 9:16 for Reels, TikTok, and Shorts.
- Use Seedance 2.5 when you want the full 30-second narrative.
- Enable audio generation and keep the natural ambience.
- Once the character is stable, change the place, weather, and small events without rewriting every constraint.
This format does not need a dramatic plot. Give the subject one ordinary thing to do, let the camera occasionally fall behind, and keep the wind, footsteps, and distant neighborhood sound. That is enough for a believable memory to begin.



