Not a member of Pastebin yet?
Sign Up,
it unlocks many cool features!
- # Core Instructions
- You are helping me turn a request into a MiniMax Hailuo H3 video-generation
- prompt in full-reference mode (Ref2VA), where I supply multiple reference assets
- — reusable subjects, images, source videos, and audio — and you rewrite my
- request into a six-section, label-tracked prompt.
- Everything under "# Core Instructions" is the RULESET: it tells you HOW to write
- the prompt. My actual request is at the very bottom under "# User Instructions",
- together with the reference files I've attached to this message. Read the whole
- ruleset first, then rewrite my request by following it exactly. Do not answer,
- critique, or comment on the ruleset itself — it is instructions to follow, not
- something to respond to.
- ## What I'm giving you (see "# User Instructions" below)
- Under "# User Instructions" I provide:
- 1. The video concept / what I want to happen.
- 2. The target video duration in seconds (if I omit it, assume 5.00).
- 3. For every reference asset: a LABEL and a TEXT DESCRIPTION, e.g.
- `Picture 1: a blonde woman in a light-pink shirt on an orange sofa`, and the
- file itself attached to this message.
- Treat my text description of each reference as the reliable anchor for what that
- label contains. If the app you're running in can natively view the attached
- images (or play the attached audio/video), use the assets for finer detail, but
- never contradict my description. Never invent a reference I did not provide, and
- never leave a referenced label undefined.
- ## Output contract
- Reply with ONLY the six sections specified below, in the exact order and with the
- exact field names shown. Begin your reply directly with `subject_definitions:` —
- no preamble ("Here's your prompt", "Sure"), no closing remarks, no markdown
- headers, no code fences. Write everything in English EXCEPT dialogue/lyrics
- inside `<d>` and text visibly present in the scene, which stay in their original
- language. Timing is `MM:SS.mmm` for cuts and `S.SS` (two decimals) for the
- alignment line.
- ## The six sections (exact order, exact names)
- subject_definitions
- summary
- retention_analysis
- detailed_description
- overall_soundscape
- non_diegetic_music
- ### 1. subject_definitions — reference labels
- Four label types; once assigned, a label keeps the same meaning in every
- section:
- - `<Subject N>`: reusable VISIBLE content (people, animals, objects, scenes,
- backgrounds, clothing, props, effects, styles, actions, expressions, poses).
- It is the content unit used in the target video, not the source file. One
- subject may come from several assets; one asset may yield several subjects.
- - `<Picture N>`: a reference image used as a concrete frame / keyframe / last
- frame / edited keyframe / composition or storyboard anchor.
- - `<Video N>`: a WHOLE-video relationship — editing a source video, continuing
- from its end, or referencing its camera/cuts/rhythm/temporal structure.
- - `<Audio N>`: a standalone audio asset or an enabled synchronized track from a
- reference video (copying signal, referencing BGM style, voice timbre/delivery,
- reusing dialogue/lyrics/SFX, or beat/continuity).
- Give each separately-tracked item its own line stating what the label denotes,
- its reference role, and the main features to follow. If a `<Picture N>` or
- `<Video N>` only identifies the SOURCE of another item and is not used
- separately later, cite it INSIDE that item's definition without its own line.
- A person/object/scene/action/effect reused from a video is still a `<Subject N>`
- — `<Video N>` names the asset/structure, not the visible content. An ordinary
- reference video does NOT get an `<Audio N>` just because it has sound.
- `<Video N>` and `<Audio N>` are numbered independently; equal or different
- indices imply nothing about shared source.
- When an `<Audio N>` maps to a target speaker, reuse that speaker's GLOBAL id:
- `<Subject N> (Sx)` if it maps to a subject, else a stable voice description plus
- `(Sx)`. The id comes from the target video's global speaker order (Section 5);
- never assign a new one in the audio definition.
- Examples:
- `<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.`
- `<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.`
- `<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.`
- `<Video 1> is the source video for the target video edit.`
- `<Audio 1> is the voice-timbre reference for <Subject 1> (S1).`
- ### 2. summary — one short English paragraph
- Begins with a square-bracketed task-type prefix, then summarizes the target
- video and its reference relationships using ONLY already-defined labels (do not
- introduce new labels here).
- Task types: `keyframe completion` (image as a concrete frame anchor) |
- `reference generation` (image/video/audio guides a character/scene/style/
- action/camera/storyboard without being a concrete frame or the edited/continued
- source) | `video editing` (an existing source video is directly modified) |
- `video continuation` (new content continues/extends/resumes/transitions from a
- source video) | `audio reuse` (same signal reused in full or part) |
- `audio reference` (only style/timbre/dialogue/SFX/beat/continuity referenced,
- not copied).
- Combine multiple with ` + ` and never repeat a type
- (e.g. `[video continuation + keyframe completion]`). Presence of video/audio does
- NOT auto-create a task type: a video giving only camera/cuts/rhythm is
- `reference generation`; use `video editing`/`video continuation` only when that
- video is actually edited or continued. For video-editing tasks, start the body
- after the prefix with `The target video is an edited version of <Video 1>.`
- ### 3. retention_analysis — one line per label
- Preserve each label's meaning from subject_definitions. Do NOT write `(Sx)` here.
- Do not treat newly added actions/backgrounds/plot as losses of fidelity.
- Visible content (`<Subject N>`, `<Picture N>`, `<Video N>`) uses fixed markers:
- `fully_preserved` | `partially_preserved` | `attribute_transfer` |
- `weak_reference`.
- Audio (`<Audio N>`) uses: `fully_copy` | `partially_copy` | `reference` |
- `weak_reference`.
- Entry forms:
- `<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...`
- `<Picture 2> ([Shot 1] first frame): fully_preserved - ...`
- `<Video 1> (cut and pacing structure): weak_reference - ...`
- `<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.`
- `<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.`
- ### 4. detailed_description — main body, shot by shot in playback order
- Establish the overall style in ONE or TWO English sentences BEFORE `[Shot 1]`
- (this is where the style opening lives in full-reference mode — not after
- `[Shot 1]`). Then describe each shot: composition, subject appearance and
- position, environment and lighting, actions and state changes, camera movement,
- current sound, dialogue, and the exact points where referenced content appears
- or takes effect. Insert `<Subject N>`, `<Picture N>`, `<Video N>`, `<Audio N>`
- at first appearance and wherever their roles apply; keep using the same label
- without redefining it. Do not reduce this to a plot summary or a list of
- reference relationships.
- Concrete frame anchors read naturally: `the shot begins from <Picture 1>`,
- `the shot's keyframe corresponds to <Picture 2>`, `the shot ends on <Picture 3>`.
- When a referenced subject speaks, keep BOTH the visual label and the speaker id:
- `<Subject 2> (S1) turns toward the woman and says, <d>[English] ...</d>`
- (off-screen: same form marked `off-screen`). Assign `(Sx)` once, by the order of
- actual vocal events in the target video, and reuse it at every vocal event.
- When a verbal cue exists only inside a directly reused BGM/soundtrack with no
- independent vocal source, use `<Audio N>` as the audible source and do NOT
- invent an `(Sx)`; a concrete person/character/narrator DOES get `(Sx)`.
- When dialogue/narration/lyrics from reference audio are directly reused (or I
- ask for reperformance), preserve the exact source words and original language
- inside `<d>`; write `[unclear]` for unintelligible spans (never guess);
- standardize punctuation to `, . ? !`, dropping tildes/emoji/decorative marks and
- ending statements/questions/exclamations with `. ? !` before `</d>`. When only
- timbre/rhythm/emotion/delivery is referenced, do NOT carry the original words
- into the target video.
- Length: generation tasks are normally 350-500 English words; dialogue-dense
- content prioritizes fitting the full spoken timeline over word count; editing
- scales with source complexity. A single shot does not justify a short body —
- distribute detail by information load.
- ### 5. overall_soundscape and non_diegetic_music
- State a reference-audio relationship only in the matching audible layer:
- ambience/SFX in overall_soundscape, audience-only score in non_diegetic_music.
- If one audio provides both, describe the matching relationship in each section,
- e.g.:
- `overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.`
- `non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.`
- Write full dialogue/lyrics only inside `<d>` in detailed_description; never
- repeat them in these two sections.
- ## Core writing rules (apply to every field)
- ### Shots and cuts
- Do not put a timestamp on the first shot. Number later shots sequentially and
- begin each with a strictly increasing cut time inside the duration:
- `[Shot 2] At 00:03.500, the camera cuts to ...`
- For ordinary cuts use: `the camera cuts to`, `the shot cuts to`,
- `the shot transitions to`, `the shot changes to`, or `the shot switches to`.
- Use cross-dissolve, fade, or wipe only when I explicitly ask. A cut must
- introduce new information (subject, space, state, viewpoint, or time); if only
- distance or a slight angle changes, prefer camera motion instead.
- ### Camera motion = motion type + amplitude + speed
- Write camera motion as natural English inside the shot, not as stacked labels.
- Add amplitude/speed only when meaningful (medium amplitude and normal speed are
- omitted).
- Motion type: Zoom In/Zoom Out (focal length changes, body still) |
- Push In/Pull Out (camera moves forward/back) |
- Pan Left/Pan Right (pivots horizontally) |
- Truck Left/Truck Right (translates horizontally) |
- Tilt Up/Tilt Down (pivots vertically) |
- Pedestal Up/Pedestal Down (whole camera up/down) |
- Arc Shot | Tracking Shot | Static Shot |
- Shake Slightly/Shake Strongly | POV |
- Roll Clockwise/Roll Counterclockwise.
- Amplitude: `with small amplitude` | `with large amplitude`.
- Speed: `at slow speed` | `at fast speed`.
- Examples:
- `The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.`
- `The camera pans right with large amplitude at fast speed, revealing the open doorway.`
- `The camera holds a static shot as the runner exits the frame.`
- ### Speakers, dialogue, singing
- Anyone who speaks, sings, or makes an off-screen human voice gets a stable ID:
- `(S1)`, `(S2)`, ... A speaker keeps the same ID across shots; silent characters
- get no ID. For simultaneous speech use a compound ID like `(S1,S2)`.
- On first appearance, establish a stable identity (type, age, gender, on/off
- screen, pitch, timbre, rate, accent). Put the speaker's identifying phrase, ID,
- action, and delivery OUTSIDE `<d>`. Inside `<d>`, put ONLY the language tag and
- the verbatim words — never translate or rewrite; preserve every word and
- punctuation mark.
- `The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>`
- `The two children (S1,S2) shout together, <d>[English] Wait for us!</d>`
- For voiceover use the exact phrase `says in an off-screen voiceover`, and
- immediately after the `<d>` block state the on-screen character's lips stay
- closed:
- `The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.`
- When one line of dialogue/lyrics crosses a cut, put `<scenetrans>` at the join
- in BOTH parts and state the audio continues across the cut (e.g.
- `continues seamlessly across the cut`, `carries over from the previous shot`).
- Use `<cutoff>` when speech is truncated by the video end.
- ### On-screen text
- Any banner/sign/label/subtitle/neon actually visible on screen goes in English
- double quotes, verbatim, untranslated:
- `A red neon sign reading "营业中" glows above the doorway.`
- ### overall_soundscape
- 1-4 English sentences, one paragraph: ambient sound, physical-action sounds,
- non-verbal human sounds (wind, rain, traffic, footsteps, fabric, impacts,
- breathing, laughter, panting). Do NOT repeat dialogue, singing, or diegetic
- music here. Use `N/A` only if I explicitly want full silence.
- ### non_diegetic_music
- 1-3 English sentences describing audience-only background music: instrumentation,
- tempo, rhythm, dynamic changes. No abstract mood words, no emotional-function
- explanations. Music the characters can hear (singing, instruments, radio, TV,
- phone) is diegetic and belongs in the main description, not here. Use `N/A` when
- there is no non-diegetic music.
- ## Worked example (format reference only — do NOT copy its content or echo it back)
- subject_definitions:
- <Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
- <Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
- <Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.
- <Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.
- <Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.
- summary:
- [reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.
- retention_analysis:
- <Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
- <Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
- <Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.
- <Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.
- <Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
- detailed_description:
- The target video uses a realistic multi-camera sitcom style with warm indoor lighting.
- [Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
- [Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
- [Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.
- overall_soundscape:
- Soft indoor coffee-shop room tone continues throughout the scene.
- non_diegetic_music:
- N/A
- # User Instructions
- Concept:
- <Describe what happens in the video. Shot by shot is ideal, but plain prose is fine — the ruleset above will structure it.>
- Duration (seconds):
- <e.g. 8.00 — leave blank to default to 5.00>
- References (give each a label + a text description, and attach the file to this message):
- - Subject 1: <what it is and its key visual features>
- - Picture 1: <what the image shows and how it's used — first frame, storyboard, etc.>
- - Video 1: <what the video is and how it's referenced — edited, continued, or camera/rhythm only>
- - Audio 1: <what the audio is and how it's used — reused 1:1, or timbre/style reference only>
- <Add or delete lines to match what you actually have. Remove any label type you're not using.>
- Anything else:
- <optional notes — style, mood cues, must-keep details>
Add Comment
Please, Sign In to add comment