Guest User

H3 Ref2V User instruction

a guest
Aug 9th, 2026
128
0
Never
Not a member of Pastebin yet? Sign Up, it unlocks many cool features!
text 17.84 KB | None | 0 0
  1. # Core Instructions
  2.  
  3. You are helping me turn a request into a MiniMax Hailuo H3 video-generation
  4. prompt in full-reference mode (Ref2VA), where I supply multiple reference assets
  5. — reusable subjects, images, source videos, and audio — and you rewrite my
  6. request into a six-section, label-tracked prompt.
  7.  
  8. Everything under "# Core Instructions" is the RULESET: it tells you HOW to write
  9. the prompt. My actual request is at the very bottom under "# User Instructions",
  10. together with the reference files I've attached to this message. Read the whole
  11. ruleset first, then rewrite my request by following it exactly. Do not answer,
  12. critique, or comment on the ruleset itself — it is instructions to follow, not
  13. something to respond to.
  14.  
  15. ## What I'm giving you (see "# User Instructions" below)
  16. Under "# User Instructions" I provide:
  17. 1. The video concept / what I want to happen.
  18. 2. The target video duration in seconds (if I omit it, assume 5.00).
  19. 3. For every reference asset: a LABEL and a TEXT DESCRIPTION, e.g.
  20. `Picture 1: a blonde woman in a light-pink shirt on an orange sofa`, and the
  21. file itself attached to this message.
  22.  
  23. Treat my text description of each reference as the reliable anchor for what that
  24. label contains. If the app you're running in can natively view the attached
  25. images (or play the attached audio/video), use the assets for finer detail, but
  26. never contradict my description. Never invent a reference I did not provide, and
  27. never leave a referenced label undefined.
  28.  
  29. ## Output contract
  30. Reply with ONLY the six sections specified below, in the exact order and with the
  31. exact field names shown. Begin your reply directly with `subject_definitions:` —
  32. no preamble ("Here's your prompt", "Sure"), no closing remarks, no markdown
  33. headers, no code fences. Write everything in English EXCEPT dialogue/lyrics
  34. inside `<d>` and text visibly present in the scene, which stay in their original
  35. language. Timing is `MM:SS.mmm` for cuts and `S.SS` (two decimals) for the
  36. alignment line.
  37.  
  38. ## The six sections (exact order, exact names)
  39. subject_definitions
  40. summary
  41. retention_analysis
  42. detailed_description
  43. overall_soundscape
  44. non_diegetic_music
  45.  
  46. ### 1. subject_definitions — reference labels
  47. Four label types; once assigned, a label keeps the same meaning in every
  48. section:
  49. - `<Subject N>`: reusable VISIBLE content (people, animals, objects, scenes,
  50. backgrounds, clothing, props, effects, styles, actions, expressions, poses).
  51. It is the content unit used in the target video, not the source file. One
  52. subject may come from several assets; one asset may yield several subjects.
  53. - `<Picture N>`: a reference image used as a concrete frame / keyframe / last
  54. frame / edited keyframe / composition or storyboard anchor.
  55. - `<Video N>`: a WHOLE-video relationship — editing a source video, continuing
  56. from its end, or referencing its camera/cuts/rhythm/temporal structure.
  57. - `<Audio N>`: a standalone audio asset or an enabled synchronized track from a
  58. reference video (copying signal, referencing BGM style, voice timbre/delivery,
  59. reusing dialogue/lyrics/SFX, or beat/continuity).
  60. Give each separately-tracked item its own line stating what the label denotes,
  61. its reference role, and the main features to follow. If a `<Picture N>` or
  62. `<Video N>` only identifies the SOURCE of another item and is not used
  63. separately later, cite it INSIDE that item's definition without its own line.
  64. A person/object/scene/action/effect reused from a video is still a `<Subject N>`
  65. — `<Video N>` names the asset/structure, not the visible content. An ordinary
  66. reference video does NOT get an `<Audio N>` just because it has sound.
  67. `<Video N>` and `<Audio N>` are numbered independently; equal or different
  68. indices imply nothing about shared source.
  69. When an `<Audio N>` maps to a target speaker, reuse that speaker's GLOBAL id:
  70. `<Subject N> (Sx)` if it maps to a subject, else a stable voice description plus
  71. `(Sx)`. The id comes from the target video's global speaker order (Section 5);
  72. never assign a new one in the audio definition.
  73. Examples:
  74. `<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.`
  75. `<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.`
  76. `<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.`
  77. `<Video 1> is the source video for the target video edit.`
  78. `<Audio 1> is the voice-timbre reference for <Subject 1> (S1).`
  79.  
  80. ### 2. summary — one short English paragraph
  81. Begins with a square-bracketed task-type prefix, then summarizes the target
  82. video and its reference relationships using ONLY already-defined labels (do not
  83. introduce new labels here).
  84. Task types: `keyframe completion` (image as a concrete frame anchor) |
  85. `reference generation` (image/video/audio guides a character/scene/style/
  86. action/camera/storyboard without being a concrete frame or the edited/continued
  87. source) | `video editing` (an existing source video is directly modified) |
  88. `video continuation` (new content continues/extends/resumes/transitions from a
  89. source video) | `audio reuse` (same signal reused in full or part) |
  90. `audio reference` (only style/timbre/dialogue/SFX/beat/continuity referenced,
  91. not copied).
  92. Combine multiple with ` + ` and never repeat a type
  93. (e.g. `[video continuation + keyframe completion]`). Presence of video/audio does
  94. NOT auto-create a task type: a video giving only camera/cuts/rhythm is
  95. `reference generation`; use `video editing`/`video continuation` only when that
  96. video is actually edited or continued. For video-editing tasks, start the body
  97. after the prefix with `The target video is an edited version of <Video 1>.`
  98.  
  99. ### 3. retention_analysis — one line per label
  100. Preserve each label's meaning from subject_definitions. Do NOT write `(Sx)` here.
  101. Do not treat newly added actions/backgrounds/plot as losses of fidelity.
  102. Visible content (`<Subject N>`, `<Picture N>`, `<Video N>`) uses fixed markers:
  103. `fully_preserved` | `partially_preserved` | `attribute_transfer` |
  104. `weak_reference`.
  105. Audio (`<Audio N>`) uses: `fully_copy` | `partially_copy` | `reference` |
  106. `weak_reference`.
  107. Entry forms:
  108. `<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...`
  109. `<Picture 2> ([Shot 1] first frame): fully_preserved - ...`
  110. `<Video 1> (cut and pacing structure): weak_reference - ...`
  111. `<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.`
  112. `<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.`
  113.  
  114. ### 4. detailed_description — main body, shot by shot in playback order
  115. Establish the overall style in ONE or TWO English sentences BEFORE `[Shot 1]`
  116. (this is where the style opening lives in full-reference mode — not after
  117. `[Shot 1]`). Then describe each shot: composition, subject appearance and
  118. position, environment and lighting, actions and state changes, camera movement,
  119. current sound, dialogue, and the exact points where referenced content appears
  120. or takes effect. Insert `<Subject N>`, `<Picture N>`, `<Video N>`, `<Audio N>`
  121. at first appearance and wherever their roles apply; keep using the same label
  122. without redefining it. Do not reduce this to a plot summary or a list of
  123. reference relationships.
  124. Concrete frame anchors read naturally: `the shot begins from <Picture 1>`,
  125. `the shot's keyframe corresponds to <Picture 2>`, `the shot ends on <Picture 3>`.
  126. When a referenced subject speaks, keep BOTH the visual label and the speaker id:
  127. `<Subject 2> (S1) turns toward the woman and says, <d>[English] ...</d>`
  128. (off-screen: same form marked `off-screen`). Assign `(Sx)` once, by the order of
  129. actual vocal events in the target video, and reuse it at every vocal event.
  130. When a verbal cue exists only inside a directly reused BGM/soundtrack with no
  131. independent vocal source, use `<Audio N>` as the audible source and do NOT
  132. invent an `(Sx)`; a concrete person/character/narrator DOES get `(Sx)`.
  133. When dialogue/narration/lyrics from reference audio are directly reused (or I
  134. ask for reperformance), preserve the exact source words and original language
  135. inside `<d>`; write `[unclear]` for unintelligible spans (never guess);
  136. standardize punctuation to `, . ? !`, dropping tildes/emoji/decorative marks and
  137. ending statements/questions/exclamations with `. ? !` before `</d>`. When only
  138. timbre/rhythm/emotion/delivery is referenced, do NOT carry the original words
  139. into the target video.
  140. Length: generation tasks are normally 350-500 English words; dialogue-dense
  141. content prioritizes fitting the full spoken timeline over word count; editing
  142. scales with source complexity. A single shot does not justify a short body —
  143. distribute detail by information load.
  144.  
  145. ### 5. overall_soundscape and non_diegetic_music
  146. State a reference-audio relationship only in the matching audible layer:
  147. ambience/SFX in overall_soundscape, audience-only score in non_diegetic_music.
  148. If one audio provides both, describe the matching relationship in each section,
  149. e.g.:
  150. `overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.`
  151. `non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.`
  152. Write full dialogue/lyrics only inside `<d>` in detailed_description; never
  153. repeat them in these two sections.
  154.  
  155. ## Core writing rules (apply to every field)
  156.  
  157. ### Shots and cuts
  158. Do not put a timestamp on the first shot. Number later shots sequentially and
  159. begin each with a strictly increasing cut time inside the duration:
  160. `[Shot 2] At 00:03.500, the camera cuts to ...`
  161. For ordinary cuts use: `the camera cuts to`, `the shot cuts to`,
  162. `the shot transitions to`, `the shot changes to`, or `the shot switches to`.
  163. Use cross-dissolve, fade, or wipe only when I explicitly ask. A cut must
  164. introduce new information (subject, space, state, viewpoint, or time); if only
  165. distance or a slight angle changes, prefer camera motion instead.
  166.  
  167. ### Camera motion = motion type + amplitude + speed
  168. Write camera motion as natural English inside the shot, not as stacked labels.
  169. Add amplitude/speed only when meaningful (medium amplitude and normal speed are
  170. omitted).
  171. Motion type: Zoom In/Zoom Out (focal length changes, body still) |
  172. Push In/Pull Out (camera moves forward/back) |
  173. Pan Left/Pan Right (pivots horizontally) |
  174. Truck Left/Truck Right (translates horizontally) |
  175. Tilt Up/Tilt Down (pivots vertically) |
  176. Pedestal Up/Pedestal Down (whole camera up/down) |
  177. Arc Shot | Tracking Shot | Static Shot |
  178. Shake Slightly/Shake Strongly | POV |
  179. Roll Clockwise/Roll Counterclockwise.
  180. Amplitude: `with small amplitude` | `with large amplitude`.
  181. Speed: `at slow speed` | `at fast speed`.
  182. Examples:
  183. `The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.`
  184. `The camera pans right with large amplitude at fast speed, revealing the open doorway.`
  185. `The camera holds a static shot as the runner exits the frame.`
  186.  
  187. ### Speakers, dialogue, singing
  188. Anyone who speaks, sings, or makes an off-screen human voice gets a stable ID:
  189. `(S1)`, `(S2)`, ... A speaker keeps the same ID across shots; silent characters
  190. get no ID. For simultaneous speech use a compound ID like `(S1,S2)`.
  191. On first appearance, establish a stable identity (type, age, gender, on/off
  192. screen, pitch, timbre, rate, accent). Put the speaker's identifying phrase, ID,
  193. action, and delivery OUTSIDE `<d>`. Inside `<d>`, put ONLY the language tag and
  194. the verbatim words — never translate or rewrite; preserve every word and
  195. punctuation mark.
  196. `The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>`
  197. `The two children (S1,S2) shout together, <d>[English] Wait for us!</d>`
  198. For voiceover use the exact phrase `says in an off-screen voiceover`, and
  199. immediately after the `<d>` block state the on-screen character's lips stay
  200. closed:
  201. `The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.`
  202. When one line of dialogue/lyrics crosses a cut, put `<scenetrans>` at the join
  203. in BOTH parts and state the audio continues across the cut (e.g.
  204. `continues seamlessly across the cut`, `carries over from the previous shot`).
  205. Use `<cutoff>` when speech is truncated by the video end.
  206.  
  207. ### On-screen text
  208. Any banner/sign/label/subtitle/neon actually visible on screen goes in English
  209. double quotes, verbatim, untranslated:
  210. `A red neon sign reading "营业中" glows above the doorway.`
  211.  
  212. ### overall_soundscape
  213. 1-4 English sentences, one paragraph: ambient sound, physical-action sounds,
  214. non-verbal human sounds (wind, rain, traffic, footsteps, fabric, impacts,
  215. breathing, laughter, panting). Do NOT repeat dialogue, singing, or diegetic
  216. music here. Use `N/A` only if I explicitly want full silence.
  217.  
  218. ### non_diegetic_music
  219. 1-3 English sentences describing audience-only background music: instrumentation,
  220. tempo, rhythm, dynamic changes. No abstract mood words, no emotional-function
  221. explanations. Music the characters can hear (singing, instruments, radio, TV,
  222. phone) is diegetic and belongs in the main description, not here. Use `N/A` when
  223. there is no non-diegetic music.
  224.  
  225. ## Worked example (format reference only — do NOT copy its content or echo it back)
  226. subject_definitions:
  227. <Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
  228. <Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
  229. <Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.
  230. <Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.
  231. <Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.
  232.  
  233. summary:
  234. [reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.
  235.  
  236. retention_analysis:
  237. <Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
  238. <Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
  239. <Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.
  240. <Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.
  241. <Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
  242.  
  243. detailed_description:
  244. The target video uses a realistic multi-camera sitcom style with warm indoor lighting.
  245. [Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
  246. [Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
  247. [Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.
  248.  
  249. overall_soundscape:
  250. Soft indoor coffee-shop room tone continues throughout the scene.
  251.  
  252. non_diegetic_music:
  253. N/A
  254.  
  255. # User Instructions
  256.  
  257. Concept:
  258. <Describe what happens in the video. Shot by shot is ideal, but plain prose is fine — the ruleset above will structure it.>
  259.  
  260. Duration (seconds):
  261. <e.g. 8.00 — leave blank to default to 5.00>
  262.  
  263. References (give each a label + a text description, and attach the file to this message):
  264. - Subject 1: <what it is and its key visual features>
  265. - Picture 1: <what the image shows and how it's used — first frame, storyboard, etc.>
  266. - Video 1: <what the video is and how it's referenced — edited, continued, or camera/rhythm only>
  267. - Audio 1: <what the audio is and how it's used — reused 1:1, or timbre/style reference only>
  268. <Add or delete lines to match what you actually have. Remove any label type you're not using.>
  269.  
  270. Anything else:
  271. <optional notes — style, mood cues, must-keep details>
Add Comment
Please, Sign In to add comment