Build a Scroll-Storytelling Website With AI Image + Video
Build HOTARU, a scroll-choreographed ink-wash firefly river: 7 gpt-image-2 plates, MiniMax H3 Max image-to-video, a canvas swarm, and the full prompt pack.

Quick Answer
What you'll build: HOTARU (蛍), a seven-section, scroll-choreographed website in a Japanese ink-wash style — a pinned hero, a poem field, four painted scenes, and a night-river call to action, with a live canvas firefly swarm over the top. Stills come from gpt-image-2. Motion comes from MiniMax H3 Max image-to-video. The build is a single self-contained HTML file with no frameworks and no CDN.
Total cost: $0.84. Every prompt, price, and bug is in this article.
Most AI-generated websites look generic for one reason: the model is asked to be creative instead of being constrained. "Design a beautiful landing page" gives the model a million mediocre answers and it picks the average of all of them.
This is the complete production record for one that doesn't look generic. It is a real build, not a hypothetical: seven generated plates, two generated video clips, a hand-rolled particle system, and roughly $0.84 of model spend. You can open the finished piece below and scroll through it.
The result, before the how
This is the actual site. Scroll inside the frame — the header nav highlights the section you're in, the fireflies drift, and the night-river section loops its own ambient clip. If the embed feels cramped, open it full-screen.
Note:
Scroll-driven pages are awkward inside an iframe — your mouse wheel scrolls the parent article until the cursor sits over the frame. If you're here for the motion, open it in its own tab.
Why this one looks designed
The site works because of what was decided before a single pixel was generated. The art direction was locked first, and the models were then forced to obey it.
The palette is seven colors, and only one of them is allowed to be loud. Near-black ink for the night, pale parchment for the paper, two green tones for the fireflies, and a single cinnabar red for the seal stamp. That's it.
| Role | Hex |
|---|---|
| Parchment (light sections) | #e8e0cc, #efeadf |
| Ink charcoal (night, text) | #211b15, #2a2521, #39332c |
| Deep indigo night (hero, CTA) | #191a24, #232533 |
| Muted water gray-green | #7f8d83, #9aa79a, #c6ccc1 |
| Firefly green (glow) | #9fd48a, #517a4f |
| Seal red (one accent) | #9c2d24 |
| Bone (text on dark) | #f1ebdf |
The type is high-contrast serif, system-only. A Didot / Bodoni 72 / Canela stack, no webfont request, no FOUT. The hero wordmark is set at clamp(64px, 15vw, 300px) with wide tracking, and it is real text — never baked into the images. That single decision is what separates a site from a picture of a site: text stays selectable, responsive, and translatable.
The composition is asymmetric on purpose. Every plate prompt asks for "generous negative space" on a named side. The hero leaves the right-of-center open for the wordmark; the bridge scene leaves the upper right for its heading. Nothing is centered and nothing is crowded.
Key Takeaway
The model doesn't need to be told to be tasteful. It needs to be told, precisely, what not to do: "no text, no UI, no people, painterly and restrained." Negative constraints are what stop AI output from looking like AI output.
This is the same discipline the Taste-Skill enforces for frontend work: a checklist of bans beats a vibe.
The recipe
Five steps, in order. The order matters — generating assets before the art direction is locked is how you end up with a folder of pretty images that don't fit together.
Source a structure, not an idea
Steal the choreography from something that already works, then swap the subject. This build re-art-directs the scroll structure of DesignCode's "Sakura River Scroll Recreation Prompt" — same pacing, same section rhythm, different world.
Lock the art direction
Palette, type stack, section list, and the negative constraints. Written down before any generation. This becomes the spec every prompt inherits.
Generate the plates with gpt-image-2
Seven 16:9 stills, batched, converted to WebP. About four cents total.
Generate motion with image-to-video
Two clips — the hero and the night CTA — animated from their approved plates with MiniMax H3 Max.
Build the page
One HTML file. A canvas-2D firefly system, a scroll-choreography loop, and a legible fixed header.
Step 1 — The plates
Seven prompts, all generated with gpt-image-2 at 16:9. The shared suffix does the heavy lifting: sumi-e dry brush, handmade-paper grain, painterly and restrained, and the three prohibitions — no text, no UI, no people.
Here is the hero, and the prompt that produced it.

Original Japanese ink-wash web background plate, 16:9, sumi-e dry brush on pale aged parchment, deep charcoal-indigo night sky, a quiet river bend with reeds, faint green fireflies (two green tones, varied brightness) drifting above the water, soft ink reflections, one dark branch entering offscreen top-left, editorial art-book composition, large open negative space right-of-center for typography, textured handmade-paper grain, muted palette (#211b15 shadows, #9fd48a/#517a4f fireflies, #e8e0cc bone parchment), no text, no UI, no people, painterly and restrained
An alternate hero was generated in the same batch — same world, different framing — so there was a real choice at sign-off rather than a single take-or-leave result.

Original Japanese ink-wash web background plate, 16:9, deep indigo night sky over pale parchment water, sumi-e dry brush texture, a wide quiet river bending through a dark reed bank, dozens of small green fireflies glowing above the water in graduated brightness, moonless summer night, generous negative space on the right half for a large serif wordmark, handmade paper grain, muted palette (#191a24 sky, #232533 deep indigo, #9fd48a/#517a4f fireflies, #e8e0cc parchment water), no text, no UI, no people, painterly and restrained
The four scene plates carry the middle of the page. Each names its own composition constraint — which side stays open for text, and how dense the fireflies should be.

Original Japanese ink-wash web background plate, 16:9, sumi-e dry brush on pale handmade parchment, dusk mood, dark silhouetted reeds and riverside grass leaning over calm water, several small green fireflies drifting low above the water surface, soft ink reflections, deep charcoal-indigo sky, editorial art-book composition, generous open parchment space on the left for section text, textured paper grain, muted palette (#211b15, #2a2521, #9fd48a, #e8e0cc), no text, no UI, no people, painterly and restrained

Original Japanese ink-wash web background plate, 16:9, sumi-e dry brush on pale aged parchment, a small dark wooden footbridge over a narrow river with a single stone lantern at its end, faint green fireflies drifting around the lantern glow and over the water, soft reflections, deep indigo summer night, editorial art-book composition, clean negative space upper-right for typography, handmade paper grain, muted palette (#211b15, #39332c, #9fd48a, #9c2d24 tiny lantern accent, #e8e0cc), no text, no UI, no people, painterly and restrained

Original Japanese ink-wash web background plate, 16:9, sumi-e dry brush on pale parchment, a quiet bamboo grove edge at dusk with a narrow earthen path, tall bamboo culms in layered ink gray-greens, faint green fireflies threading between the culms, deep indigo sky, editorial art-book composition, generous open negative space lower-right for text, textured handmade-paper grain, muted palette (#211b15, #2a2521, #7f8d83, #9fd48a, #e8e0cc), no text, no UI, no people, painterly and restrained

Original Japanese ink-wash web background plate, 16:9, sumi-e dry brush on pale parchment, distant layered mountain silhouettes fading into mist over a wide still river, small green fireflies drifting above the water with soft glowing reflections, deep indigo dusk sky, editorial art-book composition, large open negative space upper-left for typography, handmade paper grain, muted palette (#232533, #191a24, #9aa79a, #c6ccc1 water, #9fd48a fireflies, #e8e0cc), no text, no UI, no people, painterly and restrained
The final plate is the darkest one, and the only place the seal red appears in the artwork. Its prompt names the open center explicitly, because the call-to-action text has to sit there.

Original Japanese ink-wash dark final background plate, 16:9, deep brown-black handmade paper night river, warm lantern reflections rippling across dark water, faint koi shapes just under the surface, green firefly glow drifting in foreground and background, subtle wooden bridge edge and temple silhouette, one tiny cinnabar red seal accent near the bottom, center open for off-white serif call-to-action text, textured paper grain, muted palette (#160f0b, #211b15, #251d18, #9fd48a fireflies, #9c2d24 seal), no text, no UI, no people, painterly and restrained
The seven generations were run as a batch — a JSON manifest with concurrency 2 — so the whole plate set landed in one pass for $0.0387.
From PNG to WebP: the 16× that fixed a "sluggish" page
The first build shipped the plates as PNGs. It felt sluggish on first scroll, and the reason was obvious once measured: 12.85 MB of PNGs, decoded mid-scroll. Converting to WebP at quality 82 took the payload to 794 KB.

| Plate payload | Before (PNG) | After (WebP q82) |
|---|---|---|
| Total | 12,850 KB | 794 KB (16×) |
| Largest plate | 2.6 MB / ~165 ms decode | 182 KB / ~50 ms decode |
| Full-page transfer + decode | several seconds | ~300 ms |
The PNGs were kept on disk as an archive. The page only ever loads the WebP.
Step 2 — The video
Only two sections get motion, because motion is the expensive part and the page doesn't need it everywhere: the pinned hero and the night-river CTA. The four middle scenes are stills with a slow parallax drift, which is enough.
Both clips are generated by animating their approved plate. The model is minimax/h3-max/image-to-video on fal, at $0.08/second — a 5-second 768p clip is $0.40.
The video prompt is a motion brief, not a description of the scene. The plate already contains the scene; the prompt only describes what should move and how much.
Here is the hero motion prompt:
Japanese sumi-e ink-wash river at night. Faint green fireflies drift slowly
above the pale parchment water, blinking soft at different rates. The river
surface shimmers gently as reflections waver. Reeds and grass sway almost
imperceptibly. Near-still camera with a very slow, tiny push-in. Quiet,
meditative, painterly mood. No text, no UI, no people, no animals, no hard cuts.
The result — the hero clip, playing on a loop (5s, 768p, 876 KB):
And the night-river CTA prompt:
Japanese ink-wash night river, deep brown-black handmade paper. Warm
lantern reflections ripple softly across the dark water. Pale green fireflies
drift slowly in foreground and background, blinking at different rates. Faint
koi shapes move just under the surface. Very subtle sway of willow branches,
near-still camera, no cuts. Quiet final mood. No text, no UI, no people.
The night-river clip (5s, 768p, 595 KB):
Note:
Be honest with yourself about that prompt. "Near-still camera" and "imperceptibly" produced a clip with almost no motion — measured at 1.1/255 mean per-frame luminance change for the hero and 0.6/255 for the CTA, against 5–20 for a clip with real movement. The model did exactly what it was told. We come back to this in the lessons.
The API, briefly
The call is a synchronous POST to https://fal.run/minimax/h3-max/image-to-video, with the plate passed as a data URL:
minimax/h3-max/image-to-video — input parameters
Values: string
Values: string
Values: 5 (default), cap 15
Values: 768P (default) | 480P
Values: balanced (default)
Values: integer
The response returns an MP4 at body.video.url, priced per second of output.
Wiring it into mediastudio
Rather than glue together a one-off script, video support was added to the existing image-generation CLI (mediastudio), so future clips are one command. Five small changes:
types.ts
Add a per_second price kind, a video?: boolean flag on model info, and duration? / resolution? on the request type.
registry.ts
Register minimax/h3-max/image-to-video, provider falai, price $0.08/s, video: true.
router.ts
Estimate per-second cost as usd × duration, and exclude video models from automatic "cheapest" routing so an image request never gets routed to a video endpoint.
providers/fal.ts
Add a video branch: send prompt, image_url, duration, resolution; parse body.video.url.
cli.ts
Add --duration and --resolution flags, and make the file saver honour an .mp4 extension.
That turns a run into:
bun run medias edit plate.png "<motion prompt>" \
--model minimax/h3-max/image-to-video \
--duration 5 --resolution 768P \
--out out/hotaru-video
Note:
Re-encode any clip with -movflags +faststart. Without it the moov atom sits at the end of the file and the browser cannot seek — which silently breaks any currentTime manipulation. Both HOTARU clips ship faststarted.
The two finished clips are small — 876 KB and 595 KB — because a near-static 5-second 768p clip compresses well.
The design spec
This is the locked spec the build was held to. If you're adapting the recipe, write your own version of this table before you generate anything.
| Section | Plate | Text position | Notes |
|---|---|---|---|
| 1. Hero — Firefly River | 01-hero-firefly | left, vertically centered | pinned, 210vh tall, wordmark open right |
| 2. Poem Field | none (parchment) | centered | real text, no image |
| 3. Reed Bank | 02-reed-bank | left | denser fireflies low over water |
| 4. Lantern Bridge | 03-lantern-bridge | upper right | single red accent allowed |
| 5. Bamboo Grove | 04-bamboo-grove | lower right | path leads the eye |
| 6. Mountain Water | 05-mountain-water | upper left | widest, quietest frame |
| 7. Night River CTA | 06-night-cta | centered | fixed background, open center |
Motion rules: hand-rolled requestAnimationFrame + scrollY interpolation, no animation library. Text reveals are 0.9s, 16px rise, opacity and blur. Scene plates parallax at ~4vh. prefers-reduced-motion disables parallax, reveals, and the particle canvas.
One accent per page. The seal red appears exactly twice in the whole piece — the SVG seal stamps in the poem field and the CTA. That restraint is the entire trick.
Step 3 — The build
The whole site is one index.html. No bundler, no dependencies, no CDN. That's a deliberate choice for a piece like this: it loads instantly, it's trivially portable, and it can be dropped onto any static host.
The firefly swarm
The living layer is a canvas-2D particle system. Each firefly has a depth value, a blink phase, and a drift phase. Depth drives size, opacity, and how much scroll velocity it inherits — so scrolling the page drags the swarm downward like it's moving through air.
function frame() {
if (!running) { requestAnimationFrame(frame); return; }
velocity = velocity * 0.9 + (scrollY - lastY) * 0.1;
lastY = scrollY;
velocity *= 0.94;
ctx.clearRect(0, 0, W, H);
for (const f of flies) {
f.phase += f.blinkSpeed * 0.016;
f.driftPhase += 0.012;
f.x += Math.sin(f.driftPhase + f.depth * 2) * (0.15 + f.depth * 0.3);
f.y += f.vy - velocity * 0.03 * f.depth;
if (f.y < -20) { f.y = H + 20; f.x = Math.random() * W; }
const blink = 0.25 + 0.75 * (0.5 + 0.5 * Math.sin(f.phase));
const alpha = (0.22 + f.depth * 0.68) * blink;
// draw a radial-gradient glow, larger with depth
}
requestAnimationFrame(frame);
}
Counts scale with viewport: 100 fireflies above 900px, 60 above 640px, 38 on phones. Device pixel ratio is capped at 2, and the loop pauses when the tab is hidden.
The scroll choreography
One loop reads scrollY, maps it into progress values, and writes transforms. The hero plate settles from scale(1.08) to 1 and brightens slightly over the first viewport. The scene plates drift in alternating directions. Text elements get an .on class when they cross 86% of the viewport.
scenes.forEach((sec, i) => {
const plate = sec.querySelector(".plate");
const rect = sec.getBoundingClientRect();
const prog = (vh - rect.top) / (vh + rect.height);
const dir = i % 2 === 0 ? -1 : 1;
plate.style.transform =
"scale(1.12) translateY(" + (dir * (1 - prog) * 4 + (prog - 0.5) * 4) + "vh)";
});
A header you can actually read
The first version of the fixed nav used mix-blend-mode: difference with white text. It worked beautifully over the dark night sections and disappeared completely over the pale parchment ones — thin 11px glyphs inverted to a faint gray.
The fix was to stop relying on the backdrop entirely. The header now carries its own legibility: a 150px gradient scrim behind it (dense at the top, fading to nothing), bone-white text with a slight shadow, and a firefly-green underline on hover and on the active section. It reads on every section, light or dark.
Key Takeaway
mix-blend-mode: difference is a party trick that fails the moment your backdrop contains both very light and very dark regions. If text must survive an unpredictable background, give it its own scrim instead of borrowing contrast from the page.
What went wrong
The post-mortem is the useful part. Four of these were silent failures — the page looked fine and simply didn't do what it was built to do.
1. The still plate was painted on top of the video. The <video> and its fallback .plate div both sat at z-index: 0, and the plate came later in the DOM. In CSS, equal z-index means document order decides — so the opaque WebP covered the playing clip completely. The video was running the whole time, behind a picture of itself. The fix is one line of ordering: the plate must come first.
2. Scrolling a video's currentTime looks frozen. The original hero mapped scroll position to video.currentTime for a scrubbed camera push. With a near-still clip, scrubbing between two nearly identical frames changes nothing on screen, and any writer that re-asserts currentTime every frame fights the decoder. A plain autoplay loop was both simpler and obviously alive.
3. prefers-reduced-motion hid the video entirely. A well-meaning media query set .scene-video { display: none }, which meant that on any machine with Reduce Motion enabled, the video never appeared at all. Respecting the preference is right; deleting the content isn't. Pausing is the correct response, not removal.
4. The clips barely move. This is the one to internalise. The prompt asked for "near-still camera, imperceptible sway," and MiniMax delivered exactly that — so the finished page reads as a still image with a faint shimmer. The layered bug above was fixed and the video now plays, but the motion amplitude was authored away by the prompt. If you want visible movement, you have to ask for it in degrees: fireflies stream across the frame, ripples visibly, camera pushes forward.
5. A dev server without range support breaks seeking. Python's SimpleHTTP ignores Range requests — it returns 200 with the whole body and no Accept-Ranges. Browsers use ranges to seek video. Any currentTime manipulation is unreliable against it. Use a range-capable static server, or drop seeking entirely.
What it cost
| Item | Model | Quantity | Cost |
|---|---|---|---|
| Plates | gpt-image-2 | 7 images | $0.0387 |
| Hero clip | minimax/h3-max/image-to-video | 5s @ 768p | $0.40 |
| CTA clip | minimax/h3-max/image-to-video | 5s @ 768p | $0.40 |
| Total | $0.84 |
Note:
The fal endpoint carried no free tier for this model — both clips billed at the full $0.08/second. Budget on the assumption that "free tier first" won't apply.
Build your own
The whole point of this format is that it's copyable. The prompts above are complete and unedited — swap the subject, keep the structure, and you have a different site with the same bones.
Three things to carry forward:
- Lock the art direction before you generate. Palette, type, section list, and the negative constraints. Write it down.
- Let the plate carry the scene and the prompt carry the motion. They are separate jobs and mixing them produces muddy results.
- Ask for motion in degrees. "Subtle" gets you a still image. Name what moves and how far.
If you want the discipline behind decisions like the header scrim and the seven-color palette as a reusable checklist, start with Taste-Skill. For the image side, the plate prompts here are the same architecture used across our GPT image generation prompt library and the Nano Banana prompt guides.
Key Takeaway
The cost of a site like this is under a dollar. The expensive part was never the model — it's knowing what to ask for before you spend anything.
Related Articles & Guides
Building a Token Usage Tracker Plugin for OpenCode
Step-by-step guide to building an OpenCode plugin that tracks token usage per provider, reports session stats, and learns quota limits from rate-limit errors.
OpenCode Session History & Recovery
Where OpenCode stores session history, and how to query the SQLite database to recover past work, trace tool calls, find regressions, and audit model costs.
DeepSeek Harness Plugins & Extending
Understand the DeepSeek Harness plugin system: Cordis, profiles, bundles, configuration, and how to build a custom tool. Everything is a plugin.