Alias ArchiveArchive in progress
Archive in progress

Let the storyboard follow the image: keyframes and animation for a painterly pilot

Building a pipeline for a 90-second dialogue-free animated pilot: why locking composition into the storyboard failed, and where painterly texture was lost during video generation.

Filed
Field
Other
Edition
EN / Reading copy

The target was a 90-second, dialogue-free pilot with one visual direction: the brushwork and canvas texture of an old oil painting, a cool grey monochrome palette, and two characters whose faces never appear. Human work was reserved for keyframes—generated in an image model and approved by eye—while the remaining labour was assigned to generation tools.

Halfway through production, what looked like an execution detail became a directional problem: the image model offered little control, while the script was completely controllable. The statement is obvious; its consequence was not. This article records the specific failure that made the consequence unavoidable, the storyboard rewrite that followed, and a set of measurements from animation: how much painterly texture was lost, and where it disappeared.

Seven failed rounds and the conclusion they forced

One shot required a faded woman seen from behind, her arms still held in an embracing pose although they enclosed nothing, standing close to a stone wall, with no warm colour anywhere in the frame and an old-oil-painting texture. Once written into the storyboard, these became a composition specification.

Seven rounds failed review. Each round repaired the previous round’s reported defect, but every repair pushed the image further away from the actual content of the shot. To stop the wall from becoming a distant castle, its description grew increasingly forceful. The wall eventually became the grammatical subject at the beginning of the prompt; the resulting image became architectural, with the woman reduced to a dot.

The turn came from rereading the script. It required only four things: she repeats a motion she cannot stop; her arms are empty although the pose remains; she faces away and does not turn; and she can be attacked. The wall was never part of the requirement. It entered during execution, then became a reason to reject otherwise valid images.

Two production principles followed.

First, a storyboard specifies what must happen, not what it must look like. Every section now lists red lines and negotiable elements explicitly, followed by a mechanical rule: no one may reject a valid image because of an element listed as negotiable. This is not a rule of leniency. It prevents an executor’s invented constraint from overriding the script.

Second, a continuous section receives only one keyframe. Every image-model generation is an independent sample; four images from the same prompt may contain four different walls. Two independently generated keyframes within one section virtually guarantee a continuity failure. All remaining shots in that section must be derived: use the previous video’s final frame for continuous motion, and use controlled local editing for a new angle or removed element.

The storyboard was rebuilt accordingly: five large sections, one keyframe per section, beats and red lines for each, and no composition descriptions. All five passed within one to three rounds.

Approved keyframe showing the faded woman from behind and to the side as she hurries through a dead forest beside a dry-stone wall
The approved keyframe. It had already appeared in a rejected round, where it was excluded because “the wall is too low to collide with.” No collision with a wall existed in the script.

Three repeatable prompt failures

Each of the following occurred at least twice in this pipeline, through a different failure mechanism.

Unspecified body parts are filled in. An undescribed lower body becomes rags; an undescribed left hand acquires a shield; a head that was outside the crop grows a face when the frame expands. The model does not preserve emptiness. It supplies a plausible value.

A mentioned object is drawn, even when it appears in a negation. “As if carrying a child that is not there” reliably produces a child. “Cinematic” reliably produces letterbox bars. Negative parameters can act as a fallback, but they cannot carry the main constraint; in this pipeline they did nothing to remove letterboxing. The only reliable way to communicate emptiness or absence was to describe only what was present and let the absence emerge from the composition.

Some sets of constraints cannot coexist from any camera angle. “Back to camera” conflicts with “arms embracing in front of the body”: the arms are on the front and therefore hidden from a direct rear view. The model compromises by drawing a contradictory body—back-facing head, front-facing arms. This is not a capability problem; a human illustrator cannot satisfy it either. The solution is a three-quarter rear angle, which keeps the face turned away while revealing the outline of the arms.

The same mistake appeared three times because the sentence was rewritten on every attempt. The production document now stores one fixed formulation that must be copied rather than paraphrased.

One additional result followed: changing the shot scale or angle was much more effective than changing the wording. All three geometric changes succeeded—a three-quarter rear angle exposed the arms, a tighter frame stopped the environment from swallowing the subject, and a camera pullback repaired the spatial relation between two people. All four wording-only revisions failed. Framing is geometry; language cannot solve impossible geometry.

Image-to-video: inherit the style; do not rewrite it

The first animated result barely resembled the keyframe. The oil brushwork disappeared and the image became a different kind of misty forest.

The prompt contained a separate style paragraph that described the intended look again. In an image-to-video task, this was a mistake. The input image already defines composition, subject, light, and style; the prompt should describe movement only. Restating style in words introduces a new synthesis target and displaces the source image’s visual identity.

The corrected prompt declared the input image to be a strict first frame and required its style to be preserved, without restating the style. The generated result returned to the keyframe’s visual character.

The same revision fixed another problem. The first video became almost static for roughly five seconds in the middle: mean inter-frame difference fell from 8.2 in the opening section to 3.7–4.7, while the subject moved horizontally by only 9% of frame width across the entire clip. Five negative constraints at the end of the prompt—never turn, never reveal the face, never move the arms, and so on—were suppressing the performance. That language is appropriate for a seamless micro-motion loop precisely because it suppresses movement; it should not be reused for a shot that needs substantial action. Adding two positive requirements—movement begins immediately and continues, and no static pause may occur—raised mean inter-frame difference to 11.90 with a minimum of 2.50, eliminating the dead zone.

Animation test for the mother section. Playback starts only when selected; it never autoplays.

Measuring texture loss

The subjective report was “scratchy, too sharp, plastic.” That sounds contradictory: plastic imagery is usually described as overly smooth, while the observer reported excessive sharpness. The total amount of high-frequency energy therefore had to be separated from its location.

The method takes a Laplacian response from a greyscale image and partitions it by local gradient. Low-gradient areas are treated as flat surfaces, where brushwork and canvas texture live; high-gradient areas are treated as contour edges. Response strength is measured in each region, followed by their ratio.

The same keyframe, reduced to the output resolution, serves as the baseline. The keyframe was generated in Midjourney and approved manually. The two animation paths were Midjourney Animate and Seedance 2.0 through LibTV at 720p. Both received the same keyframe.

SourceFlat-surface textureEdge sharpnessEdge/surface ratioSurface texture retained
Keyframe (baseline)13.3259.864.49100%
Midjourney Animate10.1349.534.8976%
Seedance 2.05.1435.416.8939%

Seedance lost 61% of flat-surface texture but only 41% of edge response, increasing the edge-to-surface ratio by 53%. High-frequency energy concentrated around contours while surfaces were smoothed. That is the structure of the plastic appearance; the perceived “scratchiness” came from contours isolated by the smoothing. The subjective report and measurement appeared to conflict, but described the same change.

Midjourney Animate produced an edge-to-surface ratio of 4.89, close to the baseline’s 4.49. It retained not only more detail, but more of the image’s overall character.

There is an important confound: the source outputs differed in bitrate by roughly threefold (31.8 versus 10.6 Mbps), so compression loss was not controlled. The table must therefore be read as an end-to-end comparison of two complete paths, not a model benchmark under equal encoding conditions. Their operating ranges also differ. In this test, Midjourney Animate chose its own duration at roughly 5.2 seconds and offered weak directional control, while Seedance provided exact durations from 4 to 15 seconds and strict first-and-last-frame control. Texture quality and controllability belonged to opposite ends of the comparison.

The loss has a structural cause. Video generation must preserve temporal consistency, but random high-frequency textures—brushstrokes, canvas weave, film grain, stippling, hatching—cannot be matched frame by frame without flicker. Denoising therefore suppresses them. Choosing a hand-painted visual direction means choosing one of the least video-generation-friendly classes of texture.

A correction that passed the metric and failed the image

Once the missing amount was measured, the direct response was to restore it in post-production. A high-frequency residual was extracted from the keyframe by subtracting a Gaussian-blurred version, then composited over the video frames at several strengths.

At 0.6 strength, flat-surface texture rose from 5.14 to 13.29, against a baseline of 13.31—an almost exact metric match.

Four one-to-one crops of the same stone-wall region: keyframe, animation path A, animation path B, and path B with restored texture
Four-way comparison. The lower-right image matches the metric almost exactly, but dark branches and stripes appear as ghosts because the overlay no longer aligns with the moving content.

The visual result failed. The high-frequency residual was tied to the source image; once the camera advanced and the character moved, the overlay no longer aligned with the content and produced ghosting.

The experiment exposed a blind spot in the metric: flat-surface texture measures whether high frequencies exist, not whether they are correctly registered. A number that exactly matches the baseline can be a false positive. The more plausible direction is a content-independent, tileable texture layer rather than a residual taken from the source image. That approach has not yet been tested and is not presented here as a conclusion.

What currently holds

  • A storyboard should specify what must happen, not what it must look like. Negotiable elements must be explicit and cannot be used to reject an otherwise valid image.
  • A continuous section should use one keyframe; later shots derive from final frames or controlled local edits.
  • Unspecified body parts are filled in, while mentioned objects appear even when mentioned negatively.
  • Some constraint combinations are geometrically impossible. They require a new camera angle, not rewritten wording.
  • In image-to-video work, style must be inherited rather than described again.
  • Movement-suppressing prompt language cannot be reused for shots that require performance.
  • Painterly texture loss is structural: flat-surface texture is removed while contours remain relatively exposed. In this test, Midjourney Animate and Seedance 2.0 retained 76% and 39% of surface texture respectively, while offering opposing strengths in duration and movement control.

The pipeline is not complete. Tool roles for the animation stage, the trade-off between duration and texture, and a viable post-production texture-recovery method remain unresolved.

Related reading

02 / LINKS
C01

Claude Fable 5 Route: Rebuilding The Vigil with Human Keyframes

A Claude Fable 5-driven Flova restart: the creator owned visual judgment and Midjourney keyframes, while the Agent drove the CLI, per-shot review, and orchestration to produce a 14.2-second atmospheric film.

Read post ↗
C02

AI Action Direction Failure Log: Shot-by-Shot Generation and the Shift to Flova

A stone-giant combat sequence becomes a test of Niji, Image2, Seedance, Dreamina, BytePlus, and Kling—and of why action causality, spatial continuity, and three-dimensional distance ultimately broke the shot-by-shot workflow.

Read post ↗