Alias ArchiveArchive in progress
Archive in progress

Claude Fable 5 Route: Rebuilding The Vigil with Human Keyframes

A Claude Fable 5-driven Flova restart: the creator owned visual judgment and Midjourney keyframes, while the Agent drove the CLI, per-shot review, and orchestration to produce a 14.2-second atmospheric film.

Filed
Field
Other
Edition
EN / Reading copy

The previous AI action film was archived as “technically complete, unsuccessful against the objective.” Its postmortem left one hypothesis untested: perhaps the problem was not the platform itself, but the shape of the task and the discipline of the assets. This project tests that hypothesis on the same platform and with the same character world, while rebuilding everything from zero. The route was driven by Claude Fable 5, with Flova handling project orchestration and model routing while the creator retained keyframes and final visual judgment.

The result comes first: The Vigil, a 14.2-second dark-fantasy atmosphere film. It has three shots, no contact, no combat causality, and a total cost of 838 credits. Its action is simple, but every frame passed review for character identity, visual style, and atmosphere.

The Vigil · 14.2 seconds · final three-shot atmosphere film

Objective: move the task inside the capability boundary

The central lesson from the previous failure was that asking an image-to-video model to prove “the geometric overlap of a blade and a stone body” amounted to asking a visual sampler to perform collision detection. This time, the objective was deliberately reduced:

  • Three shots, roughly 15 seconds in total;
  • Only micro-motion: fog, cloak, ash, firelight, and lightning;
  • No character contact and no requirement to prove causal action;
  • Four success criteria: stable identity, style consistent with the master, intact micro-motion, and continuity between shots.

The narrative was equally simple. A knight walks alone along an ashen night road. The knight approaches a vast ruin. In a flash of lightning, a dark swordsman carrying an impossibly long blade appears among the towers. Solitary travel, awareness of the unknown, and the arrival of that unknown.

Division of labor: humans made the keyframes, Claude Fable 5 ran the process

The most important structural decision was to remove “keyframe creation” from the platform’s responsibilities and return it to the creator:

  • Human creator: confirm story and storyboard, personally create every keyframe in Midjourney, make the visual judgments, and approve or reject each shot;
  • Claude Fable 5: drive the full process through the official Flova CLI—create the project, upload assets, send directing instructions, download and sample each shot, perform objective checks for identity, objects, and continuity, and provide diagnosis and revisions after every failure;
  • Flova: organize the project through storyboard structure, project memory, inheritance between shots, asset replacement, timeline, and export;
  • Seedance 2.0 / Nano Banana: provide image-to-video generation and platform-side image generation.

Giving keyframes to the human was a quality decision rather than an efficiency decision. Midjourney keyframes share their visual source with the existing character sheets, so identity and visual language align naturally. The platform-side image model drifted under the same written constraints. The evolution of Shot 2 is direct evidence of the difference.

A parallel Codex 5.6 Sol-driven greenfield experiment, Sylas Roadblock, supplied a same-day control. That route also cleared old state and established stage gates, but it still allowed the platform to create environmentalized characters and shot-level key assets. Only two restrained single-subject shots survived, and the finished film scored 3–4/10. The contrast shows that greenfield isolation can remove historical state; it cannot replace the visual and identity authority supplied by human-authored keyframes.

Seven production rules were fixed before work began. The four most important were:

  1. Greenfield isolation: bring no assets, intermediate images, or scripts from the old project;
  2. One visual master: one image is the sole authority for visual style and character identity;
  3. First-frame gate: every new camera angle receives a still frame and passes identity review before video generation;
  4. External review: platform self-review is only a clue; the agent extracts frames and the creator makes the final decision.

Shot 1: the master image passed in one attempt

The first frame of Shot 1 was the visual master itself: a black-armored knight walking toward the camera along a foggy, ash-covered road at night.

A black-armored knight walks toward the camera on a foggy ashen road, carrying a small orange cross on the chest with faint fires among the trees behind
The film's only visual master. Every later judgment of style and identity was made against this image.

Using the master directly as the starting frame passed on the first attempt. Identity remained stable, while cloak, fog, ash, and distant firelight all moved at separate scales. The shot also revealed a useful platform behavior. When told to use an image as the first frame, the platform treated it as a compositional anchor rather than a pixel-exact starting point: the knight began farther away than in the master. That difference happened to create a complete walking beat here, but productions requiring exact first-frame continuation need to account for it.

Shot 2: three versions of a helmet, followed by deletion

Shot 2 required the knight to stop and look up from a new angle, which meant redrawing the first frame. It was the only shot not taken directly from the master and became the center of the film’s problems.

Three images compare the master, the platform's first frame with an oversized enclosed helmet, and a second frame whose pointed crown has returned
Left: master. Center: platform V1, where the helmet became a large enclosed helm. Right: the second version after written correction. The deeper issue was that the hood concealed the helmet in the master, leaving the model no structural information.

Both platform-generated first frames were stylistically close, but the helmet’s construction drifted. The problem became worse in motion. As the knight looked up and revealed the crown, the model completed it as a smooth reflective dome. Two rounds of written constraint could not lock the design because the problem was not wording: the master image did not contain the hidden structure, so the animation had no choice but to invent it.

The solution returned to the division of labor. The creator generated the new first frame in Midjourney, using the master as both image and style reference and specifying the approved helmet construction. With enough information in the first frame, the helmet remained stable in motion.

The shot was deleted anyway. Final review found perceptible character drift between Shot 1, taken directly from the master, and Shot 2, redrawn from a new angle. They no longer felt fully like the same person. Removing the shot and cutting directly from the frontal walk to a distant rear view produced a cleaner sequence. The edit added an unexpected lesson: identity remains more stable when fewer camera changes require redrawing the character, and a frontal “walking toward you” cut directly to a rear “walking away” can be narratively cleaner.

Shot 3: the failed text-only silhouette and the rebuilt impact shot

The original third shot called for lightning to reveal a black silhouette with an extremely long blade above distant ruins. To avoid contaminating the style, the first attempt gave the silhouette no visual reference at all and relied entirely on text.

The result failed cleanly. The silhouette became a small figure floating between towers, with no sense of threat. In the creator’s words, “it looks like a clown flashing into view to scare you, only to become the joke.”

The failed text-only silhouette floats as a small figure between towers on the left, while a low-angle entrance frame anchored to the character design appears beneath lightning on the right
One entrance, two methods. The left side used only a written description. The right side rebuilt the shot with the character sheet as an image reference.

The correction reused the film’s central rule—anchor every important visual with an image; use words only to control behavior—and borrowed the language of an action-film impact shot. The established dark-swordsman design became an image reference for a low-angle Midjourney entrance frame: a crown of flame, two red eyes, an electrified impossibly long blade, and a giant bolt falling behind the character. The wide travelling shot cut hard to this close entrance, with the edit landing on the opening flash.

Three rounds of convergence in animation

Animating the entrance produced three more reusable lessons:

  1. Synthetic surfaces: image-to-video processing made the thick-painted close-up look physically rendered and reduced its painterly quality. Written instructions could lock matte materials, but the deeper solution came from discovering that Midjourney’s own Animate function preserved the texture of Midjourney artwork with almost no loss. The final distant travelling shot used that route. Atmospheric micro-motion went to MJ Animate; stronger directed motion went to Seedance.
  2. A still entrance needs camera motion: the micro-motion discipline worked for atmosphere, but a completely locked entrance lost tension. An extremely slow push restored pressure immediately.
  3. Restraint creates pressure: the first version made the red eyes flicker like flames. The creator changed them to “two dim, motionless red points, like a gaze that does not move in the abyss,” while making the sword’s electrical arcs the only fast element in the frame. One contrast between stillness and movement was stronger than layered spectacle.
Four close frames show the electrical arcs around the blade changing shape, followed by two full-frame views
The approved sword arcs. Across four samples separated by 0.1 seconds, the electrical shape changes in every frame. It is the only fast element in the composition.

Sound and final assembly

Sound followed the same restrained logic. A low-frequency rumble provided the base and a sense of weight, rain established the environment, and thunder was used sparingly. The low end became much stronger during the entrance and faded cleanly with the picture. The platform’s music engine produced two candidates. The creator selected not only a version but a precise passage from it; when music is longer than the picture, choosing the section remains a creative decision.

The final timeline was: knight walking alone for five seconds, a 5.2-second distant view of the knight approaching the ruins animated with MJ Animate, and the dark swordsman’s four-second entrance.

Six frames sample the finished film: the knight approaches through fog, walks deeper into the frame, becomes a rear-view figure in distant ruins, and gives way to the dark swordsman beneath lightning before the image darkens
The approved film sampled across three sections and two hard cuts. The second cut lands on the giant flash at the beginning of the entrance.

Reusable rules

The experiment produced seven rules, ordered by importance:

  1. Task shape comes first: moving the brief inside the current capability boundary—micro-motion, no contact, and no causal proof—reduces failure discontinuously rather than gradually.
  2. Keyframes are the foundation of identity and texture: make them with the strongest image tool under human judgment, not with the video platform’s auxiliary model.
  3. The information in the first frame sets the ceiling of the animation: structures hidden in the first frame will be invented later.
  4. Anchor every important visual with an image; use text for behavior and rhythm. A major character entrance deserves its own impact shot.
  5. Assign animation engines per shot: MJ Animate preserves thick-painted atmospheric motion, while Seedance handles stronger directed movement at the cost of more synthetic close-ups.
  6. Restraint creates tension: still pressure, one fast element, and controlled light are stronger than stacked effects.
  7. External review is mandatory: platform self-review remains a clue, not a verdict. Agent-led frame extraction and human final review are what make the other rules enforceable.

The parallel failure record, Codex 5.6 Sol Route: Why the Five-Shot Sylas Greenfield Film Failed, concludes that a new project, a complete storyboard, and a successful export do not make a film successful. The Vigil adds a complementary point: creative judgment cannot be delegated to the platform either—but once humans retain keyframes and final review, much of the remaining process can safely move to Claude Fable 5 and orchestration.

Related reading

02 / LINKS
C01

Codex 5.6 Sol Route: Why the Five-Shot Sylas Greenfield Film Failed

A Codex 5.6 Sol-driven Flova experiment completed its storyboard, per-shot generation, sound, and export, yet the finished film still scored only 3–4/10.

Read post ↗
C02

Flova End-to-End Test: The Project Ran, the Film Still Failed

After shot-by-shot generation failed, the same action sequence was handed to Flova for director audit, storyboarding, keyframes, four-shot generation, sound, and export. The system delivered a 38.24-second film, but character contamination, style drift, weak contact logic, and false-positive self-review meant the workflow still failed.

Read post ↗