Alias ArchiveArchive in progress
Archive in progress

Codex 5.6 Sol Route: Why the Five-Shot Sylas Greenfield Film Failed

A Codex 5.6 Sol-driven Flova experiment completed its storyboard, per-shot generation, sound, and export, yet the finished film still scored only 3–4/10.

Filed
Field
Other
Edition
EN / Reading copy

This is a 15-second film that completed its process and failed as a work. The route was driven by Codex 5.6 Sol, with Flova serving as the orchestration and generation platform and the creator retaining final review.

The project began from scratch with new concept art for the knight and Sylas. It inherited none of the storyboard, composite frames, or video from the earlier action-film attempt. Flova completed the directing plan, five-shot storyboard, Key Elements, two rounds of video generation, sound design, timeline, and final export. Codex 5.6 Sol supplied production context, enforced stage gates, extracted frames for review, and wrote revision instructions. The experiment met the formal conditions of a greenfield restart, but it did not receive a corresponding reset in quality.

The creator’s final assessment was 3–4/10: a failure. That score is closer to the state of the project than “export successful.”

Sylas roadblock · 15 seconds · production workflow complete, final film rejected

Objective: let the directing platform plan a roadblock encounter

The scene belongs near the end of the game’s first chapter. All three braziers have been lit. The knight returns from the bell tower and finds Sylas blocking the retreat through a cold, fog-covered pass. The 15 seconds do not need to show a full fight. They only need to establish a sequence of solitary travel, detection, emergence, pressure, and confrontation.

The project deliberately imposed fewer constraints on the directing layer. Humans supplied character identity, story truth, and visual boundaries; Flova decided the number and duration of shots, camera positions, rhythm, and sound. To avoid repeating the failure of the earlier continuous-action film, the brief explicitly excluded weapon collisions, blade contact, and long multi-character action chains.

Flova’s native storyboard contained five shots:

ShotPlanned durationSingle responsibility
Knight walking aloneAbout 3 secondsEstablish the pass, cold, and isolation
Sensing dangerAbout 2 secondsThe knight stops and detects a threat
Blocking the roadAbout 4 secondsSylas emerges from the fog and bars the route with his sword
Closing pressureAbout 4 secondsA low angle and slow push strengthen Sylas’s presence
StandoffAbout 2 secondsEstablish the two-character axis before combat

The written plan was sound. It did not waste 15 seconds on exposition or assign combat beyond the models’ practical capabilities. What followed proved, however, that a storyboard can have the right causal order without the generated images preserving the right spatial relationships.

Workflow: the stage gates were complete, but visual authority remained loose

The production had six stages:

  1. Upload two Niji concepts for the knight, a frontal construction sheet, a separate longsword image, and the Sylas concept;
  2. Ask Flova to establish the Final Video Spec and five-shot storyboard;
  3. Generate Key Elements for Sylas, the knight, and the pass as a paid capability probe;
  4. Generate all five video shots in one pass, then have Codex 5.6 Sol extract frames and the creator perform final review;
  5. Keep Shots 2 and 4 while regenerating Shots 1, 3, and 5;
  6. Rescue shots that would not converge in the edit, generate a unified sound layer, and export a 15-second film.

The highest-risk probe asked whether Sylas could retain his identity, longsword, and electrical effects after being placed in the environment. The static Sylas Key Element initially passed. The first environmentalized knight, however, inherited the synthetic look of an older structural reference. A second attempt used only the two most recent Niji concepts, and the visual language improved substantially.

The first knight Key Element has hard, glossy armor surfaces and a style contaminated by the structural reference
Knight Key Element V1. Including the structural supplement pushed the image toward generic game key art and weakened the thick-painted Niji character.
The second knight Key Element, regenerated with only two Niji concept images
V2 kept only the latest Niji concepts. Identity and style recovered, but the character remained a platform reinterpretation rather than a manually approved shot-specific keyframe.
A contextual Sylas Key Element in cool blue Gothic ruins, retaining red eyes, a crown-like head, and an electrified longsword
The static Sylas capability probe passed. It proved that the platform could generate one usable concept image, not that space, motion, and identity would remain continuous across several shots.

This stage already exposed the decisive difference from the parallel Claude Fable 5 route, The Vigil. In The Vigil, the creator produced every important keyframe in Midjourney and used the platform only for orchestration and animation. This project also had gates, but Flova’s auxiliary image model still handled environmentalization and shot-level visual interpretation. The former route locked visual authority at the pixel level; the latter locked only relationships among references.

Round one: only two of five shots passed

The first round generated all five shots at once. External frame review produced the following verdict:

  • Shot 1: the knight moved too quickly and showed no weight, caution, or resistance from the environment; reject.
  • Shot 2: identity stayed stable and the restrained reaction worked over two seconds; pass.
  • Shot 3: Sylas lost his tall, slender proportions and resembled a generic soldier; reject.
  • Shot 4: the slow low-angle push preserved Sylas’s silhouette and electrical effects; pass.
  • Shot 5: the knight planted the sword vertically in the ground while Sylas resembled a distant statue; the composition had little tactical or narrative meaning; reject.
Sequential frames from Shot 1 show the knight moving rapidly away from the camera through cold ruins
Shot 1, first version. The setting and travel direction were legible, but the character moved too quickly to convey the intended weight.
Sequential frames from Shot 2 show the knight making a restrained detection movement before the ruins
Shot 2 was one of the most stable first-round results: one character, one action, and a clear composition, with no demand for complex spatial causality.
Sequential frames from Shot 4 preserve Sylas's oppressive silhouette and electrified sword in a low-angle composition
Shot 4 also worked. The character performs almost no complex action; scale, low camera position, slow movement, and electrical arcs carry the pressure.

The two successful shots shared the same structure: one subject, little movement, an explicit composition, and narrative carried by camera and sound. That pattern marks the most reliable capability boundary in this experiment.

Round two: regeneration did not solve the underlying causes

The second round regenerated only Shots 1, 3, and 5. None converged:

  • Shot 1 slowed the cadence, but the knight now walked backward in the wrong direction;
  • Shot 3 improved the proportions slightly, yet remained close to the first version and kept its stiff movement and reveal;
  • Shot 5 established a clearer confrontation axis, but the knight acquired an orange cross on his back even though that identity marker belongs only on his chest, and both figures moved stiffly.
Sequential frames from the second version of Shot 3 show Sylas appearing frontally with a sword in an archway similar to Shot 1
Shot 3, second version. A single frame contains a character and setting, but its spatial relationship to Shot 1 is never established. Sylas simply appears in a similar location without entry, occlusion, eyeline, or an editing cue to supply causality.

The reruns showed that instructions such as “slower,” “taller,” or “use a confrontation axis” constrain surface traits but do not restore information absent from the starting frame and shot design. Direction of travel, the surface carrying an identity mark, the distance between characters, and the reason they can see one another remained open to model interpretation.

Editing rescue: the timeline was completed, but the film was not rescued

The creator chose to stop spending generation credits. The timeline adopted the following compromises:

  • Shot 1 returned to the first version, slowed down and trimmed to roughly three seconds;
  • Shot 2 used the first-round pass;
  • Shot 3 used the second version despite its weaknesses;
  • Shot 4 used the first-round pass;
  • Shot 5 discarded the moving clip with the misplaced back cross and used a comparatively clean still with a subtle two-second push.
A static confrontation image with the knight holding a sword in the foreground and Sylas carrying an electrified longsword in the distance
The still used for the final Shot 5. It works as a single composition, but filling a climactic video shot with a static image does not make the shot successful.

All native audio from the generated clips was muted. Flova generated a new 15-second layer of ambience, electrical effects, and metal sounds, plus restrained dark strings and a low-frequency drone. Hard cuts and sound bridges carried most transitions. The final file ran for 15 seconds and completed the technical loop.

The problem is that editing can hide local defects, but it cannot create missing performance, spatial relationships, or narrative action. When Shot 5 regressed from a failed video to a still, the timeline became full; the film did not gain an ending.

The creator’s final scores

Final review used the creator’s viewing experience rather than platform task status:

ShotScoreFinal assessment
Shot 1Barely passingConventional and usable after slowing down, but with no distinct strength
Shot 27–8/10Stable sound, character, and style, with no obvious defect
Shot 31/10The scene resembles Shot 1, yet neither character’s appearance has spatial logic; the figure is unattractive and the weapon looks plastic. The single point is for effort
Shot 48/10Largely preserves Sylas’s identity, visual language, and pressure, although the weapon effect remains weak
Shot 50/10Only a still image, with no motion to assess

Overall: 3–4/10, a failure.

The result also shows why a simple average is misleading. Shots 2 and 4 can score seven or eight without compensating for Shot 3’s narrative break and Shot 5’s missing function. A short film is constrained by its critical structural nodes, not by the arithmetic mean of five independent shots.

Why the greenfield restart still failed

1. Clearing history is not the same as establishing visual authority

The project inherited no old state. Once multiple references with different responsibilities were uploaded, however, the platform still reinterpreted the characters. Greenfield isolation removes old contamination; it does not decide who has final pixel-level authority. The Vigil was stronger not because it opened a cleaner project, but because the creator generated and approved every important first frame in Midjourney.

2. The storyboard described causality; the shot assets did not prove it

“Walking alone—detection—emergence” reads naturally in prose. Yet when Shots 1 and 3 occupied similar spaces, the production never established where Sylas entered, why the knight saw him, or how far apart they were. The model generated two individually readable images, not one continuous event.

3. The static capability probe tested too narrow a question

The Sylas Key Element proved static identity, weapon, and style. The actual scheme-breaking risk was whether the character could appear in a specified space, move in the correct direction, and connect to adjacent shots. Approving the static probe and then generating all five shots still skipped the most dangerous question.

4. Regeneration treated symptoms without rebuilding the shots

The second round retained the original five-shot structure and asked for slower movement, taller proportions, and a better confrontation axis. When current assets cannot prove a shot’s responsibility, deletion, combination, or reconstruction around a human-authored keyframe is more effective than another iteration in the same slot.

5. Timeline completion was mistaken for project completion

The Shot 5 still allowed the file to reach 15 seconds, but it changed the final pre-conflict moment from a moving climax into a placeholder. A successful export proves only that media was encoded. It does not prove that the shot fulfilled its narrative responsibility.

Comparison with The Vigil

The parallel Vigil experiment was driven by Claude Fable 5 and also used Flova, but with a different structure. It reduced the brief to three atmospheric shots with limited motion. The creator produced every important keyframe in Midjourney. Failed shots could be deleted. Thick-painted atmospheric motion went to MJ Animate, while Seedance handled the kinds of directed motion it performed better.

Together, the experiments show that Flova is valuable for project memory, storyboarding, asset orchestration, sound, and timeline work. Their difference demonstrates that an orchestration platform cannot replace a visual anchor or the creative decision to remove a shot. When humans retain keyframes and final review, agents can safely carry much of the process. When the platform interprets the image, animates it, and evaluates itself, a greenfield project can still reliably produce a complete process and an unsuccessful film.

Rules that must change next time

  1. A greenfield definition must include asset governance: every important shot needs one uniquely approved human-authored keyframe, not merely an absence of old project state.
  2. A capability probe must test the relationship most likely to invalidate the plan—travel direction, spatial continuity, or the surface carrying an identity marker—not merely an attractive still.
  3. Before video generation, every shot must pass three gates: single-frame visual quality, spatial logic, and executable action.
  4. A rerun may change only one explicit variable. If a second attempt still fails, delete or rebuild the shot instead of filling the original storyboard slot.
  5. A still-image placeholder must remain labeled as a missing shot and cannot count as a successful video shot in final review.
  6. Flova may provide directing suggestions and orchestration; humans own keyframes, selection, and factual review. Success is always defined by watching the finished film.

The most valuable output of the experiment is not the 15-second video but a stricter conclusion: a greenfield restart clears state; a quality reset also requires resetting asset authority, capability probes, and deletion discipline.

Related reading

02 / LINKS
C01

Claude Fable 5 Route: Rebuilding The Vigil with Human Keyframes

A Claude Fable 5-driven Flova restart: the creator owned visual judgment and Midjourney keyframes, while the Agent drove the CLI, per-shot review, and orchestration to produce a 14.2-second atmospheric film.

Read post ↗
C02

Flova End-to-End Test: The Project Ran, the Film Still Failed

After shot-by-shot generation failed, the same action sequence was handed to Flova for director audit, storyboarding, keyframes, four-shot generation, sound, and export. The system delivered a 38.24-second film, but character contamination, style drift, weak contact logic, and false-positive self-review meant the workflow still failed.

Read post ↗