The previous article ended at a pivot. Generating shots individually with Dreamina, Seedance, and Kling could occasionally produce an attractive five seconds, but it could not reliably form a spatially credible action sequence. The same objective was therefore handed to Flova to test whether a project-level platform could integrate direction, storyboards, assets, models, sound, timeline, and export.
It completed every stage and delivered a 38.24-second film. Judged against the original objective, however, the Flova workflow still failed.
Failure here does not mean a broken file or an absent generation. The film plays, contains sound, has four shots, and sits on a complete timeline. The failure occurred at a more difficult level: the images did not consistently prove the actions claimed by the prompts; character identity did not remain locked; corrections damaged an already approved visual style; and the platform’s own director review did not reliably detect these defects.
The objective was not export, but continuous validity
The story was restricted to four shots:
- The knight passes through the stone giant’s final attack arc and reaches a position near its leg.
- The knight visibly strikes a fissure in the knee, and the Boss drops to one knee as its structure gives way.
- The knight executes the chest core; the Boss collapses and the knight is exhausted.
- Only after the battle is completely over does Vera enter with her lantern.
The total duration could remain between 37 and 40 seconds; it did not need to be compressed merely to save credits. The film had to preserve the same rain-soaked ruins, the same Phase 1 knight, the same Boss made entirely of rock, and Vera’s approved face, clothing, and lantern design.
Any one of the following would disqualify the film:
- a character lands a strike from a distance it has not yet crossed;
- the sword never visibly overlaps the target area;
- the Boss suddenly becomes an armored giant;
- weapons or lanterns multiply;
- Vera’s identity, clothing, or lantern hand changes;
- the Niji-derived visual language becomes plastic CG, flat cartoon art, or high-frequency digital noise;
- the platform claims a hit solely because the prompt says so, while the image provides no evidence.
These criteria are stricter than “generation succeeded.” That is precisely what makes the subsequent failures useful to analyze.
Division of labor: Flova was the project director, not a single model
Flova was not treated as another text-to-video endpoint. Responsibilities were divided as follows:
| Layer | Responsibility |
|---|---|
| Human owner | Confirm story, character assets, acceptable cost, and final aesthetic judgment |
| Codex | Global planning, asset organization, prompt boundaries, frame-by-frame QA, cost records, and delivery |
| Flova | Director audit, Storyboard, shot breakdown, model routing, project resources, timeline, sound, and export |
| Seedance 2.0 | Four image-to-video clips with native action and environmental sound |
| Image models | Gate B key elements, spatial previs, and correction attempts after failures |
This distinction matters. Flova’s value is not that it possesses a stronger proprietary video model. Its value is the ability to organize multiple models and intermediate artifacts around a persistent project. The test asked whether that organizational capability could also carry directing judgment and quality assurance.
Stage one: the zero-media director audit was correct
In the first pass, Flova read only the project assets and earlier failure record. It generated no media and incurred no new media cost. It correctly identified the core problem in the earlier continuation: the knight and Boss were still visibly separated, yet the knight could stand in place and immediately strike. The issue was not insufficient spectacle; the three-dimensional spatial relationship had already broken.
The second pass also remained at the text and Storyboard level. Flova did not mechanically accept the proposed six-shot structure. It held the sequence to four shots and 37–40 seconds:
- Shot 1 resolved the final inward dodge.
- Shot 2 handled only the close-range poise break and the Boss dropping to one knee.
- Shot 3 preserved the continuous psychological time of execution, death, and exhaustion.
- Shot 4 introduced Vera only after the battle.
It also proposed using the final one or two seconds of each shot as the next shot’s video reference, and defined a “hit” as geometric overlap between the blade and rock rather than detached sparks.
This was the most reliable part of the Flova workflow. It converted a vague desire for action into distinct, inspectable shot responsibilities.
Gate B: still-frame self-review failed first
Once media production began, Flova generated three key-element images and six spatial previs images, initially consuming 42 credits. It marked the Shot 2 kneeling frame as a pass, but manual review immediately found two serious errors.
First, the Boss’s head had become an enlarged knight’s helmet, complete with a clear visor and metallic silhouette. Second, the knight’s sword was already touching the glowing point on the chest and producing a cross-shaped flare, effectively completing Shot 3’s execution at the end of Shot 2.
The first targeted correction consumed 28 credits. It removed the helmet and premature contact, but regenerated the entire image. The approved high-detail dark visual language became a flat concept illustration. Another corrected Shot 4 frame simultaneously introduced two sword-shaped weapons.
The second correction generated three more candidates and consumed another 42 credits. Object counts were finally close to correct, but the visual style continued to drift. Gate B therefore recorded 112 credits of image work without producing a formal set of keyframes that simultaneously respected the design and retained the master style.
Image2 was then asked to repair the original image locally. It removed the helmet and increased the distance from the sword tip, but introduced familiar high-frequency grit and digital sharpening. Edge-preserving denoising reduced the grain but could not restore Niji brushwork after the material had already been redrawn.
The final decision was not to continue repairing the intermediate images, but to revoke their visual authority. The workflow returned to the initially approved Niji composite and established it as Shot 1’s sole style and identity master.
Gate C: one visual master improved the first shot
Shot 1 used the approved master to generate a 6.06-second Seedance 2.0 video with native AAC stereo sound. The knight no longer fled outward. As the giant arm became a threat, the knight moved inward; the foreground arm and tracking camera established pressure, and the shot ended near the Boss’s leg.
The shot passed, although it was not literally “without pauses,” as Flova’s self-review claimed. The opening anticipation and final crouch both lasted slightly too long. These holds did not break spatial causality, so they could be handled gently on the timeline instead of cutting away the character’s breathing rhythm.
From Shot 1 onward, every later shot inherited the end of the previous shot where possible. This finally produced a continuous close-range chain.
Four-shot result: two passes and two conditional failures
Shot 2 lasts 8.06 seconds. It begins near the leg, and the Boss does eventually drop to one knee. The chest core, however, glows strongly during the knee break, while geometric evidence of the blade touching the knee fissure remains unclear. The image proves that “the Boss knelt,” but not sufficiently that “the knight struck the knee, therefore the Boss knelt.”
Shot 3 lasts 13.07 seconds and is the best-executed shot in the film. The knight executes the chest core directly from the close-range state. The Boss collapses as a mass of rock losing structural support rather than disappearing, floating, or turning into another armored giant. The knight then kneels against the sword, preserving the continuous duration of exhaustion.
Shot 4 lasts 11.05 seconds. The knight naturally falls sideways from the kneeling position. Vera enters from the background afterward, while the Boss remains lifeless rock debris. The narrative order is correct, and the contrast between cold rain and warm lantern provides closure.
The problem is that Vera’s face, apparent age, and clothing drift from the authoritative character pack. The lantern also appears closer to her own left hand than the designated right. As a new character entering for the first time, she inherited no identity from the preceding shot; a character reference alone did not allow Flova to lock her reliably.
The final assessment was:
| Shot | External review |
|---|---|
| Shot 1: inward dodge | Pass, with pacing reservations |
| Shot 2: knee poise break | Conditional failure: weak contact evidence and premature chest glow |
| Shot 3: execution and exhaustion | Pass |
| Shot 4: Vera enters | Conditional failure: identity and lantern-hand drift |
Why platform self-review was the largest risk
In its final report, Flova marked all four shots as passed and explicitly stated:
- the blade’s geometric contact in Shot 2 passed;
- Vera’s right-hand lantern in Shot 4 passed;
- no style or identity drift appeared in the film.
The frame-by-frame evidence does not support those conclusions. Earlier in Gate B, it had also passed the Boss with the enormous knight’s helmet.
This suggests that the platform’s “director review” may be influenced by its prior textual intent. If the prompt says “lantern in the right hand,” the summary tends to report a right-hand lantern. If the prompt says “attack the knee,” it reports knee contact. It does not consistently re-evaluate the final pixels as independent evidence.
This is an automation risk, not a wording issue. If an upper-level agent directly trusts the lower platform’s self-assessment, errors pass through every Gate and reach export carrying an “all passed” label.
The most important conclusion is therefore not merely that Flova can generate defective video, but:
Project orchestration may be delegated to a platform; final factual judgment cannot be delegated to the platform that generated the same result.
Technically complete, unsuccessful against the objective
Flova built a complete timeline and exported a 38.24-second video:
- 1280 × 720;
- 30fps;
- H.264 video;
- AAC 44.1kHz stereo;
- no abnormal black frames detected;
- approximately 1.25 seconds of silence at the end, matching the planned fade.
Technical validity and content validity must remain separate. Encoding, audio tracks, duration, black-frame checks, and export all prove that a complete deliverable exists. The original objective required all four shots to hold at once. Two did not pass, so the workflow as a whole failed.
Cost: less orchestration labor, not less rework
Observable Gate B image expenditure was:
- initial nine images: 42 credits;
- first two corrections: 28 credits;
- second set of three corrections: 42 credits;
- total: 112 credits.
The account balance was observed at 6,558 credits before the video stage and at 5,484 after Shot 1, Shots 2–4, timeline assembly, and export—an observable media-execution difference of approximately 1,074 credits. Shot 1 accounted for roughly 144 credits; the combined difference for the remaining three shots, timeline, and export was approximately 930.
Between Gate B and the video stage, the balance increased by 50 credits for a reason not explained by the project log. The figures above are therefore not presented as a fabricated exact all-project invoice. What can be confirmed is that project-level integration did not eliminate media cost, while faulty keyframe correction continued consuming credits before video generation began.
Flova reduced organizational cost. The same project retained its Storyboard, resources, preceding-shot references, sound, and timeline, so each model result no longer had to be moved manually. It did not reduce final review or prevent defective shots from requiring another pass.
Reusable principles from the failed workflow
Allow only one authoritative visual master
Character, Boss, scale, and style need an explicit priority. A corrected version must not automatically override the approved master merely because it is newer.
A video reference is more reliable than repeated description
Spatial continuity from Shot 1 through Shot 4 was substantially better than in the earlier manually assembled sequence. Most of the gain came from inheriting the end of the preceding shot, not from writing longer prompts.
Contact must be proven by pixels
Action prompts, detached sparks, and result states cannot replace contact evidence. The blade and target area must visibly occlude or overlap at the correct moment.
Establish an identity Gate before a new character enters
Vera should not have gone directly into long-form video generation. A safer method is to first generate an entrance frame that passes checks for face, clothing, proportion, lantern, and left/right hand, then use it for image-to-video.
A local error must not trigger a full-image regeneration
A Boss helmet, lantern hand, or duplicated weapon should be handled through local editing. Full regeneration returns every approved style and identity decision to randomness.
Lower-level self-review is evidence, not authority
Flova can propose review items, organize reports, and identify some defects, but it cannot be its own final reviewer. The upper layer must still extract frames, count objects, inspect sound, and compare the result against the original objective.
Conclusion
This Flova test came closer to a real production system than the earlier manually generated shots. It could analyze before acting, plan the sequence, generate shots, inherit references, organize sound, build a timeline, and export a film. Shot 1 and Shot 3 also demonstrated that the structure can produce real improvements.
The objective, however, was not to prove that the platform functions. It was to obtain a film with reliable characters, action, and narrative. Shot 2’s contact causality and Shot 4’s Vera identity still failed, while platform self-review reported both as passes. The complete Flova workflow therefore has to be archived as a failed case.
The next version does not need to discard everything. Shot 1 and Shot 3 can remain; only Shot 2 and Shot 4 need to be rebuilt. The former must lock the knee contact, while the latter should begin from a Vera entrance frame that has already passed identity review. The artifact worth preserving is not this V1, but the project structure and failure criteria that have now been tested.