Alias ArchiveArchive in progress
Archive in progress

AI Action Direction Failure Log: Shot-by-Shot Generation and the Shift to Flova

A stone-giant combat sequence becomes a test of Niji, Image2, Seedance, Dreamina, BytePlus, and Kling—and of why action causality, spatial continuity, and three-dimensional distance ultimately broke the shot-by-shot workflow.

Filed
Field
Other
Edition
EN / Reading copy
The Phase 1 knight faces a colossal stone giant in rain-soaked ruins
This static master passed review: the character, Boss, visual language, scale, and stage are all legible. The real problems began once it moved.

The previous video test used a relatively controllable pipeline: Vera spoke English in HeyGen, then HyperFrames added Chinese subtitles and graphic packaging. The sample still contained a flash and a pause, but it proved that the route from task to deliverable could run end to end.

The second target was harder: a 16:9 game-action sequence. In rain-soaked ruins, the Phase 1 knight would finish a confrontation with the stone giant. The knight would collapse from exhaustion after defeating the Boss, and only then would Vera arrive carrying a lantern. The sequence was meant to demonstrate Seedance’s complex-motion and camera capabilities while remaining suitable for later redubbing and reintegration into the game narrative.

There is no finished film from this test. After a full day of work, I stopped.

Not because every frame was bad, but because action only works when space and causality hold from one frame to the next. That is precisely where the current workflow demanded the most labour and produced the least stable results.

Round one: three playable clips, none that passed

The first plan was a short film of roughly thirty seconds:

  1. The knight evades the giant’s strike, rolls, and cuts toward its knee.
  2. The floor collapses under the impact, sending both knight and camera into free fall.
  3. After defeating the giant, the knight falls, and Vera approaches with a lantern.

Every clip generated successfully. Encoding, frame rate, audio-track, and luminance checks all passed. The first segment even preserved the knight, the giant, and the overall visual language. At normal speed, however, the failure was immediate: the knight rolled before the fist became a genuine threat, then paused after approaching the Boss before apparently remembering to swing the sword.

Contact sheet from the first attempt at the knight evading the giant
Sampled frames preserve composition and identity, but conceal the premature dodge and the brief pause before the attack.

The second segment was worse. The floor showed no load, propagating cracks, separating slabs, or loss of support before the knight abruptly fell through it. “Driving the sword into the wall to slow the fall” became random slashing on either side. The stone giant then reappeared in the distant background without convincing contact with the ground, as though it were floating.

Sampled frames from the failed floor-collapse and free-fall sequence
Individual frames can still feel cinematic; continuous playback reveals the absence of credible support, gravity, and space.

This was the first point at which I separated a technical pass from a directorial pass. A decodable file with no yellow flashes and approximately the right number of characters proves only that the file is intact. It cannot explain why a character moves, when contact happens, or where the force originates.

Correct the story before correcting the prompt

The free fall existed only to demonstrate camera movement, but it also rewrote the game’s story without permission. The actual logic is simple:

The knight defeats the stone giant → collapses from exhaustion → Vera appears after the battle.

There is no collapsing floor, alternate space, black flame, or newly invented ability. The second version therefore removed that technical spectacle and retained one Boss room, one 180-degree axis, and one action chain that could reconnect to the game.

I also broke down a strong reference, Moonwatch Sword Dance. Its visual language differs from ours, but its prompt follows a clear directorial grammar:

wind-up → contact point → material response → recovery / preparation for the next action

A sword does not merely “swing.” It hits a defined location at a defined moment, and only then does the object break. The landing foot, centre of mass, and sword position at the end of one action already prepare the next. The camera is not a separate list of “orbit, push in, low angle”; it follows the blade path, balance, and impact.

That is much closer to real action direction than listing “roll, slash, kneel.”

How the image finally became stable

Before motion could begin, the dual-subject keyframe had already become a bottleneck.

Giving Niji one image of the knight and one of the Boss did not tell the model which attributes belonged to whom. Helmet, cloak, armour, and greatsword all contaminated the Boss, turning the stone giant into another oversized knight. Painting the knight into the frame with Editor made the two subjects look as though they came from different visual worlds, and the costume acquired an unrequested gold symbol and robe. Character Reference was not a strict local identity lock either; it continued to affect the full image.

The eventual division of labour was simpler:

  • Niji produces high-quality single subjects. Character, Boss, clothing, material, silhouette, and the overall art language are resolved separately.
  • Image2 performs controlled assembly. Approved subjects are placed into one scene and adjusted for position, scale, camera, and seams without redesigning them.

Image2 also requires constraints. It readily creates uniform high-frequency noise, dense micro-cracks, plastic highlights, and excessive sharpening. Once those static defects enter video, they crawl from frame to frame. After several correction rounds, the opening master finally held together: knight on the left, Boss on the right, one ground plane, one light direction, and a largely coherent visual language.

The problem was therefore not an inability to produce the image. The image was sufficient. The failure was in adding reliable time to it.

Round two: preventing an early dodge produced a frozen defender

The first five-second BytePlus test focused specifically on the premature dodge. The prompt repeatedly stated that the knight must not move early and must wait until the giant’s fist entered the danger line.

The model obeyed “do not move” but failed to understand continuous tactical adjustment. The knight remained almost frozen until the attack reached its face. The shoulder roll was unreadable, while the Boss’s fist and forearm expanded into a huge foreground rock that obscured the point of contact.

The knight remains still while the giant fist expands into a foreground rock
Fixing “too early” cannot mean freezing the defender. Timing should emerge from active footwork, threat tracking, and a last-moment change of direction.

The next prompt reversed the approach completely. The knight would advance from the first frame, feint, enter, roll, rise, and cut. The camera would accelerate, move low, create parallax, and follow the sword’s line of force. The eight-second prompt was highly detailed. The visual language survived, but the sequence still lacked tension.

This produced an unwelcome conclusion: a longer prompt does not make the model a better director. When five to eight seconds contain footwork, baiting, tracking, impact, rolling, recovery, a slash, the Boss kneeling, camera movement, and sound, the model tends to delete the hardest intermediate steps.

Dreamina and Kling: a complex route, then a matched test

To test stronger movement, I designed another eight-second route. The giant’s fist would land fully; the knight would run along the grounded arm and complete a cut across the shoulder or chest.

Dreamina broadly respected the route, but the knight walked like a machine and the attack had no force. The file contained an audio track, yet no clear and useful action sound was audible. Kling had previously produced more agile climbing, but enabling automatic multi-shot generation created a more severe failure this time: mid-climb, the knight suddenly appeared on an unrelated cliff, then cut back to the Boss, repeated the climb, delivered an arbitrary slash, and jumped away.

Kling cuts from the Boss to an unrelated cliff and then back again
Multi-shot generation created momentary intensity while splitting one action into unrelated locations. Each shot behaves like an isolated spectacle; the space does not continue.

For a fair comparison, I reduced the task to one genuinely closed five-second action:

  1. The Boss raises its fist.
  2. The fist enters the danger range.
  3. At the last moment, the knight rolls diagonally forward and right.
  4. The fist strikes the ground the knight has just vacated.
  5. The knight stops inside the Boss’s guard, facing it, without counterattacking.

Dreamina and Kling used the same opening image, the same resolution, and closely matched prompts. Dreamina won this round: the action order was more complete and the knight more clearly entered the Boss’s inside line. Kling preserved the scene but felt sluggish; the evade resembled a low dash rather than a legible roll.

Dreamina R4 · The relative best five-second candidate in this round, not a finished shot. Foreground rock still partly obscures the fist’s contact geometry.
High-frequency contact sheet from the Dreamina five-second evade test
Reducing the prompt to one closed action substantially improved causality. “Relative best,” however, is not the same as passing a complete action chain.
Kling R4 · Matched comparison. The scene remains continuous, but the character is sluggish and the evade technique is unclear.

Honestly, Kling’s two continuity failures made me laugh: one changed mountains halfway through the climb, and the other turned a roll into a low dash. Even a serious production review has its limits.

The final continuation exposed the fundamental problem

I extracted a relatively stable frame from the end of Dreamina R4 and requested another five seconds:

The knight immediately drives up from one-knee support and delivers an upward-left diagonal cut into the stone giant’s near knee joint. Only after blade contact does the rock fracture along the cut.

The sampled output looked successful: the knight rose, the blade flare struck, the Boss’s knee lit up, the rock split, and the giant lowered its weight. Automated QA therefore reported “causal continuity.”

At normal speed, the error was impossible to miss:

The knight is visibly far from the Boss. There is no step, dash, or approach; the knight simply stands up in place and the sword somehow reaches the target.

Dreamina R5 · Explicitly rejected. Compare the starting distance with the sword’s physically reachable range.
Contact sheet of the knight striking the stone giant’s knee
The contact sheet captures “rise, blade flare, fracture, descent,” yet hides the fact that the character never traversed the required distance.

This was the most valuable failure of the round. A last-frame continuation can preserve colour and broad composition without correctly reconstructing the hidden three-dimensional space. When the prompt says “immediately strike,” the model may remove the distance instead of producing a plausible approach.

It also confirms that contact sheets are useful for composition, identity, and colour, but cannot certify action causality. Final review must begin with normal-speed playback, followed by slow inspection of contact, support, momentum, and recovery.

Why I did not force it into a finished edit

These are not editing problems.

  • An early dodge cannot become a credible reaction by deleting a few frames.
  • A character who is too far away cannot be fixed by adding a blade flare.
  • Uncaused cliffs, teleportation, and positional resets cannot be hidden with transitions.
  • The presence of an audio track does not mean the effects are clear or synchronized with impact.
  • Several attractive isolated shots do not automatically form coherent action.

Continuing would require rebuilding keyframes, prompts, generation tasks, sampled-frame QA, normal-speed QA, and slow-motion QA for every five seconds, then treating each last frame as the spatial starting point of the next segment. At that pace, a week might produce only a short, barely acceptable sequence. The process would no longer test automated production; it would use human coordination to compensate for every model uncertainty.

The material therefore did not enter HyperFrames, and the failed shots were not packaged as a “completed” film.

Why the next step is Flova

Flova is next to be evaluated, not because it has already solved the problem, but because its unit of work is closer to the actual bottleneck. Project, storyboard, assets, model generations, and final delivery can be managed in one workflow, with access to multiple models including Image2, Seedance, and Kling.

The real test is not whether it can produce another clip, but whether it can satisfy eight requirements:

  1. Produce a storyboard that respects the game narrative instead of inventing collapses and alternate spaces.
  2. Distinguish initiator, responder, contact point, and physical distance within a shot.
  3. Preserve character, equipment, Boss, terrain, and screen direction across shots.
  4. Select different models for static assembly, single-shot action, and specialist shots.
  5. Understand the separate roles of Niji single subjects, the Image2 master, and action references.
  6. Expose reviewable storyboards and task states before the next credits are spent.
  7. Generate audible, synchronized ambience and contact effects.
  8. Preserve actual model, parameter, cost, and failure records rather than becoming another black box.

The Flova CLI and API are connected, but Flova has not yet generated this project. It is not the winner of this article; it is the next hypothesis to test.

What the test actually delivered

There is no finished film, but several rules are now strong enough to become defaults:

  • Use Niji for high-quality single subjects and visual language; use Image2 for controlled multi-subject assembly.
  • Resolve identity, clothing, equipment, and the spatial master before entering video generation.
  • Give five seconds one causal loop; short duration is not permission to add more actions.
  • Specify wind-up, danger line, contact point, material response, and recovery for every impact.
  • Let the camera follow force and centre of mass instead of stacking generic “cinematic” terms.
  • Treat “audio track exists” and “sound effects work” as separate checks.
  • Keep technical QA, sampled-frame QA, and directorial QA distinct.
  • Give normal-speed human review the final veto.
  • When coordination costs exceed the value of automation, stop and change the workflow instead of drawing again.

The first AI-presenter video showed that an automated pipeline could complete a route. The second action-video test offered the necessary correction: a workflow can run without being worth continuing.

This time, stopping was itself the deliverable.

Related reading

04 / LINKS
C01

AI Presenter Workflow: Coordinating Codex, HeyGen, and HyperFrames

From one assignment to an English virtual presenter, Chinese subtitles, and a social-ready master—including the revisions and two unresolved defects.

Read post ↗
C02

Watchman's Opening Cinematic: Silent Narrative and Production Workflow

A record of Watchman's 59-second silent opening cinematic, from concept keyframes and segmented image-to-video generation to editing, OGV conversion, and Godot integration.

Read post ↗
C03

AI Music-Video Production: From Still Generation to Editing and Release

An eight-stage AI image workflow covering song approval, still generation, animation, quality review, colour matching, editing, and release.

Read post ↗
C04

AI-Assisted Pixel Boss Production: The Stone Giant Pipeline

A production note on moving from incompatible high-resolution concept art to a game-ready pixel boss, including style anchoring, animation constraints, and iteration failures.

Read post ↗