Ten seconds of coffee
My first test was a golden retriever jumping over a wooden hurdle. It produced a playable file, but the anatomy and motion were unconvincing. More detail would only have made those problems easier to see. I needed to change the brief.
After experiments with quiet landscapes and space scenes, I came back to something familiar: coffee pouring from a glass carafe into a ceramic cup, lit by a window. The brief gave the model something specific to move—a connected stream and ripples on the surface—while keeping the camera and cup steady.
I generated the reference image with OpenAI image generation, then animated it locally using Wan 2.2 5B on the same NVIDIA RTX 5050 with about 8 GB of video memory. The sample is one continuous ten-second shot at 832 × 480 and 24 frames per second, with no audio. It is a manually directed image-to-video experiment; the reference-image step was not local.
The useful unit was a shot
I built the pipeline around a small division of labor. A local language model turns the request into a shot plan. A video-generation workflow renders each shot. A conventional media tool normalizes the clips and assembles the final MP4. The language model is the director in this analogy, not the source of the pixels.
The default shot length is five seconds. A sixty-second request therefore becomes twelve shot jobs in the current planner. That is the decomposition implemented in the code, not a claim that I produced a coherent one-minute film. The original dog test exercised that pipeline; the ten-second coffee sample above is a separate, manually directed image-to-video render through its ComfyUI backend.
The current loop renders shots sequentially. Independent jobs create a possible path to parallel workers later, but they do not make this version parallel. Keeping those two statements separate matters when a promising architecture starts sounding like a finished product.
| Stage | Responsibility |
|---|---|
| Request → plan | Local language model produces structured shot descriptions |
| Plan → validated jobs | Application checks shot count, assigns durations, and applies the requested style |
| Shot → generated footage | ComfyUI runs the selected video workflow; LTX in the original pipeline, Wan for the coffee sample |
| Clips → final file | FFmpeg trims, normalizes, and joins the output |
Eight gigabytes changed the design
The original workflow uses the LTX-Video 2B 0.9.8 distilled checkpoint, an FP8 text encoder, and tiled VAE decoding. The runtime log records low-VRAM operation and weight offloading. Those are specific choices in this experiment, not a claim that any video model will fit on an 8 GB card. The coffee experiment uses the installed Wan 2.2 5B workflow instead, with 20 sampling steps and the same low-VRAM approach.
Model files on disk and the memory needed while executing a stage are different quantities. The question was which weights and intermediate data had to be resident together. Offloading and tiled decoding let the workflow manage that pressure, with time and transfer costs that do not disappear just because the job completes.
I kept planning and rendering as separate steps. The original LTX workflow also uses a fixed eight-step sampling schedule. Exposing a setting named steps in an application would be misleading if changing it did not change that schedule. The interface and the engine have to agree on what a control actually controls.
The director needed a contract
Asking a language model to think like a director was not enough. The application validates the number of shots and rejects blank descriptions. It sets the shot indices and durations itself, then appends the requested visual style. These are decisions the program can enforce instead of hoping the model follows them.
The directing instructions became concrete: describe visible action before, during, and after the movement; keep the subject in frame; place an obstacle across the direction of travel; do not add another action after the requested one. The goal was a usable shot description, not increasingly cinematic prose.
If the model-backed planner fails, the application logs the failure and falls back to a deterministic planner. That keeps the workflow usable, but a fallback plan is a different kind of result. A system should preserve that distinction instead of presenting every completed job as equivalent.
A valid MP4 can still tell the wrong story
There are several independent ways for this project to succeed. The shot plan can match its schema. The renderer can return frames. The file can have the requested duration and frame rate. The final video can play in a browser. Each is worth checking, and none proves that the requested action looks convincing.
Increasing resolution gives the viewer more detail to inspect. It does not prove that the dog clears the obstacle naturally, keeps a consistent body, or moves plausibly between frames. The dog test made that limitation visible. The coffee shot made motion easier to judge: a connected stream, changing ripples, and a stable cup. Its limitation is visible too: the coffee level barely rises over ten seconds. A more convincing composition still does not prove that the model understands fluid physics.
I also kept a mock rendering mode that makes simple title-card clips. It is useful for exercising job flow and file assembly without paying for video inference. It cannot test visual quality. A test earns its value by making a specific claim, not by standing in for everything downstream.
The next problem is continuity
Short independent shots reduce the size of each generation job. They also create another problem: making the next shot belong to the same scene. A consistent character, location, and direction of motion cannot be assumed just because two clips concatenate successfully.
The coffee experiment now demonstrates a reference-image and image-to-video workflow for one shot. Automated quality scoring, selective retries, sound, and a richer editing workflow remain future work. Keeping a scene consistent across multiple shots is still unproven.
What I have now is a working slice of a local creative pipeline: a plan, a renderer, an assembly step, and an artifact I can inspect. The exciting part is how much becomes possible when the job is divided carefully. The unfinished part is teaching the pipeline to recognize when its output deserves to survive the edit.