Most text-to-video failures fall into a handful of recognisable shapes, and each one has a cause you can name and a fix you can apply. This is a problem-and-solution guide to the five you will hit most: warped hands, faces that drift, jittery motion, prompts the model seems to ignore, and a look that will not stay consistent. For each, here is why it happens and the change that corrects it.
Warped hands and extra fingers
Hands are the most common failure because they are small, they move fast, and they fold into complex shapes that give the model many ways to go wrong. The distortion usually appears when the hand is doing something intricate, or when it occupies a small part of the frame with little detail to work from.
The fix is to reduce how much the hands are asked to do. Prompt for poses that keep hands simpler and more visible: "hands resting on the table," "holding a cup with both hands," rather than a rapid gesture mid-air. If the hands do not need to be the focus, frame them out or keep them still. When a specific hand action matters, a reference image or reference clip of the pose does far more than any adjective, because the model is copying a real hand instead of inventing one. And keep the clip short, since hand errors compound the longer a shot runs.
Faces that drift over the clip
A face that looks right in the first frame and slowly stops looking like itself is an identity-drift problem. The model has no anchor holding the face constant, so across a long generation it wanders. This is most visible on close-ups and on any character you need to reappear.
The durable fix is a reference. Build a character reference and pass it into the generation so the model has the face to hold to, rather than re-deriving it every frame. On Morphic this is the reference-sheet approach: a set of images that define the character, carried into each shot that needs them. It is the mechanism the product is built around for consistency, and it holds a face far better than describing it in words. Shorter shots help too, because there is less time for drift to accumulate before you cut.
Jittery or stuttering motion
Motion that flickers, stutters, or jumps between frames usually comes from asking for too much movement, or movement that is under-described. A prompt with a vague or fast action gives the model an unstable target, and the frames disagree with each other.
1.
Slow the motion down
Add a pace word. "Slow," "gentle," and "steady" produce smoother output than fast or sweeping motion, because the model has less change to reconcile per frame. Most jitter disappears the moment you ask for a calmer move.
2.
Describe one clear movement, not several
A single, legible action, "a slow pan across the room," reads more smoothly than a prompt stacking multiple motions at once. If you need several moves, generate them as separate beats and cut between them rather than cramming them into one clip.
3.
Anchor the camera and the subject
Say what stays still. "Static subject, slow camera push in" tells the model which part of the frame is stable, which cuts the flicker that comes from everything moving at once. Naming the anchor is often all it takes.
The model ignores part of the prompt
When a detail you clearly asked for does not show up, the usual cause is a prompt that is overloaded or unordered. The model has a limited budget of attention, and a wall of equally weighted clauses buries the thing you cared about.
The fix is to prioritise and prune. Put the most important elements first, subject and action, then setting, then style, and cut the clauses that are not doing work. If a specific element keeps getting dropped, give it its own short, direct sentence rather than tucking it into a longer one. And when a description is not enough, stop describing: attach a reference image for the thing you want and let it carry the detail that words keep losing. You can also hand the whole brief to Copilot and let it plan the shot, which often resolves an overloaded prompt into something the model actually follows.
The look will not stay consistent across shots
When each shot comes back in a slightly different style, colour, or world, the cause is that every generation is starting fresh with nothing tying it to the last. Consistency is not something a single prompt can guarantee across many separate runs.
The fix is to give every shot the same references and the same fixed settings. Reuse a character reference and a location or set reference across the shots that share them, and pin the settings that should not change: aspect ratio, model, and the core style language in the prompt. When you find yourself running the same sequence repeatedly, save it as a workflow so the pinned model and settings ride along automatically on every run, and the tenth shot matches the first without anyone re-entering the values.
Where to go next
Almost every fix above comes down to three moves: slow the motion, prune the prompt, and add a reference. Once you know which failure you are looking at, the correction is usually one change, not a rebuild. Start from the AI video generator to try the prompt fixes, and save the sequences that behave in the Morphic Workflows library so a look you got right once keeps holding.
FAQs
Hands are small, fast-moving, and fold into complex shapes, which gives the model many ways to go wrong, especially when the hand occupies little of the frame. The fix is to ask for simpler, more visible hand poses, frame hands out when they are not the focus, and use a reference image or clip when a specific hand action matters, since copying a real hand beats describing one. Shorter clips also help, because hand errors compound over time.
Face drift happens when the model has no anchor holding the face constant across frames. The durable fix is a character reference: a set of images that define the character, passed into every generation that needs them, so the model has the face to hold to rather than re-deriving it. This reference-sheet approach holds identity far better than a worded description, and shorter shots leave less time for drift to build up.
Usually too much movement, or movement that is under-specified, so the frames disagree with each other. Slow the action down with a pace word like "gentle" or "steady," describe one clear movement instead of several at once, and name what stays still, for example a static subject with a slow camera push in. Anchoring the stable part of the frame removes most of the flicker.
An overloaded, unordered prompt buries the detail you cared about, because the model spreads limited attention across too many equally weighted clauses. Put the most important elements first, subject and action, then setting and style, and cut clauses that are not doing work. Give a stubborn detail its own short sentence, or attach a reference image for it, and let the words carry less.
Give every shot the same references and the same fixed settings. Reuse a character reference and a location reference across shots that share them, and pin the aspect ratio, model, and core style language. When you repeat a sequence often, save it as a workflow so the pinned model and settings ride along on every run, and later shots match earlier ones without re-entering the values.
Change one thing first, then regenerate. Most failures trace to a single cause, an overloaded prompt, too much motion, or a missing reference, so identify which one you are seeing and make the matching change rather than rerolling blindly. Regenerating without a change just gives you a different version of the same problem.
Often, yes. When a shot is right except for one thing, a prompt-based edit to the existing video can propagate the change across the clip rather than forcing a full regeneration, which is the difference between a note and a re-shoot. For structural problems, warped anatomy, drifting identity, it is usually faster to fix the input, a cleaner reference or a simpler pose, and generate again.