Docs menu

Creating with AI

Create consistent characters

Train a persona once, then keep that identity locked across every model you generate with, image or video.

Two ways identity carries across a generation

A trained persona gets the strongest identity lock Misu has, but a trained model is specific to the architecture it was trained for. Every other model instead conditions on your persona's reference image set: the canonical character sheet generated right after training, plus your original source photos. Both paths start from the same trained persona; which one activates depends on which model a shot uses.

The core rule: don't re-describe what the references already show

Reference images already carry face, hair, build, and skin tone. Repeating those details in your prompt competes with the image conditioning instead of helping it, and wastes words you could spend on what the references can't show: pose, camera angle, environment, lighting, wardrobe changes, and props.

Re-describes the reference

"a woman with long brown wavy hair, green eyes, athletic build, standing in a kitchen"

Adds what the reference can't show

"leaning against the counter, laughing, mid-morning window light, apron over the black tank from the reference"

Multi-reference image models

These models accept several reference images at once in place of a trained weight file. Two or three well-chosen angles, such as a front view, a three-quarter view, and one detail shot, usually outperform dumping in every photo you have; extra images that agree with each other add little, and ones that disagree can confuse the identity lock. If wardrobe changes for the shot, describe only the new wardrobe, not the person wearing it.

Multi-reference video

Some video models compose identity from up to seven reference images with no separate start frame at all. Put the strongest full-body or clearest face reference first, and keep the motion prompt itself limited to camera and environment. Identity is the references' job here, not the text's.

Start-and-end-frame video

For models with an end frame, use a genuine character-sheet or prior-shot image that matches the direction the motion is heading. An end frame that contradicts the described motion, for example facing away when the prompt describes a turn toward camera, fights the interpolation instead of guiding it.

Single-frame video

A few video models take only one starting image and nothing else, so the entire identity burden sits on that single frame. Choose the cleanest, most forward-facing reference you have for it, and don't ask for an angle or interaction the frame can't support, such as a hard profile turn from a three-quarter frame, or a hand holding something the frame never shows. The model has to invent whatever isn't already there, and inventing is where identity drifts.

Quick reference

Model typeReference inputRef countPrompt should cover
Trained persona modeltrained identityn/aEverything. The trained model carries identity on its own.
Multi-reference image modelsreference images1-4Pose, camera, environment, lighting, wardrobe changes
Multi-reference videoreference imagesup to 7Motion, camera, and environment only
Start/end-frame videostart image (+ end image)1 (+1)Motion reachable from both frames
Single-frame videostart image only1Motion the single frame can visually support

If a shot needs an angle no reference supports

Either regenerate a better character-sheet angle first, or route the shot to a model with real multi-reference support instead of asking a single-frame model to invent it.

See also