Docs menu
Introduction & Getting Started
Creating with AI
Products & Talent
Organize & Assets
Video Tools
Account & Workspace
Troubleshooting & FAQs
Writing Prompts
How the composer, character locking, style controls, camera settings, and Director Chat turn a description into a finished generation.
The composer isn't a text box
Every generation starts in the composer, and the composer is a chip editor, not a plain textarea. You type words, and you also insert chips: small inline blocks that carry more than text can. A finished prompt is usually a mix of both.
Two chip families do different jobs:
- Reference chips condition the model on a real image: your own upload, a character, a product, or an asset from your brand library.
- Style chips insert a written phrase into the prompt: a camera treatment, a color palette, an effect, a location, or a general style direction.
Open the + Add menu to pick from either family, or type @ anywhere in the composer for a live, filtered list of everything you can reference.
Adding characters and products
To put a specific character or product in a shot, pick it, don't type it. Open + Add, then Character (or Product) and choose from your list, or type @ and pick from the dropdown of characters, products, and brand assets that matches what you type. Typing a name alone does nothing until you select a match: the composer isn't parsing your text for names, it's waiting for you to point at a real record.
That distinction matters because it's what makes identity reliable. Every pick attaches the actual character or product record, reference photos and all, not a word that has to be re-recognized later. If that character has a trained model, a short identity tag rides along automatically. You'll never see it, write it, or need to remember it: it's just there because you picked the character.
No trained model yet? Nothing breaks.
Three style controls, three different jobs
Misu has three separate ways to steer style, and they don't overlap. Knowing which one you're touching saves a lot of confusion later.
| Control | Where you set it | Applies to | What it does |
|---|---|---|---|
| Style Lock | One toggle in brand settings (on by default) | Every generation, automatically | Appends your brand's saved style references as text, and flags anything you've marked to avoid. |
| Style / Camera / Colors / Effects / Location chips | The composer's + Add menu | This one generation | Inserts a specific written phrase, e.g. "Dark & Moody" becomes dark moody aesthetic, deep shadows, dramatic chiaroscuro lighting. |
| Brand DNA | The campaign wizard | An entire campaign | A structured style, composition, and environment profile, extracted from your own reference images or picked from a curated inspiration library. |
Style Lock is brand-wide and silent: once it's on, every generation carries your brand's look without you thinking about it. Chips are the opposite: local, visible, one generation at a time. Brand DNA sits above both, for campaigns that need dozens of shots to agree with each other, and is built either by uploading 8 to 15 real campaign images and letting Misu extract the pattern, or by starting from a reference brand in the inspiration library.
Camera and composition are closed choices
Camera, lens, focal length, and aperture are dropdowns, not free text, and the same options appear whether you're generating a single shot or building a campaign:
- Camera: None, Cinema 70mm Film, Lo-Fi Toy Camera, Camcorder, 8mm Film, 35mm Film, DSLR, Medium Format, Polaroid
- Focal length: 14mm, 24mm, 35mm, 50mm, 85mm, 135mm, 200mm
- Aperture: f/1.4, f/2, f/2.8, f/4, f/5.6, f/8, f/11, f/16, f/22
Camera defaults to None
Video works differently: instead of lens and aperture, you pick a camera movement (Static, Slow zoom in, Dolly in, Pan left / right, Orbit, Tracking, Crane up, Handheld, Drone), because what a video shot needs is motion, not focal geometry.
Director Chat: describe it instead
If you'd rather describe a shot in plain language than assemble it from chips, Director Chat is the same composer with a conversation layered on top. You can still attach characters and products with @, but you don't have to build the rest by hand.
Behind the scenes, the Director recognizes common formats and applies a tailored approach for each. A few examples from a growing library:
- Brand Bible: Interview the brand once, and every prompt after it knows the world
- UGC Ad: A creator-style video that feels like a real person posted it
- Editorial Cover: Magazine-cover framing, pose vocabulary and print-ready light
- Product Hero: Studio, lifestyle and in-hand product shots that hold up as brand assets
For a straightforward shot, the Director writes the prompt, checks its own output, picks a model, and generates. For higher-stakes formats, such as a UGC-style ad or a multi-shot film, it stops and shows you the plan first, so nothing generates (and no tokens spend) until you say go.
Field notes: the rules behind each approach
Every approach the Director can take rests on a few rules that working directors, DPs and AI-video creators arrived at the hard way. They are the same rules the in-app coach nudges you with, and the same checklist you see before generating when you pick an approach by hand — here they are in one place.
Brand Bible
Interview the brand once, and every prompt after it knows the world
- Written once, inherited everywhere. Saved to your brand's visual guidelines, the bible is read by every prompt the platform composes afterwards. An hour here is repaid on every shot for a season.
- What you refuse defines you. The 'never' list does more work than everything above it. 'Never a seamless white sweep' stops more drift than any amount of describing what you do want.
- Answer with other people's work. Three images you wish you'd made, and what you envy in each. People describe someone else's work concretely and their own in adjectives.
- Materials, not moods. 'Elevated' renders as nothing. 'Matte ceramic against raw linen, no gloss anywhere' renders as itself.
- A bible is versioned, not eternal. Re-run it when the brand shifts — a new season, a new product line, a new face. Stale canon is inherited just as faithfully as fresh canon.
Before you generate
- Three images you wish your brand had made — from anywhere, not your own work.
- One brand in your category you never want to be mistaken for.
- The one product that has to be perfect, and the detail on it that must always be readable.
- Who appears in your work repeatedly — and whether they already have a character sheet here.
UGC Ad
A creator-style video that feels like a real person posted it
- The first three seconds are their own unit. Plan the hook separately from the body — first frame, on-screen text, first words, vibe — and write three versions of it. Test hooks; keep the body.
- Write it for the sound off. Autoplay is muted. If the hook only works when you can hear it, it doesn't work. The on-screen line is the hook; the audio is the backup.
- Use a format people already recognise. Unboxing, GRWM, haul, problem–solution: each has a beat order viewers know. Follow it — a new order reads as an ad.
- The product is in their hands. A product that never touches the speaker reads as fake. Hold it, open it, use it mid-sentence.
Before you generate
- A specific, named product — with a photo if it has to look exactly right.
- Who is speaking and who they're speaking to — one sentence each.
- Any real evidence you want quoted — a result, a timing, a comparison. The speaker can't invent it.
Editorial Cover
Magazine-cover framing, pose vocabulary and print-ready light
- Attitude before product. Catalogue centres the subject and shows the clothes; editorial crops tight, holds tension in the face, and lets the product be secondary.
- Leave room for the masthead. Headroom above the crown and one clean third for coverlines. A cover that fills the frame has nowhere for the words.
- Keep the skin real. 'Flawless' renders plastic. Ask for pores, micro-imperfections and a soft specular on the cheekbone.
Before you generate
- The person is a character with reference panels or a trained model — a cover is not the place to invent a face.
- One strong garment chosen, with a clean neckline.
Product Hero
Studio, lifestyle and in-hand product shots that hold up as brand assets
- Ground the product. A surface with a finish, and a contact shadow or reflection. Without them the product floats like a cutout.
- Say the label faces camera. A beautiful shot with a turned or garbled label is a discard. State it as legible, centred on the front face.
- The detail deserves its own shot. Hands opening the box, the clasp, the label macro: an insert carries what a wide can't. Plan it as a shot, not a hope.
- One mode per image. Studio hero, lifestyle, in-hand: pick one. Blended modes read as stock.
Before you generate
- Product photos from the front and at least one angle, label readable, in even light.
- The exact colour and finish, named — 'matte forest green, brushed brass cap' — so the render doesn't drift from the real product.
- The mode is decided: a studio hero, a lifestyle scene and an in-hand shot are three different briefs.
Character Sheet
Lock a character's identity, then render the reference panels everything else inherits
- Drift usually starts in the reference. Sunglasses, a shadowed face, a busy background: the model locks onto what it can see. Fix the reference before you touch the prompt.
- The body panel has no head. A second, smaller face on the body panel bleeds into later generations and splits the identity. The face lives in one panel only.
- Change one thing per stage. Face lock, then wardrobe, then angles — one change at a time. Two at once and the identity has nowhere to anchor.
- Don't argue with a trained model. If the character has a trained model, describe framing, wardrobe and light only. Re-describing the face fights the weights.
Before you generate
- The reference photo faces the camera, evenly lit — no sunglasses, hat or hand over the face.
- The face is at least a fifth of the frame. A small face in a wide shot gives the model nothing to lock onto.
- Plain background, taken recently, at least 1024px on the short side.
- You've checked the brand's existing characters — rebuilding one creates a second, slightly different person.
Scene Plate
Cinematic scene and environment plates built on an explicit block grammar
- Write the visible. If a phrase describes how the shot feels rather than what the camera sees, cut it. 'Moody' renders as murk; 'haze at 20%, visible from 15m' renders as haze.
- One camera move per shot. A clip renders one move. A push-in and a tilt in the same shot come out as neither — make it two shots.
- Give lighting its own line. Key, fill and rim, with direction and temperature and the source that motivates it. Lighting is half of what reads as cinematic.
- Say what's there. 'Empty platform, no people' invites people. 'Bare platform, one upright bench, wet concrete' does not.
Before you generate
- A real place in mind — materials, time of day, weather — not a category like 'a street'.
- If anyone appears, their character sheet exists; the plate references it rather than re-describing the face.
- You know whether this is a standalone still or one shot in a film — in a film, the world and capture blocks must match the other shots.
Brand Film
Multi-shot narrative video with the world and cast locked across every shot
- Plan the cut, not the shot. Each clip generates on its own, so screen direction and eyelines flip between shots unless you decide them up front. Mark which shots continue the previous one and Storyboard starts them on its last frame.
- One camera move per shot. Two moves in one clip come out as neither. If you want a push-in and a tilt, that's two shots.
- Repeat the locks in every shot. Generation has no memory. 'The same woman as before' produces a different woman — the cast, world and capture locks go into every shot in full.
- Change one thing per retry. When a shot isn't right, change exactly one variable — pose, camera, light or wardrobe. Change three and you can't tell which one fixed it.
- Say what's there, not what isn't. Video models read the nouns. 'No people' puts people in the frame. Describe the empty platform instead.
Before you generate
- A character sheet exists for every recurring person — build one first if not.
- Product photos are attached for anything that has to look like the real product.
- You know the runtime and can afford the clip count — a 60s film is roughly 8–12 renders.
Multi-Angle Coverage
One filmed take, re-shot from angles that were never filmed
- Start from one clean angle. Coverage needs a take filmed as one continuous shot. A clip that already cuts gives the model two spatial layouts to reconcile, and it drifts.
- References are numbered, not named. The source take is @video1, a character sheet is @image1, the audio is @audio1. An invented name like @charsheet resolves to nothing and that reference is quietly dropped.
- Ask for a lens switch. That phrase is what produces a cut inside one generation. Timestamps on their own tend to give you a single continuous camera move.
- Attach the audio to keep the voice. Sound is generated by default, so the model re-voices the scene. Attaching the take's own audio is the only way to keep the original delivery and timing.
- Generate a few and pick. A video reference carries motion and pacing, not frame-exact detail. Two or three runs and a choice is the normal workflow, not a sign something is wrong.
Before you generate
- The take is 4–30 seconds and filmed as one continuous angle.
- You can name what must stay the same — people, wardrobe, location, light, background.
- If the voice matters, you have the take's audio to attach; otherwise it will be re-voiced.
- You have somewhere to cut the results together — the output is one clip per run, not separate angle files.
Contact Sheet
One cheap panel grid that locks beats, wardrobe and light before a single video credit
- Lock the sequence in stills first. One image showing every beat costs a fraction of one video clip. Argue about wardrobe, light and order here — not after six renders.
- Never animate the grid. A video model treats the sheet as one picture and animates the collage. Crop each panel out first, and keep the character sheet — not the storyboard art — as the identity reference.
- Keep the sheet landscape. Vertical grids miscount panels and collapse into comic layouts. The sheet stays landscape even when the panels inside it are vertical.
- Sketch boards hold identity better. Photoreal panels compete with your real references once animated. If the sheet feeds video, a loose tone or line style drifts less.
Before you generate
- The beats are decided — one line per panel.
- A character sheet exists for anyone who appears, so the sheet references a face instead of inventing one.
- You know whether this sheet is for sign-off (photoreal is fine) or will feed video (prefer sketch).
You don't need to pick a model
Misu never asks you to name a model. Auto mode looks at what you've already done, such as which character or product you attached, or whether you toggled Draft or Edit, and infers the right one. A trained character routes to its own trained model. An untrained character with reference photos routes to a model built for identity from images. Text or logo rendering routes to a model tuned for legible type. Editing an existing image routes to an edit-focused model. Draft routes to the fastest option available.
If you want to choose for yourself, switching out of Auto into "All models" hands you the full list. Most of the time, there's no reason to.
Keeping a character consistent across models
A trained model is the strongest identity lock Misu has, but it's architecture-specific: a trained weight file only loads into the model family it was trained for. Every other model instead conditions on your character's reference images directly, and how you write the prompt changes how well that holds up.
Don't re-describe what the reference images already show
Re-describes the reference
"a woman with long brown wavy hair, green eyes, athletic build, standing in a kitchen"
Adds what the reference can't show
"leaning against the counter, laughing, mid-morning window light, apron over the black tank from the reference"
Multi-reference image models
Models that accept several reference images at once work best with 2 to 3 well-chosen angles rather than every photo you have. Lead the prompt with the scene, since the references are already doing the identity work. If wardrobe changes, describe only the new wardrobe.
Multi-reference video
Some video models compose identity from up to seven reference images with no start frame at all. Put the strongest full-body or clearest face reference first, and keep the motion prompt to camera and environment: identity is the references' job here, not the text's.
Start-and-end-frame video
For models that support an end frame, use a real character-sheet or prior-shot image that matches where the motion is heading. An end frame that contradicts the described motion, such as facing away when the prompt describes a turn toward camera, fights the interpolation instead of guiding it.
Single-frame video
A small number of video models take only one starting image and nothing else, so the entire identity burden sits on that single frame. Choose the cleanest, most forward-facing reference you have, and don't ask for an angle or interaction (a profile turn, a hand holding something) the frame can't support: the model has to invent whatever isn't already there, and inventing is where identity drifts.
| Model type | Reference input | Ref count | Prompt should cover |
|---|---|---|---|
| Trained character model | trained identity | n/a | Everything. The trained model carries identity on its own. |
| Multi-reference image models | reference images | 1-4 | Pose, camera, environment, lighting, wardrobe changes |
| Multi-reference video | reference images | up to 7 | Motion, camera, and environment only |
| Start/end-frame video | start image (+ end image) | 1 (+1) | Motion reachable from both frames |
| Single-frame video | start image only | 1 | Motion the single frame can visually support |
Example prompts
Product photography
What gets sent to the model
[Product name] on a marble pedestal, soft top-down studio light. Studio camera setup. Warm neutral tones.
Trained character portrait
What gets sent to the model
[identity tag] seated on a rooftop ledge, looking past camera, late afternoon light. Cinema 70mm Film camera, 35mm lens, f/2.8 aperture.
Untrained character, first shoot
What gets sent to the model
Featuring Mara, a stylist. On a rooftop ledge, looking past camera, late afternoon light.
Cinematic scene, via Director Chat
See also
Start here
How Misu is put together — brands, products, talent, the library, and tokens.
Create your account
Set up your brand workspace in five steps.
Your dashboard
What the home screen shows and where to go next.
Generate an image
Director, Advanced, and Video modes — plus every control on the generate screen.
Generate a video
One clip at a time — frames, motion, models, and sound.
Director Chat
Describe a shot in plain language and let the Director build it.
Storyboard
Plan a multi-shot film on a timeline, render it, and export it as one video.
Campaigns
Brief a shoot, get an AI shot list, generate every shot into your Library and a Board.
Canvas
Wire models together into a repeatable node workflow.
Train a custom model
Teach Misu your product so it renders accurately every time.
Create consistent characters
Train a persona and keep an identity locked across every model.
Products
Build your catalog, import from Shopify, and categorise items.
Talent library
Platform models, Originals, and your own exclusive personas.
Wardrobe & try-on
Dress a model from your catalog and render the look.
Your library
Every generation lands here. Filter, favorite, and set status.
Boards
Group assets into collections per project or campaign.
Upload assets
Formats, size limits, and where uploads are accepted.
Editing tools overview
All ten tools, where to find them, and how edits are saved.
Upscale
Topaz models, scale factors, and face enhancement.
Color grading
Film presets, grain, halation, and colour correction.
Backgrounds & relight
Remove, replace, or relight what's behind the subject.
Viral clips
Cut a long video into scored, captioned short-form clips.
Virality predictor
Score a post before you publish it.
Brand settings
Name, guidelines, colours, and business details.
Team & roles
Add members and what each role can do.
Plans & tokens
What a token is, what each plan includes, and how spend works.
Integrations
Connect Shopify and your ad platforms.
Tokens & limits: FAQ
Costs, refunds, resets, and what happens when you run out.
Common issues
Training failures, drifting identity, and generation errors.
Contact support
How to reach us and what to include.
Acceptable use & content guidelines
What you can generate, consent rules, and ownership.