Meta Muse Image: The Essentials
Last updated: September 2, 2026
Covers Meta Muse Image and Meta Muse Image Edit
Meta Muse Image is Meta Superintelligence Labs' first image generation family on Scenario: a text-to-image model plus a matching edit and compose model. Both are built for faithful instruction-following, understanding long conversational, multi-step prompts and rendering clean, legible text directly inside the image, which makes this family a strong pick for anything that mixes visuals with real typography.
Which Model Should I Use?
Model | ID | Input | Best for |
|---|---|---|---|
| Text prompt only | Generating new images from scratch: posters, concept art, infographics, product visuals | |
model_meta-muse-image-edit | 1 to 10 reference images + a text instruction | Editing an existing photo, or composing several photos into one scene |
Start with Meta Muse Image when you have nothing to work from yet. Switch to Meta Muse Image Edit the moment you have a photo, product shot, or set of reference images you want to change or combine.
How to Use the Model
Faithful Text Rendering
Meta Muse Image renders legible, correctly spelled text directly inside the image, which makes it a strong choice for posters, menus, signage, and infographics where the words are part of the design, not an afterthought.
A hand-drawn chalkboard-style menu board for a cozy ramen shop, illustrated steaming ramen bowl with chopsticks in the corner, Japanese-inspired brush-lettering title reading "TONKOTSU RAMEN" with the price "$14" beside it, smaller chalk-script list below reading "SHOYU RAMEN $12" and "SPICY MISO $13", warm cream chalkboard texture, rustic izakaya atmosphere.More text-rendering examples across styles and layouts:
Multi-Step, Detailed Prompts
The model understands long, conversational prompts with several instructions chained together, and it rewards detail rather than getting confused by it. Avoid stacking dense keyword lists (subject, subject, subject, style, style); write it the way you would describe the shot to an illustrator.
A stop-motion claymation style scene of a tiny astronaut mouse in a handmade patchwork spacesuit exploring a candy-colored alien planet, gumdrop-shaped rock formations and lollipop trees, visible clay texture and fingerprint dents like an Aardman animation frame, soft studio lighting, whimsical miniature set design, no text.The fantasy concept art below shows the same principle with no text at all: a single dense, descriptive paragraph, no invented keyword shorthand.
Fantasy concept art of a moss-covered stone golem guardian awakening in an overgrown temple ruin, shafts of green light filtering through a dense jungle canopy above, dramatic low-angle composition looking up at the towering figure, cracked ancient carvings on its chest, painterly digital-art style with rich texture detail, no text.Precise Photo Editing
Meta Muse Image Edit takes one reference photo and a plain-language instruction, and applies the change while keeping everything else recognizable: the same face, the same pose, the same setting. It handles restoration, background swaps, object removal, outfit changes, and full mood or season changes.
Restore this vintage family photograph to modern quality: repair the torn and cracked edges, remove the scratches and creases running across the image, correct the yellowed color cast back to natural true colors, sharpen the faces of the woman in the floral blouse and the young girl in the striped shirt, keep their smiles, poses, and the porch setting with the window, wooden pillar, and potted plants exactly as they are, just clean and clear like a modern photograph.More single-image edits: removing an object, and a full season change, both from one reference photo and one instruction.
Object removal: the trash can is gone, the grass and path fill in naturally. | Season change: same path and benches, now under fresh snow. |
A full mood and lighting change on a street photo, from flat daylight to a moody, rain-soaked night:
Transform this daytime city street into a moody nighttime rain scene: wet reflective pavement, glowing streetlights and warm window light spilling from the storefronts, light rain visible in the air, deep blue-black night sky, keep the buildings, parked cars, and street layout exactly as they are.Multi-Image Composition
Feed the model 2 to 10 reference images in one call and it blends them into a single coherent scene, preserving the identity of the people, products, or objects in each source.
Combine these two photos into a single custom vacation postcard: place the smiling woman with curly brown hair and the yellow sweater from the first photo into the tropical beach scene from the second photo, standing on the white sand near the wooden longtail boats with the limestone cliffs and turquoise water behind her, golden hour lighting matching the beach photo, add a small handwritten-style postcard caption near the bottom reading "WISH YOU WERE HERE".With three references, the same approach builds a marketing composite: a hero product shot placed into an action scene built from two other photos.
Combine these three images into a single dynamic athletic footwear ad: place the black and neon-green trail shoe from the first photo as the hero product in the foreground, add the running woman from the second photo mid-stride on the mountain trail from the third photo at sunrise, golden morning light unifying the scene, dynamic action-ad composition suitable for a sports marketing campaign.Parameters
Both models share the same aspect ratio and output count controls. Meta Muse Image Edit adds the reference images input.
prompt
Required on both models (labeled Instructions on the edit model). Describe subject, style, mood, composition, and any text that should appear. Full, detailed paragraphs work better than short keyword lists, as shown in every example above.
referenceImages
Meta Muse Image Edit only. Required, 1 to 10 images. A single photo is a straight edit (restoration, background swap, outfit change, mood or season change). Two or more images let the model compose them into one scene, as in the postcard and footwear ad examples above.
aspectRatio
21:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, or 9:21. Default 1:1. Pick the ratio that matches the final placement (a poster wants 2:3 or 3:4, a game concept-art still reads better at 16:9).
numOutputs
1 to 4 images from the same prompt and source. Useful for picking the best variation of a composition; each additional image adds to the cost.
Use Cases
Marketing and social posters: product launch posters, event and festival posters, and menu boards with correctly rendered headlines and prices.
Game concept art: characters, creatures, and environment mood pieces in any art style, from painterly fantasy to 16-bit pixel art.
Education: labeled infographics and step-by-step diagrams with legible text baked into the illustration.
E-commerce and product marketing: place a product photo into a lifestyle background, or compose a hero product shot with model and location photos into one ad.
Photo restoration and archiving: repair scratches, fading, and creases on old family photos while preserving faces and poses.
Creative brainstorming: season, weather, mood, and style changes on an existing photo to explore direction before a full production shoot.
Tips for Better Results
Write full, detailed paragraphs, not keyword lists. Every example on this page is a complete sentence describing subject, style, and any text, and the model consistently rewarded that detail over short prompts.
Describe text content directly in the prompt. Put the exact words you want in quotes inside the prompt (as in the poster and menu examples) rather than describing the text abstractly.
Chain multi-step instructions in one prompt. "Make a poster, put the headline on a banner, add a price tag" style requests work in a single call.
For edits, describe what should stay the same, not just what should change. Naming the parts to preserve (face, pose, layout) measurably improved consistency in testing.
Ground compose prompts in what is actually in each reference photo. Naming the real subject of each image (hair color, clothing, landmark) produced far more coherent composites than a generic instruction.
Match aspect ratio to source images when composing. Reference images generated on another model should land on one of this model's own supported ratios (listed under Parameters above) rather than an arbitrary size, or composition quality can suffer.
Use numOutputs to explore variations before committing. Generating 2 to 4 versions of a composition is an easy way to pick the strongest layout.
Known Limitations
No negative prompt or seed control. Both models expose only prompt, aspect ratio, and output count. Fine-grained control comes from being more specific in the prompt itself.
Reference image aspect ratio matters. If a source image comes from a generator without a matching aspect-ratio control (for example a model that only accepts pixel width and height), confirm the delivered image lands on one of this model's 9 supported ratios before using it as a reference; an unusual ratio can reduce edit and compose quality.
Up to 10 reference images per call. Larger composites need to be built in stages: compose a subset, then feed the result back in as one of the next call's references.
Pinned examples are curated manually. The showcased examples on the model pages are selected and ordered by the Scenario team in the app, not through this article.