The 5-Element Frame
The prompt structure that works across major models is to write in the order subject → style → composition → lighting → details. Example: "Tokyo at sunset (subject) / ukiyo-e style (style) / diagonal overhead wide shot (composition) / warm cinematic light (lighting) / Hokusai wave patterns, cloud texture, 8 people (details)."
Subject and Object
- Narrow the lead to one (multiple subjects easily break)
- Choose concrete nouns ("30s woman, business suit" over "person")
- Add verbs/states ("brewing coffee," "concentrating")
Style Reference
- Broad categories like "photo," "watercolor," "3D render"
- Specific artist names (Midjourney references past-work URL with --sref)
- SD reuses composition with IP-Adapter / ControlNet
- Flux Kontext interactively edits like "change only the clothes to red"
Composition Keywords
- Camera: close-up, medium shot, wide shot, bird's-eye, low angle
- Lens: 35mm, 85mm, shallow depth of field, bokeh
- Composition rules: rule of thirds, symmetry, leading lines
Lighting/Color
- Time: golden hour, blue hour, midday, studio lighting
- Light source: rim light, softbox, backlight, neon
- Color tone: warm tones, muted palette, monochrome
Negative Prompts
Stating elements to avoid reduces artifacts. Like "extra fingers, deformed hands, watermark, low quality, blurry." Midjourney uses --no, SD has a negative-prompt field.
Per-Model Character and Selective Use
| Model | Strength | Weakness |
|---|---|---|
| Midjourney V7 | Aesthetic completeness, mood | Text rendering, no API |
| FLUX.2 | Realism, finger/hand accuracy | Minimal UI, little Japanese info |
| Stable Diffusion (SD4) | Customizability, local run | Hard setup |
| GPT Image 2.0 (OpenAI) | Text rendering, instruction-following | Style breadth |
| Ideogram 4.0 | Best English typography | Japanese text unstable |
As of August 2026 the current versions are GPT Image 2.0 from OpenAI, FLUX.2 from Black Forest Labs, and SD4 from Stability AI. Model numbers turn over quickly, so check each vendor's own model list before relying on a name.
Commercial-Use Check
- Adobe Firefly: training data commercially licensed, safe for enterprise
- Getty AI: opt-in material based
- Midjourney / SD / Flux: confirm terms, especially Stability's Community License
- Similarity check: confirm conflicts with existing characters/works
Practical Tips
- Generate 4-8 images with the same prompt and pick the best 1
- Keep consistency with fixed seed + tweaks on the adopted image
- When switching models, adjust prompt style too
- Record failure patterns and accumulate negative prompts
Summary
Image generation works across models when combining the 5-element frame + negatives + style reference. For commercial use, also consider license-safe Adobe Firefly or Getty AI and use them selectively by purpose.



