Intro to Generative AI: The Big Picture of Image, Video, Text, Audio

AI Navigate Original / 5/16/2026

共有:

Key Points

  • Generative AI spans text, image, video, audio, 3D—know each
  • Text free; image/video/audio mostly paid; 3D still experimental
  • Often combine multiple AIs for one work; keep your own hand
  • Start by genre: text→ChatGPT, image→Midjourney, video→Runway

The Big Picture of Generative AI

"Generative AI" spans many areas—image, video, text, audio, 3D. Before starting as a creator, organize what kinds exist and what each is good at.

5 Main Categories

1. Text Generation

  • Main: ChatGPT, Claude, Gemini
  • Uses: articles, novels, scenarios, social, copy
  • Can start sufficiently for free

2. Image Generation

  • Main: Midjourney, DALL-E, Stable Diffusion, FLUX
  • Uses: illustration, photo-style, concept art, manga
  • Mostly paid (pricing/plan names change often—check each official site)

3. Video Generation

  • Main: Google Veo, Runway, Kling, etc. (OpenAI Sora's offering changed—check the latest)
  • Uses: ads, social video, prototypes, visual expression
  • Pricier, longer generation time

4. Audio/Music Generation

  • Main: Suno, Udio (music), ElevenLabs (speech synthesis)
  • Uses: BGM, narration, voice-actor substitute
  • Confirm commercial use

5. 3D/Special

  • Main: Meshy, Luma AI, Tripo
  • Uses: 3D models, Unity / Blender integration
  • Still experimental, evolving rapidly

Positioning for Creators

Conventional workflowAfter AI integration
Ideating from blankBrainstorm/draft with AI
Make assets one by one10 variations with AI → choose
Deadline-driven mass productionCompress time with AI, humans focus on finishing
Asset cost (stock photos etc.)Generate with AI as needed

3 Things to Pin Down Before Starting

1. Copyright/Commercial Use

Rights of output and training-data issues are detailed in the next article "Copyright Guide." Free-plan output may be commercial-NG in some cases.

2. Tool Combination

Using multiple AIs for one work is common:

  • Scenario in ChatGPT → background in Midjourney → BGM in Suno → video with Runway

3. Keep "Your Own Hand"

AI 80% + human 20% shows more originality than AI 100%. Humans judge the finished form.

Where to Start

  • Text-mainly: ChatGPT / Claude
  • Illustration/photo-mainly: Midjourney
  • Video-mainly: Runway
  • Music-mainly: Suno

Next Step

The next article "How to Choose Generative AI" details the judgment axes for picking tools that fit your genre.