Ethan Mollick shows the progression of AI image and video generation with iterations of a prompt about otters using wifi on a plane. He also explains the difference between diffusion and multimodal image generation models (Midjourney vs ChatGPT). These tools get such different results because the underlying technology and approach is different.
While LLMs generate text one word at a time, always moving forward, diffusion models start with random static and transform the entire image simultaneously through dozens of steps. It is like the difference between writing a story sentence by sentence versus starting with a marble block and gradually sculpting it into a statue, every part of the image is being refined at once, not built up sequentially.
But what makes diffusion models interesting is not their increasing ability to make photorealistic images, but rather the fact that they can create images in various styles.
Unlike diffusion models that transform noise into images, multimodal generation lets Large Language Models directly create images by adding tiny patches of color one after another, just as they add words one after another. This gives AIs deep control over the images it creates.