one year on
OpenAI unveils DALL·E and CLIP, pairing text and images to give AI a grounded grasp of the world
Two new models from the maker of GPT-3 combine natural language and computer vision, letting an AI generate images from whimsical captions and recognize objects without task-specific training.
OpenAI today released two new models that bridge language and vision, extending the approach behind GPT-3 to images. DALL·E, a smaller version of GPT-3 trained on text-image pairs, generates images from natural-language captions. CLIP, a contrastive model, learns to associate images with descriptions scraped from the internet, enabling it to recognize objects without curated training sets.
The most striking result is the “avocado armchair” — a prompt the researchers used to test whether the model could combine unrelated concepts into a coherent object. DALL·E produced multiple images that look like chairs shaped from halved avocados, with the pit serving as a cushion. Other prompts, such as “a baby daikon radish in a tutu walking a dog,” also yielded plausible composites, though some researchers noted the daikon image may be heavily influenced by existing internet art.
“Text-to-image is a research challenge that has been around a while, but this is an impressive set of examples,” said Mark Riedl of Georgia Tech. CLIP, meanwhile, is less likely than other state-of-the-art image recognition models to be led astray by adversarial examples. Chief scientist Ilya Sutskever said the goal is to ground language in visual understanding: “In the long run, you’re going to have models which understand both text and images.”
The release marks a significant step toward multimodal AI, but experts caution that DALL·E still struggles with complex scenes and can memorize training data rather than generate truly novel compositions. Reactions are mixed, with researchers praising the whimsical examples while noting that DALL·E can still struggle with complex scenes and may memorize images it has seen online.
The record
Described DALL·E as an impressive set of examples and noted its ability to handle novel concepts, while expressing suspicion that some outputs may have been memorized from internet art.
Found the ability to generate synthetic images from whimsical text very interesting and praised that results obey desired semantics.
Impressed that DALL·E shows level of control drawing multiple objects and spatial reasoning not seen in existing text-to-image generators.
One year later — open only if you can handle spoilers
DALL·E and CLIP laid the foundation for the text-to-image boom of 2022-2023, inspiring a wave of models like Stable Diffusion and Midjourney. The avocado armchair became the defining meme of early generative visual AI, cited endlessly as proof that neural networks could blend concepts creatively. OpenAI would later release DALL·E 2 and DALL·E 3, each substantially improving image quality and prompt adherence.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy