one year on
OpenAI unveils DALL-E 2, a text-to-image generator with near-photorealistic output and new editing capabilities
The upgraded model produces 1024x1024 images, supports inpainting and variations, and will be released to the public through a staged rollout.
OpenAI today announced DALL-E 2, a major upgrade to its text-to-image neural network first unveiled in January 2021. The new model produces images at 1024x1024 resolution—up from 256x256—with near-photorealistic quality, and introduces editing capabilities such as inpainting and variations.
DALL-E 2 is a complete redesign from its predecessor. Instead of extending GPT-3 to predict pixels, the model works in two stages: first, OpenAI’s CLIP system translates a text prompt into an intermediate representation; then a diffusion model generates an image from random pixels guided by that representation. This produces higher-resolution images more quickly. Researchers can sign up for access today, and OpenAI plans a staged public rollout informed by feedback from initial users.
OpenAI has implemented some built-in safeguards. The model was trained on data that had some objectionable material weeded out, includes a watermark, and cannot generate recognizable faces based on a name. Users are banned from uploading or generating images that are not G-rated and could cause harm, including hate symbols, nudity, obscene gestures, or major conspiracies or events related to major ongoing geopolitical events.
AI researcher Mark Riedl says DALL-E scores well on his Lovelace 2.0 test of creative intelligence, though he says terms like “create” and “understand” are ill-defined and subject to debate. OpenAI chief scientist Ilya Sutskever described the system as “transcendent beauty as a service,” while artist Holly Herndon said it feels like working in a new medium. OpenAI’s examples include a horse-riding astronaut, teddy bear mad scientists, and a shiba inu in a beret and turtleneck.
The record
Described the leap from DALL-E to DALL-E 2 as reminiscent of the leap from GPT-2 to GPT-3.
Said she is using DALL-E 2 to create wall-sized compositions, describing the experience as 'working in a new medium'.
Noted that DALL-E scores well on the Lovelace 2.0 test of creative intelligence, but cautioned that terms like 'create' and 'understand' are subject to debate.
One year later — open only if you can handle spoilers
DALL-E 2's release ignited the generative image race; Midjourney launched in open beta within months, and Stable Diffusion followed later in 2022. The technology quickly polarized artists and attracted regulatory scrutiny over deepfakes and copyright.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy