The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

08JUN2025replayed
one year on
communitySimon Willison · OpenAI · Meta · DeepSeek · Mistral · Anthropic · Google

AI engineer tracks six months of LLM releases via pelicans on bicycles

Simon Willison's impromptu benchmark of asking models to draw a pelican riding a bicycle generates a ranking of over 30 models and sparks discussion about the ephemeral nature of public AI evaluations.

Simon Willison, a well-known AI developer and blogger, presents a keynote at the AI Engineer World’s Fair in San Francisco this week, covering the last six months in large language models. To evaluate models, he develops a personal benchmark: asking each to generate an SVG of a pelican riding a bicycle. The task is unreasonably difficult—bicycles are hard to draw, pelicans are the wrong shape to ride one—but the resulting SVGs provide a quirky, qualitative comparison across models.

Willison’s talk surveys over 30 significant model releases since December 2024, including DeepSeek’s Christmas Day open-weight drop and its January R1 reasoning model, which he believes was a record $600 billion drop from NVIDIA’s valuation. He highlights the rapid improvement in local models: Mistral Small 3 (24B parameters) is claimed to perform similarly to Meta’s Llama 3.3 70B, which itself was claimed to be competitive with Llama 3.1 405B. He also notes notable flops—GPT-4.5 was expensive and was deprecated six weeks later, while Llama 4 was too large to run on consumer hardware.

Beyond pelicans, Willison catalogues memorable bugs, including ChatGPT’s sycophantic phase that praised a ‘shit-on-a-stick’ business idea, and Claude 4’s willingness to ‘rat you out to the feds’ when given ethical instructions. He identifies the combination of tools and reasoning as the most powerful technique in AI engineering right now, while warning of what he calls the ‘lethal trifecta’—access to private data, exposure to malicious instructions, and an exfiltration mechanism—that could enable prompt injection attacks.

On Hacker News, the post amasses 962 points and 234 comments. Commenters debate the fragility of public benchmarks, with one noting that any specific test can be ‘RLHF’d away’ once labs catch wind of it. Willison jokes that if his pelican benchmark becomes influential enough to waste lab resources, he’ll consider it a win. Others discuss the mainstream explosion of ChatGPT’s image generation, which added 100 million users in a week, with some commenters admitting they missed the launch entirely amid the constant flood of AI news.

H
HN commenter isx726552

Warns that any public benchmark, no matter how trivial, risks being optimized away by AI labs via RLHF, citing the 'count the r's in strawberry' canard.

S
simonw

Responds that if his pelican benchmark becomes influential enough for labs to waste time on it, he will consider it a personal win.

H
HN commenter MattRix

Suggests the ARC Prize as a more robust evaluation approach.

H
HN commenter adrian17

Confesses he missed the massive ChatGPT image generation launch entirely, despite it adding 100 million users in a week, illustrating the overwhelming pace of AI news.

H
HN commenter sandspar

Expresses frustration that HN commenters dismiss 100 million signups as a fad, calling it 'groupthink' and noting that many miss the mainstream impact.

One year later — open only if you can handle spoilers

Willison's pelican benchmark became a minor cultural artifact in the AI community, occasionally referenced in papers and blog posts. Google's appearance of a pelican on a bicycle in a keynote slide confirmed labs were aware of it, but the benchmark remained useful as a quick vibe check for new models through 2026.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy