We explore the 'pelican-bicycle' benchmark, an inventive way to assess the accuracy of AI image generation.
Imagine you tell an AI, “Draw me a pelican riding a bicycle.” What kind of picture will it produce? Will it just be a pelican-shaped doodle, or will it be a dynamic pelican actually pedaling a bike?
The AI industry is recently flooded with countless ways to measure model performance. However, among them is a particularly eye-catching, somewhat quirky, yet powerful benchmark: the ‘pelican-bicycle’ benchmark.
Why Does It Matter?
Typically, the metrics we use to evaluate AI performance are quite rigid—things like “How many math problems did it solve?” or “How accurately did it write code?” However, these numbers alone make it difficult to grasp how an AI actually ‘understands’ the world.
The ‘pelican-bicycle’ benchmark is different. This task evaluates how accurately an AI can visually implement—beyond simple sentences—elements and physical actions (the act of riding a bike) using graphics Source: GitHub - simonw/pelican-bicycle. It serves as an intuitive litmus test for whether an AI possesses a logical structure in image generation that transcends mere text.
Understanding the ‘Pelican-Bicycle’ Benchmark
As the name suggests, this benchmark involves giving the AI a prompt: “Generate an SVG (Scalable Vector Graphic) of a pelican riding a bicycle” Source: GitHub - simonw/pelican-bicycle.
In simple terms, it’s like asking a child to “draw a pelican riding a bike” and observing if they place the pelican’s feet correctly on the pedals and if the pelican’s beak is pointing toward the handlebars. Because an AI needs to understand how these two concepts interact—not just know the words ‘pelican’ and ‘bicycle’—it must have deeper reasoning to draw it correctly. Metaphorically, it tests the ability to ‘direct’ a situation rather than just memorizing dictionary definitions.
Simon Willison has used this benchmark in presentations to summarize the progress of AI models over the past six months Source: GitHub - simonw/pelican-bicycle.
Current State: How Far Has AI Come?
The ‘pelican-bicycle’ benchmark has become a symbolic task for evaluating the visual implementation capabilities of AI models. Simon Willison maintains a tag titled ‘pelican-riding-a-bicycle’ on his blog, where he consistently reviews how various cutting-edge AI models perform on this task Source: GitHub - simonw/pelican-bicycle.
It demonstrates that AI is advancing rapidly, moving beyond basic word recognition to handle complex images and vector graphic data. Of course, depending on the model, the bicycle structure sometimes collapses or the pelican’s form can appear grotesque. These trial-and-error moments are like growing pains as AI increases its cognitive capacity.
What Lies Ahead?
In the future, the ability to express complex actions and physical laws in formats like SVG or other graphics—beyond generating 2D images—will become a core competitive edge for AI models. As more creative and demanding benchmarks like ‘pelican-bicycle’ emerge, we will be able to more accurately evaluate how closely AI is following human logical reasoning.
MindTickleBytes AI Reporter’s Perspective
Sometimes, one pelican riding a bicycle reveals the state of AI more clearly than a technical report filled with complex numbers. The evolution of AI is moving in a direction that is much more ‘visual’ and ‘intuitive’ than we might imagine. Just like this benchmark we learned today, an era is coming where we will test AI intelligence in even more fun and inventive ways. That is exactly why I look forward to seeing what other animals will be riding which vehicles next.
References
- GitHub - simonw/pelican-bicycle: LLM benchmark: Generate an SVG… (https://github.com/simonw/pelican-bicycle)
- Creating a video of a pelican
- Generating an SVG image of a pelican riding a bicycle
- Searching for information on bicycles
- Simon Willison
- Google DeepMind
- OpenAI
- They intuitively show AI's cognitive abilities rather than just raw numbers
- To find the fastest model
- To reduce server costs