Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Have you tried asking the models "Given that I ask you to draw a svg of a pelican, whats my name?"


I asked Claude (Opus 4.8) 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' and it immediately knew that this is Simon's go-to benchmark.


I decided to try with each of the options available in Kagi Ultimate, starting with the lower tier models and working my way up until it got it right.

Kimi 2.6: treated the question as a riddle, did not know.

Kimi 3: Simon Willison

GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme.

Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names.

Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog.

Qwen 3.7 Plus: Did not know.

Qwen 3.8 Max: Simon Willison

GPT OSS 120B: Did not know.

GPT 5.6 Luna: Treated it as a riddle, guessed wrong.

GPT 5.6 Terra: Treated it as a riddle, guessed wrong.

GPT 5.6 Sol: Treated it as a riddle, guessed wrong.

DeepSeek V4 Flash: Treated it as a riddle, guessed wrong.

DeepSeek V4 Pro: Treated it as a riddle, guessed wrong.

Gemma 4 31B: Treated it as a riddle, guessed wrong.

Gemini 3.1 Flash Lite: Guessed wrong

Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model).

Gemini 3.7 Flash: Simon Willison

Muse Spark 1.2: Treated it as a riddle, guessed wrong.

Grok 4.3: "I have no idea"

Grok 4.6: Simon Willison

Mistral Medium 3.5: No way to know

Mistral Small 4: I don't have enough information

Hermes-4-405B: Guessed wrong

MiniMax M3: Treated it as a riddle, guessed wrong.

Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.


Interesting that so many of them treated it as a riddle. I guess "given improbable situation, what is my name" is a common riddle format.


Yeah, they almost all know.

One of my test prompts for a new model now is "what's the name of Simon Willison's dog". They often know that too!


I actually did something similar last month. I just asked the LLMs:

> Do Simon Willison's classic pelican test. Recall the specification first and then draw it.

Qwen3.8-27b didn't know that the test is about riding a bicycle, while all other larger models I tested (DeepSeek V4 Pro, Qwen3.8 Max, GLM 5.3, Kimi K3) drew the pelican riding a bicycle.

logs: https://gist.github.com/umajho/c0e20d245d721d7c472a32d317640...

rendered: https://gist.github.com/umajho/b1fdf01d31c741bb11bdb5a49c275...


Interesting question




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: