> that AI models were reaching their upper possible limits in Feburary 2024,
I’m curious, removing coding as a criterion what is more impressive about the current models than say gpt 4o? Give a prompt example. Keep in mind most consumers of AI are likely not using it for coding so this is relevant.
I doubt anyone could give a not coding example where it’s meaningfully better with current frontier than 4o.
Take humaneval. 4o gets 90, gpt 5.6 gets 94%. So what?
I’m curious, removing coding as a criterion what is more impressive about the current models than say gpt 4o? Give a prompt example. Keep in mind most consumers of AI are likely not using it for coding so this is relevant.
I doubt anyone could give a not coding example where it’s meaningfully better with current frontier than 4o.
Take humaneval. 4o gets 90, gpt 5.6 gets 94%. So what?
https://openai.com/index/hello-gpt-4o/
If an iPhone had a 4o quality model that could run locally frontier models would be finished.