For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.
As others commenters said, it's amusing. But also the person you're replying to is the guy who created the pelican test in the first place and I appreciate the whimsy he brings to the discussion.
Well it is just a bit of fun I think. However, I also think an AGI or an extremely capable model approaching AGI would be able to paint a pelican on a bicycle fairly easily. So in that way it is a good metric.
Because of all the svg rendering stuff, I added a draw_svg tool to my harness and it has been really nice to get a quick mock-up of ui changes. And conveniently, the new DeepSeek models are really good at knowing when to use it. So I do look at pelican rendering as a small metric of useful capability.
On either side of the front wheel is a perfectly reasonable place to carry cargo. I think I'd have taken more issue with the spokes, or at least that's what stood out to me. The chain is indeed nice, however.
This is true and is only becoming more important the more they improve. I am already moving to checking so they're at least somewhat following the status quo and otherwise prioritizing price and platform. I think this will be an emerging way of viewing AI in 2027 and the winner will probably be open models and China.
Don’t think it’s happening any time soon for most people. My Mac has stayed 32gb for many years now. I don’t think I’m moving into 128gb territory any time soon with all the price hikes.
You don’t need Fable or Sol to execute tasks. However, you need them to supervise and plan. Like, Luna is cheap and is at DSV4F level, but it’s not capable of advanced reasoning.
Basket? Fish? All I see is the model recursively running itself locally on an eye-pad, which for some reason beyond our understanding is obscuring the invisible fork.
It's interesting that all three of those used roughly the same amount of tokens, and almost entirely output. Feels like the thinking level lever didn't alter cost at all for this specific task, even though it did change the output.
That raises the question of what is it actually doing?
If it isn't spending tokens on quality, is it the assumptions about the task difficulty that cause it to perform better? Or are their broader differences in the model being run.
If I live my life on the basis that some people don't share my sense of humor, and hence I should avoid doing anything funny that might be misunderstood, my life will be a lot less fun.
You're doing great. Don't let insanely low-effort (negative-effort, as in making others dumber rather than having no effect?) comments like from the above throwaway affect your actions.
Honestly they should all use their respective pelicans as their logos. Or maybe a browser plugin to do do that on the Hugging Face and OpenRouter sites.