Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...


For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.


As others commenters said, it's amusing. But also the person you're replying to is the guy who created the pelican test in the first place and I appreciate the whimsy he brings to the discussion.


It's just a little fun.


Well it is just a bit of fun I think. However, I also think an AGI or an extremely capable model approaching AGI would be able to paint a pelican on a bicycle fairly easily. So in that way it is a good metric.


I agree it's fun, no argument there.

However, it's no longer a good metric, as "drawing svg pelicans" is now showing up too much in the training data, so is not proof of generalization.


At this point I want to see some human-drawn pelicans on bicycles. I suspect the LLMs aren't doing all that bad.


Because of all the svg rendering stuff, I added a draw_svg tool to my harness and it has been really nice to get a quick mock-up of ui changes. And conveniently, the new DeepSeek models are really good at knowing when to use it. So I do look at pelican rendering as a small metric of useful capability.


I like to do something outlandish like a "cockatiel driving a UFO on it's way to austrailia"


Because seeing a pelican on a bike is always a good time. Look at it go


becuase most people don't care whether it's accurate, as long as it looks right and is funny...


It doesn't look right at all.


But it does look funny


I prefer the Browser OS test


On either side of the front wheel is a perfectly reasonable place to carry cargo. I think I'd have taken more issue with the spokes, or at least that's what stood out to me. The chain is indeed nice, however.


You don’t need the best model in 99% of cases…


This is true and is only becoming more important the more they improve. I am already moving to checking so they're at least somewhat following the status quo and otherwise prioritizing price and platform. I think this will be an emerging way of viewing AI in 2027 and the winner will probably be open models and China.


I think this likely plateaus and we all just get the smartest intelligence humans need running locally…


Don’t think it’s happening any time soon for most people. My Mac has stayed 32gb for many years now. I don’t think I’m moving into 128gb territory any time soon with all the price hikes.


You don’t need Fable or Sol to execute tasks. However, you need them to supervise and plan. Like, Luna is cheap and is at DSV4F level, but it’s not capable of advanced reasoning.


Might be time to also have it try more three dimension pelican rendering. Or a short animated version (even just a few frames) still in standard SVG


these links never work for me. always "Error: Enter a valid URL" when opening in Firefox. maybe a URL escape issue with Glider?


Can confirm. Also use glider which seems to double encode. I have to open the comment in Firefox/browser and then click on the link


Is this really the right place to bring up an issue with a 3rd party frontend for HN?


to be honest, i hadn't affirmatively diagnosed, but i'm glad i mentioned it because another Glider user corroborated!


my firefox works


we hitting singularity levels of bicycle chain here


Basket? Fish? All I see is the model recursively running itself locally on an eye-pad, which for some reason beyond our understanding is obscuring the invisible fork.


Your tool is giving "Error: Gist API returned 403"


I think I saw a better overall composition out of Flash 0731

Effort on this one?


Default effort for OpenRouter. I'll try a grid of efforts...

Wow, the low, medium, and high pelicans came out in surprisingly different styles: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...


It's interesting that all three of those used roughly the same amount of tokens, and almost entirely output. Feels like the thinking level lever didn't alter cost at all for this specific task, even though it did change the output.


I never trust OpenRouter to forward parameters correctly and would only ever conduct benchmarks with the official api, personally.


should use an open source one you can audit like my site TrustedRouter


That raises the question of what is it actually doing?

If it isn't spending tokens on quality, is it the assumptions about the task difficulty that cause it to perform better? Or are their broader differences in the model being run.


[flagged]


If I live my life on the basis that some people don't share my sense of humor, and hence I should avoid doing anything funny that might be misunderstood, my life will be a lot less fun.


You're doing great. Don't let insanely low-effort (negative-effort, as in making others dumber rather than having no effect?) comments like from the above throwaway affect your actions.


I’m always excited to see “simonw” on my screen.

More likely than not to be interesting & accessible for those of us outside e.g. compsci.

Hope you stay happy so you stay nerdy and keep sharing both with us!


exciting. it's almost like 3 models in one. that variety would matter when trying to solve a creative problem.


Wonder if we'll ever see optimizations for pelican riding a bicycle svg, make it in to model training runs.


Wondering about this in simonw pelican threads is part of the tradition.


I assume this isn't watermarked...


Not bad but the left foot is still in the wrong place


Looks like a belt driven bicycle to me. :D


Honestly they should all use their respective pelicans as their logos. Or maybe a browser plugin to do do that on the Hugging Face and OpenRouter sites.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: