Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Sorry, "this" referred to the parent comment's claim.

> models starting becoming "moody" due to their proprietors arbitrarily modifying their performance capabilities

The tokenizer changes are measurable, the above is quite difficult to quantify.

There are a few sites floating around that purport to, but all of them have fatal flaws in their methodology.



Unfortunately, LLM performance isn't an exact science and some observations are going to be subjective. Observations like ChatGPT being "lazy" in the Winter. Wanting to form opinions based on hard data, aka science, and not vibes is entirely reasonable but doesn't make the vibes a figment of imagination. Or as Jeff Bezos put it, "When the data and the anecdotes disagree, the anecdotes are usually right." And while he's not a scientist, his success does put some weight behind that quote. (as does digging deeper in what he meant by that.)




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: