Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

So the benchmark is : Two models with different harness produced very different results .

Glm game was completely broken Opus game was at first glance ok but also with bugs

Different models with different cost produced different non perfect results . How is it “close” ? :)

Also on costs : glm burns more tokens on average vs opus . Gpt5.5 burns less surprisingly



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: