Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder where that's from? Maybe they're overindexing on coding even more than others? I can't say I've noticed it in my (coding) usage so far, has anyone seen it make up potential root causes or other speculative stuff more than other models?


0731 is definitely tuned for coding. I mean: https://gertlabs.com/rankings?mode=agentic_coding

But it is also a decent translator from English to Czech in my experience.


Yeah. I stopped ising deepseek v4 flashbecause it is awful (even worse than my local qwen3.6 35B model) at multilingual prose.


Working on a language related app makes me realize that all these supposed language models don't have many good language benchmarks




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: