The past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder where that's from? Maybe they're overindexing on coding even more than others? I can't say I've noticed it in my (coding) usage so far, has anyone seen it make up potential root causes or other speculative stuff more than other models?