已过期75% 置信事实时间未知
两款模型的最终测试总比分为平局
1
来源数
75%
置信度
中期 (~90 天)
时效性
2026/5/31
首次发现
有效期至:2026/8/29(已过期)
来源
Gemini 3.1 Pro vs Opus 4.6:前端编程能力实测对比
bilibili伊莱文思帕
涉及实体
相关事实
待验证The tester assessed that O3 and Gemini 2.5 Pro performed the best in the first round of the snake battle game test.58% 相似待验证在Agent Test: Last Exam测试中,Soul得分53.6分,Anthropic的Opus 4.5得分40.5分54% 相似待验证top-1%准确率是评估量化模型与原始全精度模型输出一致性的指标,关注模型置信度最高的1%的token是否与全精度版本一致53% 相似待验证DeepSeek V4 Pro在Humanity's Last Exam基准测试中得分37.753% 相似待验证LFM2.5在多项基准测试中展现出与体量高达其4倍模型相当的竞争力53% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/9566API
curl https://kongchang.com/api/v1/knowledge/claims/9566MCP
get_claim(id=9566)