待验证75% 置信事实时间未知
GPT 5.5在TerminalBench 2.0基准测试中以87%的成绩大幅领先,超过了Anthropic尚未发布的Mythos模型
1
来源数
75%
置信度
中期 (~90 天)
时效性
2026/5/31
首次发现
有效期至:2026/8/29
来源
GPT 5.5 vs Claude Code vs DeepSeek V4:三大编码模型实测对比
bilibili量子菠萝_
涉及实体
相关事实
待验证GPT 5.5 leads on TerminalBench 2.0 with a score of 87%, surpassing even Anthropic's unreleased Mythos model84% 相似待验证在DeepSWE基准测试中,GPT 5.5以约70%的通过率高居榜首,比Opus 4.7领先超过15个百分点80% 相似待验证GPT-5在HumanEval、SWE-bench等主流编程基准测试上相比前代模型表现大幅提升80% 相似已验证GPT-5.4在SWE-Bench Pro测试中拿下57.7%准确率,超越了GPT-5.3-Codex的56.8%79% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/2464API
curl https://kongchang.com/api/v1/knowledge/claims/2464MCP
get_claim(id=2464)