待验证50% 置信事实精确时间
Gemini 2.5 Pro's visual score drops to 11.7 and functionality score drops to 22.6 on full-stack web development tasks.
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
Tsinghua and Zhipu Release Full-Stack Web Dev Benchmark: Top AI Models Fail Spectacularly
bilibili跟我学一辈子AI吧2026/5/22
相关事实
待验证Gemini 2.5 Pro experienced a cliff-like score drop from approximately 63 on static pages to 11.7 visual score on full-stack tasks, representing a drop of over 80%.84% 相似待验证Gemini 2.5 Pro在全栈网站任务中视觉得分仅为11.7,功能分仅为22.683% 相似待验证Gemini 2.5 Pro从静态网页到全栈任务的得分降幅超过60%72% 相似待验证本月 Gemini Drops 更新为小微企业提供新的 AI 支持方式66% 相似待验证Gemini 3.5 Flash在Humanities Last Exam中得分40.2%,落后于Opus的46.9%和Pro的44.4%65% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/50414API
curl https://kongchang.com/api/v1/knowledge/claims/50414MCP
get_claim(id=50414)