已验证65% 置信事实精确时间
FrontierCode is an evaluation benchmark for AI models' programming capabilities focusing on real engineering tasks rather than simple algorithm problems
3
来源数
65%
置信度
长期有效
时效性
2026/7/23
首次发现
来源
涉及实体
相关事实
待验证Frontier Code的核心理念是评估代码是否能被项目维护者合并(PR可合并性),而非仅仅通过测试75% 相似待验证Frontier Code从行为正确性、回归安全性、代码规范性、测试正确性、影响范围、代码质量六个维度评估代码质量67% 相似待验证Frontier Code使用名为Mutagen的工具动态调整参考测试以适配智能体的不同实现方式67% 相似待验证AI Guardrails Index项目基于开源数据和代码构建,评估方法论、测试数据集、评分逻辑全部公开64% 相似待验证The coding domain provides clear evaluation criteria for AI alignment research because code either compiles/passes tests or it does not62% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/680220API
curl https://kongchang.com/api/v1/knowledge/claims/680220MCP
get_claim(id=680220)