📦 FeatureBench Dataset V1.1 Released
We released FeatureBench Dataset V1.1, primarily correcting inconsistencies and omissions in task problem statements.
We released FeatureBench Dataset V1.1, primarily correcting inconsistencies and omissions in task problem statements.
We added lite split evaluation results for frontier models including GPT-5.5, Claude Opus 4.7, DeepSeek-V4, GLM-5.1, Kimi-2.6, Mimo-V2.5-Pro, and more to the leaderboard.
We released the fast split containing 100 instances (a subset of the full split). These instances require no GPU and are optimized for rapid evaluation. On an Intel Xeon Platinum 8457C with 944GB RAM, the average evaluation time per instance using gold patches is 57.2 seconds.
We now support one-click inference for mainstream agent frameworks, including OpenHands, Claude Code, Codex, Gemini CLI, and mini-swe-agent. All supported agent frameworks can be found here. We have also open-sourced the FeatureBench data pipeline.
FeatureBench is an execution-based benchmark for evaluating agents on end-to-end development of complex real features.