Technology

SWE-Bench Verified

SWE-bench Verified is a human-validated coding benchmark of 500 real-world GitHub issues designed to reliably test AI agents on software engineering tasks.

Developed in collaboration with OpenAI's Preparedness team, SWE-bench Verified isolates a high-fidelity subset of 500 tasks from the original SWE-bench dataset. Human annotators manually vetted each instance to guarantee clear problem descriptions, correct test patches, and solvable tasks (eliminating the noisy, underspecified, or overly specific unit tests that plague automated benchmarks). By focusing on 500 verified issue-pull request pairs across popular Python repositories, it serves as a rigorous, containerized standard for evaluating how effectively LLMs and AI coding agents can resolve actual software engineering bugs.

https://www.swebench.com

Recent Talks & Demos

Showing 1-0 of 0

Members-Only

Sign in to see who built these projects

No public projects found for this technology yet.