Technology
SWE-Bench Verified
SWE-bench Verified is a human-validated coding benchmark of 500 real-world GitHub issues designed to reliably test AI agents on software engineering tasks.
Developed in collaboration with OpenAI's Preparedness team, SWE-bench Verified isolates a high-fidelity subset of 500 tasks from the original SWE-bench dataset. Human annotators manually vetted each instance to guarantee clear problem descriptions, correct test patches, and solvable tasks (eliminating the noisy, underspecified, or overly specific unit tests that plague automated benchmarks). By focusing on 500 verified issue-pull request pairs across popular Python repositories, it serves as a rigorous, containerized standard for evaluating how effectively LLMs and AI coding agents can resolve actual software engineering bugs.
Recent Talks & Demos
Showing 1-0 of 0