# SWE-Bench Projects at AI Tinkerers

> Canonical HTML: https://aitinkerers.org/technologies/swe-bench
> Markdown URL: https://aitinkerers.org/technologies/swe-bench.md
> Technology record last updated: 2026-03-19T06:09:04Z
> Generated: 2026-09-20T16:41:30Z

An evaluation framework that tests LLMs on their ability to resolve real-world GitHub issues through autonomous software engineering.

SWE-bench benchmarks large language models by tasking them with fixing 2,294 functional bugs sourced from popular open-source repositories like django/django and scikit-learn/scikit-learn. Unlike static coding tests, it requires models to navigate complex codebases, modify multiple files, and verify solutions using unit tests. By measuring the percentage of issues successfully resolved (Pass@1), it provides a rigorous metric for the practical autonomy of AI coding agents.

- Official technology site: https://www.swebench.com
- Public AI Tinkerers demos and talks: 0
- Result page: 1 of 1

## Recent Public Talks and Demos

No public projects are currently indexed for this technology.
