Technology

SWE-Bench Verified and Terminal-Bench v2 for benchmarking

SWE-bench Verified and Terminal-Bench v2 serve as the industry-standard gauntlets for testing AI agents on real-world software engineering and sandboxed command-line operations.

Evaluating AI agents requires testing them on practical, messy engineering tasks rather than isolated code snippets. SWE-bench Verified solves this by offering a human-vetted subset of 500 real-world GitHub issues (drawn from major Python repositories like Django and pandas) to ensure tasks are clear and solvable. To complement this, Terminal-Bench v2 tests end-to-end command-line competence by requiring agents to execute complex system administration, compilation, and environment-setup tasks inside sandboxed containers. Together, these benchmarks provide the most reliable, reproducible metrics for measuring how effectively modern LLMs can operate as autonomous software engineers.

https://github.com/princeton-nlp/SWE-bench
1 project · 1 city

Related technologies

Recent Talks & Demos

Showing 1-1 of 1

Members-Only

Sign in to see who built these projects