Technology
harbor environment/evals
Harbor is an open-source framework for evaluating, optimizing, and training AI agents in sandboxed container environments.
Built by the creators of Terminal-Bench, Harbor standardizes how developers test and optimize AI agents (such as Claude Code and OpenHands) inside secure, isolated containers. The platform solves the infrastructure headache of agent evaluation by orchestrating parallel trials across cloud runtimes like Daytona and Modal. By decoupling tasks, environments, and agents into modular configurations, Harbor allows teams to run complex benchmarks (including SWE-Bench Verified and Terminal-Bench 2.0), generate high-quality trajectories for reinforcement learning, and implement robust CI/CD pipelines for agentic workflows.
Recent Talks & Demos
Showing 1-0 of 0