Technology
Harbor for containerized agent environments
Harbor is an open-source framework for running, evaluating, and optimizing AI agents inside secure, containerized sandbox environments.
Built by the creators of Terminal-Bench, Harbor provides a unified harness to evaluate and optimize AI agents across thousands of isolated sandboxes. The framework ships with out-of-the-box support for popular agents (including Claude Code, OpenHands, and Codex CLI) and standard benchmarks like SWE-Bench and Terminal-Bench-2.0. By decoupling the agent from its execution layer, Harbor allows developers to run massive parallel evaluations locally via Docker or scale horizontally using cloud providers like Daytona, Modal, and E2B. It is the go-to infrastructure for generating clean rollout data, optimizing prompts, and training agents through reinforcement learning.
Related technologies
Recent Talks & Demos
Showing 1-1 of 1