Harbor Framework Team
Harbor Framework Team develops Harbor, an open-source framework for evaluating and optimizing AI agents and language models in isolated container environments. It provides a common task, agent, sandbox, rollout, and verification interface so researchers and developers can run reproducible evaluations and large-scale training rollouts.
At a Glance
- AI model developers and frontier AI labs
- Agent developers and CLI-agent builders
- Benchmark and evaluation researchers
- Academic research groups
- +2 more
AI Tools by Harbor Framework Team
(2)Terminal-Bench-Science
AI Agent Science Benchmark
terminal-bench
AI Agent Terminal Benchmark
Discussions
No discussions yet
Be the first to start a discussion about Harbor Framework Team
Latest News
Harbor x LangChain: a unified stack for evaluating agents, connecting Harbor with LangGraph/Deep Agents, LangSmith Sandboxes, and LangSmith Observability.
Harbor Hub adds job-result sharing so users can share run results with team members or customers.
Snorkel’s Benchtalks interview details Harbor’s adoption, benchmark-factory vision, and expansion from coding into finance, law, and general automation tasks.
Laude Institute names Harbor among its Slingshots // Two projects and describes it as an agent evaluation framework for environment-based tasks.
Products & Services
Open-source evaluation and optimization harness for agents and language models. It defines environment-based tasks, installs arbitrary container-runnable agents inside sandboxes, runs trials and verifiers, records results and trajectories, and scales rollouts through local or cloud container providers.
A Harbor-compatible benchmark for evaluating AI agents on research workflows across life, physical, earth, mathematical, and engineering sciences. It is distributed under Apache License 2.0 and can be run through Harbor with Modal or Daytona.
A sharing and registry surface associated with Harbor for publishing datasets and sharing job results; the official Harbor news page specifically documents job-result sharing for team members and customers.
Market Position
Harbor positions itself as a simple, flexible, open-source execution and evaluation layer that abstracts container orchestration while retaining configurable tasks and verifiers. Its differentiation is a unified interface spanning arbitrary agents, benchmarks, isolated sandboxes, repeated trials, and training rollouts; adjacent alternatives and complements include bespoke evaluation harnesses, LangSmith, and benchmark-specific runners, with Harbor integrating directly with LangSmith rather than requiring a single agent framework.
Leadership
Founders
Alex Shaw
Co-creator of Terminal-Bench and Harbor; Founding Member of Technical Staff at Laude Institute. Previously worked at Google on ad recommendations and conversion modeling.
Mike Merrill
Co-creator of Terminal-Bench and Harbor; postdoctoral researcher at Stanford University working on agents, evaluations, and autonomy.
Andy Konwinski
Co-founder of Databricks and Perplexity, and a founder of Laude Institute; provided the research-to-artifact and early-user-feedback vision behind the Terminal-Bench/Harbor work.
Ludwig Schmidt
Stanford University researcher and collaborator on Terminal-Bench and Harbor; worked with Mike Merrill on promoting terminal-based computer use and helped connect the project to Laude Institute.
Executive Team
Alex Shaw
Co-creator and Harbor project lead
Founding MTS at Laude Institute; previously worked at Google on ad recommendations and conversion modeling.
Mike Merrill
Co-creator and Harbor project lead
Stanford postdoctoral researcher focused on agents, evaluations, and autonomy; co-created Terminal-Bench and Harbor with Alex Shaw.
Board of Directors
Founding Story
Harbor grew out of the Terminal-Bench team’s experience building and operating containerized agent evaluations. The team saw that evaluating in containers was slow, scaling to thousands of cloud environments was difficult, and the same infrastructure could support not only evaluation but also SFT, reinforcement learning, and prompt optimization; Harbor was started as an experimental package to make task definition and large-scale rollouts simple and reusable across benchmarks.
Business Model
Revenue Model
Harbor is distributed as open-source software under Apache License 2.0. The public materials describe a self-managed package and integrations with paid cloud sandbox and observability providers, but do not document a Harbor subscription or usage-based price.
Target Markets
- AI model developers and frontier AI labs
- Agent developers and CLI-agent builders
- Benchmark and evaluation researchers
- Academic research groups
- Companies testing agents in CI/CD and production-like workflows
- Researchers performing SFT, RL, prompt optimization, and data-generation experiments
- Benchmarking CLI and computer-use agents
- Evaluating model and agent changes during development
- Running large-scale agent rollouts for reinforcement learning and supervised fine-tuning
- Prompt optimization and automated agent experimentation
- CI/CD testing of agents against reproducible environment-based tasks
- Building and sharing custom benchmarks and datasets
- Stanford University
- AfterQuery
- Tensorlake
- Snorkel AI