EveryDev.ai
Subscribe
Home
Developers

3,472+ AI companies

  • Radar
  • Trending
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Developers
    3. Harbor Framework Team

    Harbor Framework Team

    Harbor Framework Team develops Harbor, an open-source framework for evaluating and optimizing AI agents and language models in isolated container environments. It provides a common task, agent, sandbox, rollout, and verification interface so researchers and developers can run reproducible evaluations and large-scale training rollouts.

    Visit Website

    At a Glance

    2Tools Listed
    3Products
    39Tool Views
    9Capabilities
    Discussions
    San Francisco, CAHeadquarters
    2025Est.
    20Employees
    $100MRaised
    Focus Areas
    LLM Evaluations
    Autonomous Systems
    Agent Frameworks
    Academic Research
    Connect
    Latest News
    Harbor x LangChain: a unified stack for evaluating agents, connecting Harbor with LangGraph/Deep Agents, LangSmith Sandboxes, and LangSmith Observability.Jun 30, 2026
    Harbor Hub adds job-result sharing so users can share run results with team members or customers.May 27, 2026
    Markets
    • AI model developers and frontier AI labs
    • Agent developers and CLI-agent builders
    • Benchmark and evaluation researchers
    • Academic research groups
    • +2 more

    AI Tools by Harbor Framework Team

    (2)
    View Terminal-Bench-Science
    Terminal-Bench-Science tool icon

    Terminal-Bench-Science

    AI Agent Science Benchmark

    LLM EvaluationsAcademic ResearchAgent Harness
    View terminal-bench
    terminal-bench tool icon

    terminal-bench

    AI Agent Terminal Benchmark

    LLM EvaluationsAutonomous SystemsAgent Frameworks

    Discussions

    No discussions yet

    Be the first to start a discussion about Harbor Framework Team

    Latest News

    06/30/2026

    Harbor x LangChain: a unified stack for evaluating agents, connecting Harbor with LangGraph/Deep Agents, LangSmith Sandboxes, and LangSmith Observability.

    langchain.com
    05/27/2026

    Harbor Hub adds job-result sharing so users can share run results with team members or customers.

    harborframework.com
    03/31/2026

    Snorkel’s Benchtalks interview details Harbor’s adoption, benchmark-factory vision, and expansion from coding into finance, law, and general automation tasks.

    snorkel.ai
    02/26/2026

    Laude Institute names Harbor among its Slingshots // Two projects and describes it as an agent evaluation framework for environment-based tasks.

    laude.org

    Products & Services

    3
    Harbor
    November 7, 2025

    Open-source evaluation and optimization harness for agents and language models. It defines environment-based tasks, installs arbitrary container-runnable agents inside sandboxes, runs trials and verifiers, records results and trajectories, and scales rollouts through local or cloud container providers.

    Terminal-Bench-Science
    2026

    A Harbor-compatible benchmark for evaluating AI agents on research workflows across life, physical, earth, mathematical, and engineering sciences. It is distributed under Apache License 2.0 and can be run through Harbor with Modal or Daytona.

    Harbor Hub
    2026

    A sharing and registry surface associated with Harbor for publishing datasets and sharing job results; the official Harbor news page specifically documents job-result sharing for team members and customers.

    Market Position

    Harbor positions itself as a simple, flexible, open-source execution and evaluation layer that abstracts container orchestration while retaining configurable tasks and verifiers. Its differentiation is a unified interface spanning arbitrary agents, benchmarks, isolated sandboxes, repeated trials, and training rollouts; adjacent alternatives and complements include bespoke evaluation harnesses, LangSmith, and benchmark-specific runners, with Harbor integrating directly with LangSmith rather than requiring a single agent framework.

    Leadership

    Founders

    AS

    Alex Shaw

    Co-creator of Terminal-Bench and Harbor; Founding Member of Technical Staff at Laude Institute. Previously worked at Google on ad recommendations and conversion modeling.

    MM

    Mike Merrill

    Co-creator of Terminal-Bench and Harbor; postdoctoral researcher at Stanford University working on agents, evaluations, and autonomy.

    AK

    Andy Konwinski

    Co-founder of Databricks and Perplexity, and a founder of Laude Institute; provided the research-to-artifact and early-user-feedback vision behind the Terminal-Bench/Harbor work.

    LS

    Ludwig Schmidt

    Stanford University researcher and collaborator on Terminal-Bench and Harbor; worked with Mike Merrill on promoting terminal-based computer use and helped connect the project to Laude Institute.

    Executive Team

    AS

    Alex Shaw

    Co-creator and Harbor project lead

    Founding MTS at Laude Institute; previously worked at Google on ad recommendations and conversion modeling.

    MM

    Mike Merrill

    Co-creator and Harbor project lead

    Stanford postdoctoral researcher focused on agents, evaluations, and autonomy; co-created Terminal-Bench and Harbor with Alex Shaw.

    Board of Directors

    AK
    Andy Konwinski
    Founder/Primary Advisor

    Founding Story

    Harbor grew out of the Terminal-Bench team’s experience building and operating containerized agent evaluations. The team saw that evaluating in containers was slow, scaling to thousands of cloud environments was difficult, and the same infrastructure could support not only evaluation but also SFT, reinforcement learning, and prompt optimization; Harbor was started as an experimental package to make task definition and large-scale rollouts simple and reusable across benchmarks.

    Business Model

    Revenue Model

    Harbor is distributed as open-source software under Apache License 2.0. The public materials describe a self-managed package and integrations with paid cloud sandbox and observability providers, but do not document a Harbor subscription or usage-based price.

    Harbor Framework Team is presented as an open-source research/software project rather than a publicly traded company; no IPO plan is reported.

    Target Markets

    Industries & Segments
    • AI model developers and frontier AI labs
    • Agent developers and CLI-agent builders
    • Benchmark and evaluation researchers
    • Academic research groups
    • Companies testing agents in CI/CD and production-like workflows
    • Researchers performing SFT, RL, prompt optimization, and data-generation experiments
    Use Cases
    • Benchmarking CLI and computer-use agents
    • Evaluating model and agent changes during development
    • Running large-scale agent rollouts for reinforcement learning and supervised fine-tuning
    • Prompt optimization and automated agent experimentation
    • CI/CD testing of agents against reproducible environment-based tasks
    • Building and sharing custom benchmarks and datasets
    Notable Customers
    • Stanford University
    • AfterQuery
    • Tensorlake
    • Snorkel AI

    Quick Facts

    Headquarters
    San Francisco, CA
    Founded
    2025
    Entity Type
    Non-profit
    Employees
    20
    Total Funding
    $100M (Laude Institute Endowment)
    Investors
    Andy Konwinski, Hugging Face (Grant provider)
    Office Locations
    San Francisco

    Funding History

    Founding Grant$100,000,000
    2024
    N/A valuation
    Andy Konwinski

    History & Milestones

    January 2026

    Harbor’s software citation was published with Harbor Framework Team as author and Harbor described as a framework for evaluating and optimizing agents and models in container environments.

    February 26, 2026

    Laude Institute announced Harbor as one of its Slingshots // Two research projects, describing it as an agent-evaluation framework for environment-based tasks.

    May 27, 2026

    Harbor Hub introduced job-result sharing, allowing users to share results from a run with team members or customers without manually zipping results.

    June 30, 2026

    LangChain announced a unified Harbor/LangGraph/LangSmith stack, including LangSmith sandboxes and observability integrations for isolated, parallel agent evaluations.

    May 2025

    Terminal-Bench 1.0 launched, establishing the benchmark and the containerized task format from which Harbor evolved.

    Key Capabilities

    9
    Sandboxed, reproducible execution of agent trials in containers
    Task format based on an instruction, environment, tests/verifier, and optional solution
    Support for arbitrary agents that can be installed and run in a container, including Claude Code, OpenHands, Codex CLI, Aider, Gemini CLI, Cursor, and others
    Parallel execution across thousands of environments through pluggable cloud sandbox providers
    Evaluation of models and agents with repeated attempts, deterministic verifiers, rewards, and job-level metrics
    Rollout and trajectory generation for supervised fine-tuning and reinforcement learning

    Integrations & Partnerships

    Platform Integrations

    • Local Docker environments
    • Daytona
    • Modal
    • E2B
    • Runloop
    • Tensorlake
    • LangSmith Sandboxes and LangSmith Observability
    • Blaxel

    Key Partnerships

    Laude Institute and Stanford University lead the broader Terminal-Bench/Harbor work.
    Snorkel AI collaborated on Terminal-Bench tasks and evaluation-data analysis.
    LangChain/LangSmith integrated Harbor with LangGraph/Deep Agents, LangSmith Sandboxes, and LangSmith Observability.

    Connect

    Website
    harborframework.com/
    GitHub
    harbor-framework
    X / Twitter
    jeffwpli
    Discord
    6xWPKhGDbA

    AI Topics

    5

    Harbor Framework Team focuses on these topics:

    LLM Evaluations(2)
    Autonomous Systems(1)
    Agent Frameworks(1)
    Academic Research(1)
    Agent Harness(1)
    Back to all developersSuggest an edit