EveryDev.ai
Subscribe
Home
Tools

3,804+ AI tools

  • New
  • Trending
  • Featured
  • Compare
  • Arena
Categories
  • Agents2782
  • Coding1973
  • Infrastructure825
  • Projects603
  • Marketing598
  • Research520
  • Analytics468
  • Design462
  • MCP419
  • Testing346
  • Security323
  • Data305
  • Integration224
  • Prompts220
  • Communication210
  • Extensions196
  • Learning179
  • Voice175
  • Commerce160
  • DevOps135
  • Web95
  • Finance31
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. Terminal-Bench-Science
    Terminal-Bench-Science icon

    Terminal-Bench-Science

    LLM Evaluations

    An open-source benchmark for evaluating AI agents on expert-curated research workflows across life, physical, earth, mathematical, and engineering sciences.

    Visit Website

    At a Glance

    Pricing
    Open Source

    Fully free and open-source under Apache 2.0. Use, modify, and distribute freely.

    Engagement

    Available On

    Web
    CLI
    API

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    LLM EvaluationsAcademic ResearchAgent Harness

    Alternatives

    terminal-benchEnterpriseRAG-BenchProgramBench
    Developer
    Harbor Framework TeamSan Francisco, CAEst. 2025$100M raised

    Listed Sep 2026

    About Terminal-Bench-Science

    Terminal-Bench-Science is an open academic benchmark designed to measure the frontier of AI agent capabilities on challenging, expert-curated scientific research workflows. It is hosted by Stanford University and the Laude Institute, and built by the Terminal-Bench and Harbor Framework team. The project is licensed under Apache 2.0 and released on GitHub, with version v0.1.0 published in August 2026.

    What It Is

    Terminal-Bench-Science is a continuous, community-driven benchmark that evaluates AI agents on real scientific research tasks drawn from five broad domains: life sciences, physical sciences, earth sciences, mathematical sciences, and engineering sciences. Tasks are authored and reviewed by domain experts, and the benchmark evolves alongside frontier AI to create a feedback loop between scientific needs and AI development. It is part of the Terminal-Bench franchise and uses the Harbor Framework as its execution and evaluation infrastructure.

    Task Coverage and Structure

    The benchmark currently contains 70 expert-curated tasks and is growing toward 100+. Each task is designed to be genuinely challenging for frontier AI agents and must produce outcomes that can be objectively verified in a terminal environment. The contribution flow follows a three-stage process:

    • Propose: Submit a task idea via a structured proposal form for feedback and approval.
    • Build: Implement the task and open a pull request following the contributing guide.
    • Review: Tasks undergo automated checks (static validation, Docker build, oracle/nop validation, agent trials, cheat trials), parallel domain and technical review, and final bar-raiser approval before merge.

    A public task dashboard tracks every proposal, pull request, and review status alongside domain coverage.

    Leaderboard and Evaluation

    The benchmark publishes a live leaderboard ranking AI models and agents by resolution rate. As of the latest data, Anthropic's Opus 5 running via Claude Code leads with a 30.0% resolution rate, followed by GPT-5.6 Sol via Codex at 22.4%. The leaderboard also tracks token usage and cost per run, providing a Pareto view of performance versus efficiency. Evaluations are run using the Harbor CLI tool against sandboxed environments such as Modal or Daytona.

    Open-Source Architecture and Running the Benchmark

    The benchmark is built on the Harbor Framework, an open-source agent harness. Users install Harbor via uv tool install or pip install and run tasks against the published dataset on Harbor Hub. The oracle solution can be run 5x to confirm task correctness in a given sandboxing environment. Agent and model selection is handled via CLI flags (--agent, --model), and concurrent runs are supported with --n-concurrent.

    Community and Research Partners

    Terminal-Bench-Science is a large-scale scientific community effort. The project lists contributors from institutions including Stanford University, MIT, Caltech, Princeton, Oxford, ETH Zurich, and many others. Research partners include the Stanford AI Lab (SAIL), Stanford HAI, the Allen Institute for AI, and the NSF AI Institute for Foundations of Machine Learning (IFML). Industry sponsors listed by the project include Anthropic, Google, Modal, Moonshot AI, and others. The community coordinates via a dedicated #tb-science Discord channel and weekly project calendar meetings.

    Update: v0.1.0 Release

    Version v0.1.0 was published on August 26, 2026, with a concept DOI (10.5281/zenodo.22110253) that always resolves to the latest release. The project is actively accepting pull requests for v0.2, with the README noting that early contribution is encouraged due to multi-round review requirements. The GitHub repository shows active development with the last push on August 29, 2026.

    Terminal-Bench-Science - 1

    Community Discussions

    Be the first to start a conversation about Terminal-Bench-Science

    Share your experience with Terminal-Bench-Science, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source

    Fully free and open-source under Apache 2.0. Use, modify, and distribute freely.

    • Full access to all benchmark tasks
    • Live leaderboard access
    • Harbor Framework CLI integration
    • Community Discord access
    • Contribution workflow (Propose → Build → Review)

    Capabilities

    Key Features

    • Expert-curated scientific research tasks across 5 domains
    • Live leaderboard with resolution rate, token usage, and cost tracking
    • Continuous benchmark that evolves alongside frontier AI
    • Automated task review pipeline (static checks, Docker build, oracle/nop, agent trials, cheat trials)
    • Community contribution workflow (Propose → Build → Review)
    • Public task dashboard for tracking proposals and reviews
    • Harbor Framework integration for sandboxed agent execution
    • Support for multiple AI agents and models via CLI flags
    • DOI-versioned releases on Zenodo
    • Apache 2.0 open-source license

    Integrations

    Harbor Framework
    Modal
    Daytona
    Claude Code
    OpenAI Codex
    Anthropic API
    OpenAI API
    Gemini API
    Harbor Hub
    Zenodo
    Airtable (task proposal form)
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate Terminal-Bench-Science and help others make informed decisions.

    Developer

    Harbor Framework Team

    The Harbor Framework Team builds open-source infrastructure for evaluating and optimizing AI agents and language models. Their flagship products are Terminal-Bench, a benchmark suite for terminal-use agents, and Harbor, the evaluation harness that powers it. The project is described as a Stanford × Laude collaboration, with contributors including researchers and engineers from those institutions. Harbor supports parallel cloud-based evaluations and RL environment generation, making it a general-purpose platform for agent developers.

    Founded 2025
    San Francisco, CA
    $100M raised
    20 employees

    Used by

    Stanford University
    AfterQuery
    Tensorlake
    Snorkel AI
    Read more about Harbor Framework Team
    WebsiteGitHubX / Twitter
    2 tools in directory

    Similar Tools

    terminal-bench icon

    terminal-bench

    Terminal-Bench is an open-source benchmark suite for evaluating AI agents' ability to complete complex tasks in terminal environments, built on the Harbor framework.

    EnterpriseRAG-Bench icon

    EnterpriseRAG-Bench

    An open-source benchmark dataset of 500,000+ enterprise documents and 500 questions for evaluating RAG systems on realistic company internal data.

    ProgramBench icon

    ProgramBench

    A benchmark that tests whether AI agents can rebuild real-world programs from scratch given only a compiled binary and its documentation, with no access to source code.

    Browse all tools

    Related Topics

    LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    121 tools

    Academic Research

    AI tools designed specifically for academic and scientific research.

    62 tools

    Agent Harness

    Infrastructure, orchestrators, and task runners that wrap around LLM coding agents — covering session management, context delivery, worktree isolation, architecture enforcement, and issue-to-PR pipelines.

    161 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions