EveryDev.ai
Subscribe
Home
Topics

209 topics

  • Trending
AI Topics
  • Agents2849
  • Coding2023
  • Infrastructure848
  • Projects607
  • Marketing604
  • Research532
  • Analytics472
  • Design469
  • MCP435
  • Testing352
  • Security327
  • Data308
  • Prompts228
  • Integration225
  • Communication213
  • Extensions197
  • Learning180
  • Voice178
  • Commerce161
  • DevOps137
  • Web95
  • Finance31
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Topics
    3. Testing
    4. LLM Evaluations

    AI Tools & Discussions in LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    LLM Evaluations Tools (122)

    View Spanda
    Spanda tool icon

    Spanda

    LLM Hallucination Detection Library

    LLM EvaluationsAI InfrastructureObservability
    View Revalvo
    Revalvo tool icon

    Revalvo

    Local Prompt Eval Workbench

    LLM EvaluationsPrompt EngineeringPrompt Management
    View Preseason
    Preseason tool icon

    Preseason

    AI Tool Recommendation Benchmark

    LLM EvaluationsTool DirectoriesVibe Coding
    View Terminal-Bench-Science
    Terminal-Bench-Science tool icon

    Terminal-Bench-Science

    AI Agent Science Benchmark

    LLM EvaluationsAcademic ResearchAgent Harness
    View Lenz
    Lenz tool icon

    Lenz

    Featured

    AI Hallucination Detection API

    AI InfrastructureLLM EvaluationsContent Analysis
    View LittleLearner
    LittleLearner tool icon

    LittleLearner

    LM Trained on K5 Curriculum

    Academic ResearchAI Dev LibrariesLLM Evaluations
    View inferock-bench
    inferock-bench tool icon

    inferock-bench

    Featured

    Local LLM API Cost Tracker

    LLM EvaluationsObservabilityAI Infrastructure
    View Oqoqo
    Oqoqo tool icon

    Oqoqo

    Featured

    AI Agent Evaluation Platform

    LLM EvaluationsAgent FrameworksMCP Tools
    View PromptWright
    PromptWright tool icon

    PromptWright

    Structured Prompt Engineering Platform

    Prompt EngineeringPrompt ManagementLLM Evaluations
    View Armature
    Armature tool icon

    Armature

    Featured

    MCP Server Analytics Platform

    LLM EvaluationsMCP ToolsObservability

    Top Tools in LLM Evaluations

    Highest trending score

    Design Arena

    A crowdsourced benchmark platform that pits top AI models against each other on design tasks and lets users vote to power live leaderboards.

    Arena (LMArena)

    A community-powered platform for evaluating and comparing frontier AI models through real-world human feedback, featuring a public leaderboard and battle mode.

    BridgeBench

    BridgeBench ranks AI coding models across UI generation, security, refactoring, hallucination, debugging, and speed benchmarks.

    New in LLM Evaluations

    Spanda11m agoRevalvo3d agoPreseason3d ago

    Featured Tool

    Design Arena screenshot
    Design Arena

    A crowdsourced benchmark platform that pits top AI models against each other on design tasks and lets users vote to power live leaderboards.

    Last 7 Days

    4
    New Tools
    45
    Featured
    28
    Upvotes

    Related Topics

    Automated Testing145 tools
    Bug Detection48 tools
    Test Generation19 tools
    Visual Testing15 tools
    Performance Testing3 tools

    LLM Evaluations Discussions

    Start a new discussion about LLM Evaluations
    Seth Zhai's avatar
    Seth Zhai
    September 13, 2026

    Spark KNE Verify 0.1.8: testing evidence applicability for AI claims

    Spark KNE Verify 0.1.8 is now live. This release improves how the extension handles claim context, current events, risk, review, freshness, and source applicability. The goal is to avoid treating generic authoritative links as evidence for a specific claim. You can try 10 verifications without signi…

    0

    Weekly Newsletter

    One weekly email. New AI dev tools, news, and trends.

    No spam — unsubscribe anytime