EveryDev.ai
Subscribe
Home
Topics

209 topics

  • Trending
AI Topics
  • Agents2189
  • Coding1574
  • Infrastructure698
  • Marketing534
  • Projects498
  • Research456
  • Design416
  • Analytics389
  • Testing296
  • MCP290
  • Security286
  • Data262
  • Integration197
  • Prompts189
  • Communication183
  • Extensions173
  • Learning170
  • Voice151
  • Commerce135
  • DevOps123
  • Web86
  • Finance26
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Topics
    3. Testing
    4. LLM Evaluations

    AI Tools & Discussions in LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    LLM Evaluations Tools (115)

    View Oqoqo
    Oqoqo tool icon

    Oqoqo

    Featured

    AI Agent Evaluation Platform

    LLM EvaluationsAgent FrameworksMCP Tools
    View PromptWright
    PromptWright tool icon

    PromptWright

    Structured Prompt Engineering Platform

    Prompt EngineeringPrompt ManagementLLM Evaluations
    View Armature
    Armature tool icon

    Armature

    Featured

    MCP Server Analytics Platform

    LLM EvaluationsMCP ToolsObservability
    View BenchGen
    BenchGen tool icon

    BenchGen

    Featured

    AI Agent Simulation Platform

    LLM EvaluationsAutomated TestingAgent Frameworks
    View Scenario
    Scenario tool icon

    Scenario

    AI Agent Simulation Testing Framework

    Automated TestingAgent FrameworksLLM Evaluations
    View Prefactor
    Prefactor tool icon

    Prefactor

    Featured

    AI Agent Runtime Enforcement SDK

    LLM EvaluationsObservabilityCompliance & Gov
    View Supabase Evals
    Supabase Evals tool icon

    Supabase Evals

    Featured

    AI Agent Benchmark for Supabase

    LLM EvaluationsAgent HarnessAI Infrastructure
    View Next.js Evals
    Next.js Evals tool icon

    Next.js Evals

    Next.js AI Coding Benchmark

    LLM EvaluationsAI Coding Asst.Automated Testing
    View ContextQA
    ContextQA tool icon

    ContextQA

    AI Testing Platform for QA Teams

    Automated TestingAI InfrastructureLLM Evaluations
    View G0DM0D3
    G0DM0D3 tool icon

    G0DM0D3

    AI Red Teaming Chat Interface

    Conversational AILLM EvaluationsPrompt Engineering

    Top Tools in LLM Evaluations

    Highest trending score

    Design Arena

    A crowdsourced benchmark platform that pits top AI models against each other on design tasks and lets users vote to power live leaderboards.

    Arena (LMArena)

    A community-powered platform for evaluating and comparing frontier AI models through real-world human feedback, featuring a public leaderboard and battle mode.

    Artificial Analysis

    Independent AI benchmarking platform that evaluates and compares AI models across intelligence, speed, cost, and capabilities to help users choose the best model and provider for their use case.

    New in LLM Evaluations

    Oqoqo5m agoPromptWright5d agoArmature5d ago

    Featured Tool

    Design Arena screenshot
    Design Arena

    A crowdsourced benchmark platform that pits top AI models against each other on design tasks and lets users vote to power live leaderboards.

    Last 7 Days

    3
    New Tools
    43
    Featured
    28
    Upvotes

    Related Topics

    Automated Testing137 tools
    Bug Detection43 tools
    Test Generation19 tools
    Visual Testing12 tools
    Performance Testing2 tools

    LLM Evaluations Discussions

    No discussions yet

    Be the first to start a discussion about LLM Evaluations

    Weekly Newsletter

    One weekly email. New AI dev tools, news, and trends.

    No spam — unsubscribe anytime