EveryDev.ai
Subscribe
Home
Developers

3,795+ AI companies

  • Radar
  • Trending
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Developers
    3. Preference Model

    Preference Model

    Preference Model is a superintelligence data research company building robust reinforcement-learning environments for training capable, better-aligned AI systems. It focuses on reward functions, secure harnesses and graders, and environments for AI research and ML engineering.

    Visit Website

    At a Glance

    1Tool Listed
    2Products
    10Capabilities
    Discussions
    San Francisco, CaliforniaHeadquarters
    14Employees
    $16MRaised
    Focus Areas
    AI Development Libraries
    Agent Frameworks
    LLM Evaluations
    Connect
    Latest News
    Published 'Building RL Environments for Superintelligence,' outlining the company's approach to robust environments for increasingly capable models.Oct 8, 2026
    Emerged from stealth and announced a $16 million seed round led by a16z.Oct 7, 2026
    Markets
    • Frontier AI research labs
    • AI model developers and post-training teams
    • ML research and engineering organizations
    • Teams building agent benchmarks and evaluation environments

    AI Tools by Preference Model

    (1)
    View Karotte
    Karotte tool icon

    Karotte

    Open Source RL Environment Framework

    AI Dev LibrariesAgent FrameworksLLM Evaluations

    Discussions

    No discussions yet

    Be the first to start a discussion about Preference Model

    Latest News

    10/08/2026

    Published 'Building RL Environments for Superintelligence,' outlining the company's approach to robust environments for increasingly capable models.

    preferencemodel.com
    10/07/2026

    Emerged from stealth and announced a $16 million seed round led by a16z.

    wsgr.com
    10/07/2026

    Open-sourced Karotte, its framework for building robust RL environments.

    preferencemodel.com
    10/07/2026

    Andreessen Horowitz published its investment announcement describing Preference Model's focus on AI research and ML-engineering environments.

    a16z.com

    Products & Services

    2
    Karotte
    October 7, 2026

    An open-source framework for building robust reinforcement-learning environments. Developers write tasks in Python; Karotte runs agents in sandboxes, provides tools and graders, scores outcomes, and records transcripts. It uses secure defaults including unprivileged execution, resource limits, firewalling, process cleanup, and protected grading artifacts.

    Custom RL environments and training data

    Research and development of RL environments and associated training tasks for frontier AI labs, especially environments for ML research, software engineering, GPU kernels, data curation, post-training, and computer-use-style work.

    Market Position

    Preference Model positions itself as a specialist in robust, adversarially tested RL environments for AI research and ML engineering, rather than static labeling datasets. Its differentiation is security-minded infrastructure designed for models that actively optimize against graders, with defenses hardened through more than one million evaluation runs. Relevant alternatives include labs' in-house environment teams and RL/evaluation frameworks such as Harbor, HUD, AgentEnv, verifiers, Inspect, Habitat, DeepTune, Fleet, Vmax, Turing, Mechanize, and Bespoke.

    Leadership

    Founders

    JZ

    Jennifer Zhou

    Co-founder and CEO. Previously worked on Anthropic's data team, building data infrastructure, tokenizers, and datasets, and earlier worked at Stripe.

    NC

    Ning Cao

    Co-founder, leading strategy, business development, and recruiting. Previously an early employee at DatologyAI, where he helped build the company from 0 to 1.

    Executive Team

    JZ

    Jennifer Zhou

    Co-founder and Chief Executive Officer

    Former Anthropic data-team member who worked on data infrastructure, tokenizers, and Claude pretraining datasets; previously worked at Stripe.

    NC

    Ning Cao

    Co-founder; Strategy, Business Development, and Recruiting

    Early DatologyAI employee who helped build that company from 0 to 1.

    Founding Story

    The founders started Preference Model around the view that alignment and capability depend heavily on the quality of the reward signals used in training. After seeing data and training infrastructure up close at Anthropic and DatologyAI, they set out to build robust, secure RL environments that prevent reward hacking and help frontier models learn useful behavior rather than loopholes.

    Business Model

    Revenue Model

    The company builds and sells custom RL environments and training data to frontier AI labs; its Karotte framework is open source and functions as a public framework around the company's environment-development work.

    Target Markets

    Industries & Segments
    • Frontier AI research labs
    • AI model developers and post-training teams
    • ML research and engineering organizations
    • Teams building agent benchmarks and evaluation environments
    Use Cases
    • Training and evaluating AI models on ML research and engineering tasks
    • Writing and optimizing GPU/CUDA kernels
    • Debugging training runs and designing experiments
    • Post-training and reasoning-token efficiency
    • Data curation and labeling
    • Software engineering and coding environments

    Quick Facts

    Headquarters
    San Francisco, California, United States
    Employees
    14
    Total Funding
    $16 million
    Investors
    Andreessen Horowitz (a16z), SignalFire
    Office Locations
    San Francisco

    Funding History

    Seed$16 million
    October 7, 2026
    Andreessen Horowitz (a16z)

    History & Milestones

    October 7, 2026

    Preference Model emerged from stealth and announced a $16 million seed financing led by Andreessen Horowitz (a16z), with SignalFire, South Park Commons, Scale Angel Group, and prominent researchers participating.

    October 7, 2026

    The company open-sourced Karotte, its framework for building robust RL environments with secure defaults and defenses against reward hacks.

    October 8, 2026

    Preference Model published its perspective on building RL environments for the superintelligence era, emphasizing deeper tasks, outcome-based grading, and robust reward signals.

    2025-2026

    The team built RL environments for several frontier AI labs over the course of a year and hardened them through more than one million internal evaluation runs and controlled red-teaming.

    Key Capabilities

    10
    Sandboxed, unprivileged agent execution
    Resource limits for memory, processes, disk, and runtime
    Firewalling and checks that prevent sandbox internet access
    Persistent bash and other agent tools with timeouts and output limits
    Protected root-only data and expected answers
    Process cleanup and safe submission collection before grading

    Integrations & Partnerships

    Platform Integrations

    • Python task authoring
    • VM runtimes using Apple container on macOS or Firecracker on Linux
    • Docker and Podman runtime support
    • LiteLLM-backed inference endpoints
    • Installation through uv/uvx
    • GitHub source repository and karotte.dev documentation

    Key Partnerships

    Andreessen Horowitz partnered with Preference Model and led its seed financing.
    Wilson Sonsini Goodrich & Rosati advised Preference Model on its seed financing.
    The company has built RL environments for several leading/frontier AI labs, though the labs were not named publicly.

    Connect

    Website
    preferencemodel.com
    GitHub
    preferencemodel

    AI Topics

    3

    Preference Model focuses on these topics:

    AI Development Libraries(1)
    Agent Frameworks(1)
    LLM Evaluations(1)
    Back to all developersSuggest an edit