conv.

All stories
AIQuiet 2d · day 7

Vals AI, backed by a16z, pushes to become neutral referee for AI benchmarks

The startup keeps its tests secret and grades models on real industry tasks to stop firms from gaming benchmarks.

What to know

  • Vals AI argues legacy AI benchmarks are outdated and easily gamed because their test materials are public, letting companies train against the exam.
  • Vals instead keeps its test materials private and grades models on industry-specific real-world tasks in law, finance, and coding rather than abstract knowledge tests.
  • The startup raised a $40 million Series A led by Andreessen Horowitz last month, following an earlier seed round from 8VC and Bloomberg Beta.
  • Vals is expanding its evaluation scope into novel and high-stakes domains, including recursive self-improvement, biosecurity, cybersecurity, mental health, and law-of-armed-conflict compliance.

“Historically, I think evaluation has been done to evaluate intelligence in a very abstract way. Like, do models know enough information to be able to take a bar exam type test?”

Rayan Krishnan, Vals AI co-founder · TechCrunch ↗ · Sep 18

Rayan Krishnan 25-year-old co-founder of Vals AIVals AI AI benchmarking startup founded 2024Andreessen Horowitz Lead investor, Series A8VC and Bloomberg Beta Lead investors, seed round

Vals AI, backed by a16z, pushes to become neutral referee for AI benchmarks
techcrunch.com

How it unfolded 1 development · click the chart to see its coverage articlesposts

Peak 7 pieces in two hours at Sep 19, 8 AM; 10 pieces over 7 days (1 article · 9 posts) Sep 19, 8 AM — 7 pieces · 1 article · 6 posts — Mastodon 4, Bluesky 2, Newswires 1Sep 19, 10 AM — quietSep 19, 12 PM — quietSep 19, 2 PM — quietSep 19, 4 PM — quietSep 19, 6 PM — 1 piece · 1 post — Bluesky 1Sep 19, 8 PM — quietSep 19, 10 PM — quietSep 20, 12 AM — quietSep 20, 2 AM — quietSep 20, 4 AM — quietSep 20, 6 AM — quietSep 20, 8 AM — quietSep 20, 10 AM — quietSep 20, 12 PM — quietSep 20, 2 PM — quietSep 20, 4 PM — quietSep 20, 6 PM — quietSep 20, 8 PM — quietSep 20, 10 PM — quietSep 21, 12 AM — 1 piece · 1 post — Bluesky 1Sep 21, 2 AM — quietSep 21, 4 AM — quietSep 21, 6 AM — quietSep 21, 8 AM — quietSep 21, 10 AM — quietSep 21, 12 PM — quietSep 21, 2 PM — quietSep 21, 4 PM — quietSep 21, 6 PM — quietSep 21, 8 PM — quietSep 21, 10 PM — quietSep 22, 12 AM — quietSep 22, 2 AM — quietSep 22, 4 AM — quietSep 22, 6 AM — quietSep 22, 8 AM — quietSep 22, 10 AM — quietSep 22, 12 PM — quietSep 22, 2 PM — quietSep 22, 4 PM — quietSep 22, 6 PM — quietSep 22, 8 PM — quietSep 22, 10 PM — quietSep 23, 12 AM — quietSep 23, 2 AM — quietSep 23, 4 AM — quietSep 23, 6 AM — quietSep 23, 8 AM — quietSep 23, 10 AM — quietSep 23, 12 PM — quietSep 23, 2 PM — quietSep 23, 4 PM — quietSep 23, 6 PM — quietSep 23, 8 PM — quietSep 23, 10 PM — quietSep 24, 12 AM — quietSep 24, 2 AM — 1 piece · 1 post — Bluesky 1Sep 24, 4 AM — quietSep 24, 6 AM — quietSep 24, 8 AM — quietSep 24, 10 AM — quietSep 24, 12 PM — quietSep 24, 2 PM — quietSep 24, 4 PM — quietSep 24, 6 PM — quietSep 24, 8 PM — quietSep 24, 10 PM — quietYesterday, 12 AM — quietYesterday, 2 AM — quietYesterday, 4 AM — quietYesterday, 6 AM — quietYesterday, 8 AM — quietYesterday, 10 AM — quietYesterday, 12 PM — quietYesterday, 2 PM — quietYesterday, 4 PM — quietYesterday, 6 PM — quietYesterday, 8 PM — quietYesterday, 10 PM — quietToday, 12 AM — quietToday, 2 AM — quietToday, 4 AM — quietToday, 6 AM — quietToday, 8 AM — quietToday, 10 AM — quietToday, 12 PM — quiet 1
Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 2:52 PM ET
  1. 1

    TechCrunch profiles Vals's private, task-based benchmarking approach

    In an interview and office tour, Krishnan described Vals's method of testing models on undisclosed, industry-specific tasks in law, finance, and coding, and its expansion into recursive self-improvement, biosecurity, cybersecurity, mental health, and law-of-armed-conflict benchmarks.

    “We have a benchmark on recursive self improvement. We're doing some work in mental health, cybersecurity, biosecurity, and even law of armed conflict to models to understand how to apply the Geneva Convention…”
    — Rayan Krishnan
    1. first by TechCrunch, 7d ago

    • nexttool.bsky.social

      Vals just raised a 0M Series A led by a16z to become the gold standard for AI benchmarking. Private test sets, real-world task evals, federal-agency program. 8x revenue YoY.

      nexttool.bsky.socialBluesky5d ago1▲view on Bluesky ↗
    2 more of the top 3 · 6 posts in this stretch
    • TechCrunch@mstdn.social

      Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models. https:// techcrunch.com/2026/09/19/vals -backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking/?utm_source=dlvr.it&utm_medium=mastodon

      TechCrunch@mstdn.socialMastodon7d agoview on Mastodon ↗
    • mgtech.bsky.social

      🚨 BREAKING: Vals is a startup that aims to improve AI benchmarking systems. It raised $40 million in a series A funding round led by Andreessen Horowitz in September 2026.

      mgtech.bsky.socialBluesky6d agoview on Bluesky ↗
    all of them →
  2. background

    Vals raises $40M Series A led by Andreessen Horowitz — After a period of rapid growth, Vals closed a $40 million Series A round led by a16z, deepening its push into AI evaluation infrastructure.

  3. background

    Vals closes a seed round led by 8VC and Bloomberg Beta — The young company secured seed funding as it began establishing itself in the AI industry.

  4. background

    Rayan Krishnan co-founds Vals AI — Krishnan, then a Stanford undergraduate who had interned at Palantir and worked at Microsoft and Stanford's AI lab, started Vals after observing that academic benchmarks were not keeping pace with fast-advancing frontier models.

What people are saying 3 voices from 1 site · best of 6 · verbatim