<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Vetto — Research &amp; Blog</title>
    <link>https://vetto.ai/companies/research-blog/</link>
    <description>Research and engineering notes from Vetto — benchmarks, evaluations, and learnings on how post-training data improves frontier AI models.</description>
    <language>en</language>
    <lastBuildDate>Tue, 18 Aug 2026 00:00:00 GMT</lastBuildDate>
    <atom:link href="https://vetto.ai/companies/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Plot Twist Bench</title>
      <link>https://vetto.ai/companies/plot-twist-bench/</link>
      <guid isPermaLink="true">https://vetto.ai/companies/plot-twist-bench/</guid>
      <pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate>
      <description>Plot Twist Bench asks whether frontier models watch a video or recognise it: 48 multiple-choice questions on 26 recordings filmed to order, each built so the usual thing is not what happened. Twelve models, no rubric and no judge in the grading, and the strongest averages 75% against a human 85%.</description>
      <category>Benchmarks</category>
    </item>
    <item>
      <title>ProgramBench Vetted</title>
      <link>https://vetto.ai/companies/programbench-vetted/</link>
      <guid isPermaLink="true">https://vetto.ai/companies/programbench-vetted/</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>ProgramBench Vetted remediates the original reverse-engineering benchmark with calibrated tasks, stronger verification gates, and evidence-backed analysis.</description>
      <category>Benchmarks</category>
    </item>
    <item>
      <title>GulliBench</title>
      <link>https://vetto.ai/companies/gullibench/</link>
      <guid isPermaLink="true">https://vetto.ai/companies/gullibench/</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>GulliBench measures AI gullibility: on tiny, easy tasks, a convenient stored field lies while the truth sits one derivation away in the same data. Does the model check, or grab the number it was handed? Gullibility is a behavioral trait, largely orthogonal to intelligence — and no model solves even half once two lies are chained.</description>
      <category>Benchmarks</category>
    </item>
    <item>
      <title>Terminal Tasks v1.0</title>
      <link>https://vetto.ai/companies/computer-anthology-terminal-tasks/</link>
      <guid isPermaLink="true">https://vetto.ai/companies/computer-anthology-terminal-tasks/</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>Introducing Computer Anthology</description>
      <category>Benchmarks</category>
    </item>
  </channel>
</rss>
