<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>futurehouse on tomrochette.com</title>
    <link>https://tomrochette.com/tags/futurehouse/</link>
    <description>Recent content in futurehouse on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Tue, 06 Oct 2026 04:48:21 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/futurehouse/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>FutureHouse Robin</title>
      <link>https://tomrochette.com/agents/automated-research/futurehouse-robin/</link>
      <pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/automated-research/futurehouse-robin/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>automated-research</category><category>futurehouse</category><category>biology</category><category>drug-discovery</category><category>multi-agent</category>
      <description>&lt;p&gt;Robin is FutureHouse&amp;rsquo;s open-source multi-agent system that automates the intellectual half of biological discovery, hypothesis generation, experimental strategy, and data analysis, around human-executed laboratory experiments.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Robin is this category&amp;rsquo;s wet-lab proof point, and its paper&amp;rsquo;s own ablations make the engineering argument: strip out the specialized agents and the base model hallucinates nearly half its references, so the harness, not the model, is what carries the loop.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A Python workflow, Apache-2.0 at github.com/Future-House/robin, that orchestrates three specialized agents: Crow and Falcon for concise and deep literature search (both built on PaperQA2) and Finch for data analysis, which runs eight parallel analysis trajectories and merges them into a consensus.&#xA;It was built by FutureHouse, a San Francisco 501(c)(3) nonprofit founded in 2023 by Sam Rodriques and Andrew White, funded philanthropically (the Schmidts, OpenPhilanthropy, the National Science Foundation, and others), with the for-profit Edison Scientific spinning out in 2025 to commercialize the underlying tools.&#xA;Given a disease target, Robin proposes disease mechanisms, ranks drug candidates in an LLM-judged tournament, proposes the assays, analyzes the raw flow-cytometry and RNA-seq data, and proposes the next round.&#xA;Humans write the protocols, run every physical experiment, and review the ranked candidates before testing: the paper is explicit that human researchers executed the experiments while the intellectual framework was AI-driven.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Published in Nature on May 19, 2026 (volume 655, pages 497 to 505, open access), first announced May 20, 2025, with concept to submission taking 2.5 months, as of 2026-10-06.&#xA;The demo: pointed at dry age-related macular degeneration, Robin proposed enhancing RPE phagocytosis, and its rounds of candidates identified ripasudil, an approved glaucoma drug never before proposed for the disease, plus KL001, a circadian modulator, both confirmed in vitro and revalidated in primary human RPE stem cells from a donor over 60.&#xA;Nature reports about 268,000 accesses and 117 citations in the five months since publication, as of 2026-10-06.&#xA;The repository carries the loop and example trajectories: 730 stars, 122 forks, Apache-2.0, last push 2026-04-21, as of 2026-10-06.&#xA;A paper-configured workflow run costs about US$11 of API calls (45 Crow and 30 Falcon calls), and Robin read 551 papers in 30 minutes against an estimated 294 human hours.&#xA;&lt;strong&gt;The developer-community footprint is nearly absent: the announcement drew an 18-point Hacker News thread with 4 comments as of 2026-10-06, so adoption runs through academia and Nature readers rather than the harness ecosystem.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The category&amp;rsquo;s only head-to-head against OpenAI Deep Research: given the same candidate-generation task, Deep Research produced 17 candidates, none were hits in the assay, and it never suggested ROCK inhibition.&lt;/li&gt;&#xA;&lt;li&gt;Published ablations: replacing Crow and Falcon with o4-mini leaves 44.5 percent of references hallucinated per assay proposal, evidence the specialized agents are necessary rather than decoration.&lt;/li&gt;&#xA;&lt;li&gt;The loop ran continuously across rounds, and every hypothesis, experiment choice, analysis, and main-text figure in the paper came from the system.&lt;/li&gt;&#xA;&lt;li&gt;Open source with example trajectories, so the workflow is copyable, unlike most lab science programs in this category.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The bench stayed human: physical experiments, protocol execution, and candidate review are all human work, and the 200-fold speedup estimate covers cognitive labor only, modeled from surveys rather than measured end to end.&lt;/li&gt;&#xA;&lt;li&gt;The result is cell-culture validation, with in vivo work explicitly still required; no therapy exists, and generality beyond one disease is asserted, not shown.&lt;/li&gt;&#xA;&lt;li&gt;Finch scored 22.8 percent on an expert panel of 170 BixBench questions (47.9 percent on statistics, 15.3 percent on multi-step bioinformatics), so the analysis agent needs tight task framing.&lt;/li&gt;&#xA;&lt;li&gt;The claims come from the nonprofit&amp;rsquo;s own peer-reviewed paper, and Edison&amp;rsquo;s Kosmos, the commercialized sibling, self-estimates 80 percent of its findings as accurate, a number nobody has checked from outside.&lt;/li&gt;&#xA;&lt;li&gt;Peer review covers the paper, not the pace narrative; the Asimov Press profile records the founders themselves saying it is too early to know how good these systems are.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Robin is free and open source under Apache-2.0; FutureHouse is a nonprofit and sells nothing.&#xA;Edison Scientific commercializes the underlying agents (the FutureHouse platform became Edison&amp;rsquo;s), with no public prices on the fetched pages.&#xA;The only concrete cost in the record is the paper&amp;rsquo;s: about US$11 of API calls per workflow run, plus human bench time.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/agon/&#34; &gt;Agon&lt;/a&gt;: both open-source research loops; Agon runs software experiments toward papers with producer-critic loops, Robin runs biology around people at the bench.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/anthropic-claude-math/&#34; &gt;Anthropic Claude mathematical research&lt;/a&gt;: Anthropic&amp;rsquo;s life-sciences agents found a CRISPR-like enzyme with about 950 agents; Robin is the peer-reviewed, open-code counterpart with the discovery published as a Nature paper.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/openai-deep-research/&#34; &gt;OpenAI Deep Research&lt;/a&gt;: the prose-only literature loop; Robin&amp;rsquo;s paper used it as a control group and it scored zero hits.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for engineers building agent workflows around domain experts: it is the best-documented public demonstration that a multi-agent loop can carry the cognitive half of discovery while people keep the bench.&lt;/strong&gt;&#xA;Not for anyone expecting autonomous laboratories; the loop plans and analyzes, and every experiment still needs hands.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-06 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/agon/&#34; &gt;Agon&lt;/a&gt; - the open-source research loop on the software side of the same idea&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/anthropic-claude-math/&#34; &gt;Anthropic Claude mathematical research&lt;/a&gt; - the rival lab loop whose life-sciences wing made the enzyme discovery&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/openai-deep-research/&#34; &gt;OpenAI Deep Research&lt;/a&gt; - the prose loop Robin&amp;rsquo;s paper benchmarks against&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/automated-research-feature-matrix/&#34; &gt;Automated Research Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.nature.com/articles/s41586-026-10652-y&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.nature.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.nature.com/articles/s41586-026-10652-y&lt;/a&gt; - the peer-reviewed paper: architecture, ripasudil and KL001 results, ablations, the Deep Research control, and the BixBench scores (fetched 200, 2026-10-06)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.futurehouse.org/research/demonstrating-end-to-end-scientific-discovery-with-robin-a-multi-agent-system&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.futurehouse.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.futurehouse.org/research/demonstrating-end-to-end-scientific-discovery-with-robin-a-multi-agent-system&lt;/a&gt; - the announcement: agent roster, the 2.5-month build, and the humans-execute-the-experiments disclosure (fetched 200, 2026-10-06)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/Future-House/robin&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/Future-House/robin&lt;/a&gt; - repository license, stars, forks, and last push date for the status line (fetched via GitHub API, 2026-10-06)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.futurehouse.org/about&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.futurehouse.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.futurehouse.org/about&lt;/a&gt; - org structure: nonprofit, philanthropic funding, the Edison Scientific spinout, and the Kosmos 80-percent self-estimate (fetched 200, 2026-10-06)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.asimov.press/p/futurehouse&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.asimov.press&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.asimov.press/p/futurehouse&lt;/a&gt; - the independent profile: founders on evaluation limits and on engineering, not AI, being the hard part (fetched 200, 2026-10-06)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/search?query=FutureHouse&amp;amp;tags=story&amp;amp;hitsPerPage=10&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/search?query=FutureHouse&amp;tags=story&amp;hitsPerPage=10&lt;/a&gt; - the thin developer-community footprint: an 18-point announcement thread and a 71-point team profile (fetched 200, 2026-10-06)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
