Inherent’s Faraday Model Outperforms Claude Opus 4.8 and GPT-5.5 at Replicating Science
London AI lab Inherent says its 27-billion-parameter Faraday agent beat Claude Opus 4.8 and GPT-5.5 at reproducing published research results, using reinforcement learning instead of raw model scale.
London AI lab Inherent says its AI agent, Faraday, has outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at a specific scientific task: independently reproducing results from published research papers without being told the answer in advance. Inherent detailed the work in a research blog post and paper published August 14, and cofounder and chief scientist Edward Hughes discussed the result with TechCrunch this week.
Table Of Content
A much smaller model wins on a scientist’s task
What makes the result notable is the size gap. Faraday is built by post-training a 27-billion-parameter model, Qwen3.6-27B, using a turn-level credit variant of the reinforcement learning algorithm GRPO. Claude Opus 4.8 and GPT-5.5 are both far larger, general-purpose frontier systems: in the accompanying paper, Inherent’s researchers cite an external estimate that puts GPT-5.5 Codex, the specific system Faraday is measured against, at roughly 5 trillion parameters.
Hughes told TechCrunch that beating the larger systems wasn’t really the point. “What was most interesting to us about this was not so much the result of beating those frontier agents (which of course we liked) but was actually the way we went about building this,” he said.
How Replica tests for “research taste”
To train and evaluate Faraday, Inherent built a task suite it calls Replica: 310 replication tasks drawn from 100 machine learning and AI-for-science papers published between 1990 and 2026. Each task strips a results figure out of a paper and asks an agent to reproduce it under a limited time and compute budget, without ever seeing the original plot. The papers cover a range of fields, including natural language processing, materials science, weather forecasting, meta-learning and structural biology.
Scoring the result is harder than it sounds. A faithful replication needs more than a matching chart; it also has to reflect strong experimental design and effective use of available resources, a quality Inherent calls “research taste.” The company built an automated, per-task rubric judge to grade this, validated against a human study, after finding that scoring with a raw LLM judge alone produced too noisy a reward signal for reinforcement learning to work well.
Run against baselines that Inherent tested inside the Claude Code and Codex harnesses at extra-high thinking effort, Faraday beat both Claude Opus 4.8 and GPT-5.5 on 73 percent of in-distribution replication tasks and 60 percent of held-out, out-of-distribution tasks, according to the paper. Faraday’s own base model, before reinforcement learning, was the weakest agent tested, and degraded fastest on more recent papers.
Directing a much larger model, not replacing it
Faraday doesn’t write or run its own code. Instead, it calls OpenAI’s GPT-5.5 Codex as a tool, the same way Inherent says human scientists lean on existing coding assistants rather than build their own from scratch. According to the paper, Faraday was trained while directing a smaller model, GPT-5.4-mini, and generalized at test time to directing the larger GPT-5.5 Codex without additional training, suggesting its research-taste layer can work with whichever coding agent is available, including ones many times its own size.
These are Inherent’s own, self-reported numbers on a benchmark it designed and controls; they have not been independently reproduced by a third party.
A quiet startup with a bigger goal
Inherent emerged from stealth in late May with a $50 million seed round led by Index Ventures, with participation from Radical Ventures. It was founded by Google DeepMind alumni Edward Hughes, Tantum Collins and Louis Kirsch, along with Kaloyan Aleksiev, previously of Reka AI and Microsoft. Former UK government AI tsar Matt Clifford, a cofounder of Entrepreneurs First, sits on the company as an adviser.
The team, about a dozen people, works in person out of King’s Cross in London, and plans to grow to about 20 to 25 people by the end of the year. Hughes has also used his public platform to argue against “garden leave,” the UK practice of barring departing employees from joining or starting a rival company for months after they resign, a restriction that American researchers generally do not face. “This is a personal view rather than a company view, but I was affected by the garden leave problem,” he told TechCrunch.
Inherent’s longer-term goal is more ambitious than replication: building AI that can discover new scientific results, not just verify old ones. The company frames paper replication as a training ground for that goal, the same way many PhD students start their own research careers by trying to reproduce existing results before running original experiments of their own.








No Comment! Be the first one.