TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Articles/Five Startups Turn Transformer Compute Costs Into an Architecture Race
Articles

Five Startups Turn Transformer Compute Costs Into an Architecture Race

Five AI startups are betting that sparse attention, recurrent retention, hybrid neural networks, diffusion, and brain-inspired design can outrun the transformer's ballooning compute bill.

August 10, 2026 7 Min Read
45

MIT Technology Review’s What’s Next series set out to answer a specific question: is the transformer, the architecture behind every major large language model since Google researchers introduced it in a 2017 paper titled “Attention Is All You Need,” starting to run into a wall it cannot afford to keep hitting? Its answer comes in the form of five startups, each betting on a different fix. Subquadratic wants to make the transformer’s attention mechanism sparser. Manifest AI wants to replace attention with a recurrent mechanism called power retention. Liquid AI wants to blend transformers with liquid neural networks. Inception wants to keep the transformer but abandon word-by-word generation for diffusion. Pathway wants an architecture that traces its lineage more to neuroscience than to the original transformer paper. What unites them is the same cost problem, and the same bet that whoever solves it first gets to define what a large language model looks like for the next decade.

Table Of Content

  • Why Attention Got So Expensive
  • Five Different Bets on What Comes After Attention
  • Subquadratic Bets on Sparse Attention, and Draws Skeptics
  • Manifest AI Trades Attention for Retention
  • Liquid AI Keeps a Fifth of the Transformer and Replaces the Rest
  • Inception Keeps the Transformer but Skips Sequential Generation
  • Pathway Builds Something Closer to a Brain Than a Transformer
  • The Problem Every One of These Startups Shares

Why Attention Got So Expensive

The transformer’s defining feature, the attention mechanism, is also its biggest liability at scale. Attention lets a model weigh every token in its context against every other token, which is what makes transformers so good at capturing long-range relationships in text. It is also what makes them expensive: the computation required grows quadratically as context length grows, meaning a document twice as long does not cost twice as much to process: it costs roughly four times as much.

That math now shows up directly on balance sheets. OpenAI expects to spend $50 billion on computing in 2026, a figure cofounder and president Greg Brockman disclosed while testifying in the company’s court battle with Elon Musk, Bloomberg reported. The International Energy Agency projects global data center electricity consumption will more than double by 2030, from 415 terawatt-hours in 2024 to roughly 945 terawatt-hours, with AI-optimized data centers as the single biggest driver of that growth. Every major AI lab is absorbing some version of that cost curve, which is why a wave of startups now sees an opening to sell a cheaper alternative to an architecture almost everyone assumed was permanent.

Five Different Bets on What Comes After Attention

Subquadratic Bets on Sparse Attention, and Draws Skeptics

Miami-based Subquadratic, led by CEO Justin Dangel, raised $29 million in seed funding and launched its SubQ model with a 12 million token context window, roughly the length of 120 books. Rather than discarding attention outright, the company built what it calls Subquadratic Selective Attention: a content-dependent routing system that computes exact attention only against the tokens it judges relevant, instead of every token in the context. SiliconANGLE reported that Subquadratic’s own benchmarks show SubQ running more than 50 times faster and 50 times cheaper than leading frontier models at 1 million tokens of context. “The entire AI industry is built on transformers,” Dangel told MIT Technology Review. “They are one of the most important innovations in the history of computer science.”

The claims have not gone unchallenged. In an earlier MIT Technology Review report on Subquadratic, AI engineer Dan McAteer summed up the skepticism bluntly: “SubQ is either the biggest breakthrough since the Transformer, or it’s AI Theranos.” Former OpenAI researcher Will Depue was more measured but no less pointed, comparing the claimed breakthrough to “running a four-minute mile” and concluding that “the public evidence does not yet justify the stronger claim that they have solved the quadratic attention bottleneck.” Part of the friction is access: SubQ has a long waitlist, and independent researchers have had little opportunity to reproduce the company’s numbers. Subquadratic also built SubQ by adapting pretrained weights from Alibaba’s open-source Qwen models rather than training from scratch, a detail that shortens the path to a working model but complicates claims of a from-the-ground-up architectural reinvention.

Manifest AI Trades Attention for Retention

Manifest AI, cofounded by CEO Jacob Buckman and CTO Carles Gelada, replaced attention entirely in its Brumby-14B-Base model with a mechanism it calls power retention. Instead of comparing every token against every other token, power retention behaves like a true recurrent network: it keeps a fixed-size internal state that updates as new tokens arrive, so cost stops scaling with context length at all. Manifest’s own writeup describes the goal as building a model that identifies which parts of its past matter and retains only those, rather than one that reprocesses every detail of everything it has ever seen.

The company demonstrated the approach by retraining, not pretraining, an existing model: it took Qwen3-14B-Base’s weights as a starting point and converted them to power retention, matching Qwen3’s original training loss after just 3,000 steps. The total cost was $4,000, run over 60 hours on 32 Nvidia H100 GPUs, against Manifest’s own estimate of roughly $200,000 to train a comparably sized model from scratch. Gelada frames the payoff as unlocking workloads no transformer can handle affordably today: hours-long video analysis and long-running autonomous agents that need to hold far more context than current context windows allow.

Liquid AI Keeps a Fifth of the Transformer and Replaces the Rest

Cambridge, Massachusetts-based Liquid AI, an MIT spinout led by cofounder and CEO Ramin Hasani, took a hybrid approach instead of a clean break. Hasani is also a machine learning scientist at MIT’s Computer Science and Artificial Intelligence Laboratory, and Liquid AI’s foundation models combine transformers with liquid neural networks, a technique inspired by the compact nervous systems of roundworms, in roughly a 20-to-80 split favoring the liquid component. Hasani told MIT Technology Review the company’s models have been downloaded almost 34 million times, and that they can match the performance of rival models four times their size, including versions of Alibaba’s Qwen and Google’s Gemma, while running on hardware as modest as a $50 Raspberry Pi. “Your brain is an AGI system,” Hasani said, framing the pitch around efficiency rather than raw scale. “It operates with 20 watts of power.” Liquid AI offers its models free to organizations under $10 million in annual revenue, a distribution strategy that helps explain the download numbers even before the architecture’s efficiency claims are independently stress-tested at scale.

Inception Keeps the Transformer but Skips Sequential Generation

Palo Alto-based Inception, cofounded by CEO Stefano Ermon along with Aditya Grover and Volodymyr Kuleshov, took a narrower approach than the other four: it kept the transformer’s underlying architecture and instead replaced its word-by-word generation process with diffusion, the same class of technique that powers image generators like Stable Diffusion. “You’re still using a big transformer model, but you can predict many tokens at the same time,” Ermon told MIT Technology Review. Instead of predicting one token at a time, Inception’s Mercury models start with a rough draft of an entire response and refine it in parallel across multiple passes; Ermon has called standard autoregressive generation “fancy autocomplete” by comparison. Inception’s own published figures claim Mercury delivers higher quality than Anthropic’s Claude 4.5 Haiku at one-fifth the latency and less than one-quarter the price, pricing input at $0.25 per million tokens and output at $1.00 per million tokens. The company raised $50 million in funding led by Menlo Ventures, with individual backing from Andrew Ng and Andrej Karpathy. Ermon’s framing of the endgame is unambiguous: “Ultimately, the currency is going to be intelligence per dollar.”

Pathway Builds Something Closer to a Brain Than a Transformer

Pathway, led by cofounder and CEO Zuzanna Stamirowska, took the most conceptually distinct approach of the five. Its architecture, called BDH or Dragon Hatchling, drops attention entirely in favor of a design its own researchers describe as bridging deep learning and neuroscience. The model relies on a large internal “latent reasoning space” combined with sparse activations, with roughly 5 percent of the network’s neurons active at any given moment, far closer to biological brains than the dense activations inside a standard transformer. Pathway tested BDH against roughly 250,000 Sudoku Extreme puzzles, a benchmark that requires holding many candidate solutions in mind at once rather than generating a single left-to-right answer. BDH solved 97.4 percent of them. Leading reasoning models it compared against, including OpenAI’s o3-mini, DeepSeek’s R1, and Anthropic’s Claude 3.7 at an 8K context window, scored close to zero, according to Pathway’s published results. Pathway attributes the gap to what happens inside a transformer during generation. “Each decision gets locked in as text is generated,” the company argues, so transformers “cannot hold multiple candidate strategies in parallel” the way BDH can.

The Problem Every One of These Startups Shares

None of these five approaches has been independently reproduced at the scale its backers claim, and that is the pattern skeptics keep pointing to across all of them, not just Subquadratic. Waitlists, cherry-picked benchmarks, and single-company evaluations are common to early-stage AI startups generally, and extraordinary efficiency claims against transformers, an architecture with roughly a decade of production hardening and tooling behind it, invite an unusually high bar of proof. Two of the five, Subquadratic and Manifest AI, sidestepped part of that burden by converting existing open-weight Qwen models rather than training their new architectures from scratch, a legitimate engineering shortcut that also means their headline results are partly inherited from the very transformer-based models they are trying to replace.

What is not in dispute is the incentive driving all of it. As long as OpenAI is spending $50 billion a year on compute and the IEA is projecting data center power demand to double before the end of the decade, there is a commercial opening for anyone who can cut that bill, whether the winning idea turns out to be sparser attention, a recurrent replacement for it, a hybrid with brain-inspired components, diffusion, or something that looks nothing like a transformer at all.

Tags:

AIai-model-architectureai-startupsLLM BenchmarksLLM Infrastructure

Share

An open, empty slot with tangled cabling inside a server rack, exposed between two mounted appliances
Previous Post

CISA Orders Federal Agencies to Patch a Critical Kemp LoadMaster Flaw Under Active Attack

A Royal Navy autonomous surface vehicle, fitted with a mast-mounted camera and sensor array, underway during a 2020 trial exercise in Norway
Next Post

Royal Navy Strips Internet Access From K3 Scout Drone Cameras Caught Signaling China

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Blue-lit server racks in a modern data center, illustrating the compute infrastructure behind the AI boom.
Articles

The AI Boom Is Spending Real Money Before Proving Real Returns

June 7, 2026
Technician working with a laptop beside server racks, representing enterprise AI retrieval infrastructure
Articles

Google’s Agentic RAG Push Makes Enterprise AI Less of a One-Shot Guess

June 7, 2026
A person with a laptop and smartphone, representing digital attention and AI-assisted work
Articles

AI Chatbots Are Making Attention a Design Problem

June 7, 2026
A customer-support representative wearing a headset against a dark studio background.
Articles

The Meta AI Support Hack Was a Plain Old Authorization Failure

June 7, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026