The $2,000 Superintelligence, and the Crack It’s Pointing At
A reasoning engine that runs frontier intelligence on a CPU for two grand a month would end the GPU arms race overnight...if it existed.
A reasoning engine that runs frontier intelligence on a CPU for two grand a month would end the GPU arms race overnight.
If it existed.
Last week somebody forwarded me a write-up of ARIA — “Artificial Recursive Intelligence Architecture,” a 238-million-edge reasoning core that allegedly runs on commodity CPUs at roughly $2,000 a month, hallucinates at under 0.0001%, and routes attention across 4,751 cognitive agents mapped to the wiring diagram of a fruit fly’s brain. The piece called it “the superintelligence Silicon Valley is desperately trying to ignore.” It claimed 1,200+ production deployments across banking, aerospace, and Fortune 500 environments.
I wanted it to be real. I’ve spent the better part of a year arguing that the industry is overpaying for intelligence, and here was a document arguing the same thing, louder. So I did the thing I do before anything goes out under my name: I went looking for the primary source.
Here is what I found. And here is the crack the ARIA story is pointing at, which is real, even though ARIA almost certainly is not.
Disclosure up front, because it matters for what follows: I build and operate sovereign AI compute infrastructure. I have a direct commercial interest in where the “cheaper intelligence” argument lands. Read me with that in mind, and check my links.
What the source actually is
The entire factual basis for ARIA is a single self-deposited preprint on Zenodo — record 20590584, credited to a Nicola Vettorato of “Counsellor Service Holding LTD.” Zenodo is a legitimate, CERN-operated open repository. It is also a place where anyone can upload anything. There is no peer review. There is no gatekeeper. The “citable source” the article congratulates itself on having is a document that cites only itself.
So I searched for independent corroboration. Anyone replicating the benchmark. A repository. A named client. A single hostile engineer poking at the <0.0001% hallucination claim. There are at least six unrelated projects called “ARIA” in AI research right now — a data-analysis framework out of Tsinghua, a math auto-formalization agent, an augmented-reality museum tool. The $2,000-superintelligence ARIA appears in exactly two places on the entire internet: its own Zenodo file, and the SEO article written to promote it. That article, by the way, ships with a “Meta Title” and “Meta Description” baked into the top of the document. It was built to rank, not to inform.
Then there’s the architecture itself. The claim is that you can map an “Economic Attention Network” onto the Drosophila melanogaster connectome and get verifiable reasoning out the other end. The fly connectome is real, beautiful science — FlyWire mapped roughly 140,000 neurons and published it in Nature in 2024. Bolting a commercial reasoning engine onto that name is the oldest move in the book: borrow the credibility of a verified result to launder an unverifiable one. The “No comment” when a reporter asked Vettorato to name one of the 1,200 deployments isn’t intrigue. When you ask a company for one verifiable customer and the answer is silence, the silence is the finding.
A hallucination rate of “structurally impossible.” A trillion-dollar problem solved by one man and a fruit fly. No code, no benchmark, no client. I’ve seen this pattern in maybe thirty pitch decks over twenty years. It has a name, and the name is not “superintelligence.”
So: ARIA goes in the bin. But don’t throw out the question it’s exploiting, because the question is the most important one in this industry.
The crack is real, and it isn’t in the model layer
ARIA is selling a true anxiety. The cost of intelligence at the model layer is collapsing, and everyone who built a business on the opposite assumption is quietly terrified.
The numbers are not in dispute, and they don’t come from a Zenodo PDF. Stanford’s 2025 AI Index found that the inference cost of a GPT-3.5-level system fell more than 280-fold between November 2022 and October 2024. Open-weight models closed the gap to closed models from 8% to 1.7% on some benchmarks in a single year. NVIDIA’s own analysis cites the same Index: hardware cost per unit down about 30% a year, energy efficiency up about 40% a year. Independent trackers put comparable-capability inference at roughly a 10x annual decline — from about $0.06 per thousand tokens in early 2025 to about $0.006 by mid-2026.
The model layer is commoditizing from below. That part of ARIA’s pitch is correct.
What ARIA gets wrong — what the entire “clever architecture beats the GPU monopoly” genre gets wrong — is where the escape hatch is. They think the moat you break is the model. It isn’t. The model is the thing dissolving on its own. A 13-billion-parameter model now hits 95% of GPT-3’s MMLU score. DeepSeek undercut incumbent pricing by roughly 90% and the sky did not fall on NVIDIA; Blackwell allocation is still constrained. Cheaper serving doesn’t shrink compute demand. It expands the set of products worth building, which pulls more compute, not less. That’s the part the CPU-miracle crowd never models.
The thing that stays scarce when intelligence gets cheap is not a better algorithm. It’s a megawatt you can actually plug into, and an operator who can turn it into revenue before the hardware goes obsolete.
Where the real bottleneck lives
Here is the number that should be on the wall of every AI infrastructure investor, and it has nothing to do with a fruit fly.
As of the end of 2024, roughly 2,290 gigawatts of generation and storage capacity sat stuck in U.S. interconnection queues — close to twice the entire installed U.S. generating fleet — per Lawrence Berkeley National Laboratory’s Queued Up report. The median project reaching commercial operation in 2024 spent about 55 months in the queue. Historically, only around 13% of the capacity that entered the queue from 2000 to 2019 had been built by the end of 2024. Most of it withdraws.
Now hold that against the clock the technology runs on. GPU performance roughly doubles every 18 to 24 months. A traditional data center takes 48 to 72 months to build. Grid interconnection averages five years. By the time the centralized megacampus is energized, the silicon it was poured for is two generations stale. PJM’s capacity auction told the same story in price: clearing rates jumped from $28.92/MW-day in 2024–25 to $329.17/MW-day for 2026–27. That is not a market signal. That is a market screaming.
This is the crack. Not the model. The interconnect. The mismatch between how fast compute demand compounds and how slowly grid-dependent infrastructure gets built is structural, and it does not get solved by building the old thing faster, or by a CPU graph engine nobody can reproduce.
It gets solved by moving the compute to where the power already is.
The answer has scars, not a preprint
This is where I disclose my hand fully, because the rest is what I actually do for a living.
Moving a megawatt-hour roughly a thousand miles costs about $41.50. Moving the data costs essentially nothing. So you stop dragging energy to compute and you drag compute to energy: retired power-plant campuses, industrial brownfields, stranded generation, substations with spare headroom. You manufacture the intelligence factory in a shop, truck it to the power, and run a live workload in roughly eight months instead of five years.
Our first unit is one ~1 MW modular intelligence factory: 392 GPUs, around 450 kW of IT load, closed-loop liquid cooling at a design PUE at or below 1.12, co-located at a power-plant site in Europe. Under contract, in construction, go live targeted for Q4/2026. One ~$45M asset. Four people on the operating side.
That last number is the part the ARIA story and the hyperscaler story both miss from opposite ends. ARIA fantasizes that you can make the compute almost free. The real lever is making the operator almost free — software agents carrying reconciliation, reporting, capacity pricing, network operations, commissioning paperwork, while humans hold the four things machines must never hold: money movement, signatures, client relationships, physical witness. Every machine action lands in an approval queue. Nothing gets sent, signed, or wired without a person. The agent that can’t beat 20% of a loaded human’s cost within ninety days gets retired. That’s not a slogan; it’s a kill rule we run against ourselves.
I’ll own the obvious objection: this is a hard thing to do, and most of who tries it will get the human-interlock part wrong and either move too slow or hand a machine a checkbook. The discipline is the product. The infrastructure is buildable. The operating model is where companies will actually die.
The tell that separates the two stories
ARIA published a hallucination rate and hid the engine. We publish the operating metrics and hide nothing that matters: fleet revenue per orchestrator, agents per human, human-touch minutes per closed deal, cost per deliverable against the human alternative. When a number and a derivation disagree, the derivation wins. When we miss, the miss gets more words than the win.
That’s the difference between a claim and a build. One asks you to trust a fruit fly. The other shows you the queue data, the auction prints, the contract date, and says: check it yourself.
The cheaper intelligence gets, the more the advantage migrates away from whoever owns the cleverest model and toward whoever owns the megawatt and can operate it without a thousand people. ARIA is right that the old assumption is dying. It’s just selling you the wrong funeral.
The scarce thing in 2027 won’t be intelligence. It’ll be a trustworthy operator standing next to a power plant.
_______________________________________________________________________
Sources: LBNL Queued Up (2025 edition, data through 2024); Stanford HAI 2025 AI Index; NVIDIA inference economics; FlyWire connectome; the ARIA Zenodo record — judge it for yourself.




