- A 150M model reached 29.5% while costing honest $0.0007 per job
- ChatGPT scored increased, yet its comparable reasoning runs mark considerably more
- BDH-CQ performs reasoning internally as a replace of generating lengthy intermediate textual snort material
Pathway, an AI lab centered on constructing Post-Transformer architectures, has released new benchmark outcomes for its BDH-CQ reasoning model.
Constant with the researchers, their 150M-parameter model scored 29.5% pass@2 on the final public ARC-AGI-1 analysis space.
It performed this at a computed inference mark of $0.0007 per job, roughly eleven cases more cost-effective than ChatGPT’s underlying GPT 5.6 Luna (Low) model.
A inexpensive technique to motive
On the composed time, many AI tools raze computing energy resulting from how they’re designed, now not because deep reasoning requires it.
“Today’s AI pays a steep token cost for reasoning, but that cost is imposed by architecture, not by any law of intelligence,” stated Zuzanna Stamirowska, CEO and co-founder of Pathway.
“We gift that a particular structure adjustments the sport and opens up a total new space by process of how significant intelligence per greenback.”
Amazon Web Companies and products believes that BDH-CQ’s outcome is a promising step against the exhaust of advanced AI reasoning in trusty products more affordably.
Sign as much as the TechRadar Pro e-newsletter to ranking the total top news, thought, parts and steering your on-line commercial desires to prevail!
“Customers are increasingly exploring how to move advanced reasoning from experimentation into production, where performance, efficiency, and scalability all matter,” stated Nicolas Tarducci of AWS.
ARC-AGI-1, a widely used reasoning benchmark for AI systems, checks whether or now not a machine can infer an underlying rule from diminutive examples and apply it accurately to new inputs.
On this test, OpenAI’s Luna model scored easiest rather of increased at 34.2%, yet operating it detached charges deal more ($0.008 per job).
That mark gap already entails OpenAI’s most in vogue 80% mark minimize on Luna, which began on July 30th of this one year.
Extra up the chart, Claude Opus 5 and Gemini 3.1 Pro reach 97–98% nonetheless mark around $0.5 – $0.6 per job, that suggests the frontier’s very top charges discontinuance to a thousand cases more than BDH-CQ for the highest scores.
On a worth range discontinue, Qwen3 235B charges over three cases more than BDH-CQ while scoring worse than even its Low variant, so it is now not undoubtedly a trusty competitor on both mark or efficiency.
“Pathway shows that model architecture, not just scale, can drive the next leap in AI reasoning,” stated Łukasz Kaiser, co-creator of the long-established 2017 Transformer paper.
Why it charges so significant much less
The effectivity gap stems primarily from a structural incompatibility in how every machine genuinely performs reasoning for the length of inference computations.
Many reasoning AI systems generate intermediate textual snort material, at the side of one token after one more sooner than producing their closing solutions.
The longer that written reasoning turns into, the more it charges to bustle and the slower the AI responds to every request.
BDH-CQ works slightly differently, quietly solving complications internal its possess memory as a replace of writing the total lot down first as visible textual snort material.
Pathway additionally confirmed that early experiments already observe recurring Transformer-enjoy scaling legal guidelines all over model sizes from 1B to 600B parameters.
The firm additionally plans to enhance this methodology against tougher benchmarks, at the side of mathematical reasoning, ARC-AGI-2, and at final fleshy ARC-AGI-3 opinions.
If these effectivity features assign all over larger and more complex initiatives, mark in space of raw ability would possibly perhaps perhaps an increasing number of separate rival reasoning systems.

Apply TechRadar on Google News and add us as a most smartly-liked source to ranking our expert news, opinions, and thought in your feeds.


