11X cheaper than ChatGPT: Tiny 150M model just proved AI doesn’t need to “think out loud” to be smart
- A 150M model reached 29.5% while costing just $0.0007 per task
- ChatGPT scored higher, yet its comparable reasoning runs cost substantially more
- BDH-CQ performs reasoning internally instead of generating lengthy intermediate text
Pathway, an AI lab focused on building Post-Transformer architectures, has released new benchmark results for its BDH-CQ reasoning model.
According to the researchers, their 150M-parameter model scored 29.5% pass@2 on the public ARC-AGI-1 evaluation set.
It achieved this at a computed inference cost of $0.0007 per task, roughly eleven times cheaper than ChatGPT’s underlying GPT 5.6 Luna (Low) model.
Latest Videos FromTechRadar
A cheaper way to reason
Today, many AI tools waste computing power because of how they are designed, not because deep reasoning demands it.
“Today’s AI pays a steep token cost for reasoning, but that cost is imposed by architecture, not by any law of intelligence,” said Zuzanna Stamirowska, CEO and co-founder of Pathway.
“We show that a different architecture changes the game and opens up a whole new space in terms of how much intelligence per dollar.”
Amazon Web Services believes that BDH-CQ’s result is a promising step toward using advanced AI reasoning in real products more affordably.
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!
“Customers are increasingly exploring how to move advanced reasoning from experimentation into production, where performance, efficiency, and scalability all matter,” said Nicolas Tarducci of AWS.
ARC-AGI-1, a widely used reasoning benchmark for AI systems, checks whether a system can infer an underlying rule from limited examples and apply it correctly to new inputs.
In this test, OpenAI’s Luna model scored only slightly higher at 34.2%, yet running it still costs significantly more ($0.008 per task).
That price gap already includes OpenAI’s recent 80% price cut on Luna, which began on July 30th of this year.
Further up the chart, Claude Opus 5 and Gemini 3.1 Pro reach 97–98% but cost around $0.5 – $0.6 per task, meaning the frontier’s very top costs close to a thousand times more than BDH-CQ for the highest scores.
On the cheap end, Qwen3 235B costs over three times more than BDH-CQ while scoring worse than even its Low variant, so it isn’t a real competitor on either price or performance.
“Pathway shows that model architecture, not just scale, can drive the next leap in AI reasoning,” said Łukasz Kaiser, co-author of the original 2017 Transformer paper.
Why it costs so much less
The efficiency gap stems mainly from a structural difference in how each system actually performs reasoning during inference computations.
Many reasoning AI systems generate intermediate text, adding one token after another before producing their final answers.
The longer that written reasoning becomes, the more it costs to run and the slower the AI responds to each request.
BDH-CQ works quite differently, quietly solving problems inside its own memory instead of writing everything down first as visible text.
Pathway also confirmed that early experiments already follow standard Transformer-like scaling laws across model sizes from 1B to 600B parameters.
The company also plans to extend this approach toward harder benchmarks, including mathematical reasoning, ARC-AGI-2, and eventually full ARC-AGI-3 evaluations.
If these efficiency gains hold across larger and more difficult tasks, cost rather than raw capability could increasingly separate rival reasoning systems.

Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
A 150M model reached 29.5% while costing just $0.0007 per task ChatGPT scored higher, yet its comparable reasoning runs cost substantially more BDH-CQ performs reasoning internally instead of generating lengthy intermediate text Pathway, an AI lab focused on building Post-Transformer architectures, has released new benchmark results for its BDH-CQ reasoning…
Recent Posts
- ‘The best Shark vacuum cleaner overall’ is now down 48% — just in time for spring cleaning
- 11X cheaper than ChatGPT: Tiny 150M model just proved AI doesn’t need to “think out loud” to be smart
- Russia’s Microsoft Office challenger forced to shut down offices and fire workers as reports claim it is losing millions
- A rare two-player Computer Space arcade machine is up for auction
- Quordle hints and answers for Thursday, August 13 (game #1662)
Archives
- August 2026
- July 2026
- June 2026
- May 2026
- April 2026
- March 2026
- February 2026
- January 2026
- December 2025
- November 2025
- October 2025
- September 2025
- August 2025
- July 2025
- June 2025
- May 2025
- April 2025
- March 2025
- February 2025
- January 2025
- December 2024
- November 2024
- October 2024
- September 2024
- August 2024
- July 2024
- June 2024
- May 2024
- April 2024
- March 2024
- February 2024
- January 2024
- December 2023
- November 2023
- October 2023
- September 2023
- August 2023