‘That is where the machine starts winning on cost’: Expert pits AMD Radeon AI PRO R9700 rig against ChatGPT and gives surprising verdict
- Two AMD cards cost $18,775 yet beat GPT-5.6 Sol within hours weekly
- Multi-Token Prediction nearly doubled throughput to 320.2 tokens every second
- Twenty million monthly tokens save a team $11,738 yearly against Sol pricing
A hardware reviewer compared a dual-GPU AMD workstation against cloud subscription pricing to determine which option delivers cheaper AI inference over time.
Two AMD Radeon AI PRO R9700 cards, each carrying 32 GB of memory, were installed inside a workstation costing roughly $18,775 as tested.
The evaluation measured electricity draw, token throughput, and amortized hardware cost, then set those figures against several cloud subscription tiers for comparison.
Latest Videos FromTechRadar
How cloud pricing sets the bar
Cloud AI providers charge customers per million tokens generated, with prices ranging from $1.20 for GPT-5.6 Luna up to $30 for GPT-5.6 Sol.
Mid-tier models sit in between, with Claude Sonnet 5 at $10 and Claude Opus 5 at $25 per million output tokens generated.
The more expensive the cloud model a team would otherwise use, the sooner owned hardware pays for itself.
Testing both AMD cards together, the workstation generated 156.2 tokens every second while serving eight simultaneous users during this test.
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!
A speed technique called Multi-Token Prediction nearly doubled that figure, pushing throughput up to 320.2 tokens every second, with identical output quality.
At that 320.2 token-per-second speed, the machine only needs 3.5 hours of weekly use to beat GPT-5.6 Sol on cost, 4.2 hours to beat Claude Opus 5, and 8.8 hours to beat Gemini 3.1 Pro.
A team generating above 20 million tokens per month against GPT-5.6 Sol pricing gains real savings using this owned hardware setup.
At that volume, running the workstation costs about $6,262 yearly in electricity and amortized hardware, against roughly $18,000 yearly in matching Sol fees.
That $11,738 yearly gap is the actual evidence behind the claim that heavy monthly usage makes AMD’s rig worthwhile.
If a company instead relies on GPT-5.6 Luna, priced at just $1.20 per million tokens, that math flips entirely in the other direction.
The workstation would then need 94.3 hours of weekly use just to match that far cheaper cloud subscription’s total cost.
Since a single week only contains 168 hours total, reaching that particular break-even point remains genuinely difficult without near constant, saturated usage.
Electricity itself was a minor factor throughout, since both cards together drew between 310 and 510 watts under sustained load conditions.
Where the economics tip in AMD’s favor
A smaller team producing only five million tokens monthly, priced against Gemini 3.1 Pro, would spend $6,262 yearly to displace just $720 in cloud costs, a clear loss.
That comparison shows the hardware only makes financial sense once usage climbs high enough to close a large yearly cost gap.
A single R9700 card handled an eight billion parameter AI model alone, processing 34.5 tokens every second without help from a second card.
Running just one card also lowers the effective break-even point further, since a lone card draws far less power under equivalent load.
That single card pulled between 221 and 283 watts depending on simultaneous user count, well under the 310 to 510 watts both cards drew together.
Smaller models therefore offer a second path into positive economics, letting lighter workloads justify a $1,880 single-card purchase long before a team can justify the full $3,760 dual-card upgrade.
Price alone still does not settle the question, since these locally run models measurably trail top cloud systems on complex reasoning benchmarks.
On one independent intelligence index, the 27 billion parameter AMD model scored 37 points against Google‘s Gemini 3.1 Pro scoring 46 points instead.
Therefore, a team chasing the cheaper token count may be trading away real reasoning quality, not just cloud subscription fees.
Via Puget Systems

Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
Two AMD cards cost $18,775 yet beat GPT-5.6 Sol within hours weekly Multi-Token Prediction nearly doubled throughput to 320.2 tokens every second Twenty million monthly tokens save a team $11,738 yearly against Sol pricing A hardware reviewer compared a dual-GPU AMD workstation against cloud subscription pricing to determine which option…
Recent Posts
- The NBN speed upgrades launched one year ago — here’s how they’ve changed the face of Australian broadband
- Lanterns episode 5 ending explained: who dies, why does John Stewart [spoiler], and more big questions answered about the DC comic book show’s latest entry
- Warlock: Dungeons & Dragons will have a smaller set of spells with multiple uses rather than a large variety to make ‘everything accessible on a controller,’ Invoke says
- Sony denies claims it is stopping disc production — but admits its last plant will drop production by 10%
- MagSafe vs. USB-C: Which is better for charging your MacBook?
Archives
- September 2026
- August 2026
- July 2026
- June 2026
- May 2026
- April 2026
- March 2026
- February 2026
- January 2026
- December 2025
- November 2025
- October 2025
- September 2025
- August 2025
- July 2025
- June 2025
- May 2025
- April 2025
- March 2025
- February 2025
- January 2025
- December 2024
- November 2024
- October 2024
- September 2024
- August 2024
- July 2024
- June 2024
- May 2024
- April 2024
- March 2024
- February 2024
- January 2024
- December 2023
- November 2023
- October 2023
- September 2023