Ai

Why the Price of an AI Token Doesn't Tell You What AI Actually Costs

Kasun Illankoon

By: Kasun Illankoon

5 min read

A new whitepaper from an Abu Dhabi infrastructure company argues the AI industry has been measuring cost wrong, and that fixing it means treating speed, spend, and data sovereignty as a single design problem.

[For more news, click here]

Every enterprise buying artificial intelligence today fixates on one number: the price of a token. It shows up on vendor pricing pages, in procurement spreadsheets, and in boardroom arguments over which provider deserves next year's technology budget. A new whitepaper from Core42, a G42 company built around sovereign cloud and AI infrastructure, argues that number is close to meaningless on its own.

Titled “Efficiency Wins the Inference Era,” the paper focuses on inference, the everyday act of a trained model answering a prompt after it has already been built. Every chatbot reply, every AI agent completing a multistep task, and every automated workflow humming quietly in the background of a company's operations runs on inference, and each of those interactions consumes tokens, the units a model uses to read a request and produce a response. As generative AI moves from pilot projects into daily operations, that consumption compounds quickly. Prompts get longer, context windows get denser, and agentic systems that once made a single model call now make several, most of them invisible to the person who typed the original request.

The Metric Executives Have Been Missing

Core42's argument is that the sticker price of a token obscures the real question a business needs answered: how much useful work a system delivers, at what cost, and how quickly. The whitepaper frames that as tokens per second per dollar, a single measure that ties together speed, output, and spend. A cheap token that arrives slowly, or that requires several follow-up calls to produce a usable answer, is not actually cheap. The shift the paper is pushing for is a move away from shopping for the lowest sticker price and toward measuring whether AI reliably produces business outcomes at the speed a given task demands.

One Chip Doesn't Fit Every Job

That shift carries a practical consequence: no single model or piece of hardware suits every workload. A real-time customer service agent needs low latency above almost everything else. A nightly batch job processing months of documents cares more about raw throughput than speed. Core42's answer to that mismatch is Compass, a platform that routes each workload to a specific combination of model, accelerator, and deployment path across a mix of silicon that includes NVIDIA, AMD, Qualcomm, and Cerebras hardware. On production numbers, the platform operates with a 99.5 percent availability commitment, serves more than seven million API requests, and processes upward of 100 billion tokens every week, according to Core42. Its fastest path, built on Cerebras chips, reports throughput gains of up to 20 times over standard configurations.

That heterogeneous approach echoes what the company has said about its philosophy before. When Core42 and Cerebras expanded global access to OpenAI's open-weight gpt-oss-120B model in 2025, delivering inference at roughly 3,000 tokens per second, Kiril Evtimov, who was serving as Core42's chief executive at the time and now leads technology strategy across G42 as Group Chief Technology Officer, said the deployment was “setting a new benchmark for performance, flexibility, and compliance in AI.” The same logic runs through the new whitepaper: speed and governance are framed not as competing priorities but as two outputs of the same underlying design.

Sovereignty Built Into the Architecture, Not Bolted On

That governance piece is where the paper gets more ambitious. Efficient inference and data sovereignty are usually treated as separate conversations, one about cost engineering, the other about compliance. Core42 argues the two have to be managed together, because the same infrastructure decisions that determine speed and cost also determine where data physically sits, who can reach it, and how its use is tracked. Compass builds in centralized visibility over usage and spend, budget and access controls, in-country data residency, encryption, zero logging of customer data, and what the company counts as more than 170 distinct security policies.

A Gulf Blueprint With a North American Audience

The timing matters beyond Abu Dhabi. North American enterprises have spent the past two years racing to deploy generative AI, and many are now running into the same sticker shock that appears to have prompted this paper: pilots that looked inexpensive on a per-token basis turned costly once they scaled into production, particularly as agentic systems multiplied calls behind the scenes. Core42's pitch, that efficiency and sovereignty are the same design problem rather than two separate ones, hands North American and European technology buyers a framework for that conversation, while giving Gulf governments that have made sovereign AI a strategic priority a working example to point to. Under Interim Chief Executive Talal M. Al Kaissi, who stepped into the role in December, Core42 has continued expanding that infrastructure across the United States, Europe, and the Middle East, positioning the company less as a regional cloud provider and more as a reference point for how enterprises anywhere might scale AI without losing control of either the bill or the data.

The bet underlying “Efficiency Wins the Inference Era” is straightforward: the organizations that treat AI economics and AI governance as one problem, rather than two, will be the ones still standing once the novelty of generative AI wears off and the accounting begins.

Related Articles:

Core42 Is Building the AI Compute Network That Doesn't Pick Sides Between AMD and NVIDIA

The AI Chip Squeeze Is Redrawing Enterprise Tech Budgets, and Speed Now Wins

Exclusive: How Edge AI Is Reshaping Data Sovereignty Across Saudi Arabia and the UAE

Share this article

Related Articles