OpenAI's chief financial officer, Sarah Friar, published a framework she calls a scorecard for the AI age, arguing that the right way to measure return on AI is not seats, licenses, or cost per token but what she terms useful intelligence per dollar: the amount of successfully completed work a system produces for a given spend. The piece proposes four questions for buyers to ask. How much useful work actually gets done. What is the cost per successful task, including retries and failures rather than the headline per-token price. How dependable is the system, meaning how often it gets the work right without supervision. And does each additional dollar of compute buy proportionally more completed work as usage scales.
What makes the essay more than a finance memo is the concrete evidence it marshals from OpenAI's newest models. Friar references GPT-5.6, which she says shipped the prior week in three tiers named Sol, the flagship, Terra, the balanced option, and Luna, the cheapest and fastest. She reports that GPT-5.6 Sol at maximum reasoning set a state-of-the-art result on the Artificial Analysis Coding Agent Index while using fifty-four percent fewer output tokens than the prior best, and that on the DeepSWE v1.1 benchmark it reached 72.7 percent, edging past Claude Fable 5's 69.9 percent while costing an estimated thirty-six percent less per task in API terms. The token-efficiency claim is the load-bearing one: a model that matches or beats rivals while emitting far fewer tokens directly lowers the cost per successful task, which is exactly the metric the scorecard is built around.
The framing lands the same week that NVIDIA published its own intelligence-per-dollar argument about post-training hardware, and it arrives against a backdrop of investor anxiety about whether AI spending is converting into returns. Read that way, the scorecard is as much a market-positioning document as a measurement proposal, tying OpenAI's efficiency numbers to a rubric that happens to favor efficient models. The benchmark deltas are self-reported and await independent replication, and cost-per-task comparisons depend heavily on task mix and reasoning settings. Still, the shift in vocabulary is notable: the frontier conversation is moving from raw capability toward the economics of getting verified work done.
- OpenAI frames the metric around completed work and cost per successful task rather than per-token price.
- The essay doubles as positioning for GPT-5.6's token-efficiency numbers, which are self-reported and unreplicated.
- Echoes NVIDIA's same-week 'intelligence per dollar' post-training argument, signaling an industry-wide pivot to cost accounting.