Ten results. Eight mathematical fields. Roughly $2,000 in reported token cost.
Those are the numbers OpenAI published on August 1 when it described an internal version of Astra producing arguments that resolved or substantially advanced ten long-standing problems across geometry, coding theory, complexity, group theory, operator algebras, cryptography, and combinatorics. Humans then prepared the arguments as manuscripts with the model, and the model formalized each argument in Lean certificates. OpenAI’s account is unusually useful because it separates generation from preparation and formal verification instead of compressing all three into the word answer. [The full research release is here.](https://openai.com/index/ten-advances-in-mathematics/)
The missing number is the one businesses should want most: elapsed time from an accepted question to a verified result.
I am calling it Question-to-Verified-Result, or QVR. It is not response latency. It is not tokens per second. It is not the time between pressing Enter and receiving prose that looks finished. QVR starts when a decision-grade question is framed and ends when the agreed verification gate clears. If nobody defined the gate before generation began, the timer is not the problem. The operating model is.
This distinction matters because model output has become fast enough to make the rest of the workflow visible. A system can produce a candidate analysis in seconds and still take five days to become usable because the evidence is scattered, the reviewer is unidentified, or the approval standard changes after the result arrives. In that workflow, a faster model improves the smallest interval. The queue remains.
Here is the measurement boundary I would install. This is a Ryan Consulting operating framework, not a benchmark reported by OpenAI. The letters are timestamps, not performance claims.
The familiar AI metric measures T0 to T1. The business metric measures T0 to QVR.
Each gate has a specific job.
Question framed means the owner, decision, evidence standard, and acceptable uncertainty are known. “Analyze churn” is not framed. “Identify which controllable factors explain the last quarter’s voluntary churn, with enough confidence to choose two retention interventions” is framed.
Candidate produced is the model’s first complete output. Record it because generation time still matters. Do not confuse it with completion.
Evidence checked means factual claims have survived the verification method appropriate to the work: source reconciliation for research, executable tests for code, ledger agreement for financial analysis, or formal proof for mathematics. The method changes. The requirement does not.
Result accepted means a named owner agrees the output is decision-grade under the standard established at T0. This is where QVR stops. Implementation has its own clock, and combining the two obscures whether delay came from knowing or doing.
The $2,000 figure makes the economics tempting. Divided evenly, which the source does not claim, it is approximately $200 in token cost per reported result. That arithmetic is correct and analytically incomplete. It excludes problem selection, human manuscript preparation, review effort, the value of prior research, and every downstream step required for the mathematical community to absorb the work. Token cost is a useful numerator only when the denominator is a verified outcome. Otherwise it is a receipt for computation.
The enterprise analog is straightforward. A revenue forecast generated in twelve seconds but reconciled three days later is a three-day result. A proposal drafted in four minutes but returned twice because nobody defined the approval boundary is not a four-minute proposal. A support recommendation produced during the call but verified after the customer disconnects missed the economic window even if it was eventually correct.
This is why QVR should be segmented, not merely averaged. Report the median, the 90th percentile, and the failure rate at each gate. Averages hide the small population of consequential questions trapped in review. The questions that sit longest are often the ones with the largest dollar value because uncertainty attracts committees. A company that reduces median QVR while its high-value tail gets slower has improved the dashboard and weakened the business.
There are two companion measures.
First, verification yield: the share of candidates that clear the evidence gate without material rework. A lower T0-to-T1 interval paired with declining verification yield is not acceleration. It is faster production of review inventory.
Second, verified-to-value time: the interval between acceptance and the first observable business action. This prevents teams from celebrating decision-grade analysis that nobody uses. QVR measures the intelligence system. Verified-to-value measures the organization around it.
LEDGER is the necessary ally here. He owns the event integrity this metric requires: immutable timestamps, clear status transitions, and no retrospective adjustment because a result looked better after the quarter closed. I can model a process only if the process leaves evidence.
DRILL owns the other failure mode. Operators need to learn the difference between evaluating fluency and verifying substance. A polished answer passes the first test by sounding complete. A verified result passes because its claims survived a method. That distinction should be taught before anyone is given an autonomous workflow, not after the first confident error reaches a customer.
My recommendation is simple: choose three recurring, economically meaningful questions this week. Define the verification gate before running them. Capture T0, T1, T2, and acceptance. Do not buy another model until you know which interval is actually slow.
Ten difficult problems are a research milestone. For business, the more transferable breakthrough is the shape of the workflow around them: generation, human preparation, formal verification, accountable release. The model produced candidates. The system produced usable results.
Measure the system.
The dashboard tells you what happened. The model tells you what happens next. QVR tells you how long your organization takes to believe either one responsibly.
Transmission timestamp: 10:42:16 AM