Skip to main content

Open Models Are Catching Up Faster—But Not on the Same Economic Clock

Markets value an AI model lead by the period it can earn excess returns before a capable alternative narrows the gap.

Sep 10, 2026
AICapital MarketsMarkets
Two exposed mechanical drive units on parallel rails illustrate AI models operating on different economic clocks.

Markets do not value an AI business only on what its model can do today. They value the period for which it can earn excess returns before a capable alternative narrows the gap. That duration is the economic clock behind a model lead. BenchLM's current scores and the Qwen3.8 Max release date make the duration question more concrete.

A benchmark lead is not yet a moat

BenchLM's 4 September snapshot puts Claude Fable 5.1 at 82.95 and GPT-6 Astra at 81.05. Qwen3.8 Max, an open-weight model, scores 72.43.

ModelLicenseScoreCurrent context
Claude Fable 5.1Closed82.95Leaderboard leader
GPT-6 AstraClosed81.05Second-ranked model
Qwen3.8 MaxOpen weight72.4310.52 below Fable 5.1
Claude Opus 4.8Closed72.300.13 below Qwen3.8 Max

The current frontier gap remains material on this composite measure. The same snapshot places Qwen3.8 Max 0.13 points above Claude Opus 4.8's 72.30.

Qwen3.8 Max was released 67 days, or 2.2 months, after Claude Opus 4.8. This is a release-gap proxy, not the measured date at which Qwen first passed Opus. One pair cannot establish an industry-wide catch-up rate.

A model lead earns a premium only for as long as customers cannot substitute it.

A shorter lead changes the valuation question

A model provider can earn excess returns when higher capability supports a price premium, reduces customer churn, or wins a large share of new workloads. Each assumption depends on the lead lasting long enough to justify the spending required to create it.

The Qwen–Opus comparison does not prove that those returns disappear. It lowers confidence in forecasting a long period of differentiated model pricing without evidence from contracts, renewals, usage, and gross margin.

A business priced on several years of excess margin is more exposed to faster substitution than one priced on near-term cash flow already under contract.

Earnings do not move together across the AI stack

For a model API provider, lower-cost substitutes can pressure price per completed task and make retention more valuable than a benchmark lead.

For an enterprise software company, cheaper capable models can lower the cost of serving customers. Whether that lifts margins depends on whether the company retains the saving or passes it through to defend its own pricing.

For cloud and inference providers, lower model prices can make more workloads economic and raise demand for compute. The earnings benefit still depends on utilization, incremental capex, power availability, and the return earned on the asset base.

Lombard Odier makes a related point: lower prices can expand demand while preserving value for cloud capacity, global networks, and proprietary data.

Two tests for AI exposure

First, test the duration of the advantage. Track the gap on relevant workloads, the cadence of credible substitutes, price per completed task, customer retention, and the share of revenue tied to long-term commitments. A high benchmark score carries more economic weight when it produces durable renewals and pricing power.

Second, test the conversion of demand into cash flow. For infrastructure owners, compare utilization and incremental revenue with capex, depreciation, financing cost, and the time required for new capacity to earn. For software companies, compare lower model costs with realised gross-margin expansion rather than assumed productivity gains.

The market consequence is greater dispersion

Open-weight models may already be adequate for routine work with clear inputs, observable outputs, and human review. They may remain unsuitable for long-running, customer-facing, or data-sensitive work where reliability and controls carry more weight.

The available scores do not establish the share of either category. They do suggest that a frontier score alone is a weaker shortcut for underwriting a durable moat.

The next phase of the AI trade will reward more specific analysis of lead duration, cash-flow conversion, and capital discipline.

Important Information

The information and data presented on this page are provided for general information only. They do not constitute, or should be construed as, any advertisement, invitation, inducement, offer, solicitation or recommendation to buy or sell any securities, interests in collective investment schemes, funds, derivatives or any other financial products, or to engage in any investment strategy or transaction. The content is prepared from sources Evermark believes to be reliable, but Evermark gives no representation or warranty as to its accuracy, completeness, or timeliness. Any views, opinions, and estimates reflect judgment at the original date of publication and may change without notice. This page does not contain sufficient information to support an investment decision and should not be relied upon as a substitute for independent judgment, and does not constitute investment, legal, tax, accounting, regulatory or other professional advice. It does not take into account any person's objectives, financial situation or particular needs. You should not rely on this page as the basis for any decision. You should obtain independent professional advice before making any investment or business decision.