Open-Weight AI Models: What Morgan Stanley’s ROIC Numbers Mean
Morgan Stanley says open-weight AI models could push token prices down but keep hyperscaler ROIC at 23-39%. Here's what the numbers actually mean.

Morgan Stanley analysts published estimates on August 15, 2026, projecting that model providers running one gigawatt of owned Nvidia GB300 infrastructure could still earn 20% to 60% returns on invested capital, even as cheaper open-weight AI models push token prices lower. For hyperscalers renting out GPU capacity, the bank puts ROIC at 23% to 39%. The key argument: falling model prices should expand usage volumes enough to offset tighter margins, and cloud providers can layer on higher-margin services to the same customers.
What happened
| Data point | Detail |
|---|---|
| Report date | August 15, 2026 |
| Model provider ROIC range | 20% to 60% (owned Nvidia GB300, 1 GW) |
| Assumed token price | ~$1.75 per million tokens |
| Assumed throughput | 2,000 to 3,500 tokens per second per GPU |
| Hyperscaler GPU rental ROIC | 23% to 39% (varies by hourly pricing) |
| Morgan Stanley ratings | Overweight on Amazon (AMZN) and Alphabet (GOOGL) |
Open-weight AI models (models whose trained parameters can be downloaded, modified, and run on any infrastructure the user controls) are spreading quickly. Morgan Stanley analysts say the trend will squeeze pricing for proprietary model developers, but the story for cloud infrastructure providers looks different.
The bank’s core thesis: even with token prices falling, the companies that own or rent the compute still win, because cheaper access creates more demand, and more demand fills more capacity.
Why it matters
Four reasons hyperscaler economics hold up
Morgan Stanley identified four factors that should keep returns attractive despite pricing pressure on models:
- Compute scarcity persists. Enterprises running open-weight inference workloads still need cloud providers to supply the GPUs. Owning the hardware remains the leverage point.
- Volume offsets margin compression. Lower model costs accelerate adoption, especially among smaller businesses. Hyperscalers can adjust capacity prices with supply and demand, so total profit growth can outpace any drop in percentage margins.
- Throughput keeps improving. Better chips, faster interconnects, more efficient model architectures, and request-batching software all raise the number of tokens each GPU can process per second. Amazon uses its Trainium chips and Alphabet uses tensor processing units (TPUs) to reduce their own cost of supplying compute.
- Open-weight workloads pull in connected revenue. Managed APIs, GPU rentals, databases, storage, and security tools all attach to the same customers. Low-priced model access becomes a loss leader that draws more profitable cloud spending.
Which models to watch
The analysts specifically called out Meta’s Muse models and Alphabet’s Gemini Flash products as the ones most worth tracking for signs of pricing pressure, adoption rates, and throughput gains. For businesses evaluating AI integration, these two product lines are likely to set the floor on token pricing across the industry.
Our take
Morgan Stanley’s framing is essentially: the model layer commoditises, the infrastructure layer does not. That is a reasonable read, and it matches what we see when helping clients choose where to run AI workloads.
The $1.75 per million token assumption is worth scrutinising. Token prices have moved fast, and the 2,000 to 3,500 tokens-per-second throughput range is a wide band. ROIC at the bottom of that range and the bottom of the pricing range looks a lot less comfortable than the headline 20-60% figure suggests.
For business operators, the practical signal here is not which stock to buy. It is that open-weight models are going to keep getting cheaper and more capable, which means the cost argument for avoiding AI tools is weakening. The switching cost is also lower with open-weight models: you are not locked into one vendor’s API. That is worth factoring into any procurement decision you make now. We cover the broader automation platform landscape in our comparison of n8n, Make, Zapier, and similar tools, which touches on how model access costs affect workflow tooling choices.
The “loss leader into cloud spending” dynamic Morgan Stanley describes is real. If you are using cheap model inference from a hyperscaler, check what else you are being nudged to buy from that same vendor before assuming the total bill is low.
What to do about it
- Audit your current AI spend: note whether you are paying per token via a proprietary API and at what rate.
- Test an open-weight alternative (such as a self-hosted model on rented GPU capacity) against your current provider on cost and latency.
- Watch Gemini Flash and Meta Muse pricing announcements as reference points for where the market floor is heading.
- Before committing to a hyperscaler’s managed AI stack, itemise the connected services (storage, databases, APIs) that get bundled in, and price them separately.
- If you want help mapping this to your specific stack, talk to the Lumien team about a no-jargon audit of your current AI tooling costs.
The cheapest model is not always the cheapest solution once you count the infrastructure around it.
Frequently asked questions
What is an open-weight AI model?
An open-weight model is one where the trained parameters (weights) are publicly released, allowing anyone to download, customise, and run the model on their own infrastructure rather than being locked into a vendor's API.
What ROIC does Morgan Stanley forecast for AI model providers?
Morgan Stanley estimates model providers could earn 20% to 60% returns on invested capital from one gigawatt of owned Nvidia GB300 infrastructure, assuming a token price of roughly $1.75 per million tokens and throughput of 2,000 to 3,500 tokens per second per GPU.
Will cheaper AI models hurt cloud companies like Amazon and Google?
According to Morgan Stanley, not significantly. Hyperscalers are expected to maintain 23-39% ROIC on GPU rentals because lower model costs drive higher usage volumes, and cloud providers can attach higher-margin services like managed APIs, storage, and security tools to the same customers.
Which AI models should investors watch for pricing pressure?
Morgan Stanley specifically flagged Meta's Muse models and Alphabet's Gemini Flash products as the key products to monitor for signs of token price compression, adoption growth, and throughput improvements.


