HomeArticles

AI Capex: Bubble or Opportunity? The Insiders Are Betting on Scaling Laws, and 2027 Is the Checkpoint

Written in 2024. Charts are the originals from the time of publication. · Collected in AI main line, 2023 to 2025

By Picaca · 2024-07-11 · Read the Chinese original

Insiders expect inference to fall from $10 to $0.1 per million tokens in two to five years, and models to beat humans by 2027. 2027 is the checkpoint.

The argument over whether AI investment is a bubble keeps getting louder, and the market has split into two camps. One side sees no business model yet and reads the spending as overinvestment. The other side expects exponential gains to keep making models smarter until applications actually land. Over the past two weeks we worked through three accounts from people inside the industry, wanting to understand the optimists: one projecting AI's exponential path in orders of magnitude, one on inference costs falling over time, and one from the Anthropic CEO. We wrote each up as a memo at the time. Our comments are in those notes, but this post pulls the whole argument together.

Key takeaways

  • Insiders are running the Scaling Laws playbook. Until the laws break, every vendor keeps sprinting for exponential growth in effective compute, which makes time the biggest cost in the race.
  • Inference prices are already falling fast: GPT-4 to GPT-4o took the price from $37.5 to $7.5 per million tokens, and Claude 3 Opus to Claude 3.5 Sonnet from $30 to $6.
  • The application sweet spot arrives when inference gets from roughly $10 per million tokens today down to $0.1, which insiders expect within two to five years.
  • 2027 is the checkpoint: by then models should beat humans on most tasks, and if they do not, the industry drops into years of algorithm work and a short AI winter.

Three points carry most of the weight in what the insiders are saying.

  • They follow the Scaling Laws. As long as the laws have not broken, the job at every vendor is to keep pushing effective compute up an exponential curve, and if you accept that, then for all of them time is money. Any gain in compute, whether it comes from hardware or from algorithms, gets recycled straight back into model performance, because whoever ships the better model first sits in the best position to take share once applications go mainstream.
  • History says the cost of missing this one is too high. AI may be the next wave after the internet, mobile internet and cloud that resets who competes in an industry, and missing it can mean having no seat at the table in the next technology cycle. AI already raises productivity, and each new model release lowers the barrier to using it. Real end applications are still not visible, but no vendor can miss this round, so as long as the rest of the business can carry the spending, they will spend.
  • Now to 2027 is the window that matters. The insiders expect an improved model every few months, with performance gains and cost reduction arriving together, and they expect AI inference to fall from $10 per million tokens today to $0.1 per million tokens over the next two to five years, at which point the application sweet spot should appear. The trend is already clear in 2024: GPT-4 to GPT-4o went from $37.5 to $7.5, and Claude 3 Opus to Claude 3.5 Sonnet from $30 to $6. We are still waiting on Claude 3.5 Opus, and hoping the price on it is aggressive.

Several risks are obvious enough that we will keep tracking them.

  • The Scaling Laws break. GPT-5 has been pushed out by one to one and a half years, which we read as negative. Claude 3.5 Sonnet landed the same week and covered the gap. As long as some vendor keeps shipping to the Scaling Laws, this is manageable.
  • Power. The recent news that the Oracle and xAI deal fell apart looks like it came down to doubts that the power site xAI preferred could deliver enough capacity, plus a build schedule xAI wanted that was too aggressive. Oracle also picked up an OpenAI training order this quarter, which we think may reflect data center construction running behind plan. Getting land and getting power could become the binding limit on exponential growth from here.
  • The financial health of the companies doing the spending. Public cloud growth has recovered at every vendor over the past two quarters, as AI restarted the wave of data moving to the cloud. When we checked last quarter their finances still looked normal, so there is no immediate problem, but this needs watching from here (Meta in particular).
  • Total revenue from end applications. This is what the market worries about most. Revenue at the vendors supplying new models is in fact growing quickly, and the real question is whether other applications develop. If the insiders are right, then as prices move toward the sweet spot over the next year we should see more applications appear.

That is the short version. The rest of this post walks through each piece. Speaking as people who use these tools every day, we genuinely want the next models and the price sweet spot to arrive, and it is easy to see why the big vendors treat this as a race they cannot afford to lose. As smarter models show up, many of today's application problems may simply go away, so we will keep watching what these companies build.

This post condenses three memos we published earlier this month on how insiders see GenAI developing. Our notes did not name the authors of the first two, so we describe each memo by its subject.

  • Part one, July 3 2024: using orders of magnitude to project the exponential path of AI, and forecasting a major breakthrough in 2027.
  • Part two, July 5 2024: AI inference costs falling sharply over time, and whether that creates the sweet spot for applications to land.
  • Part three, July 5 2024: the Anthropic CEO on the future of AI.

The insiders are following the Scaling Laws

Scaling Laws are a central idea in AI, and in large language model development in particular. Researchers at OpenAI set them out formally in a landmark 2020 paper, Scaling Laws for Neural Language Models.

The core claims are these.

  • Model performance keeps improving as scale goes up, counting parameters, training data and compute. The relationships hold across different model architectures and different tasks.
  • The improvement is not only a change in quantity. It produces jumps in kind, giving models new capabilities.
  • No ceiling on this scaling has been observed yet, which means adding scale may keep producing striking breakthroughs.

Those findings gave AI research a clear direction, that scale buys performance, and laid the groundwork for the large language model work that followed. They also explain why the large technology companies and research labs are pouring resources into training ever larger models.

The author of the first memo argues it this way. Assume computing capability, meaning raw compute, algorithmic efficiency and unhobbling gains (what you get from removing the constraints that keep a model from using what it already has), keeps growing exponentially. Then look at the record: from 2019 to 2023, going from GPT-2 to GPT-4 added 5 to 6 orders of magnitude (OOM) and produced a leap in capability. Over the next four years, 2023 to 2027, with all three of those inputs still on exponential paths, we can expect another leap on the scale of the last one. That memo sits at the aggressive end of the range: its author thinks AGI is possible, and that it would bring one more OOM acceleration at that point.

Stacked contribution in orders of magnitude from the three drivers of AI progress: raw compute, algorithmic efficiency and unhobbling gains.
Figure 1: Figure 1: Orders of magnitude contributed by the key drivers of AI progress: raw compute, algorithmic efficiency and unhobbling gains

The second memo argues that AI inference costs will fall sharply over time, building on the record from 2022 to 2024: language model accuracy on MMLU (Massive Multitask Language Understanding, the standard multi-subject knowledge benchmark) kept climbing while cost per million tokens kept falling. The memo puts that down to better chips, to Mixture of Experts (MoE) for efficient compute, and to parameter training and inference techniques, which lines up with the three growth drivers in the first memo.

The author expects MMLU accuracy above 90% within the next 1.5 years, with cost falling far enough to lower the barrier to adoption, and in two to five years MMLU accuracy of 95% to 100% at a price sweet spot below $0.1 per million tokens. If that holds, applications could start to take off within roughly the next two years.

Scatter plot of language models from 2022 to 2024 showing MMLU accuracy against inference cost per million tokens.
Figure 2: Figure 2: MMLU accuracy versus inference cost across language models, 2022 to 2024

That fits the third memo, in which the Anthropic CEO puts 80% of company spending into compute supply and predicts, optimistically, that the scale of AI training could rise to $10B or $100B between 2025 and 2027, at which point AI models have a chance to beat humans on most tasks.

If models really do reach that level of performance at a relatively cheap price, then applications should take off, and the companies that got there first take the market that comes with it.

That is why every large vendor is racing to release newer, larger, more usable models. Anthropic, to take one example, plans a stronger version of Claude every few months. It echoes what the NVIDIA CEO said on the last earnings call: time is money. When the customer needs compute right now to train a stronger model, time is their largest cost, and they cannot wait.

History says this is a paradigm shift nobody can afford to lose

AI is infrastructure. It could change how almost every industry operates. The potential impact of this shift may be larger than any before it, because AI can change how we work, how we innovate and how we solve problems.

Facing an industry revolution this size, the large technology companies on the field today, Microsoft, Alphabet, Meta, Amazon and the rest (the hyperscalers), have watched the internet era, then mobile, then cloud produce clear winners and clear losers. Whether they won or lost in those earlier shifts, we believe every one of them understands in detail what missing a major technology transition costs.

Look back at the previous waves. Each one reset the whole technology ecosystem.

  • The internet era: Google rose on its search engine and became the giant of the information age.
  • Mobile internet: Apple and Google (Android) caught the smartphone revolution and came to dominate the mobile ecosystem.
  • Cloud: Amazon (AWS), Microsoft (Azure) and Google Cloud invested ahead of the market and innovated, and became the leaders of the cloud market, while some traditional IT companies struggled through the transition.

So staying ahead through a technology change is critical. The giants are committing enormous sums precisely so they do not fall behind in this race, and three things are at stake.

  • Market leadership. Falling behind in the AI era can mean losing competitiveness across several key areas at once.
  • Data and scale advantage. AI development depends heavily on huge data sets and heavy compute, which is exactly where these giants are strong, and that can create a first-mover advantage in which the big get bigger.
  • The ecosystem. A successful AI platform can attract large numbers of developers and enterprise users and build an ecosystem that is hard to dislodge, much like the fight over mobile operating systems in the smartphone era. We are in the middle of the next wave of enterprise cloud migration right now, so tying enterprises into your own services, and giving them more productivity there, can create more business and possibly a moat.

Put the insider view together with what the giants have at stake, and the mindset behind the investment becomes clear, along with why they expect both spending and compute to grow exponentially through 2027.

Bottom line: the window that matters runs from now to 2027

Take the insiders at their word and the whole thesis rests on the Scaling Laws holding. In that case the job in front of them is to keep the exponential growth in AI investment going until 2027.

2027 is a critical checkpoint, because by then AI models should be beating humans on most tasks. If that happens, applications take off because prices have reached the sweet spot, and smarter models may deliver one more acceleration in effective compute on a log scale. If it does not happen, after this many resources and this much compute have gone in without a large breakthrough, the industry may face years of algorithm tuning and a short AI winter.

For our own part, when generative AI first appeared in 2023 we were not convinced applications would land. The algorithms are a black box, and there were still plenty of limits in practice: we spent time training models, working with vector databases, which store text as numeric vectors so it can be searched by meaning, and with Langchain, a toolkit for chaining model calls together, and still could not properly solve the problems we had. Back then we leaned toward the view that the technology needed a long development period and that overinvestment could produce a short-term bubble.

In 2024, with OpenAI's competitors out in the open, the wave of model releases in the first half alone has cut the barrier to building on them. GPT-4o and Claude 3.5 in particular, smarter models at falling prices, have already improved our own workflow substantially. So we now lean toward the view that applications will land as models get smarter. Even with applications not here yet, the power of generative AI is obvious in personal use, and many problems we could not solve in 2023 have been solved as smarter models arrived, which is reason enough to expect a lot from the next releases.

We spent this stretch compiling insider views to find out what the people doing the spending are thinking. As long as everyone still believes in the Scaling Laws, meaning that adding scale (parameters and training data) keeps improving what a model can do, then at this stage no large cloud platform can afford to stay out. Once someone makes a model good enough for applications to take off, any large cloud platform that is not in it could be out of the next generation.

One caution: the Scaling Laws hold over the range observed so far, but there is no certainty they continue forever. Some researchers think scale alone will not get to artificial general intelligence (AGI). The environmental and economic cost of scaling has also become a live debate.

So if you buy the insider view, here is what to track as it moves.

  • Do the Scaling Laws hold forever? OpenAI pushing the GPT-5 release out to late 2025 or early 2026 is a seriously negative signal, and our worry is that something has gone wrong that stops the hypothesis from playing out on schedule. Claude 3.5 shipped the same week and covered for it: if OpenAI has run into internal problems, then a competitor needs to pass it to show the hypothesis has not broken.
  • As models get smarter, are applications landing? The investment numbers keep rising, so whether the application market grows has to be watched closely. So far, as models get more usable, revenue at the large model startups has grown several times over. The contribution at the public cloud vendors looks small, but it is pushing enterprises toward the cloud in the short run, so the picture is acceptable for now. Over the medium term the revenue sources have to get clearer. The optimists have very high expectations for 2027, so if this is working, the data should build step by step with each model release, and something clearly application-shaped should show up in 2025 to 2026. Alongside that, check whether the hyperscalers' financial performance is being dragged down by the size of the investment. Watch Apple's edge AI after the fall 2024 iOS update.
  • Do obstacles show up that nobody expected, such as power shortages or government intervention? As things stand, power shortage is the more immediate risk, and regulation is more of a medium-term one.

That is the current round of notes. Next we will work through the arguments of the people who think AI is a bubble.