Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

55% Positive

Analyzed from 3238 words in the discussion.

Trending Topics

#compute#more#demand#prices#https#shortage#cost#price#money#scarce

Discussion (88 Comments)Read Original on HackerNews

khursabout 5 hours ago
This is written by a financial company that is pouring billions into data centres.

https://www.apollo.com/insights-news/pressreleases/2026/01/a...

https://www.apollo.com/insights-news/pressreleases/2025/11/a...

etc

antrabout 4 hours ago
So Apollo believes in the thesis strongly enough to put billions of its own capital behind it. That sounds more like putting their money where their mouth is than a rebuttal.
TehCorwizabout 4 hours ago
Or they invested billions and it behooves them to generate justification for that investment at a time when everyone else is also building railroads, I mean data centers.
antrabout 4 hours ago
That's possible, but it's still an argument about Apollo's incentives rather than the thesis itself. By all means scrutinize anyone talking their book, but then show where the analysis is wrong. Otherwise we've moved from “Apollo is conflicted” to “Apollo must be wrong because it invested,” which is not much of an argument.
reticulatesabout 5 hours ago
Yes but how much of that compute shortage is from demand that is subsidized? We’ve seen companies like Uber drastically cut how much they are willing to spend on AI because they are paying actual usage costs, while at the same time OpenAI and Anthropic increase the limits on their fixed cost plans for individuals meaning people not paying usage costs are using it more and more… doesn’t this show that the compute shortage is because OpenAI and Anthropic are paying for it, not their customers? And the moment OpenAI and Anthropic stop paying for it, demand will collapse.
loolhahalmaoabout 5 hours ago
"subsidizing" aka making 30% gross margin instead of 90%.
arijunabout 5 hours ago
Do they actually have net positive income (excluding research, I guess)? I assumed no but I’ve never seen number one way or another.
tristanjabout 4 hours ago
Anthropic is currently profitable, generating around $1B/quarter and ~$50B in ARR. About 75% to 85% of Anthropic's revenue comes from its usage-based API business, which has a gross margin that exceeds 80%. https://www.tradingkey.com/analysis/stocks/us-stocks/2620181...

Meanwhile, OpenAI is at ~$25B ARR, but is likely not yet profitable.

reticulatesabout 5 hours ago
no, even if we assume their margins are 90% (they are not) they are still losing money because the $200 plans allow for tens of thousands of dollars worth of inference and a huge number of users are milking every cent across multiple accounts. Every “reset” OpenAI and Anthropic do is setting money on fire.

If it were true that they’re making money hand over fist they wouldn’t need to raise tens of billions of dollars every few months.

https://xcancel.com/i/article/2076078865060151465

OGWhalesabout 4 hours ago
> even if we assume their margins are 90% (they are not)

How do you know they are not?

It will be curious to see the cost of inference for these newly released open weight models and will help give an idea of the actual cost of inference. But for now, I think saying the $200 plans allows for "tens of thousands of dollars worth of inference" provides very little insight when you are measuring the inference cost in API pricing with an unknown margin.

dash2about 4 hours ago
In Amodei’s Dwarkesh podcast he says that they are constantly estimating the increase in demand for the next leg up, then going out to raise money to build for it. So it’s not necessarily that they are raising money because they are unprofitable.
rayinerabout 5 hours ago
The compute demand is not fake. Non-coding industries have barely begun to deploy this technology. In the legal sector, I’ve been a tech pessimist my entire career, because it was uniformly quite bad. I’ve spent the last few months demoing legal tools backed by frontier models, and we’re definitely going to buy one of them. They’re real and they work and they address a bunch of needs.
reticulatesabout 5 hours ago
You’re talking across the issue. The demand is real because it is cheap. The demand is being generated by OpenAI and Anthropic selling inference below cost on fixed price plans. If everyone was paying the actual costs then demand would fall through the floor. The legal tools you’re looking at use barely any compute. They’re not driving the compute demand. You can validate this by asking how much they are spending on API usage. A company spending $100,000 per month on a frontier model via an API is the equivalent of… 10 or so OpenAI and Anthropic fixed price plan customers. Are these legal tools spending hundreds of millions per year on the frontier models?
dash2about 4 hours ago
Compute would have to be very expensive not to be cheaper than a junior lawyer.
rayinerabout 3 hours ago
My point is that you’re projecting forward into the future where providers need to raise prices, but overlooking the demand that will arise when the rest of the economy starts using AI.
sjsdaiuasgdiaabout 4 hours ago
I'll believe you when I stop seeing lawyers getting busted for hallucinated references in their AI-generated filings.
mutkachabout 4 hours ago
True, and even selling token price may not be illustrative of what it actually costs them to provide the service. Tokens may be sold at a loss if majority of spenders are running agents 24/7

The industry and users moves on from single chat-based to more and more "agentic" workflows that may generate longer workloads with multiple simultaneous agents (separate agents - separate contexts - separate KV caches).

My estimation is based on, say, running a Kimi3 on a 24 B200 GPUs - it is very easy to lose money when selling tokens at "market" prices.

OleksandrCabout 5 hours ago
Recent Kimi K3 release is supposedly frontier-grade, so would be interesting to see what it really costs to run inference for models of this class, when the weights get released and other independent providers pick it up (then we could reasonably expect competitive pricing with low-ish margins).
wkjagtabout 4 hours ago
When gas prices went up in the 70s because of fuel shortage, smaller (more fuel efficient) cars became more popular to use less fuel to do the same thing. I wonder if the same will happen with compute, by making software more efficient, and do the same thing with less compute.
maxericksonabout 2 hours ago
For the companies doing massive build outs, there's already significant pressure to improve effectiveness and efficiency.
zehaevaabout 4 hours ago
The only thing I really hope for if this continues for an extended period of time is everyone optimizing their programs to use a minimal amount of ram.
sdsdssweew213about 4 hours ago
I bet everyone will soon use a thin client with 4GB of RAM, and all the compute happens in the cloud of some American corporation you pay subscription fees to.
Ekarosabout 2 hours ago
And that will probably fail. I don't think on average companies are capable anymore to deliver software that can operate on only 4GB... Even if everything but presentation layer is cloud based...
OGWhalesabout 4 hours ago
I can see that happening, though that sounds pretty bleak
RunSetabout 5 hours ago
> In 2026, compute is becoming the spice of our era.

I think it more resembles the "content" of our era.

https://refactoringenglish.com/blog/why-i-stopped-creating-c...

oeziabout 5 hours ago
At least one graph showing global total datacenter compute over time and projected to come online would have been helpful.
no-name-hereabout 5 hours ago
https://epoch.ai/data-insights/ai-chip-production from the beginning of 2026 claimed global AI compute capacity was doubling every 7 months.
incognito124about 5 hours ago
> from the beginning of 2026 claimed global AI compute capacity was doubling every 7 months.

That's definitely one way to say "it doubled this year"

computerphageabout 5 hours ago
I read it as "was posted earlier this year. It claimed that compute has been doubling every seven months for quite some time, on average"
no-name-hereabout 2 hours ago
No, it was "<linked article> from the beginning of 2026 claimed <x>." Also, if you open the article, you'll see the data goes back multiple years, but does not cover 2026 (as it was posted at the beginning of 2026).
post-itabout 4 hours ago
Not what it says.

> We fit a trend to the quarterly quantities of AI compute sold, as measured in H100-equivalents. We find that computing capacity has been growing by 3.3x per year, equivalent to a doubling time of 7 months.

This is just a graph of cash changing hands. I transferred a loonie from one hand to the other one trillion times this morning and have the largest data centre in the world.

Mistletoeabout 5 hours ago
Graphs like this are all I can think of when I see unsustainable euphoria like above.

https://www.bigtrends.com/education/lessons-from-the-past-10...

jcattleabout 5 hours ago
Ah I thought you would link https://xkcd.com/605/
droobyabout 4 hours ago
This article is about silicon. But the other shortage that matters is meat-compute.

AI is trained on human intelligence. The hyper-scalers are squeezing every last drop of automatically verified reward, and that may get us very far. A compiler passes or fails in milliseconds for free, forever.

But.. "good design taste" has no compiler.

Of course, taste isn't unverifiable. But it's is expensively verifiable. Noisy, slow, and orders of magnitude lower throughput. People with deep domain knowledge often can't articulate well _why_ one design works and the other doesn't. So, judgment arrives as a verdict, and not a crisp rationale. I guess we'll see if sample efficiency outpaces the cost of human judgement.

In the mean time, leverage will sit with whoever holds this tacit knowledge (incumbents). I.e., hospital systems, law firms, chip designers, studios, SaaS that are dominating their niche.. and not with the labs training on it. To me, this is why valuations of companies like Palantir could potentially make sense.

ACCount37about 3 hours ago
"A compiler" isn't the bottleneck. The bottleneck of RLVR is the rollout.

Before you even get the code for a compiler to do its thing, you need an AI to ingest and emit thousands of tokens autoregressively. That's what makes RLVR so expensive.

Likewise - I don't think your "design taste" argument holds? Areas where "taste" exists are far easier for AI to tackle than areas where no data exists. It's easier to teach an AI how to make an image that looks good than to teach an AI how to de-solder a BGA chip. "Tacit knowledge" yes, "taste" - not really?

cmiles8about 4 hours ago
Four things are currently correct:

1. There is a huge demand for compute, specifically GPU compute

2. Infrastructure providers are building like crazy, including taking on massive debt to fund this because their own cash flow can’t cover the bills

3. The demand for that compute is broadly being paid for with investor dollars pumping up the valuation of AI companies, not cash flow from said companies. If those subsidies go away these companies can’t pay for the compute they’re buying.

4. Those that own a lot of compute are starting to offload it, looking for interested buyers (e.g., Meta looking to build a cloud biz or SpaceX selling its excess compute to others).

All while advances in open weight models are making it appear that the major labs truly have no model moat.

Put together those 4 things paint a very ugly business and financial picture that seems unlikely to just correct itself naturally. History tells us, very clearly, that “the way out” of such a scenario is a series of events that is likely to leave some of the current players severely damaged if not simply out of business.

user43928about 4 hours ago
About 4:

Does it not make sense to rent out your compute if competitors have a better model and demand at higher prices?

About the moat, Mythos was first made available to customers at the beginning of April. Kimi K3 is still behind this.

Both OpenAI and Anthropic are expected to deploy significant upgrades in August.

cmiles8about 3 hours ago
Being first rarely matters in tech. Fast follow that’s “good enough” and cheaper eats “first” for lunch all day long, and that’s the pattern starting to play out.
senordevnycabout 4 hours ago
Eh, I don’t think there’s any reason to think #3 is true, and the whole thing being a house of cards is predicated on that one.

If the massive demand is still present for compute at market rates (which I believe it is), then your second point is just investors spending cap ex to build out valuable and profitable assets, no problem there.

Time will tell.

cmiles8about 4 hours ago
What evidence says otherwise? OpenAI is projected to have massive losses for years to come and all indications are that’s driven by the cost of compute being far higher than the revenue generated by said compute.
isodevabout 4 hours ago
I feel this post is blind to many of the secondary side effects of this "shortage". The rapid increase in prices and delivery times is having deleterious effects on all things tech - everything from phones to smart appliances and all sorts of gadgets has moved into unreachable price levels.

Just imagine, if the rumours are true and Apple's 'foldable' phone costs 2500€ or more - who is going to buys this? So much of Apple's ecosystem depends on enough users consuming services, buying apps and using their phones to facilitate digital interactions. I've been waiting for 2 months to get a new mac mini for our office lab, and the Studios have moved into "we can't really justify this expense" price range.

So the framing that AI is inevitable or projections that put the world into a year or more of this "scarcity", where people can no longer afford non-entry level gadgets, are just naive IMO.

RiverCrochetabout 3 hours ago
> Just imagine, if the rumours are true and Apple's 'foldable' phone costs 2500€ or more - who is going to buys this?

Phones are a bad example in the US at least-not sure how it works in the EU.

I think most people in the US simply finance new phones through their provider who is already providing them the plan, so it just looks like an expensive phone plan to them for a few years. Apple could offer a deal where a $3000 phone is simply a $200/mo. add-on to your bill over 3 or 4 years.

I bet this causes more things to be tied to a cellular provider and offered the same way, like game consoles and maybe even TVs. Imagine signing on to your cellular provider for a 10 year contract to get a few new phones, a game console or two, and a couple TVs.

user43928about 4 hours ago
I for one am going to buy Apple's foldable at the expected price between $2000-$2500.

As long as they sell an iPhone 18 at a more accessible price point, I do not see the issue.

isodevabout 4 hours ago
The issue is that this will be another vision pro. The issue is that every year, we take device and services price increase (which of course nobody can do anything about because Apple-Google is a cartel based in the US where consumer protection is not even in the dictionary but that's a different topic probably) and nobody keeps track of just how much these things cost for a "not an influencer / person who works in tech / friend of the president" household.
user43928about 4 hours ago
I think it is likely the foldable will sell well.

While it will be at a premium price point, people use their phones for hours every day, while the Vision Pro is a niche gadget.

With videos, photos, web browsing, and reading being much better on a foldable, I can see this have mass appeal.

Normal_gaussianabout 5 hours ago
What on earth is that H100 price trend graph. The equally spaced x-axis points are 2x6 months, 6x3 months, 9x1 month. The whole visualisation of the trend is ruined on the back of that.

There are liars, damned liars, and people who play silly buggers with scales.

Diogenesianabout 4 hours ago
I think this is a severely compromised visualization but not actually misleading, either by design or in effect. Seems like they wanted to include the full picture - namely, that the 2026 spike is still less than the 2023 price when new, but doing all the data monthly would have been hard to read. Note that doing it all monthly would have made the derivative of 2026 more visually stark, so the way the graph is presented actually weakens their argument (albeit inconsequentially).

I think the best way to fix the graph would be shading or line texture to indicate the scale. The biggest problem is that the inconsistent scale is a surprise you only see upon close reading, after you've visually digested the trend. So the scale needs to be apparent in this first visual digest. (Sort of like how many logarithmic charts include thin axis marker lines in the body of the graph itself, so as to immediately inform a quick glance.)

nh23423fefeabout 5 hours ago
i dont see how compressing the past ruins the present day extrapolation? How does the conclusion change for you if the graph was 3x wider on the left? The title is "rates are rising" is that not true?
Normal_gaussianabout 4 hours ago
A graph is much more than one conclusion; in fact, almost the entire point of graphing is to allow the comparison of "shapes" and to easily hypothesise about associations across datasets.

This graph misrepresents the rate the price declines and the length of time it has been stable for, which throws off nearly all non-trivial conclusions.

nh23423fefeabout 3 hours ago
It doesn't mis-represent it. If you think that, then you think log plots misrepresent rates.
danlittabout 5 hours ago
That specific conclusion is unaffected, but that doesn't make it okay. If they used fake numbers, but the conclusion were still the same, you presumably wouldn't think that was okay either?
danlittabout 5 hours ago
I had to look 3 times to see what you meant. That's pure evil!
sghiassyabout 5 hours ago
Isn’t the compute shortage temporary?

I don’t believe in 5-10yrs we’ll be in a shortage anymore

gobdovanabout 5 hours ago
We're in such a bull market, even shortages grow.
Advertisement
lifeisstillgoodabout 5 hours ago
Meh. That’s like saying we have a parking shortage in cities. No we have more car journies than we need.

We have societies designed by the default choice. If that choice is walkable streets, low cost electric buses, dense neighbourhoods (mostly I mean you can walk for miles along streets and parks and shops without crossing a car park)

Then you get far less car use. People aren’t stupid, but living in downtown Houston means you have far less choice about driving everywhere than living in suburban Amsterdam

It’s all choices

We just are making bad ones mostly

The choices now are would you like ads or more ads. Shall I waste compute seeing if the web page you clicked on can be summarised ?

The simple answer to AI is to charge it at cost - not subsidised. Then the market will start to shake out.

It might take the US stock market with it …

twoodfinabout 5 hours ago
If instead of us training the models to be human utility maximizers, the models had figured out how to train us to maximize their own utility, could we tell the difference?
GolfPopperabout 5 hours ago
I suspect that utility may not enter the picture at all, and what we have are LLMs that have been optimized for getting the sort of humans who decide what to spend money on to maximize spending on LLMs and related infrastructure. The LLMs in turn could be anything from Skynet to a very spicy cold-reading autocomplete.
redwoodabout 2 hours ago
While building Fabs takes time you can only imagine that the amount of capital flooding in means there will be an explosion in infrastructure supply over the next decade and eventually prices will crater
kittikittiabout 4 hours ago
To preface, this article presents valid points that I agree with. I would have also liked to know their findings on computational memory (CXL) and not just HBM or DDR. Also, a theoretical background on the Von Neumann and Harvard architectures would be helpful.

Many people designing these systems understand that there are vast shortages in almost all sectors in computing hardware. I'm more interested in AI strategy of a future where there's an oversupply. I speculate that by 2030 there will be both an overproduction in the factories that make these chips and a burgeoning second hand market of AI GPU's. The differences in the compute requirements for training versus inference of AI will explain the hindsight in oversupply.

zer00eyzabout 4 hours ago
> Compute Is Becoming a Competitive Moat

No, it is not a moat.

The story here changes dramatically when you start to look at the facets of the industry that the poster is skipping over.

> At the semiconductor level, TSMC’s advanced-node capacity—particularly N3, which underpins much of the AI accelerator ecosystem—is approaching full utilization through at least 2027.

MS bought more GPU's than they had rack space for: https://www.datacenterdynamics.com/en/news/microsoft-has-ai-...

Open AI bought out the memory: https://x.com/kwharrison13/status/2029248559388746168 but they have no means to consume anywhere close to their order.

Meanwhile both google and amazon are consuming a bunch of TSMC capacity to build their own ai chips, bypassing NVIDIA ... And as for them, they seem to be addicted to burning power to keep scaling, and that is a massive problem - if the next gen chips burn more watts for the same amount of work that is only going to exacerbate the power issues were having not help them.

Tokens are just Gacha for business. https://en.wikipedia.org/wiki/Gacha_game - It is software you dont control and you are going to pay for a cache miss. That isnt a model that is sustainable (even more so in the authors multi agent flows).

At the point that prices come down, (and they will) you're going to see a lot of corporations move from the cloud to on premise or back into colocation.

sneakabout 5 hours ago
GPUs aren’t scarce. There is just a lot of demand, so prices have gone up.

Anyone can get GPUs right now and build out what they need; it’s just a matter of paying more for them than the next company. I’ve been looking at building a multi-TB HBM system lately and they’re readily available - they just cost $400k.

hibikirabout 4 hours ago
Without realizing it, you are coming close to the economics definition of relative scarcity. If we went by uour definition, basically nothing is scarce, ever, because it's possible to provide it, but the prices are set in a way that is completely unrelated to the cost of production + risk related margin. Instead prices resemble what would be an auction. And it all only works out because demand is elastic, so there are prices that people refuse to pay, because for their use case, they'd lose money.

Housing, collectors items, access to the time of a top of the line doctor... all relatively scarce. And You'd be laughed out pf most rooms if you claimed those things aren't scarce.

StilesCrisisabout 5 hours ago
"Shortage" means exactly that. It doesn't mean "there's no GPUs at all." In a gasoline shortage, prices go up and there are lines at the pump, but there are still cars on the road.
eternauta3kabout 5 hours ago
The lines at the pump only happened because of rationing rules, which led people to hoard and buy gas when they otherwise wouldn't have.
fhdkweigabout 4 hours ago
If you are talking about Russia, hoarding or not, if you burn enough refineries, you will have shortages.
StilesCrisisabout 4 hours ago
Sure. Rationing is a common response to shortage in general even though it isn't very effective.
hackeraccountabout 3 hours ago
If prices go up there aren't lines. You get lines when demand goes up and supply doesn't or supply goes down and demand doesn't follow. Either way the "fix" is either to raise prices or have lines.
gsprabout 5 hours ago
> GPUs aren’t scarce. There is just a lot of demand, so prices have gone up.

What does "scarce" mean to you?

I'd wager that to most reasonable people it means "available in a supply that, measured against demand, is low". Market forces typically react to such a state by pushing prices up. Saying "they're not scarce, they're just expensive" is just silly.

The exception to this would be during e.g. supply chain hiccups (lots of compute sitting there, unable to get to its operational destination) or when other forces artificially push up prices. But that's not what we're seeing here. Compute is much scarcer than it's been for years.

fhdkweigabout 5 hours ago
There is the type of scarce when you can't get something for love or money. Look at the gasoline lines in Russia. Those people sit in lines miles long for gas stations that aren't even open. No matter how much money they are willing to spend, they still can't get the gas.

I'm just grateful that I don't need to upgrade my computer for a while, and cross my fingers that I don't have a hardware failure in the next couple of years. I had an unpleasant laptop failure in 2021 that I don't wish to repeat.

alt227about 5 hours ago
The definition of scarce is rare, or insufficient to meet demand. Not just that supply is low in relation to demand, it means that you physically have trouble getting something.
gsprabout 4 hours ago
Fine, replace "low" with "too low".
antonvsabout 4 hours ago
In economics, scarcity applies to almost all goods and services, because they’re not infinitely available - see https://en.wikipedia.org/wiki/Scarcity

As such, saying “GPUs are scarce” is an empty statement from a strictly economic perspective - so are all other physical goods. The correct economic phrasing would be “there is a shortage of GPUs” or “GPUs are in high demand”.

Of course, in colloquial English we interpret “GPUs are scarce” as meaning the same thing.

s08148692about 5 hours ago
It will be fascinating to watch this play out. Particularly if AI-compute satellite constellations become a reality - We're trending towards a matrioshka brain and I'll be happy if I live to see the beginnings of that future
post-itabout 4 hours ago
They'll become a reality on the same day as solar roadways.