Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 350 words in the discussion.

Trending Topics

#run#hour#bubble#pop#cost#cloud#amd#tokens#users#consumer

Discussion (24 Comments)Read Original on HackerNews

majkeabout 2 hours ago
I don't think you can buy a single "MI300X" unit, right? Only the box with x8 of these at a cost of ~250K EUR.
zhoutongabout 2 hours ago
It’s available on demand from a few cloud providers. Seems like the cheapest is AMD Developer Cloud (https://www.amd.com/en/developer/resources/cloud-access/amd-...) powered by Digital Ocean at $1.99/hour.

Edit: Now I think about it, this might be the cheapest way to run the DeepSeek V4 Flash 0731 on a dedicated inference server at original weights. I haven’t run mixed load benchmarks but I guess it’s possible to generate $3-$4 worth of tokens per hour and still maintain a usable per-user throughput.

WASDxabout 2 hours ago
At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.
lnenadabout 1 hour ago
830t/s is burst aggregate. ~500 is sustained and it's for 8 concurrent users. Meaning for $1.99/hour if you serve 8 users it's 8*$0.54, not just $0.54.

You shouldn't rent one out if you're just serving it for yourself, but from a financial standpoint if you sell to users you can take a 100% margin.

Almondsetat40 minutes ago
You get privacy for 4 times the cost
krisknezabout 1 hour ago
How is that economically viable? They are selling at a loss?
thrownaway561about 1 hour ago
This is exactly what I came to say. The price of Flash is so cheap that trying to run it locally or with your own hardware is pointless. I was using it about a month ago to program some stuff and ran it for 4 days non-stop and it cost me about $2.
langsabout 1 hour ago
You need to optimize the KVCache part(save to disk to save compute) to achieve this goal.
Lwerewolfabout 1 hour ago
The MI350p exists and should run a decent quant (say, the ~96GB antirez mix) well, but you can get two rtx pro 6000s for one of these, or 8x (actually more) r9700 + probably the gear to run them, etc.

Otherwise, you can probably buy one of these second hand from somewhere (SXM A100s are available that way) and run it in an adapter board.

_joel40 minutes ago
I thought it was a consumer grade GPU until I saw the 192GB of HBM and 256GB or RAM.
varispeed30 minutes ago
To be fair the development of GPUs have stalled over the years. If they kept up with the progress instead of focusing on enterprise market, likely 256GB consumer GPU would be a norm today.
baalimagoabout 2 hours ago
Give it an AI-bubble pop and these will be flooding the market.
_factorabout 2 hours ago
They will be instantly bought out by companies, not individuals. The consumer bubble won’t pop for quite a while yet. Production also won’t ramp up while lack of real competition keeps the demand high.
aurareturnabout 1 hour ago
When is it popping? Is the AI bubble in the room with us now?
amrit312823 minutes ago
Tomorrow? Next year? In 5 years? Nobody can say. But we do know that AI is overvalued, so it WILL pop.
baalimagoabout 1 hour ago
Next month perpetually
slaw20 minutes ago
The AI bubble will pop when China gets access to EUV, so the earliest it could happen is 2030