Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

57% Positive

Analyzed from 1578 words in the discussion.

Trending Topics

#cuda#nvidia#gpu#amd#compute#open#hip#more#https#sure

Discussion (71 Comments)Read Original on HackerNews

linuxhanslabout 9 hours ago
Off-topic and somewhat of a rant, but I'd far prefer us all focusing on open standards like HIP, SYCL, OpenCL, etc.

It's unbearable that most LLM inference happens on closed H/W, closed drivers, and closed SDKs.

mistercowabout 8 hours ago
On the SDK front, you tried just having an agent reimplement the model you're interested in and just use the weights? I've taken to treating off the shelf implementations as reference implementations anyway, because I can often squeeze out significantly better performance for my configuration and use case by having Codex hammer at it for a few hours.
srousseyabout 8 hours ago
Hugging face is working on something like this where well known models get fused into a single implementation.
drivebyhootingabout 8 hours ago
Could you share the prompts and workflow? I’ve tried this too, but with mixed success when it comes to creating custom tile kernels.

I would really appreciate your input!

mistercowabout 8 hours ago
It's been pretty ad hoc, but my prompts are nothing special. Things I generally do:

1. Top end model on high/xhigh thinking (last time I did it it was Sol xhigh I think)

2. Make sure it creates some representative fixtures of different sizes and sets up a good testing, profiling and benchmarking loop that doesn't require my input.

3. Make sure it has access to reference implementation code

Edit: Oh and one obvious pitfall that for some reason I still have to remind even smart models of from time to time: make sure it knows not to try to parallelize its benchmark runs. I've occasionally had an agent struggle to figure out absolutely nonsensical data because it tried to run multiple tests on the same compute hardware simultaneously.

mschuetzabout 7 hours ago
The problem with the open standards is that their dev UX is absolutely horrible. You can't neglect usability, and then be surprised that there are no users.
the__alchemistabout 5 hours ago
This is at the core of the matter for me, and my knowledge is too weak to understand why this is the case. I don't enjoy the idea of relying on Nvidia's stack for GPU compute, but the alternatives I've tried (e.g. Vulkan compute) are higher friction to use. I am trying to reconcile why; Nvidia shouldn't have this moat.

My software is labeled "CPU only unless using an nVidia GPU". I would prefer to strikethrough "nVidia". Incidentally, this means no more Mac support.

swernerabout 4 hours ago
Vulkan compute is not really designed or intended to be a CUDA competitor, its feature set is much more restricted, and Vulkan host side code is much more verbose than CUDA. OpenCL or SYCL are much closer in features to CUDA. I found that when using SYCL on Nvidia, debugging symbols etc can be passed through and you can use tools like NSight Compute to profile it as if it were CUDA.
HeavyStormabout 6 hours ago
Commercially it's better to have AMD support CUDA which can help break Nvidia soft monopoly.
bigyabaiabout 9 hours ago
I would too, but sadly that's Khronos' job to organize, and they've had trouble getting American vendors to work together.

It's likely that CUDA will continue dominating until they put aside their differences. The current MLX/MPS/ROCm ecosystems are too fractured to threaten Nvidia.

boredatomsabout 8 hours ago
Maybe a not-Khronos org should try
high_na_euvabout 9 hours ago
Wdym American Vendors?

Intel uses SPIRV iirc

my123about 8 hours ago
> Intel uses SPIRV iirc

They're migrating away from SPIR-V to their own, Intel PISA: https://discourse.llvm.org/t/rfc-upstreaming-the-pisa-backen...

bigyabaiabout 9 hours ago
I'm talking about holistic efforts like OpenCL, and standards that would be equivalent to Nvidia's "Compute Capability" versioning.

The basic underlying tech can be agreed on, but Apple/AMD/Intel all have different GPU priorities that limit their ability to agree on a CUDA-adjacent hardware platform.

kiiciaabout 6 hours ago
it will only get worse now when hughingface was bough by nvidia, not immediately, but in span of year or years...
anon291about 6 hours ago
It's impossible to have an open standard. Hardware accelerators are nothing alike and have different perf characteristics. Each kernel is tuned to the hardware. The idea of writing a performant kernel in opencl is a fantasy.

Source: worked at a bunch of accelerator companies in the kernels or equivalent team. They're nothing alike.

swernerabout 6 hours ago
For a performant portable language, we’d have to go to a higher level where you describe what to do and leave the how to do to the compiler. It would then need to be able to adjust memory layout, access patterns, data type choice to the underlying hardware. I’m not sure if this is possible to do reliably - the closest we have right now may in fact be highly detailed plain English descriptions of the algorithms fed to an LLM prompted to produce assembly.
anon291about 5 hours ago
It's not possible to do reliably. None of the models in current use today use any esoteric math. It's extremely easy to implement the math behind both the inference and learning of all modern models.
swernerabout 8 hours ago
AI will take down Nvidia’s moat. When it becomes trivial to translate CUDA/PTX to HIP, SYCL or Metal, CUDA is no longer the moat, it becomes the intermediate representation.
larodiabout 6 hours ago
trivial to translate (or transpile) - okay. trivial to understand the result - not so much. trivial to then evolve it - hm... perhaps a different story. still, it seems very likely now, that such "quick rewrites" are viable, not sure if an open approach to them is viable. a newly born open project that was LLM-derived, and not by a credible author, which spans hundreds of files no human eye has ever looked at, can only work for a closed organization, but will never be trusted by the general audience... just like that.
bayindirhabout 8 hours ago
> When it becomes trivial to translate CUDA/PTX to HIP,...

ZLUDA is already doing that, no?

swernerabout 7 hours ago
I don't think we're at a point yet where anyone would trust ZLUDA enough to ship commercial products that rely on it. I would be delighted though, if anyone can prove me wrong.
bayindirhabout 7 hours ago
No, but we can go there. This is an open source project. Anyone can put some more effort behind it and push it further. It's improving, AFAICS.

Src: https://github.com/vosen/ZLUDA

Keyframeabout 8 hours ago
yeah yeah, "when" an often keyword with AI it seems. As Mr. E. Nigma put it - what always comes but never arrives? Meanwhile the moat deepens and it's build on inertia and laziness and Nvidia knows this really REALLY well.
swernerabout 6 hours ago
Oh, absolutely. Nvidia is the modern day “nobody gets fired for buying IBM”. The reason we’re still using Unix is not because it’s the best, but because it had to much inertia to let any alternative become its successor. Similarly, C and HTML are maybe the most terrible yet extremely useful languages we have.
mathisfun123about 8 hours ago
i swear people who are outsiders here have only clickbait takes; if you've never had to ship GPU code professionally you should just not comment on these things.

the source language has never been the moat. Nvidia sells to hyperscalers. Hyperscalers have armies of kernel authors who have no issue translating shaders by hand (or now with claude). Nvidia's moat is (and will remain for the foreseeable future) the entire stack. you cannot fathom the pain and misery of working on literally any other stack. if you've never debugged a GPU synchronization error or kernel panic due to some GPU firmware bug or fought absolute shit profilers hunting for perf you really have no idea what you're talking about.

EDIT: i can't believe this really requires saying but graphics and compute are not the same domain at all. if you work in graphics for GPU but not compute then you are still way out of your depth commenting. to wit: graphics people do not (and cannot) write CUDA kernels/shaders.

swernerabout 7 hours ago
If that is your standard, I do have an idea what I’m talking about.
mathisfun123about 7 hours ago
Ya? do tell us about your experience that leads you to believe mere translation is the bottleneck in the market...
swernerabout 6 hours ago
Most modern graphics is compute. Pixar, Dreamworks, Sony, etc do not use Vulkan to render their movies. It’s CPUs or CUDA.

“graphics people do not (and cannot) write CUDA kernels/shaders” is just not true at all. All it would take to verify that would be things like reading the introduction of the OptiX documentation, a small sample of SIGGRAPH GPU papers or the Blender/Cycles source code.

triwatsabout 2 hours ago
Interesting option for CDNA architecture chips. I wonder if this moves to an open standard?

AMD GPUs build for AI specs for reference: https://flopper.io/gpus?vendor=AMD&page=1

lulzxabout 9 hours ago
I made cuda-metal btw (for mac kek), https://github.com/lulzx/cuda-metal
srousseyabout 8 hours ago
What models can it run?
KennyBlankenabout 3 hours ago
To save everyone a click: No cuDNN and based off an ancient version of ROCm for windows (7.1 has been out for ages, 7.2 is current.)
Nuryssoabout 7 hours ago
man i can't say how much i used to like cuda when i had a nvidia gpu it made ml so much fun and on amd its a war especially on rdna 2 cards which i have. hopefully one day we will be able to properly translate cuda for its amd counter parts
system2about 10 hours ago
I wish there were a way to use RDNA1 cards with CUDA for AMD. My 5700XTs are sitting in a drawer.
monster_truckabout 9 hours ago
RDNA1 isn't good for a whole lot, even flagship RDNA2 cards are a stretch for many things. The lack of WMMA/matrix multiply/BF16 is too severe of a penalty.

The FP16 throughput on RDNA1 is both shader reliant and requires everything to be packed first. Even with 2 or 4 or 1000 cards, you would be consuming all of the available memory and memory bandwidth just packing and unpacking values, and if you really want to dump a hundred billion tokens into making it work anyways, you're only going to find out that even if you bother to sit there ferrying packed values to ram or disk before then issuing the instructions, paying that already severe penalty again when the values then have to be unpacked is so steep of a cost that the 256 BF16 flops/cu/clock's effective throughput is outright lower than simply doing it on a Zen 2 processor. You also don't have INT8 (or really INT4) on RDNA1 so the other RNS/CRT tricks aren't viable.

Sadly RDNA1's VCN2 also lacks actually good x264 bframe encoding support, or even P010 for 10 bit color, so what I'm saying is you should sell them. Used Radeon VII's are like $260, you'll go a lot further with those especially if you throw in a 7900XTX, and then augment that further with a 9070 CRE (you only want it for its int8 cores), and of course 128GB of ram.

E: And sure, that's 3, or ideally 4 GPUs, and a good bit of extra work. But that gets you up to more than halfway to the naive performance of a $15,000 MI300x in a surprising amount of cases, with additional strengths that it lacks. For far less than half of the cost

system2about 8 hours ago
I agree, I should've sold them last year when they hit $500 each. I have like 8 of them from my mining days. What a silly mistake I made.
dracotomesabout 7 hours ago
I still have a 5700XT I bought in 2019 (i think) for 300€ in my gaming rig. Crazy that they were going for $500 6 years later.
latchkeyabout 6 hours ago
There are also interesting efforts like:

https://github.com/Zaneham/Booth

https://scale-lang.com/

chiassedu80about 11 hours ago
CUDA for AMD on Windows

I’ve been working on a Windows setup that lets CUDA-targeted applications run on AMD GPUs using ZLUDA + ROCm/HIP.

Repo: https://github.com/Speedstu/CUDA-for-AMD-Windows

So far, it has only been tested on my RX 9060 XT (gfx1200), where I’ve used it with CUDA-enabled LibTorch workloads, including long ai training and use.

I also added a GPU scanner / auto-detection system that detects:

AMD GPU model

gfxXXXX architecture

ROCm/HIP installation

driver info

whether the GPU has already been validated by the project

Example:

RX 9060 XT → gfx1200 → RDNA4 → HIP detected → validated

The goal now is to test it on more hardware, especially RX 6000 / 7000 / 9000 cards.

If you have an AMD GPU on Windows and want to try it, I’d really appreciate compatibility reports working or broken. There’s a dedicated GPU compatibility issue template in the repo.

If this is useful to you, a star would also help the project get more testers.

nine_kabout 9 hours ago
(As a side note, I love the name ZLUDA; it very aptly means "delusion" or "deception" in Polish.)
swernerabout 6 hours ago
TIL, I didn’t know that. I always assumed it came from level zero”, the Intel computer layer that ZLUDA was translating to before its developer was hired by AMD to target HIP.