Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

67% Positive

Analyzed from 701 words in the discussion.

Trending Topics

#rust#pytorch#python#architecture#should#gpu#transformer#don#said#version

Discussion (26 Comments)Read Original on HackerNews

JPLeRouzic•12 minutes ago
I have read somewhere that Transformer architecture has a quadratic cost (which explains the high costs associated with LLMs and the difficulty for constant improvement without state size pockets).

For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need the full KV in memory to generate a single token:

https://arxiv.org/abs/2503.00392

https://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b1...

janalsncm•about 3 hours ago
OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.

You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.

Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?

In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.

yjftsjthsd-h•about 2 hours ago
I dunno, I could probably be convinced to try a new tool purely on the basis of not having to deal with installing pytorch
IshKebab•about 1 hour ago
Yeah likewise. Pytorch needs to die.

That said I don't know what's wrong with using a Rust AI framework like Candle.

intoXbox•about 1 hour ago
I’m curious, what’s the criticism for PyTorch?
jeroenhd•19 minutes ago
I don't think it's caused by PyTorch on its own, but every AI-related Python project I try out locally manages to depend on a version of PyTorch that isn't in my disk cache yet. Having to download a gigabyte of dependencies for every project gets tiresome.

The Rust compile cycle will probably generate a gigabyte of files locally as well, but at least they can be `rm`'d out of `target/` once it's done.

It should be said that for this project that's entirely irrelevant of course, but seeing PyTorch has made me skip over projects on the HN homepage before and probably will again in the future.

tomtom1337•about 1 hour ago
One criticism is that you have to install the same package, torch, but from different Python indexes in order to install the cpu version or gpu version, on Linux. On windows, `pip install torch` gets you the cpu version. On linux, that gets you a ton of Nvidia extras that take a lot of space.

GPU support should really be a optional extra eg `torch[gpu]` or `torch[nvidia]`.

alightsoul•about 3 hours ago
Apparently op cares a lot about speed, which is fine, but ML researchers care about correctness first, speed second. And it makes sense, because they are not as resource constrained as OP.
janalsncm•about 2 hours ago
Most PyTorch tensor operations are cython not python. So imo rewriting in rust is not going to have an enormous speed up. If that really was the concern we should see a throughput comparison vs PyTorch or something.
nicman23•about 2 hours ago
yeah that is why they compute in fp32 lol
prospero_•about 1 hour ago
Apparently none of the people complaining about rust use read past the title because it's the 2nd section of the readme and impossible to miss.
kasumispencer2•about 1 hour ago
That part is newly added after people complained.

Edit: the claimed reason of "updates" also seems to not exist in code. If the author is going to use LLM for this, the very least they can do is to ask it to check properly before publishing.

hashar•about 1 hour ago
To quote the readme:

> The implementation language is a detail, and a Python port is welcome.

I guess the post title could have dropped "in rust"

mllev15•about 4 hours ago
You can just tell when the idea itself was generated
Marcuss2•about 2 hours ago
How would it compare to current state of the art tested state space layers like Kimi Delta Attention?
meredithbloom•about 1 hour ago
> The architecture is the claim here.

Weird turn of phrase very typical of AI.

saglogog•about 2 hours ago
Nice idea actually, I always wondered why there was no actual programming in ML!
a19486•about 3 hours ago
Homie just had some tokens to burn at the end of the month and “in Rust” is pure HN clickbait.
kasumispencer2•about 3 hours ago
Isn't this just RNN and nearly nothing about this is actually new?
purple-leafy•about 4 hours ago
“In Rust” man of all the cliche title baits, I hate this one the most.

Why did you choose Rust? Why does that matter?

Nothing wrong with Rust. Lots wrong with the bandwagon that “in Rust” somehow adds value.

janalsncm•about 3 hours ago
Rust has its place but python is default for this kind of thing. By doing it in rust, they have now changed two things: the implementation of the transformer, and this new model.
echelon•about 3 hours ago
Counterpoint: I only clicked on it because it said "in Rust".
IshKebab•about 1 hour ago
Rust matters because it means you can actually deploy it without going insane. I wonder if the anti-Rust zealots have actually ever used pip, especially on Windows.

He could have said "... not written in Python" - would that have been acceptable?

anon291•about 4 hours ago
Yeah I read the title and did a double-take. It's an uninteresting choice for systems like these. The more pressing concerns, which go undescribed, are the exact mathematical choices behind the actual model. Rust provides almost zero value here because tensor stuff is all just 2d-arrays of floats for the most part.
anon291•about 4 hours ago
Seems similar to a neural turing machine.
Advertisement
octoberfranklin•about 4 hours ago
vibe coded vibe coder!