Back to News
Advertisement
Advertisement

⚑ Community Insights

Discussion Sentiment

85% Positive

Analyzed from 783 words in the discussion.

Trending Topics

#simd#let#shuffle#portable#gpu#evaluator#rust#amp#executor#gpus

Discussion (32 Comments)Read Original on HackerNews

O3marchnativeβ€’41 minutes ago
The author mentions Rust's portable SIMD library [0]. The only issue with portable SIMD is it's only available on nightly. I used it in my FFT crate, but we had to switch to the fearless_simd crate in order to get a portable SIMD solution that works on stable [1].

[0] https://doc.rust-lang.org/std/simd/index.html

[1] https://github.com/linebender/fearless_simd

camel-cdrβ€’7 minutes ago
I love how ever example of portable SIMD isn't portable.

They specifies a constant SIMD width so it's non-portable. Well, not performance portable, but why are we using SIMD again?

6r17β€’about 1 hour ago
My heard hurts - i was stupid enough to think that SIMD was a CPU only thing - I don't understand why it would be ported to GPU - huge kudos to managing to surprise me
ismailmajβ€’7 minutes ago
There is something very SIMD-coded in GPU programming which is coalesced stores/loads, if a warp (32 threads) handles contiguous memory, it will create ~4 transactions instead of 32.
chlorionβ€’about 1 hour ago
GPUs work on vectors and matrices very often, that's what they are good at, so it makes a lot of sense that they can operate with SIMD I think!
hingler36β€’about 1 hour ago
Welcome to the lucky 10,000! SIMD is actually a pretty integral part of how GPUs are able to work efficiently, it's part of why there's such a strong focus on branchless programming in the field.
nynxβ€’39 minutes ago
Do you have examples of complex algorithms running on the gpu with rust with competative performance? Radix sort might be a good one to start with
LegNeatoβ€’about 2 hours ago
Author here, AMA.
lbhdcβ€’about 1 hour ago
What is vectorware's business model? Are you planning to sell support/consulting to companies using your stack? Or are you looking to sell licenses to your tool? Or something else?
LegNeatoβ€’about 1 hour ago
The tentative plan is to open source all the compiler and `std` bits with our products built on top (compilers are not good businesses). More about our products coming in the next couple of months!
lbhdcβ€’18 minutes ago
Looking forward to reading more about it. Good luck on the launch :)
jcranmerβ€’about 2 hours ago
The post is kind of vague on the IR you're targeting. Can you give some examples of what the SIMD-ized IR looks like, and how it maps to the target PTX?
the__alchemistβ€’about 1 hour ago
I'm confused too. How does this fit between these approaches for paraellization:

  - CUDA kernels and Tiles (e.g. Cudarc, cuda-oxide, rust-gpu etc) - SIMD on the GPU. (E.g. as in the title...)
  - CPU SIMD using avx or SSE instructions (And probably thin wrappers for vectors so you can have sane syntax). Or the maybe-upcoming core simd which should abstract over architecture-specific instructions. Magic floats etc which do 4-16 computations at once, but are a bit clumsy to work with
  - Rayon thread pools - arbitrary parallel computations, including SIMD, one per CPU core.
It looks like from the code samples like maybe a cleaner syntax for writing code on the GPU than CUDA kernels? E.g. without mucking with serialization, host and device by abstracting over it? And inspired by core::simd. (Good choice if so, in the interest of standardizing on syntax; I did this for my x86 SIMD vector/quaternion lib as well)
LegNeatoβ€’about 1 hour ago
Didn't want to go into crazy detail in the post.

Each family of operations is a trait parameterized by the operation itself:

  pub trait EvaluateReduction<Operation, T>: LaneEvaluator {
      /// Reduce one distributed definition to an ordinary uniform scalar.
      fn evaluate_reduction(&self, value: LaneValue<Self, role::Distributed, T>) -> T;
  }

Call sites name the operation:

  let one   = evaluator.splat::<Splat, _>(1_u32);
  let two   = evaluator.splat::<Splat, _>(2_u32);
  let three = evaluator.binary::<Add, _>(one, two);

  let total   = evaluator.reduce::<Sum, u32>(three);   // a uniform u32
  let running = <Executor as EvaluateScan<Scan<Sum, Exclusive>, u32>>::scan(&evaluator, three);

Operations like Sum, Max, ReduceXor, Inclusive, and Exclusive are all distinct types.

As mentioned in the post, execution shape is typed too. A static shuffle takes its control as a type-level constant, and the shuffle mode constrains which controls are expressible:

  // Shift down one lane, keeping our own value where the source is inactive.
  let down  = <Executor as EvaluateShuffle<Shuffle<Down>, DownOrSelf<1>, u32>>::shuffle(&ev, v);
  // Broadcast from lane zero.
  let bcast = <Executor as EvaluateShuffle<Shuffle<Broadcast>, WarpLane<0>, u32>>::shuffle(&ev, down);
  // Butterfly exchange with the neighbor one bit away.
  let bfly  = <Executor as EvaluateShuffle<Shuffle<Xor>, Butterfly<1>, u32>>::shuffle(&ev, bcast);

For an example of errors caught, a warp-scoped executor for a device-scoped barrier is a compile error:

  <ScopedWarpExecutor<'_, WarpUniform> as EvaluateBarrier<Barrier<Device>>>::barrier(evaluator)
  // error[E0277]: the trait bound `Device: NvptxBarrierScope` is not satisfied
  //               help: the trait `NvptxBarrierScope` is implemented for `Warp`

Strip mining is typed on the amount of work and the lane capacity, and it hands back one chunk at a time along with the predicate saying which lanes live in that chunk:

  // Six work items across four active lanes: two chunks, based at 0 and 4.
  <Executor as EvaluateStripMine<StripMine, (WorkItems, ActiveLanes<StripMined<4>>), i32>>::
      for_each_strip_mined(
          &evaluator,
          (WorkItems::new(6)?, ActiveLanes::new(4)?),
          |index, active| {
           // ...
          },
      );

Hopefully that gives the flavor of it.
lbhdcβ€’about 1 hour ago
This is really cool! It sounds like y'all have a compiler fork that you are using to make this work. I wanna tinker with this, is your compiler available?
LegNeatoβ€’38 minutes ago
It is not currently available but we intend to make it available after we launch our products.
Eridrusβ€’about 1 hour ago
Given the massive demand for GPUs for LLMs, what sorts of work do you expect to economically benefit from utilizing GPUs more?
LegNeatoβ€’about 1 hour ago
Part of our thesis is that decent GPUs are in every shipping device and most software doesn't use them and should.
Eridrusβ€’29 minutes ago
I guess you're looking at consumer hardware then since servers have exactly what you pay for.

Can you say more about the application space you're targeting?

efnxβ€’about 2 hours ago
Congrats to the Rust-GPU folks! Nice to see the good work flowing.
rust-langβ€’about 1 hour ago
Good job!
the__alchemistβ€’about 1 hour ago
Hey - this is probably off-topic/meta, but what is going on with the comments here? Is it bots?
dev_l1x_beβ€’about 1 hour ago
No idea, but it seems HN needs POW challenges.
lukanβ€’about 1 hour ago
Could also just be trolls attracted by the Rust topic.
donald-trumpβ€’2 minutes ago
100% tariff on trolls