https://arxiv.org/pdf/2608.10605: Our new paper folds the systems stage into the scaling-law stage. Price every candidate architecture on what the cluster actually delivers, and the answer changes: the sparsity an MoE 'should' have depends on the cluster you train it on.
Advertisement
Advertisement

Discussion (0 Comments)Read Original on HackerNews
No comments available or they could not be loaded.