d-Matrix Claims 20x Bandwidth Density Over NVIDIA Rubin In New Chip

In AI inference, you have two phases of the workload: prefill and decode. In most cases, decode is overwhelmingly the more time-consuming portion of the workload because it’s strictly memory-bandwidth bound. You can have all the compute in the world, and that’s great for prefill, but decode doesn’t care. There have been various strategies
