p.enthalabs

SambaNova’s SN50 RDU for AI at Hot Chips 2026

![Image 1: SambaNova SN50 Hot Chips 2026 SN50 Dataflow](https://www.servethehome.com/wp-content/uploads/2026/08/SambaNova-SN50-Hot-Chips-2026-SN50-Dataflow.png)

SambaNova SN50 Hot Chips 2026 SN50 Dataflow

For as new as the dedicated AI accelerator field is, SambaNova is one of the older and more established hardware vendors. The company is now in the fifth generation of their reconfigurable dataflow unit (RDU) technology, with the SN50 that was launched earlier this year. As with the other major AI vendors at this year’s Hot Chips conference, the company has come to present new technical details on SN50, and outline what makes it competitive in the burgeoning field of dedicated AI accelerators.

SambaNova’s hardware has taken on an increased prominence in the industry thanks to the company’s connection to Intel. While Intel itself is still trying to catch up on AI accelerators, the company has become increasingly attached to the hip to SambaNova, whose RDUs provide the dedicated, high-efficiency and low-latency AI accelerators that round out Intel’s hardware stack. Thus the company’s progress with the SN50 (and future RDUs) is material not just for SambaNova, but for Intel as well.

![Image 2: SambaNova SN50 Hot Chips 2026 Inference Decode](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-inference-decode/)

SambaNova SN50 Hot Chips 2026 Inference Decode

Setting the stage, agentic inference is all the rage right now. Where does all the execution time go? SambaNova has a breakdown of it. Most time is spent in decode, especially on DeepSeek V3 where it’s 97% of the time, versus 3% for prefill.

![Image 3: SambaNova SN50 Hot Chips 2026 Decode Bandwidth Bound](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-decode-bandwidth-bound/)

SambaNova SN50 Hot Chips 2026 Decode Bandwidth Bound

And decode, in turn, is bandwidth-bound. The FLOPS-per-byte ratio is quite low, even for large batch sizes.

![Image 4: SambaNova SN50 Hot Chips 2026 HBM Bandwidth Utilization](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-hbm-bandwidth-utilization/)

SambaNova SN50 Hot Chips 2026 HBM Bandwidth Utilization

Bandwidth utilization is often misunderstood. SambaNova is laying out what they mean for this talk. In short, they aren’t talking about just how much HBM bandwidth is being used, but rather the Model Bandwidth Utilization (MBU) model. And specifically, what fraction of that is being used to cache data and otherwise handle data usage.

![Image 5: SambaNova SN50 Hot Chips 2026 Model Bandwidth Utilization](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-model-bandwidth-utilization/)

SambaNova SN50 Hot Chips 2026 Model Bandwidth Utilization

Looking at the current state of tech, GPUs offer low bandwidth usage, even with highly optimized GPU-friendly benchmarks.

![Image 6: SambaNova SN50 Hot Chips 2026 GPU Scaling](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-gpu-scaling/)

SambaNova SN50 Hot Chips 2026 GPU Scaling

Things get worse for GPUs when you scale up the number of them; performance does go up, but MBU drops significantly.

![Image 7: SambaNova SN50 Hot Chips 2026 Frontier Models](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-frontier-models/)

SambaNova SN50 Hot Chips 2026 Frontier Models

Meanwhile frontier models require being able to scale up.

![Image 8: SambaNova SN50 Hot Chips 2026 Power Capacity](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-power-capacity/)

SambaNova SN50 Hot Chips 2026 Power Capacity

Again with a GPU example, a GPU can get to around 30TB/second of model bandwidth. But they can’t get past that.

![Image 9: SambaNova SN50 Hot Chips 2026 SN50 Dataflow](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-sn50-dataflow/)

SambaNova SN50 Hot Chips 2026 SN50 Dataflow

Enter SambaNova’s SN50 dataflow RBU. They have doubled-down on what worked well from SN40, such as the large on-chip SRAM. 5x as many FLOPS as SN40, and it is designed to scale-up to a much larger domain of 256+ chips. And there is a separate scale-out network using 400Gb networking.

Notably, there are no I/O dies or similar here. Instead it is just two max reticle dies for the logic, and then HBM stacks for the memory.

Though it is interesting that SambaNova’s choice of HBM here is quite dated; SN50 still uses HBM2e here (which is going to be a problem in the future as production of the memory is already ramping down).

![Image 10: SambaNova SN50 Hot Chips 2026 SN50 Rack](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-sn50-rack/)

SambaNova SN50 Hot Chips 2026 SN50 Rack

Moving up to the SN50 rack architecture, there are 16 RDUs in a single air-cooled rack, split over two nodes.

![Image 11: SambaNova SN50 Hot Chips 2026 Chip Architecture](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-chip-architecture/)

SambaNova SN50 Hot Chips 2026 Chip Architecture

Diving a bit deeper into SN50 and the dataflow architecture. The core element of the SN50 is the sea of compute cores (PCUs) and memory cores (PMUs). There is no hardware memory management; this is all software managed. Every unit operates when it has input and sends it to the outputs.

![Image 12: SambaNova SN50 Hot Chips 2026 Transformer Structure](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-transformer-structure/)

SambaNova SN50 Hot Chips 2026 Transformer Structure

To better illustrate how the dataflow architecture works, here is an example of how it maps to a transformer.

![Image 13: SambaNova SN50 Hot Chips 2026 GPU Transformers](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-gpu-transformers/)

SambaNova SN50 Hot Chips 2026 GPU Transformers

Here is what a GPU looks like.

![Image 14: SambaNova SN50 Hot Chips 2026 SN50 Transformers](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-sn50-transformers/)

SambaNova SN50 Hot Chips 2026 SN50 Transformers

And how it look on the SN50.

![Image 15: SambaNova SN50 Hot Chips 2026 Compute Comms Overlap](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-compute-comms-overlap/)

SambaNova SN50 Hot Chips 2026 Compute Comms Overlap

For compute, data from the HBM is fed into the AGCU portal that control off-chip access, and from there into the PCUs and PMUs.

![Image 16: SambaNova SN50 Hot Chips 2026 Double-Buffered PMUs](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-double-buffered-pmus/)

SambaNova SN50 Hot Chips 2026 Double-Buffered PMUs

The SRAM amount used is not a function of the size of the model.

![Image 17: SambaNova SN50 Hot Chips 2026 No Global Synchronization](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-no-global-synchronization/)

SambaNova SN50 Hot Chips 2026 No Global Synchronization

![Image 18: SambaNova SN50 Hot Chips 2026 SN50 Scaling](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-sn50-scaling/)

SambaNova SN50 Hot Chips 2026 SN50 Scaling

That was one RDU. How do things scale up for multiple RDUs? SambaNova employs both scale-up and scale-out networking. The Scale-up network is based on 800GbE, while scale-out is 400GbE. And then there is a front-end network.

![Image 19: SambaNova SN50 Hot Chips 2026 Scale-up and Scale-Out](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-scale-up-and-scale-out/)

SambaNova SN50 Hot Chips 2026 Scale-up and Scale-Out

SambaNova uses an all-to-all topology for an 8 socket configuration.

![Image 20: SambaNova SN50 Hot Chips 2026 Scaling Beyond One Node](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-scaling-beyond-one-node/)

SambaNova SN50 Hot Chips 2026 Scaling Beyond One Node

To go above 8 sockets, then the scale-up network is employed using Ethernet switches. Links are ganged, and every node is connected to each of two switches in this 64 chip socket configuration.

![Image 21: SambaNova SN50 Hot Chips 2026 512 Socket Example](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-512-socket-example/)

SambaNova SN50 Hot Chips 2026 512 Socket Example

Then things can be scaled out further, in this case employing both scale-up and out for a 512 socket configuration.

![Image 22: SambaNova SN50 Hot Chips 2026 Model Parallelism](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-model-parallelism/)

SambaNova SN50 Hot Chips 2026 Model Parallelism

The key to performance on SN50 is overlap. SN50 supports all forms of model parallelism, and the collective communication forms that these models are built on.

![Image 23: SambaNova SN50 Hot Chips 2026 GEMM Benchmarking](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-gemm-benchmarking/)

SambaNova SN50 Hot Chips 2026 GEMM Benchmarking

Here is a brief look at performance with GEMM benchmarking. The utilization is consistently 70% or higher even at 32 sockets.

![Image 24: SambaNova SN50 Hot Chips 2026 Importance of Overlap](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-importance-of-overlap/)

SambaNova SN50 Hot Chips 2026 Importance of Overlap

If you are able to overlap, you can do the compute and communications in parallel. That kind of overlap is not something GPUs can do.

![Image 25: SambaNova SN50 Hot Chips 2026 HBM and Compute Overlap](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-hbm-and-compute-overlap/)

SambaNova SN50 Hot Chips 2026 HBM and Compute Overlap

The building block for SambaNova is collective communication, which is the purple boxes in these diagrams. And the SRAMs can stream from one to another without having to go through a higher layer (e.g. HBM).

![Image 26: SambaNova SN50 Hot Chips 2026 MoEs with TP](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-moes-with-tp/)

SambaNova SN50 Hot Chips 2026 MoEs with TP

Here’s a look at parallelism with tensor parallel.

![Image 27: SambaNova SN50 Hot Chips 2026 DeepSeek TP-32](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-deepseek-tp-32/)

SambaNova SN50 Hot Chips 2026 DeepSeek TP-32

Here is a look at the bandwidth utilization that SN has achieved with DeepSeek.

![Image 28: SambaNova SN50 Hot Chips 2026 MoEs with TP and EP](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-moes-with-tp-and-ep/)

SambaNova SN50 Hot Chips 2026 MoEs with TP and EP

Meanwhile they can also use expert parallel (EP) as an additional form of parallelism. This relies on broadcast-dispatch as well as all-to-all dispatch-combine.

![Image 29: SambaNova SN50 Hot Chips 2026 Token Dispatch](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-token-dispatch/)

SambaNova SN50 Hot Chips 2026 Token Dispatch

With all-to-all, one way is to dynamically send everything to the target RDUs.

![Image 30: SambaNova SN50 Hot Chips 2026 Dispatch Broadcast](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-dispatch-broadcast/)

SambaNova SN50 Hot Chips 2026 Dispatch Broadcast

Alternatively, you can just blast everything to all of the RDUs and then filter out things afterwards.

![Image 31: SambaNova SN50 Hot Chips 2026 Dispatch All-to-All](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-dispatch-all-to-all/)

SambaNova SN50 Hot Chips 2026 Dispatch All-to-All

The all-to-all method requires a group-by operation at the end of the router to collect (group) the tokens before transmitting them SRAM-to-SRAM. All-to-all also means allowing dynamic traffic.

![Image 32: SambaNova SN50 Hot Chips 2026 Dispatch Broadcast and Filter](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-dispatch-broadcast-and-filter/)

SambaNova SN50 Hot Chips 2026 Dispatch Broadcast and Filter

Now here’s the other method of broadcast + filter. That is still an SRAM-to-SRAM operation, but with a filter operation on the PCUs of the receiving RDU. This keeps the network traffic parallel; though it does increase it a bit. And by not depending on the router, the transfer can be started early.

![Image 33: SambaNova SN50 Hot Chips 2026 EP on 64-Socket SN50](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-ep-on-64-socket-sn50/)

SambaNova SN50 Hot Chips 2026 EP on 64-Socket SN50

Here is another DeepSeek example, with SambaNova getting close to 80% bandwidth utilization for loading the experts in MoE.

![Image 34: SambaNova SN50 Hot Chips 2026 High MBU at Scale](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-high-mbu-at-scale/)

SambaNova SN50 Hot Chips 2026 High MBU at Scale

As a result of this, SN50 achieves a high MBU value even at scale, with MBU holding at 45% even with 256 SN50s. And this makes it possible to keep adding RDUs to scale up things even further. This, in turn, means that models don’t have to give up bandwidth.

![Image 35: SambaNova SN50 Hot Chips 2026 SN50 Power Scaling](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-sn50-power-scaling/)

SambaNova SN50 Hot Chips 2026 SN50 Power Scaling

Going back to SambaNova’s original chart about power scaling, here is what SN50 clusters of different sizes look like. A 512 RDU configuration is able to scale up to an aggregate model bandwidth capacity of over 350 TB/second. The systems can strongly scale, with MBUs still in the 40% range at 512 sockets.

![Image 36: SambaNova SN50 Hot Chips 2026 SN50 Heterogeneous Disaggregation](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-sn50-heterogeneous-disaggregation/)

SambaNova SN50 Hot Chips 2026 SN50 Heterogeneous Disaggregation

Ultimately SambaNova is promoting a very similar picture as other dedicated inference chip firms, using one type of chips for prefill (and midfill), while using separate accelerators (i.e. SN50) for decode. Specifically, they’ve been using NVIDIA H200 + SN50, with RoCE for transferring between them.

![Image 37: SambaNova SN50 Hot Chips 2026 SN50 Heterogeneous Disaggregation In Action](https://www.servethehome.com/sambanovas-sn50-rdu-for-ai-at-hot-chips-2026/sambanova-sn50-hot-chips-2026-sn50-heterogeneous-disaggregation-in-action/)

SambaNova SN50 Hot Chips 2026 SN50 Heterogeneous Disaggregation In Action

Finally, taking a look at that performance in action, based on an Artificial Analysis benchmark of SN50. The hardware achieves over 750 tokens-per-second in MiniMax M2.7.