| HN Mirror

Y	Hacker News new \| ask \| show \| jobs

by panarky 4276 days ago

The 100 terabyte benchmark used 206 Spark nodes, compared with 2100 Hadoop nodes.

Going up to 1 petabyte, the Hadoop comparison adds more nodes, 3800, while the Spark benchmark actually reduced the number of nodes to 190.

Does Spark scale well beyond ~200 nodes, or does the network become the bottleneck?

In any case, it's an impressive result considering that they didn't use Spark's in-memory cache.

2 comments

Lanzaa 4276 days ago

I believe the network had become a bottleneck. As per the article:

> [O]ur Spark cluster was able to sustain ... 1.1 GB/s/node network activity during the reduce phase, saturating the 10Gbps link available on these machines.

If the network is the bottleneck it makes sense to reduce the number of nodes to reduce the network communications.

link

rxin 4276 days ago

The job is actually very linearly scalable. i.e. running it on 200 nodes roughly doubles the throughput of 100 nodes.

link

rxin 4276 days ago

It is mainly the cost of getting nodes from EC2 at that point. It becomes hard to get a huge number of i2.8xl instances.

Spark runs fine on thousands of nodes.

link