Hacker News new | ask | show | jobs
by rsync 23 days ago
“They can't realistically do anything else to estimate it …”

Of course they can.

How do you think they get the MTBF for hard drives?

What they do is deploy thousands of them and wait for the first one to fail. Then they wait for the second and third ones to fail… And then they can construct a statistical abstraction for the entire population based on the very first degradations … and that doesn’t take long at all.

1 comments

If they want to avoid confounding variables to real world use, this could take years. Nobody wants to wait years after an innovation to sell their product, so they develop ways to speed up the testing and get a number.

A machine that presses a keyboard switch 500 times a second for several days straight is obviously not indicative of someone actually using the keyboard. But it'll get the "absolutely beat the snot out of it" number, which is usually good enough for marketing.

No, it doesn't take years and the testing is performed, roughly, as I just described it.

They build a test rig with 10,000 drives which they run continuously producing a total "drive hours" count and then wait for the first one or two or ten drives to fail.

X failures / 10,000,000 drive-hours will give you your MTBF, etc. ... and only takes ~40 days given a 10k drive test rig.

Some pictures from Seagate lab in Longmont:

https://cdn.mos.cms.futurecdn.net/yHqDagRDz9H5koWBefi24K-120...

https://cdn.mos.cms.futurecdn.net/EiZ9tA6p3UbaJFTaW75JdN.jpg

Yes, this is one of the hacks they use that are fundamentally inaccurate due to confounding variables like bathtub curves and different issues causing early failures and long-term failures...

Because, otherwise, they would need to run those drives for years and years with reasonable cycling and read/write pressure to report accurate numbers based on real world failure rates in real conditions.