>The number traces to Meta’s Llama 3 technical report, which documented 419 unforeseen disruptions across 16,384 H100s over 54 days of training, of which 148 were GPU failures and 72 were HBM3 memory failures.
From an annualized number on the llama 3 training report. would be interesting to see if we have a better idea given that we're already on rubin.
It depends entirely on the time-to-failure distribution though whether used cards are a good deal or not. Often this kind of hardware has a bathtub shaped hazard rate, actually getting burned in cards may mean you get the weeded out solid specimens, and forgo the lemons.
A bathtub curve is also memoryless though, and the back portion can be exponential. Consider two sub-populations of the same hardware. The first is small, evaded QA, and will fail early. Failures follow an exponential fall off. The second is much larger and failures rise exponentially with age.
Also consider mixing in "was dropped during shipping" or "was stored improperly".
My mistake, it seems I misunderstood what OP meant by that. Ignore that semantic detail and my position still stands. You've got a mix of subpopulations and many of the failure modes are definitely time dependent in some manner at least in the real world.
From an annualized number on the llama 3 training report. would be interesting to see if we have a better idea given that we're already on rubin.