|
|
|
|
|
by eso_logic
17 days ago
|
|
This initial round of benchmarking was to understand if there was any usecase here at all and I think there is. In a follow up, I'll be trying to answer questions like this. How big of a model can you fit on 4x M60, 4x P100, 4x V100? What are the tok/second when varying context length? Do you have a set of models you'd like me to look at? |
|