Hacker News new | ask | show | jobs
by mark_l_watson 4 days ago
This seems like a smart move, given their ability to host efficiently. I approve of efforts to make the cost of inference for smaller useful models slowly approach 'close to zero' and there are many good paths for getting there. It is useful for companies to get fast hosting for the class of smaller models they may end up hosting in house.
2 comments

I really hope they dont stop at the small models though! The bigger ones that dont fit on a single GPU are more interesting I think
Problem is right now the biggest GPU boxes they have is single rtx pro 6000s.