| I respect Knaup's opinion but disagree with the claim he makes here. Kubernetes took off because everyone could run it on pretty much anything. Bigger hardware meant bigger clusters, but developers could spin up a cluster on their laptops and deploy their apps into it. More importantly, companies could repurpose their decommissioned servers as k8s clusters, a massive unlock seeing how huge companies had heaps of these in their racks. This isn't possible with open weights models. Not in the same way. First, you're out of the game if you don't have a data center class GPU (or it's sort of prosumer equivalent). Model servers support CPU inference, but you might as well watch paint dry as you wait for results...and you'll still have to run super quantized low-parameter models that aren't as good. Realistically, companies will need to purchase millions of dollars of nVIDIA gear (through suppliers) to serve agentic-capable models at scale, an activity that is being made more expensive and complicated by the day as the hyperscaleds slurp up the demand. It also really is a huge problem that all of the open weights models are coming out of one country that also happens to be a superpower. I don't think this can be handwaved away, and it's concerning to see so many folks here minimize this. American companies running on Chinese intelligence. As a country that prides itself on being the knowledge capital of the world, the optics manifested by this are horrible, not to mention the absolutely massive supply chain risks (the counterfeit Cisco devices comes to mind, except worse because the rangers are in the weights and might not be possible to distill out). This is a threat even if you focus solely on the individual developer. Recall how AWS became...AWS. They "fanatically" focused on the developer experience. They designated this as key to their growth strategy, and rightfully so. How people talk about using Qwen or Kimi and the like on here feels like that (ignoring how these models are ALSO from huge for-profits). The only solution here is for the big labs to make some of their most capable models open-weights. There is a lot of secret sauce around routing, inference, hosting and other stuff that makes them work as well as they do, but the community can figure that out. Of course, this is basically a death sentence to those companies, but that's what they get for playing with monkey paws, I guess. |