| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by skrebbel 152 days ago
	How does this work? Do they buy lots of openai credits and then hit their api billions of times and somehow try to train on the results?

1 comments

g-mork 152 days ago

dont forget the plethora of middleman chat services with liberal logging policies. i've no doubt there is a whole subindustry lurking in here

link

skrebbel 152 days ago

i wasn't judging, i was asking how it works. why would openai/anthrophic/google let a competitor scrape their results in sufficient amounts that it lets them train their own thing?

link

victorbjorklund 152 days ago

I think the point is that they can't really stop it. Let's say that I purchase API credits, and I let the resell it to DeepSeek.

That's going to be pretty hard for OpenAI to figure out and even if they figure it out and they stop me there will be thousands of other companies willing to do that arbitrage. (Just for the record, I'm not doing this, but I'm sure people are.)

They would need to be very restrictive about who is allowed to use the API and not and that would kill their growth because because then customers would just go to Google or another provider that is less restrictive.

link

skrebbel 151 days ago

Yeah but are we all just speculating or is it accepted knowledge that this is actually happening?

link

sally_glance 151 days ago

Speculation I think, because for one those supposed proxy providers would have to provide some kind of pricing advantage compared to the original provider. Maybe I missed them but where are the X0% cheaper SOTA model proxies?

Number two I'm not sure if random samples collected over even a moderately large number of users does make a great base of training examples for distillation. I would expect they need some more focused samples over very specific areas to achieve good results.

link

skrebbel 151 days ago

Thanks I that case my conclusion is that all the people saying that these models are "distilling SOTA models" are, by extension, also speculating. How can you distill what you don't have?

link

mike_hearn 151 days ago

OpenAI implemented ID verification for their API at some point and I think they stated that this was the reason.

link