|
|
|
|
|
by kianN
18 days ago
|
|
I see a fair number of comments here advocating for either codex to hand-roll this themselves, or to simply punt to SQL. I do want to advocate for the difficulty of the problem, even if I can't speak to the company itself. At the scale of a few hundred to a few thousand documents, especially short documents, there are a few out of the box methods that can yield reasonable results, whether it be embedding clustering or leveraging LLMs for tagging. However as your (1) datasets gets larger (2) documents expand from tweets and text messages to 30+ minute conversations and (3) you build downstream analytics on top of the learned semantic units, you really start to feel the limitations of LLMs and embedding for reliable annotation. That doesn't even get into the nuances associated with taxonomy management, seasonality, and model drift. TLDR; this problem solved effectively has a lot of value and is a lot harder than it seems. |
|
How is it easier to sign up and manage a different service, implement a different API, etc.
And from the company side the fatal flaw is that these types of tools rely upon 1% of their users having huge spend. Nobody is going to be a huge spender here because it's easier to hand roll than navigate procurement on this (not to mention impossible to justify the spend, additional security/privacy risk, etc.)
It feels approximately impossible for this company to have large accounts.