|
|
|
|
|
by Diogenesian
23 days ago
|
|
This shouldn't be ignored in the discussion here: The job performed by the humans was broader than what was requested of the model in this benchmark: humans also had to find the relevant invoices (searching through mailboxes, or requesting them from providers) and reason through any circumstances which cannot be inferred from the bank feed and invoices/receipts on their own. In the benchmark these circumstances are presented to the model as “user notes."
This is precisely the kind of fine print on white-collar AI capability that companies keep running into: pretty much any non-entry office job worth having involves a lot of undocumented (even undocumentable) problems requiring judgment and experience.And I would be pretty nervous about asking any of the frontier LLMs to retrieve invoices: "cool, Claude logged that it found the May 6th bill from the paper supplier, I am sure it didn't just make something up arbitrary, then compound on the error by agentically iterating over the made-up invoice lurking in its reasoning traces. I checked the first 30 times and there were no problems!" |
|
edit: There's already a number of LLM which are intended for outgoing data loss protection to redact or prevent PII from escaping. Is anyone specifically working on a training set and agent that is specialized in reviewing "is this legit to pay", as a sub-task or filtering step in an AP workflow? I suppose it's a GIGO problem, as it would work best only if you have suppliers enrolled in some kind of existing db, with a specific contracted format for invoices, and correlating with project numbers/cost codes.