|
|
|
|
|
by piterrro
20 days ago
|
|
For a while now, everytime I see „for AI agents” Im looking for benchmarks with comparable solutions. The only „for AI agent” solution that will gain adoption is/should be measured by two dimensions: token usage and correctness. If the solution doesnt use less tokens than generating chart.js code for example - why should I use it? Same for correctness, if the generated chart spec is correct only 90% of the time - why should I use it? It’s still early but I think this is the direction we should be thinking about „for AI agents” libraries and projects. |
|