|
|
|
|
|
by russlan
24 days ago
|
|
generic model benchmarks are less useful than workflow benchmarks. Nowadays it's all about workflows , think prompt injections, tool permissions is a big one, logs and lately what anthropic claims is the model distillation risks |
|