|
|
|
|
|
by ltononro
17 days ago
|
|
Worth noting/understanding, for those that are not familiar with it, that these 33k tokens are not a single 33k batch of tokens, they persist and increment in every request.. so if you have 10 requests that is 330k tokens, 100 requests 3.3M requests, 3.3M that could be ~5x less if used another harness.
These 33k tokens are mostly because the harness is completely bloated with mcp, skils, plugin, loop, basic alignment and all sort of explanations needed in the system prompt for it to work the way anthropic wants it to...
imo this is sub-optimal and if the model needs that much orientation, it says a lot about its intelligence too.. Good models should be as harness agnostic as possible |
|