|
|
|
|
|
by theturtletalks
1 day ago
|
|
I’ve heard of people using DeepSeek to reverse engineer those thinking tokens. Even though they don’t tell you the thinking content, they do tell you how many tokens were used. You can ask DeepSeek how we got from A + x thinking tokens = B. Claude and Codex flag this as a distillation attempt so you have to use an open weight model and the results are, of course, just a guess. |
|