Have you used this method yourself for long workflows and with contexts approaching 1 million tokens? It doesn't work very well. LLM context is nearly half unusable.
I always let GPT 5.5/5.6 sol reach compatction. I'm only offered a 258K window, but I think once they release 1M to normal plan users, it will be usable across the whole 1M. At least with Opus I could use the whole 1M. I'd argue that the more context I use, the better performance gets. It has more info already available. I don't notice degradation.