Then you need longer contexts, which is proving to a much more stubborn problem than general knowledge compression.