Even in chinese trained models too. the amount of tokens it requires to communicate the same ideas is higher. it is just less efficient to tokenize the Chinese language because logographic languages are archaic, ridiculous, and inefficient compared to latin languages.