If a model can improve at drawing "X doing Y" and that prompt wasn't in the training set then it means it has improved its internal mapping from text-to-spatial-to-text.
The improvement would be limited to the template "X doing Y". It would not be a general improvement unless hundreds or thousands of diverse templates were used.