How did you assemble your training dataset, especially since you mention some of the considerations with different training mixes in your controlled eval?
Recapping here: Our first train was on 4,000 images from the MaskFactory dataset alone. This improved some benchmarks but regressed on others. We took this as a sign of narrow datasets causing unintended specialization.
In our next run, we assembled 26.1K images from 10 different datasets. We capped the amount of images that could come from one source, to prevent a single type of example from dominating. This composite set covered several cases like crowded scenes, camouflage, high-res subjects, fine objects like hair, blurred backgrounds, etc. We then shuffled everything together and trained FeyNoBg.
Recapping here: Our first train was on 4,000 images from the MaskFactory dataset alone. This improved some benchmarks but regressed on others. We took this as a sign of narrow datasets causing unintended specialization.
In our next run, we assembled 26.1K images from 10 different datasets. We capped the amount of images that could come from one source, to prevent a single type of example from dominating. This composite set covered several cases like crowded scenes, camouflage, high-res subjects, fine objects like hair, blurred backgrounds, etc. We then shuffled everything together and trained FeyNoBg.