How is changing history and making a a group of objectively white people black "removing bias" ˋ? What it is is literally bias. Like thinking Pi could be 4. Removing bias ends with truth, not these crazy wonky results.
If you're aruging about historical accuracy, but still want accurate looking generated images, I don't know what to say.
But to the technical point, A large part of the training corpus has biases that if left unchecked would cause PR based disasters for the company hosting it. ie the classic black teenager/white teenager.
Now as training of models is not an exact science, and neither is the fine tuning, its analogous to forcing a water balloon into a square box. Its possible but it has odd side effects when you get to the corners.
When making a _product_ you need to choose the least worse failure case. For grok it was for a long time, pandering to the ego of the owner. For Google, who is an advertising company, its about trying not to scare advertisers. This means everthing must be vanilla
So you have a huge number of photos of white people in the training data set, but other ethnicities exist. So to make the otherwise white-biased dataset less biased, you try to e.g. add a hidden system prompt that whenever the user asks for a group of people (unspecified ethnicity), it may instead ask for "mixed ethnicities" or whatever.
Ask for a group of Nazis, and that's it - this is how models work. No "LGBTQ liberal" propaganda is needed to explain it. Unlike what Musk is doing.
The training data and the end result should reflect the real world. No you don't need to "add bias to fight bias", just keep it truthful. Don't force SF-brained forced diversity/inclusion/whatever, we know well some models do.
> OpenAI invented a technique in July 2022 whereby its system would insert terms reflecting diversity (like “Black,” “female,” or “Asian”) into image-generation prompts in a way that was hidden from the user.
> Google’s Gemini system seems to do something similar, taking a user’s image-generation prompt (the instruction, such as “make a painting of the founding fathers”) and inserting terms for racial and gender diversity, such as “South Asian” or “non-binary” into the prompt
It seems then that their objective is to superficially increase the diversity of results, to avoid bad PR, rather than actually neutralising harmful and untruthful biases of the input data.
A model where asking for a math professor results in an image of an asian man, asking for an engineer in an image of a white man, and asking for a criminal results in an image of a black man
If you try to remove that in the name of "diversity" or being "less bigoted" you quickly end up with racially diverse nazis