|
|
|
|
|
by jhhh
29 days ago
|
|
The article presents two possible methods to protect privacy in a dataset. It then attacks a theoretical weakness in a contrived scenario in the old method, which intends to incline us to choose the other, newer solution. The article does not describe in detail the newer solution beyond its name. I have have questions that the article doesn't cover in enough detail: 1) Has the coarsening failed in the way described in the article in practice which leaked information? 2) How does the 'other' solution we're expected to desire work? 3) What is an example of the difference in level of detail offered by the newer solution that was not possible when the data needed to be coarsened in practice? |
|
(2) By adding carefully tuned Gaussian noise. In the last 6 years we have also figured out how to add much less Gaussian noise: "The 2020 Census Disclosure Avoidance System TopDown Algorithm" https://arxiv.org/abs/2204.08986
(3) This one is harder to answer, since the Census Bureau aimed to release the same style of statistics as in previous decades. So the goal of 2020 was to release the same statistics with the same error bounds. Evidence suggests they succeeded in doing this. "Evaluating Bias and Noise Induced by the U.S. Census Bureau's Privacy Protection Methods" http://arxiv.org/abs/2306.07521, "Evaluating the Impacts of Swapping on the US Decennial Census" http://arxiv.org/abs/2502.01320