Abstract I consider the problem of learning from data corrupted by underrepresentation bias, where positive examples are filtered out at different, unknown rates for a fixed number of sensitive groups. I show that with a small amount of unbiased data, I can efficiently estimate the group-wise drop-out rates, even in settings where intersectional group membership makes learning each intersectional rate computationally infeasible. Using these estimates, I construct a reweighting scheme that allows me to approximate the loss of any hypothesis on the true distribution and present an algorithm encapsulating this process. Finally, I define a bespoke notion of PAC learnability for the underrepresentation and intersectional bias setting and show that my algorithm allows efficient learning for model classes of finite VC dimension.
Alexander Williams Tolbert (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: