Women earn less than men in every labour force survey India has run. The Oaxaca–Blinder decomposition uses two wage regressions to split that gap into the part that comes from differences in education, experience, sector and location, and the part that comes from the same characteristics being paid differently. This page runs it in the browser on your own microdata, with the choices that change the answer made explicit: which group's coefficients are the reference, whether to use the two-fold or three-fold form, and how much of the “unexplained” part depends on how categorical variables are coded.
The gap
Generate or load data, then decompose.
Two-fold decomposition
What is being computed
With ȳA − ȳB the mean gap in log wages (men minus women by default), and separate OLS fits βA, βB, the two-fold decomposition is (X̄A − X̄B)′β* + [X̄A′(βA − β*) + X̄B′(β* − βB)]. The first term is the part of the gap due to differences in characteristics valued at the reference prices β* (the “explained” or endowments part); the second is due to differences in returns, including the intercept (the “unexplained” part, sometimes read as discrimination, which it is only if the model has every productive characteristic in it). β* from the pooled regression with a group dummy is Jann's (2008) recommendation, because omitting the dummy lets the group difference leak into the slope coefficients. Detailed contributions are shown for each variable; for dummies the unexplained detail depends on which category is the base (Oaxaca and Ransom 1999), which is why only the totals are invariant.
Three-fold decomposition from the disadvantaged group's viewpoint
What is being computed
Endowments E = (X̄A − X̄B)′βB, coefficients C = X̄B′(βA − βB), and the interaction I = (X̄A − X̄B)′(βA − βB): what women would gain with men's characteristics at women's prices, what they would gain with men's prices at their own characteristics, and the cross term that belongs to neither.