The selection bias is a major issue when using non-probability samples for inference about finite populations. Minimizing it requires sufficient auxiliary power from external sources and proper modelling. This article examines the potential bias of estimates of income distribution parameters using real-life data from a household survey (European Union Statistics on Income and Living Conditions, EU-SILC) and computer simulations. It investigates the change in the performance of the estimates depending on data defectiveness, the sample size, the pool of auxiliary variables, sources of auxiliary information, and the estimation method. The non-probability samples had non-ignorable bias intentionally introduced by designing selection mechanisms directly tied to income or the access to a computer and the Internet. The results indicated that the bias of the estimates behaved predictably for simple parameters, such as the median and mean. However, for more complex parameters, the bias patterns were less stable. The use of sociodemographic variables for weighting did not consistently reduce the bias; on the contrary, in certain cases, it led to the aggravation of the bias. The number of auxiliary variables and the sample size were as important for the bias reduction as the actual estimation methods, of which the doubly robust estimation performed best.
selection bias, non-probability sampling, non-ignorable mechanism, income inequality, data weighting
Baker, R., Brick, J. M., Bates, N. A., Battaglia, M., Couper, M. P., Dever, J. A., Gile, K. and Tourangeau, R., (2013). Report of the AAPOR Task Force on Non-probability Sampling. Technical report. Deerfield: The American Association for Public Opinion Research.
Beaumont, J.-F., (2020). Are probability surveys bound to disappear for the production of official statistics? Survey Methodology, 46(1), pp. 1–29.
Beaumont, J.-F., Bosa, K., Brennan, A., Charlebois, J. and Chu, K., (2024). Handling non-probability samples through inverse probability weighting with an application to Statistics Canada’s crowdsourcing data. Survey Methodology, 50(1), pp. 77–106.
Beaumont, J.-F., Rao, J. N. K., (2021). Pitfalls of making inferences from non-probability samples: Can data integration through probability samples provide remedies. The Survey Statistician, 83, pp. 11–22.
Brzeziński, M., Sałach, K. and Wroński, M., (2020). Wealth inequality in Central and Eastern Europe: Evidence from household survey and rich lists’ data combined. Economics of Transition and Institutional Change, 28(4), pp. 637–660. Available at: https://doi.org/10.1111/ecot.12257.
Buelens, B., Burger, J. and Brakel, J. A. van den, (2018). Comparing Inference Methods for Non-probability Samples. International Statistical Review, 86(2), pp. 322–343. Available at: https://doi.org/10.1111/insr.12253.
Carranza, R., Morgan, M. and Nolan, B., (2023). Top Income Adjustments and Inequality: An Investigation of the EU-SILC. Review of Income and Wealth, 69(3), pp. 725–754. Available at: https://doi.org/10.1111/roiw.12591.
Chen, J. K. T., Valliant, R. L. and Elliott, M. R., (2019). Calibrating non-probability surveys to estimated control totals using LASSO, with an application to political polling. Journal of the Royal Statistical Society: Series C (Applied Statistics), 68(3), pp. 657–681. Available at: https://doi.org/10.1111/rssc.12327.
Chen, Y., Li, P. and Wu, C., (2020). Doubly Robust Inference With Nonprobability Survey Samples. Journal of the American Statistical Association, 115(532), pp. 2011– 2021. Available at: https://doi.org/10.1080/01621459.2019.1677241.
Cochran, W.G. (1963) Sampling techniques. New York: Wiley.
Cornesse, C., Blom, A. G., Dutwin, D., Krosnick, J. A., Leeuw, E. D. D., Legleye, S., Pasek, J., Pennay, D., Phillips, B., Sakshaug, J. W., Struminskaya, B. and Wenz, A., (2020). A review of conceptual approaches and empirical evidence on probability and nonprobability sample survey research. Journal of Survey Statistics and Methodology, 8(1), pp. 4–36. Available at: https://doi.org/10.1093/jssam/smz041.
Deville, J.-C., Särndal, C.-E., (1992) Calibration estimators in survey sampling. Journal of the American Statistical Association, 87(418), pp. 376–382.
Elliott, M. R., Valliant, R., (2017) Inference for Nonprobability Samples. Statistical Science, 32(2), pp. 249–264.
European Commission, (2022). Methodological guidelines and description of EU-SILC target variables. Eurostat. Available at: https://ec.europa.eu/eurostat/documents/203647/16993001/Methodological+guidelines+2022+operation+v7.pdf/ec6bc779-6462-34aa-a8bd-256f8af34d31?t=1703153300474 (Accessed: 16 July 2024).
Eurostat, (2024). EU-SILC Microdata. Available at: https://ec.europa.eu/eurostat/web/microdata/european-union-statistics-on-income-and-living-conditions.
Ferri-García, R., Castro-Martín, L. and Rueda, M. del M., (2021). Evaluating Machine Learning methods for estimation in online surveys with superpopulation modeling. Mathematics and Computers in Simulation, 186, pp. 19–28. Available at: https://doi.org/10.1016/j.matcom.2020.03.005.
Ferri-García, R., Rueda-Sánchez, J. L., Rueda, M. del M. and Cobo, B., (2024). Estimating response propensities in nonprobability surveys using machine learning weighted models. Mathematics and Computers in Simulation, 225, pp. 779–793. Available at: https://doi.org/10.1016/j.matcom.2024.06.012.
Gelman, A., Little, T. C., (1997). Poststratification into Many Categories Using Hierarchical Logistic Regression. Survey Methodology, 23(2), pp. 127–135.
GUS, (2023). Incomes and living conditions of the population of Poland – report from the EU-SILC survey of 2021. Warszawa: GUS.
Hansen, M.H. and Hurwitz, W.N. (1943) On the theory of sampling from finite populations. Annals of Mathematical Statistics, 14(4), pp. 333–362. Available at: https://doi.org/10.2307/2235923.
Horvitz, D. G., Thompson, D. J., (1952). A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47(260), pp. 663–685.
Kalton, G., (2023). Probability vs. Nonprobability Sampling: From the Birth of Survey Sampling to the Present Day. Statistics in Transition new series, 24(3). Available at: https://doi.org/10.59170/stattrans-2023-029.
Kim, J. K., Kwon, Y., (2024). Comments on “Exchangeability assumption in propensityscore based adjustment methods for population mean estimation using non-probability samples”. Survey Methodology, 50(1), pp. 57–63.
Kim, J. K., Morikawa, K., (2023). An empirical likelihood approach to reduce selection bias in voluntary samples, arXiv Available at: https://doi.org/10.48550/arXiv.2211.02998.
Kim, J. K., Park, S., Chen, Y. and Wu, C., (2021). Combining non-probability and probability survey samples through mass imputation. Journal of the Royal Statistical Society. Series A: Statistics in Society, 184(3), pp. 941–963. Available at: https://doi.org/10.1111/rssa.12696.
Kish, L., (1995). Survey sampling. New York: John Wiley & Sons.
Lavrakas, P. J., Pennay, D., Neiger, D. and Phillips, B., (2022). Comparing Probability- Based Surveys and Nonprobability Online Panel Surveys in Australia: A Total Survey Error Perspective. Survey Research Methods, 16(2), pp. 241–266. Available at: https://doi.org/10.18148/srm/2022.v16i2.7907.
Lee, S., Valliant, R., (2009). Estimation for Volunteer Panel Web Surveys Using Propensity Score Adjustment and Calibration Adjustment. Sociological Methods & Research, 37(3), pp. 319–343. Available at: https://doi.org/10.1177/0049124108329643.
Li, Y., (2024). Exchangeability assumption in propensity-score based adjustment methods for population mean estimation using non-probability samples. Survey Methodology, 50(1), pp. 37–55.
Liu, A.-C., Scholtus, S. and De Waal, T., (2023). Correcting Selection Bias in Big Data by Pseudo-Weighting. Journal of Survey Statistics and Methodology, 11(5), pp. 1181–1203. Available at: https://doi.org/10.1093/jssam/smac029.
Liu, Z., Wang, D. and Pan, Y., (2025). Superpopulation model inference for non probability samples under informative sampling with high-dimensional data. Communications in Statistics – Theory and Methods, 54(5), pp. 1370–1390. Available at: https://doi.org/10.1080/03610926.2024.2335543.
Lohr, S. L., (2022). Sampling. Design and analysis. Third edition. London, New York: CRC Press.
Marella, D., (2023). Adjusting for Selection Bias in Nonprobability Samples by Empirical Likelihood Approach. Journal of Official Statistics, 39(2), pp. 151–172. Available at: https://doi.org/10.2478/jos-2023-0008.
Meng, X. L., (2018). Statistical paradises and paradoxes in big data (I): Law of large populations, big data paradox, and the 2016 us presidential election. Annals of Applied Statistics, 12(2), pp. 685–726. Available at: https://doi.org/10.1214/18-AOAS1161SF.
Mercer, A. W., Kreuter, F., Keeter, S. and Stuart, E. A., (2017). Theory and Practice in Nonprobability Surveys: Parallels between Causal Inference and Survey Inference. Public Opinion Quarterly, 81(S1), pp. 250–271. Available at: https://doi.org/10.1093/poq/nfw060.
Mercer, A. W., Lau, A., (2023). Comparing Two Types of Online Survey Samples, Pew Research Center Methods, 7 September. Available at: https://www.pewresearch.org/methods/2023/09/07/comparing-two-types-of-online-survey-samples/ (Accessed: 15 August 2024).
Neyman, J., (1934). On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection. Journal of the Royal Statistical Society, 97(4), pp. 558–625. Available at: https://doi.org/10.2307/2342192.
Rivers, D., (2006). Understanding People. Sample Matching. YouGovPolimetrix. Available at: http://www.websm.org/db/12/16528/Web%20Survey%20Bibliography/Understanding_people_Sample_matching/ (Accessed: 26 April 2023).
Rueda, M. del M., Pasadas-del-Amo, S., Rodríguez, B. C., Castro-Martín, L. and Ferri- García, R., (2023). Enhancing estimation methods for integrating probability and nonprobability survey samples with machine-learning techniques. An application to a Survey on the impact of the COVID-19 pandemic in Spain. Biometrical Journal, 65(2), p. 2200035. Available at: https://doi.org/10.1002/bimj.202200035.
Särndal, C.-E., Swensson, B. and Wretman, J. H., (1997). Model assisted survey sampling. New York: Springer.
Szreder, M., Kozłowski, A., (2024). Wnioskowanie na podstawie prób losowych i nielosowych. Gdańsk: Wydawnictwo Uniwersytetu Gdańskiego.
Tillé, Y., (2006). Sampling algorithms. New York: Springer.
Valliant, R., (2020). Comparing Alternatives for Estimation from Nonprobability Samples. Journal of Survey Statistics and Methodology, 8(2), pp. 231–263. Available at: https://doi.org/10.1093/jssam/smz003.
Valliant, R., Dever, J. A., (2011). Estimating Propensity Adjustments for Volunteer Web Surveys. Sociological Methods & Research, 40(1), pp. 105–137. Available at: https://doi.org/10.1177/0049124110392533.
Valliant, R., Dever, J. A. and Kreuter, F., (2013). Practical tools for designing and weighting survey samples. New York: Springer.
Valliant, R., Dorfman, A. H. and Royall, R. M., (2000). Finite population sampling and inference: A prediction approach. New York: John Wiley & Sons.
Wang, L., Valliant, R. and Li, Y., (2021). Adjusted logistic propensity weighting methods for population inference using nonprobability volunteer-based epidemiologic cohorts. Statistics in Medicine, 40(24), pp. 5237–5250. Available at: https://doi.org/10.1002/sim.9122.
Wang, W., Rothschild, D., Goel, S. and Gelman, A., (2015). Forecasting elections with non-representative polls. International Journal of Forecasting, 31(3), pp. 980–991. Available at: https://doi.org/10.1016/j.ijforecast.2014.06.001.
Wiśniowski, A., Sakshaug, J. W., Ruiz, D. A. P. and Blom, A. G., (2020). Integrating probability and nonprobability samples for survey inference. Journal of Survey Statistics and Methodology, 8(1), pp. 120–147. Available at: https://doi.org/10.1093/jssam/smz051.
Yang, S., Kim, J. K. and Song, R., (2020). Doubly Robust Inference when Combining Probability and Non-Probability Samples with High Dimensional Data. Journal of the Royal Statistical Society Series B: Statistical Methodology, 82(2), pp. 445–465. Available at: https://doi.org/10.1111/rssb.12354.
Zhang, L. C., (2019). On valid descriptive inference from non-probability sample, Statistical Theory and Related Fields, 3(2), pp. 103–113.