Equidad en aprendizaje automático: comparación de intervenciones de preprocesamiento y entrenamiento en un caso colombiano

Autores/as

DOI:

https://doi.org/10.61799/2216-0388.2338

Palabras clave:

aprendizaje automático, destilación de conocimiento, equidad algorítmica, inteligencia artificial, toma de decisiones

Resumen

Esta investigación tuvo como objetivo evaluar si la equidad alcanzada por clasificadores tradicionales mediante intervenciones de preprocesamiento puede transferirse a un perceptrón multicapa mediante destilación de conocimiento, y si el uso del regularizador MinDiff aporta una mejora adicional en la reducción de la disparidad de falsos negativos entre sexos en un clasificador de acceso al Beneficio de Inserción Económica en Colombia. Para ello se aplicó un diseño experimental por fases con registros administrativos de personas desmovilizadas: con una arquitectura de perceptrón multicapa fija se evaluaron diferentes intervenciones que combinaron reponderación, análisis de componentes principales, MinDiff, adversarial debiasing y seis estrategias de destilación desde dos maestros mitigados, con cinco semillas, tres particiones estratificadas, intervalos de confianza BCa, prueba de Wilcoxon con corrección de Holm y frontera de Pareto. Los experimentos mostraron que la destilación de los modelos maestros previamente analizados produjo una disminución en la diferencia entre tasas de falsos negativos, en comparación con el perceptrón multicapa de referencia. La destilación redujo la diferencia entre las tasas de falsos negativos respecto al MLP de referencia. Los menores valores se obtuvieron al combinar las probabilidades de los maestros. Con ponderación por exactitud y MinDiff, FNR_DIF pasó de 0.000982 a 0.000502, mientras el recall se mantuvo en 0,999598. La reducción de FNR_DIF fue descriptiva entre las ejecuciones; en DPD y recall sí se obtuvieron diferencias estadísticamente significativas. Accuracy presentó una disminución moderada.

Descargas

Los datos de descarga aún no están disponibles.

Referencias

[1]

T. Hellstrom, V. Dignum y S. Bensch, «Bias in Machine Learning What is it Good for?,» arXiv, pp. 1-8, 2020. doi:https://doi.org/10.48550/arXiv.2004.00686

[2] K. Makhlouf, S. Zhioua y C. Palamidessi, «Machine learning fairness notions: Bridging the gap with real-world applications,» Information Processing & Management, pp. 1-32, 2021. doi:https://doi.org/10.1016/j.ipm.2021.102642

[3] S. Caton y C. Haas, «Fairness in Machine Learning: a Survey,» arXiv, pp. 1-33, 2020. doi:https://doi.org/10.48550/arXiv.2010.04053

[4] N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman y A. Galstyan, «A Survey on Bias and Fairness in Machine Learning,» arXiv, pp. 1-31, 2019. doi:https://doi.org/10.48550/arXiv.1908.09635

[5] M. Hort, Z. Chen, J. Zhang, M. Harman y F. Sarro, «Bias Mitigation for Machine Learning Classifiers: A Comprehensive Survey,» ACM Journal on Responsible Computing, pp. 1-52, 2024. doi: https://doi.org/10.1145/3631326

[6] A. Agarwal y H. Agarwal, «A seven-layer model with checklists for standardising fairness assessment throughout the AI lifecycle,» AI and Ethics, pp. 299-314, 2024. doi:https://doi.org/10.1007/s43681-023-00266-9

[7] A. Olteanu, C. Castillo, F. Diaz y E. Kıcıman, «Social Data: Biases, Methodological Pitfalls, and Ethical Boundaries,» Frontiers in Big Data, pp. 1-33, 2019. doi:https://doi.org/10.3389/fdata.2019.00013

[8] S. Verma y J. Rubin, «Fairness Definitions Explained,» de International Conference on Software Engineering, Gotemburgo, 2018 doi:https://doi.org/10.1145/3194770.3194776.

[9] R. Burke, «Multisided Fairness for Recommendation,» de Fairness, Accountability, and Transparency in Machine Learning, Halifax, 2017. doi:https://doi.org/10.48550/arXiv.1707.00093

[10] S. Barocas, E. Bradley, V. Honavar y F. Provost, «Big Data, Data Science, and Civil Rights,» Computing Community Consortium (CCC), pp. 1-8, 2017. doi:https://doi.org/10.48550/arXiv.1706.03102

[11] I. Zliobaite, «Fairness-aware machine learning: a perspective,» ArXiv, pp. 1-10, 2017. doi:https://doi.org/10.48550/arXiv.1708.00754

[12] B. Richardson y J. Gilbert, «A Framework for Fairness: A Systematic Review of Existing Fair AI Solutions,» arXiv, pp. 1-28, 2021. doi:https://doi.org/10.48550/arXiv.2112.05700

[13] D. Danks y A. London, «Algorithmic Bias in Autonomous Systems,» de Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, Melbourne, 2017. doi:https://doi.org/10.24963/ijcai.2017/654

[14] T. Salazar, M. Seoane, H. Araújo y P. Henriques, «FAWOS: Fairness-Aware Oversampling Algorithm Based on Distributions of Sensitive Attributes,» IEEE Access, pp. 81370-81379, 2021. doi:https://doi.org/10.1109/ACCESS.2021.3084121

[15] A. Rosado, M. L. Calderón y O. Espinosa, «Data preprocessing to improve fairness in machine learning models: An application to the reintegration process of demobilized members of armed groups in Colombia,» Applied Soft Computing, pp. 1-14, 2024. doi:https://doi.org/10.1016/j.asoc.2023.111193

[16] G. Hinton, O. Vinyals y J. Dean, «Distilling the Knowledge in a Neural Network,» arXiv, pp. 1-9, 2015. doi:https://doi.org/10.48550/arXiv.1503.02531

[17] J. Gou, B. Yu, S. Maybank y D. Tao, «Knowledge Distillation: A Survey,» International Journal of Computer Vision, p. 1789–1819, 2021. doi:https://doi.org/10.1007/s11263-021-01453-z

[18] A. Mohammadshahi y Y. Ioannou, «What is Left After Distillation? How Knowledge Transfer Impacts Fairness and Bias,» arXiv, pp. 1-22, 2024. doi:https://doi.org/10.48550/arXiv.2410.08407

[19] H. Tian, B. Liu, T. Zhu, W. Zhou y P. S. Yu, «Distilling Fair Representations From Fair Teachers,» IEEE Transactions on Big Data, pp. 1419-1433, 2025.

[20] P. Lahoti, K. Gummadi y G. Weikum, «Operationalizing individual fairness with pairwise fair representations,» Proceedings of the VLDB Endowment, pp. 506-518, 2019. doi:https://doi.org/10.48550/arXiv.1907.01439

[21] J. Atwood, T. Tian, B. Packer, M. Deodhar, J. Chen, A. Beutel, F. Prost y A. Beirami, «Towards A Scalable Solution for Improving Multi-Group Fairness in Compositional Classification,» arXiv, pp. 1-10, 2023. doi: https://doi.org/10.48550/arXiv.2307.05728

[22] M. Wan, D. Zha, N. Liu y N. Zou, «In-Processing modeling techniques for machine learning fairness: A survey,» ACM Trans. Knowl. Discov. Data, pp. 1-27, 2023. doi:https://doi.org/10.1145/3551390

[23] A. K. Veldanda, I. Brugere, J. Chen, S. Dutta, A. Mishler y S. Garg, «Fairness via In-Processing in the Over-parameterized Regime: A Cautionary Tale,» arXiv, pp. 1-14, 2022. doi:https://doi.org/10.48550/arXiv.2206.14853

[24] Z. Chen, J. Zhang, F. Sarro y M. Harman, «A comprehensive empirical study of bias mitigation methods for machine learning classifiers,» arXiv, pp. 1-30, 2023. doi:https://doi.org/10.48550/arXiv.2207.03277

[25] X. He, K. Zhao y X. Chu, «AutoML: A survey of the state-of-the-art,» Knowledge-Based Systems, pp. 1-36, 2020. doi:https://doi.org/10.1016/j.knosys.2020.106622

[26] M. Zhou, V. Abhishek, T. Derdenger, J. Kim y K. Srinivasan, «Bias in Generative AI,» arXiv, pp. 1-21, 2024. doi:https://doi.org/10.48550/arXiv.2403.02726

[27] H. Suresh and J. Guttag, “A framework for understanding sources of harm throughout the machine learning life cycle,” ACM Conf. Equity and Access in Algorithms, Mechanisms, and Optimization, pp. 1-9, 2021, doi: 10.1145/3465416.3483305

[28] A. Nielsen, Practical Fairness Achieving Fair and Secure Data Models, Sebastopol: O'Reilly, 2020.

[29] F. Kamiran y T. Calders, «Data preprocessing techniques for classification without discrimination,» Knowledge and Information Systems, vol. 33, p. 1–33, 2012. doi:https://doi.org/10.1007/s10115-011-0463-8

[30] M. Kamani, F. Haddadpour, R. Forsati y M. Mahdavi, «Efficient fair principal component analysis,» Machine Learning, pp. 3671-3702, 2022. doi:https://doi.org/10.1007/s10994-021-06100-9

[31] S. Masís, Interpretable machine learning with Python: Learn to build interpretable high-performance models with hands-on real-world examples, Birmingham: Packt, 2021.

[32] A. Tawakuli y T. Engel, «Make Your Data Fair: A Survey of Data Preprocessing Techniques that Address Biases in Data Towards Fair AI,» Journal of Engineering Research, pp. 2307-1877, 2025. doi:https://doi.org/10.1016/j.jer.2024.06.016

[33] X. Wang, Y. Zhang y R. Zhu, «A Brief Review on Algorithmic Fairness,» Management System Engineering, 2022.

[34] J. Atwood, N. Scherrer, P. Lahoti, A. Balashankar, F. Prost y A. Beirami, «Inducing Group Fairness in Prompt-Based Language Model Decisions,» arXiv, pp 1-11, 2024. doi:https://doi.org/10.48550/arXiv.2406.16738

[35] A. Tifrea, P. Lahoti, B. Packer, Y. Halpern, A. Beirami y F. Prost, «FRAPPE: A Group Fairness Framework for Post-Processing Everything,» arXiv, pp. 1-23, 2023. doi:https://doi.org/10.48550/arXiv.2312.02592

[36] P. Dantas, W. Sabino da Silva, L. Cordeiro y C. Carvalho, «Distilling Fair Representations From Fair Teachers,» Applied Intelligence, pp. 11804–11844, 2024. doi:https://doi.org/10.1109/TBDATA.2024.3460532

[37] S. Wu, X. Luo, J. Liu y Y. Deng, «Knowledge distillation with adapted weight,» Statistics, pp. 470-497, 2025. doi:https://doi.org/10.1080/02331888.2025.2451944

[38] S. Zhao, R. Duan, X. Wang y W. Xingxing, «Improving adversarial robust fairness via anti-bias soft label distillation,» de NeurIPS, Vancouver, 2024.

[39] F. Sikder, R. Ramachandranpillai, D. de Leng y F. Heintz, «Promoting intersectional fairness through knowledge distillation,» IOS Press, pp. 3431-3434, 2025. doi:https://doi.org/10.3233/FAIA251214

[40] M. Masroor, T. Hassan, Y. Tian, K. Wells, D. Rosewarne, T.-T. Do y G. Carneiro, «Fair Distillation: Teaching Fairness from Biased Teachers in Medical Imaging,» arXiv, pp. 1-16, 2024. doi:https://doi.org/10.48550/arXiv.2411.11939

[41] X. Yue, N. Mou, Q. Wang y L. Zhao, «Revisiting adversarial robustness distillation from the perspective of robust fairness,» de International Conference on Neural Information Processing Systems, New Orleans, 2023. https://doi.org/10.52202/075280-1323

[42] S. a. H. F. García, «An Extension on Statistical Comparisons of Classifiers over Multiple Data Sets for All Pairwise Comparisons,» Journal of Machine Learning Research, pp. 2677-2694, 2008.

[43] J. Demšar, «Statistical Comparisons of Classifiers over Multiple Data Sets,» Journal of Machine Learning Research, pp. 1-30, 2006.

[44] S. Holm, «A Simple Sequentially Rejective Multiple Test Procedure,» Scandinavian Journal of Statistics, vol. 6, no. 2, pp. 65–70, 1979.

[45] R. Wasserstein y N. Lazar, «The ASA Statement on p-Values: Context, Process, and Purpose,» The American Statistician, pp. 129–133, 2016. doi:https://doi.org/10.1080/00031305.2016.1154108

[46] D. S. Kerby, «The Simple Difference Formula: An Approach to Teaching Nonparametric Correlation,» Comprehensive Psychology, pp. 1-9, 2014. doi:https://doi.org/10.2466/11.IT.3.1

[47] F. Wilcoxon, «Individual Comparisons by Ranking Methods,» de Breakthroughs in Statistics: Methodology and Distribution, New York, Springer, 1992, pp. 196-202.

[48] R. Schwartz, L. Down, A. Jonas y E. Tabassi, A Proposal for Identifying and Managing Bias in Artificial Intelligence, Gaithersburg: National Institute of Standards and Technology, 2021. doi:https://doi.org/10.6028/NIST.SP.1270-draft

[49] J. Li, X. Feng, T. Gu y L. Chang, «Dual-Teacher De-Biasing Distillation Framework for Multi-Domain Fake News Detection,» de International Conference on Data Engineering (ICDE), Utrecht, 2024. https://doi.org/10.1109/ICDE60146.2024.00279

Publicado

2026-09-01

Número

Sección

Artículo Originales

Cómo citar

[1]
Rosado Gomez, A. 2026. Equidad en aprendizaje automático: comparación de intervenciones de preprocesamiento y entrenamiento en un caso colombiano. Mundo FESC. 16, 36 (Sep. 2026). DOI:https://doi.org/10.61799/2216-0388.2338.