Dissertation (August 25, 2026): Felipe de Ávila Tavares
Student: Felipe de Ávila Tavares
Title: Efeito da Cardinalidade e da Associação entre Atributos na Imputação de Dados Categóricos
Advisor: Jorge de Abreu Soares
Committee: Jorge de Abreu Soares (Cefet/RJ), Diego Nunes Brandão (Cefet/RJ), Ronaldo Ribeiro Goldschmidt (IME)
Day/Hours: August 25, 2026 / 2:30 p.m.
Room: Pavilhão I, sala P1-202 (laboratório 6)
Abstract: This dissertation investigates the problem of categorical data imputation, focusing on the comparative analysis of different imputation methods and on how dataset structural characteristics influence their performance. The study considers three distinct scenarios: simple datasets composed predominantly of numerical attributes, mixed datasets containing both numerical and categorical variables, and fully categorical datasets. The proposed methodology artificially generates missing values under the Missing Completely at Random (MCAR) mechanism at rates of 10%, 20%, and 30%, using multiple random seeds to evaluate the stability of the analyzed methods. Approaches based on KNN, Multi-Layer Perceptron (MLP), and Random Forest were investigated, considering both imputation accuracy and the performance of classifiers trained on the imputed data. In addition, Cramer’s V coefficient was used to analyze the influence of correlation between categorical attributes in the imputation process. The obtained results indicate that global model-based methods, especially Random Forest, achieve greater robustness and stability against increasing missingness rates and dataset structural complexity, while KNN shows greater sensitivity to data heterogeneity and information loss. The findings reinforce the importance of considering characteristics such as cardinality and attribute correlation when selecting the most appropriate imputation method