Multı-parameter clusterıng of cıtızens ın an electronıc demography platform: a mıxed-type dıstance approach to the dıgıtal demographıc portraıt
Abstract
Electronic demography platforms describe each citizen by many administrative variables of different kinds, and grouping citizens into a small number of interpretable segments is a natural basis for digital-inclusion policy. The objective of this study is to develop and validate a clustering method for the mixed-type digital demographic portrait, a ten-feature description (five continuous, one ordinal, one nominal and three binary variables) produced by an earlier feature-selection stage of this research. Because the portrait mixes meas-urement types, the standard k-means algorithm is inappropriate; we therefore base the analysis on the Gower distance and compare partitioning around medoids, the k-prototypes algorithm, average-linkage hierarchical clustering, and a one-hot k-means baseline. The number of segments is chosen with the silhouette width and the Calinski-Harabasz and Davies-Bouldin indices, and stability is assessed by bootstrap resampling using the adjusted Rand index. On 8,620 adults from a synthetic dataset, the validity indices and the stability analysis support a three-segment solution. Partitioning around medoids on the Gower distance gives the most bal-anced, interpretable and stable segmentation, whereas the k-prototypes and one-hot k-means baselines sepa-rate the segments less clearly. The three segments, an urban digitally-engaged group, a rural moderate-digital group and an older low-digital group, show a clear gradient in e-service engagement (67, 56 and 45 per cent). Because the data are synthetic, the results illustrate the method rather than describe a real population. The pipeline provides a reproducible, mixed-type segmentation method for e-demography.
Keywords
e-demography; mixed-type clustering; Gower distance; partitioning around medoids; citizen segmentation; digital divide.