A Cascading Hybrid Strategy for Address Normalization using Fine-Tuned LLMs to Strengthen Cancer Registry Geospatial Analysis
DOI:
https://doi.org/10.24054/rcta.v2i48.4512Keywords:
data normalization, geocoding, large language models, natural language processing, unsupervised learningAbstract
The accuracy of geospatial analysis within the Population-based Cancer Registry of the Municipality of Pasto (RPCMP) is severely compromised by the heterogeneous and ambiguous nature of physical addresses collected manually. These source data inconsistencies limit the analytical reach of the YACHAY-SIG infrastructure for generating multiscale territorial intelligence. This paper describes the development and implementation of a hybrid methodological framework designed to optimize spatial data integrity through a two-phase processing workflow.The initial phase deploys a deterministic geocoding engine specifically tailored to the local urban morphology. For records exhibiting high semantic complexity such as non-conventional "block and lot" (manzana y lote) nomenclatures a second stage based on Large Language Models (LLMs) is incorporated. This component integrates Natural Language Processing (NLP) techniques, utilizing TF-IDF and Word2Vec vectorization for the discernment and correction of structural anomalies in the addresses.To maximize efficiency in resource-constrained computational environments, a Supervised Fine-Tuning (SFT) protocol was applied using the Low-Rank Adaptation (LoRA) technique on several open-source models. Experimental results reveal that the integration of the Llama-3.2-3B-Instruct model achieved superior performance, reaching a normalization accuracy of 99.40%. System evaluation over an effective test dataset of 2,324 addresses (extracted from a total corpus of 15,492 unique instances allocated into 70% for training, 15% for validation, and 15% for testing) using standard metrics including precision, recall, and F1-score demonstrates a substantial improvement in spatial recovery, increasing from an initial baseline of 33.02% to 55.13%, effectively mitigating the semantic gap regardless of local cartographic constraints. This architecture stands as a scalable data engineering model for epidemiological registries, with future work focusing on integration with spatial providers such as OpenStreetMap and the Agustín Codazzi Geographic Institute (IGAC) to strengthen national geographic coverage.
Downloads
References
R. Timarán, L. Bravo, A. Hidalgo, R. Suárez, A. Chaves, and F. Vidal, “YACHAY-SIG: Sistema georreferenciado para el registro poblacional de cáncer del municipio de Pasto”, in Avances en Investigaciones de Sistemas Inteligentes y Gestión de Conocimiento: Libro de Resúmenes del CISIGESCO 2025. Popayán, Colombia: Editorial Institución Universitaria Colegio Mayor del Cauca, 2025, p. 11. [Online]. Available: https://sired.udenar.edu.co/17622/1/17622.pdf
R. Timarán, A. Torres, F. Vidal, and A. Hidalgo. “YACHAY: un sistema inteligente para la gstión de la incidencia de cáncer en el municipio de Pasto", in Proc. Encuentro Internacional de Educación en Ingeniería ACOFI (EIEI 2025), Cartagena, Colombia, Sep. 16-19, 2025.
Departamento Administrativo Nacional de Estadística (DANE), “Geoportal DANE: Marco Geoestadístico Nacional,” 2026. [Online]. Available: https://geoportal.dane.gov.co/
J. D. Viveros, R. Timarán and A. Calderón, "Diagnóstico de la Calidad de Datos para la Geocodificación Inteligente en el Registro Poblacional de Cáncer de Pasto", in Avances en Investigaciones de Sistemas Inteligentes y Gestión de Conocimiento: Libro de resúmenes del CISIGESCO 2025, Popayán, Colombia: Editorial Institución Universitaria Colegio Mayor del Cauca, 2025, p. 19. [Online]. Available: https://sired.udenar.edu.co/17622/1/17622.pdf
R. Timarán, G. Hernández and N. Quemá, "Geocodificador de eventos delictivos georreferenciados a nivel de direcciones urbanas en el municipio de Pasto", en Memorias del 5° Congreso Internacional de Gestión Tecnológica y de la Innovación (COGESTEC), Bucaramanga, Colombia: Universidad Industrial de Santander, 2016, pp. 314-327.
J. Hu, et al., "GeoAI agent: To empower LLMs using geospatial tools for address standardization and spatial analysis," International Journal of Geographical Information Science, vol. 39, no. 4, pp. 612–635, 2025.
M. Goodchild, "GIScience and systems: Past, present, and future trajectories in geographic automation," Computers, Environment and Urban Systems, vol. 102, p. 101954, 2023.
A. Zhai, et al., "Transformer-based models for spatial entity extraction and address matching: A comparative review," ISPRS International Journal of Geo-Information, vol. 12, no. 7, p. 284, 2023.
K. Yao, et al., "Contextual semantic understanding in street address parsing using pre-trained language models," Transactions in GIS, vol. 27, no. 5, pp. 1420–1440, 2023.
R. Mai, et al., "Mapping with words: Integrating large language models into geospatial practice and location intelligence," ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XI-4/W2-2026, pp. 115–122, 2026.
Y. J. McDonald, M. Schwind, D. W. Goldberg, A. Lampley y C. M. Wheeler, "An analysis of the process and results of manual geocode correction," Geospatial Health, vol. 12, no. 1, 2017, doi: 10.4081/gh.2017.526.
S. Gupta y K. Nishu, "Mapping Local News Coverage: Precise location extraction in textual news content using fine-tuned BERT based language model," en Proc. of the Fourth Workshop on Natural Language Processing and Computational Social Science, 2020, pp. 155-162.
Y. Guermazi, S. Sellami y O. Boucelma, "A RoBERTa Based Approach for Address Validation," en New Trends in Database and Information Systems, Springer, 2022, doi: 10.1007/978-3-031-15743-1_31.
W. X. Zhao et al., "A Survey of Large Language Models", CoRR, vol. abs/2303.18223, 2023 [Online]. Available: https://arxiv.org/abs/2303.18223
Meta AI, "Llama 3.2: Model Cards and Prompt formats," 2025. [En línea]. Disponible en: https://ai.meta.com/blog/llama-3-2-connect-2024/
M. Batty, "The computational city: Big data and geographical information systems in urban modeling," Progress in Human Geography, vol. 47, no. 2, pp. 210–228, 2024.
P. A. Longley, M. F. Goodchild, D. J. Maguire, and D. W. Rhind, Geographic Information Science and Systems, 4th ed. Hoboken, NJ: Wiley, 2023.
T. Vaswani et al., "Attention is all you need," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 5998–6008, 2017.
J. L. Pérez et al., "A review of hybrid geocoding architectures for health informatics: Challenges in scalability and data governance," International Journal of Health Geographics, vol. 23, no. 1, p. 14, 2024.
M. Open-Weights Consortium, "Efficient parameter fine-tuning of open-source language models for geospatial parsing in resource-constrained environments," Computers, Environment and Urban Systems, vol. 118, p. 102271, 2025.
Gemma Team et al., "Gemma 3 Technical Report," arXiv preprint arXiv:2503.19786, 2025, doi: 10.48550/arXiv.2503.19786.
A. Paszke et al., "PyTorch: An imperative style, high-performance deep learning library," in Advances in Neural Information Processing Systems (NeurIPS), vol. 32, pp. 8024–8035, 2019.
Y. Zheng, R. Zhang, J. Zhang, Y. Ye, and Z. Luo, “LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models,” in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (ACL), Vol. 3: System Demonstrations, Bangkok, Thailand, Aug. 2024, pp. 400-410, doi: 10.18653/v1/2024.acl-demos.38.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2022.
T. Wolf et al., "Hugging Face's transformers: State-of-the-art natural language processing," in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2020, pp. 38–45.
G. Gerganov, "llama.cpp: Port of Facebook's LLaMA model in C/C++," GitHub Repository, 2023. [En línea]. Available: https://github.com/ggerganov/llama.cpp
I. Loshchilov and F. Hutter, "SGDR: Stochastic gradient descent with warm restarts," International Conference on Learning Representations (ICLR), 2017.
L. Prechelt, "Early stopping—but when?", in Neural Networks: Tricks of the Trade, Berlin, Heidelberg: Springer, 2012, pp. 53–67.
C. M. Bishop, Pattern Recognition and Machine Learning, New York: Springer, 2006.
J. Hoffmann et al., “Training Compute-Optimal Large Language Models,” in Proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS 2022), New Orleans, LA, USA, 2022.
Y. Wang, Z. Yu, W. Yao, Z. Zeng, L. Yang, C. Wang, H. Chen, C. Jiang,2R. Xie, J. Wang, X. Xie, W. Ye, S. Zhang, and Y. Zhang,3“PandaLM: An Automatic Evaluation Benchmark for LLM Instruction4Tuning Optimization,” in International Conference on Learning Representations (ICLR), 2024.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Silvio Ricardo Timarán Pereira, Jonathan David Viveros Córdoba

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




