A Cascading Hybrid Strategy for Address Normalization using Fine-Tuned LLMs to Strengthen Cancer Registry Geospatial Analysis

Authors

DOI:

https://doi.org/10.24054/rcta.v2i48.4512

Keywords:

data normalization, geocoding, large language models, natural language processing, unsupervised learning

Abstract

The accuracy of geospatial analysis within the Population-based Cancer Registry of the Municipality of Pasto (RPCMP) is severely compromised by the heterogeneous and ambiguous nature of physical addresses collected manually. These source data inconsistencies limit the analytical reach of the YACHAY-SIG infrastructure for generating multiscale territorial intelligence. This paper describes the development and implementation of a hybrid methodological framework designed to optimize spatial data integrity through a two-phase processing workflow.The initial phase deploys a deterministic geocoding engine specifically tailored to the local urban morphology. For records exhibiting high semantic complexity such as non-conventional "block and lot" (manzana y lote) nomenclatures a second stage based on Large Language Models (LLMs) is incorporated. This component integrates Natural Language Processing (NLP) techniques, utilizing TF-IDF and Word2Vec vectorization for the discernment and correction of structural anomalies in the addresses.To maximize efficiency in resource-constrained computational environments, a Supervised Fine-Tuning (SFT) protocol was applied using the Low-Rank Adaptation (LoRA) technique on several open-source models. Experimental results reveal that the integration of the Llama-3.2-3B-Instruct model achieved superior performance, reaching a normalization accuracy of 99.40%. System evaluation over an effective test dataset of 2,324 addresses (extracted from a total corpus of 15,492 unique instances allocated into 70% for training, 15% for validation, and 15% for testing) using standard metrics including precision, recall, and F1-score demonstrates a substantial improvement in spatial recovery, increasing from an initial baseline of 33.02% to 55.13%, effectively mitigating the semantic gap regardless of local cartographic constraints. This architecture stands as a scalable data engineering model for epidemiological registries, with future work focusing on integration with spatial providers such as OpenStreetMap and the Agustín Codazzi Geographic Institute (IGAC) to strengthen national geographic coverage.

368 95

Downloads

Download data is not yet available.

Author Biographies

  • Silvio Ricardo Timarán Pereira, Universidad de Nariño, San Juan de Pasto, Nariño, Colombia

    Ricardo Timarán Pereira es Ingeniero de Sistemas y Computación y Master of Science en Ingeniería de la Universidad Politécnica de Donetsk (Ucrania), Especialista en Multimedia Educativa de la Universidad Antonio Nariño y Doctor en Ingeniería con énfasis en Ciencias de la Computación de la Universidad del Valle. Actualmente, es Profesor Titular del Departamento de Sistemas de la Universidad de Nariño e Investigador Senior reconocido por Minciencias, donde dirige el Grupo de Investigación en Sistemas Aplicados (GRIAS). Con una destacada trayectoria en la coordinación de la Maestría en Gestión de Tecnologías de la Información y el Conocimiento (MGTIC), su labor investigativa se centra en la minería de datos, inteligencia de negocios y modelos de lenguaje de gran escala aplicados a la salud pública y la educación. Es autor de diversos libros especializados que abordan el descubrimiento de patrones en el rendimiento académico y la gestión del conocimiento en registros oncológicos; además, ha participado como ponente en múltiples eventos nacionales e internacionales, consolidándose como un referente en la transferencia tecnológica y la soberanía digital en el contexto latinoamericano.

  • Jonathan David Viveros Córdoba, Universidad de Nariño, San Juan de Pasto, Nariño, Colombia

    Ingeniero de Sistemas de la Universidad de Nariño. Integrate del grupo de investigacion GRIAS del departamento de Sistemas de la Facultad de Ingeniería de la Universidad de Nariño. Asistente de investigación en el proyecto MINCIENCIAS 82288 denominado "Yachay un sistema Inteligente de genaración de Conocimiento para el Registro Poblacional de Cáncer del Municipio de Pasto. Experiencia en desarrollo de software para Inteligencia de Negocios bajo Metabase y Modelos de Lenguaje a Gran Escala (LLMs). 

References

R. Timarán, L. Bravo, A. Hidalgo, R. Suárez, A. Chaves, and F. Vidal, “YACHAY-SIG: Sistema georreferenciado para el registro poblacional de cáncer del municipio de Pasto”, in Avances en Investigaciones de Sistemas Inteligentes y Gestión de Conocimiento: Libro de Resúmenes del CISIGESCO 2025. Popayán, Colombia: Editorial Institución Universitaria Colegio Mayor del Cauca, 2025, p. 11. [Online]. Available: https://sired.udenar.edu.co/17622/1/17622.pdf

R. Timarán, A. Torres, F. Vidal, and A. Hidalgo. “YACHAY: un sistema inteligente para la gstión de la incidencia de cáncer en el municipio de Pasto", in Proc. Encuentro Internacional de Educación en Ingeniería ACOFI (EIEI 2025), Cartagena, Colombia, Sep. 16-19, 2025.

Departamento Administrativo Nacional de Estadística (DANE), “Geoportal DANE: Marco Geoestadístico Nacional,” 2026. [Online]. Available: https://geoportal.dane.gov.co/

J. D. Viveros, R. Timarán and A. Calderón, "Diagnóstico de la Calidad de Datos para la Geocodificación Inteligente en el Registro Poblacional de Cáncer de Pasto", in Avances en Investigaciones de Sistemas Inteligentes y Gestión de Conocimiento: Libro de resúmenes del CISIGESCO 2025, Popayán, Colombia: Editorial Institución Universitaria Colegio Mayor del Cauca, 2025, p. 19. [Online]. Available: https://sired.udenar.edu.co/17622/1/17622.pdf

R. Timarán, G. Hernández and N. Quemá, "Geocodificador de eventos delictivos georreferenciados a nivel de direcciones urbanas en el municipio de Pasto", en Memorias del 5° Congreso Internacional de Gestión Tecnológica y de la Innovación (COGESTEC), Bucaramanga, Colombia: Universidad Industrial de Santander, 2016, pp. 314-327.

J. Hu, et al., "GeoAI agent: To empower LLMs using geospatial tools for address standardization and spatial analysis," International Journal of Geographical Information Science, vol. 39, no. 4, pp. 612–635, 2025.

M. Goodchild, "GIScience and systems: Past, present, and future trajectories in geographic automation," Computers, Environment and Urban Systems, vol. 102, p. 101954, 2023.

A. Zhai, et al., "Transformer-based models for spatial entity extraction and address matching: A comparative review," ISPRS International Journal of Geo-Information, vol. 12, no. 7, p. 284, 2023.

K. Yao, et al., "Contextual semantic understanding in street address parsing using pre-trained language models," Transactions in GIS, vol. 27, no. 5, pp. 1420–1440, 2023.

R. Mai, et al., "Mapping with words: Integrating large language models into geospatial practice and location intelligence," ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XI-4/W2-2026, pp. 115–122, 2026.

Y. J. McDonald, M. Schwind, D. W. Goldberg, A. Lampley y C. M. Wheeler, "An analysis of the process and results of manual geocode correction," Geospatial Health, vol. 12, no. 1, 2017, doi: 10.4081/gh.2017.526.

S. Gupta y K. Nishu, "Mapping Local News Coverage: Precise location extraction in textual news content using fine-tuned BERT based language model," en Proc. of the Fourth Workshop on Natural Language Processing and Computational Social Science, 2020, pp. 155-162.

Y. Guermazi, S. Sellami y O. Boucelma, "A RoBERTa Based Approach for Address Validation," en New Trends in Database and Information Systems, Springer, 2022, doi: 10.1007/978-3-031-15743-1_31.

W. X. Zhao et al., "A Survey of Large Language Models", CoRR, vol. abs/2303.18223, 2023 [Online]. Available: https://arxiv.org/abs/2303.18223

Meta AI, "Llama 3.2: Model Cards and Prompt formats," 2025. [En línea]. Disponible en: https://ai.meta.com/blog/llama-3-2-connect-2024/

M. Batty, "The computational city: Big data and geographical information systems in urban modeling," Progress in Human Geography, vol. 47, no. 2, pp. 210–228, 2024.

P. A. Longley, M. F. Goodchild, D. J. Maguire, and D. W. Rhind, Geographic Information Science and Systems, 4th ed. Hoboken, NJ: Wiley, 2023.

T. Vaswani et al., "Attention is all you need," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 5998–6008, 2017.

J. L. Pérez et al., "A review of hybrid geocoding architectures for health informatics: Challenges in scalability and data governance," International Journal of Health Geographics, vol. 23, no. 1, p. 14, 2024.

M. Open-Weights Consortium, "Efficient parameter fine-tuning of open-source language models for geospatial parsing in resource-constrained environments," Computers, Environment and Urban Systems, vol. 118, p. 102271, 2025.

Gemma Team et al., "Gemma 3 Technical Report," arXiv preprint arXiv:2503.19786, 2025, doi: 10.48550/arXiv.2503.19786.

A. Paszke et al., "PyTorch: An imperative style, high-performance deep learning library," in Advances in Neural Information Processing Systems (NeurIPS), vol. 32, pp. 8024–8035, 2019.

Y. Zheng, R. Zhang, J. Zhang, Y. Ye, and Z. Luo, “LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models,” in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (ACL), Vol. 3: System Demonstrations, Bangkok, Thailand, Aug. 2024, pp. 400-410, doi: 10.18653/v1/2024.acl-demos.38.

E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2022.

T. Wolf et al., "Hugging Face's transformers: State-of-the-art natural language processing," in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2020, pp. 38–45.

G. Gerganov, "llama.cpp: Port of Facebook's LLaMA model in C/C++," GitHub Repository, 2023. [En línea]. Available: https://github.com/ggerganov/llama.cpp

I. Loshchilov and F. Hutter, "SGDR: Stochastic gradient descent with warm restarts," International Conference on Learning Representations (ICLR), 2017.

L. Prechelt, "Early stopping—but when?", in Neural Networks: Tricks of the Trade, Berlin, Heidelberg: Springer, 2012, pp. 53–67.

C. M. Bishop, Pattern Recognition and Machine Learning, New York: Springer, 2006.

J. Hoffmann et al., “Training Compute-Optimal Large Language Models,” in Proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS 2022), New Orleans, LA, USA, 2022.

Y. Wang, Z. Yu, W. Yao, Z. Zeng, L. Yang, C. Wang, H. Chen, C. Jiang,2R. Xie, J. Wang, X. Xie, W. Ye, S. Zhang, and Y. Zhang,3“PandaLM: An Automatic Evaluation Benchmark for LLM Instruction4Tuning Optimization,” in International Conference on Learning Representations (ICLR), 2024.

Published

2026-07-28

How to Cite

[1]
S. R. Timarán Pereira and J. D. Viveros Córdoba, “A Cascading Hybrid Strategy for Address Normalization using Fine-Tuned LLMs to Strengthen Cancer Registry Geospatial Analysis”, RCTA, vol. 2, no. 48, pp. 166–180, Jul. 2026, doi: 10.24054/rcta.v2i48.4512.

Similar Articles

1-10 of 633

You may also start an advanced similarity search for this article.