METAEXTRACT: A HYBRID METADATA EXTRACTION SYSTEM COMBINING RULE- BASED AND STATISTICAL APPROACHES
DOI:
https://doi.org/10.5281/zenodo.22825400Abstract
The issue of automatic extraction of structural metadata from documents is a widely studied topic in the field of document analysis, in which two main paradigms — universal models based on deep learning and rule-based systems — compete. Universal models have high generalization capabilities, but they require large amounts of annotated training data, which is a serious limitation for low-resource languages and domains. In this paper, a hybrid architecture combining a rule-based geometric-textual approach with a statistical named entity recognition (NER) model is presented through the example of extracting bibliographic metadata from scientific journal articles. The proposed system is based on the principle of creating rules from a single sample document and applying them to other documents sharing the same template, which makes it practically applicable even without large amounts of training data. The paper discusses design principles, the justification of architectural decisions, implementation details, and the role of this approach in the broader field of document extraction.Keywords
document metadata extraction, rule-based systems, hybrid architecture, named entity recognition (NER), low-resource languages, software designReferences
Romary, L., & Lopez, P. (2015). GROBID — Information extraction from scientific publications. ERCIM News, 2015(100).
Tkaczyk, D., Szostek, P., Fedoryszak, M., Dendek, P. J., & Bolikowski, Ł. (2015). CERMINE: Automatic extraction of structured metadata from scientific literature. International Journal on Document Analysis and Recognition (IJDAR), 18(4), 317–335. https://doi.org/10.1007/s10032-015-0249-8
Duc, D. M., Truong, Q. X., Hong, V. T., Anh, L. H., Tra, M. T. M., Van Thuy, N., ... & Van, V. N. (2026). A hybrid method for low-resource named entity recognition. arXiv preprint arXiv:2605.04489.
Ahmed, M. W., & Afzal, M. T. (2020). FLAG-PDFe: Features oriented metadata extraction framework for scientific publications. IEEE Access, 8, 99458–99469. https://doi.org/10.1109/ACCESS.2020.2996112
Skluzacek, T. J., Chard, K., & Foster, I. (2022). Automated metadata extraction: Challenges and opportunities. In 2022 IEEE 18th International Conference on e-Science (e-Science) (pp. 495–500). IEEE. https://doi.org/10.1109/ eScience56278.2022.00078
Ubewikkrama, P. T. Automatic invoice data identification with relations [Doctoral dissertation].
Skluzacek, T. J., Chen, M., Hsu, E., Chard, K., & Foster, I. (2022). Models and metrics for mining meaningful metadata. In International Conference on Computational Science (pp. 417–430). Springer, Cham. https://doi.org/10.1007/978-3- 031-08757-8_35
Raval, P., & Bhaidasna, H. (2025). Metadata extraction from scholarly document using deep learning. In 2025 3rd International Conference on Inventive Computing and Informatics (ICICI) (pp. 1–5). IEEE.
Boukhers, Z., Beili, N., Hartmann, T., Goswami, P., & Zafar, M. A. (2021). MExPub: Deep transfer learning for metadata extraction from German publications. In 2021 ACM/IEEE Joint Conference on Digital Libraries (JCDL) (pp. 250–253). IEEE. https://doi.org/10.1109/JCDL52503.2021.00038
Skluzacek, T. J. Automated metadata extraction can make data swamps more navigable [Doctoral dissertation, The University of Chicago].

