Application of Large-Scale Subgraph Matching Algorithms in Natural Language Processing

Authors

  • Edmund Kapuściński Faculty of Computer Science, AGH University of Science and Technology, Krakow, 30-059, Poland
  • Marcel Barczyk Faculty of Computer Science, AGH University of Science and Technology, Krakow, 30-059, Poland

DOI:

https://doi.org/10.64972/jaat.2023v1.283p5e:62-75

Keywords:

Large-scale Subgraph Matching, Natural Language Processing, Typed Textual Graph, Neural Re-ranking, Evidence Retrieval

Abstract

Large-scale natural language processing increasingly needs to have structured evidence that can link entities, predicates, events and discourse relations in massive corpora. This paper investigates the application of large-scale subgraph matching algorithms in NLP and proposes LSM-NLP, a hybrid framework that combines typed textual graph construction, structural indexing, candidate verification and neural re-ranking. The framework regards subgraph matching as an engineering problem of symbolic linguistic structure and neural semantic representation. It has been shown through experiments that it can perform relation extraction, semantic retrieval, question-answer evidence selection and document clustering. According to the above planned results, LSM-NLP will increase the average F1 score to 80.7 from 75.4 and lower the median latency to 612ms at the 10-million-node graph scale, keeping shard-level score drift below 0.033 after calibration. Based on the above studies, a scalable subgraph matching can be used to improve both the interpretability and efficiency of precise relational evidence in NLP systems.

Downloads

Published

2023-02-19

How to Cite

Kapuściński, E., & Barczyk, M. (2023). Application of Large-Scale Subgraph Matching Algorithms in Natural Language Processing. Journal of Applied Automation Technologies, 1, 5e:62–75. https://doi.org/10.64972/jaat.2023v1.283p5e:62-75

Issue

Section

Articles