Semantic Annotation Model and Method Based on Internet Open Dataset

Xin Gao et al.

International Journal of Intelligent Information Technologies2025https://doi.org/10.4018/ijiit.370966article
ABDC C
Weight
0.37

What the paper says

Traditional semantic annotation faces the problem of dataset diversity. Different fields and scenarios need to be specially annotated, and annotation work usually requires a lot of manpower and time investment. To meet these challenges, this paper deeply studies the semantic annotation model and method based on internet open datasets, aiming to improve annotation efficiency and accuracy and promote data resource sharing and utilization. This paper selects Common Crawl dataset to provide sufficient training samples; methods such as removing stop words and deduplication are used to preprocess data to improve data quality; a keyword extraction model based on heuristic rules and text context is constructed. In terms of semantic annotation model, this paper constructs a model based on Bidirectional Long Short-Term Memory (BiLSTM), which can make full use of the part-of-speech information of the corpus context, capture the part-of-speech features of the corpus, and generate semantic tags through supervised learning.

1 citation

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.4018/ijiit.370966

Or copy a formatted citation

@article{xin2025,
  title        = {{Semantic Annotation Model and Method Based on Internet Open Dataset}},
  author       = {Xin Gao et al.},
  journal      = {International Journal of Intelligent Information Technologies},
  year         = {2025},
  doi          = {https://doi.org/https://doi.org/10.4018/ijiit.370966},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Semantic Annotation Model and Method Based on Internet Open Dataset

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.37

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.16 × 0.4 = 0.06
M · momentum0.53 × 0.15 = 0.08
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.