Key facts
Contents & fields
The dataset contains 5 files: data_pubmed_all (all articles), data_pubmed (filtered articles), data_pubmed_train, data_pubmed_val, data_pubmed_test (random 60/20/20 split per journal). Each article record includes the following fields:
- pubmed_id——PubMed unique identifier
- title——Article title
- keywords——Keywords
- journal——Journal name
- abstract——Abstract
- conclusions——Conclusions
- methods——Methods
- results——Results
- copyrights——Copyright information
- doi——Digital Object Identifier
- publication_date——Publication date
- authors——Authors
Research uses
Suitable for journal recommendation systems, academic literature classification, text mining, and other NLP and information retrieval research.
This card was drafted from the source page; institution, coverage, time span, scale, fields and license are subject to the official page (pending human review).
Keywords
Why this is hard to get on your own
When obtaining structured PubMed article data for journal recommendation model training, issues such as incomplete fields or unclear splits often arise.
