DeepSData
Dataset guide · Machine learning & corpora

Pubmed Journal Recommendation System dataset

This dataset is based on PubMed Open Access articles, including title, abstract, keywords, and journal information, for journal recommendation system research.

← Back to dataset library · 中文版

Machine learning & corporaFree to access

Key facts

InstitutionSee official page
CoveragePubMed Open Access articles
Time span2018-01-01 to 2022-12-13
Scale262,870 articles, 469 journals, 5 files
LicenseLicense statement is still being checked
AccessZenodo

Contents & fields

The dataset contains 5 files: data_pubmed_all (all articles), data_pubmed (filtered articles), data_pubmed_train, data_pubmed_val, data_pubmed_test (random 60/20/20 split per journal). Each article record includes the following fields:

  • pubmed_id——PubMed unique identifier
  • title——Article title
  • keywords——Keywords
  • journal——Journal name
  • abstract——Abstract
  • conclusions——Conclusions
  • methods——Methods
  • results——Results
  • copyrights——Copyright information
  • doi——Digital Object Identifier
  • publication_date——Publication date
  • authors——Authors

Research uses

Suitable for journal recommendation systems, academic literature classification, text mining, and other NLP and information retrieval research.

Information comes from the source page. Please check that page for current details and terms.

Keywords

PubMedjournal recommendationdatasettext classificationmachine learningopen access

Access & license

License: License statement is still being checked | Free to access

Why this is hard to get on your own

When obtaining structured PubMed article data for journal recommendation model training, issues such as incomplete fields or unclear splits often arise.

Related datasets

Same domain

Machine learning & corporaFree to access

music_genre

A music genre classification dataset containing 1700 audio clips (mp3 format, each about 270-300 seconds), div…

Need this data retrieved and prepared?

Tell us your hard requirements. We first assess availability, then retrieve for real — and if it truly cannot be obtained, we say so plainly.