Key facts
Contents & fields
The dataset includes scripts, input, and output data to reproduce results from the ICPP 2023 paper. Input data are simulated X-ray Free Electron Laser (XFEL) protein diffraction patterns at low, medium, and high beam intensities, each containing 63,508 training images and 15,876 test images (80/20 split). Output data include neural network models and metadata generated with the A4NN workflow on different GPU distributions, with 100 NN models per experiment, each trained for up to 25 epochs, totaling approximately 72,900 model-related files.
- README.txt——Documentation file with usage instructions
- Training images——Simulated XFEL diffraction images for training
- Test images——Simulated XFEL diffraction images for testing
- NN models——Neural network model files, 100 per experiment
- Metadata——Training history and other metadata files
Research uses
This dataset is suitable for research on protein classification, neural architecture search, and reproducible computational workflows.
This card was drafted from the source page; institution, coverage, time span, scale, fields and license are subject to the official page (pending human review).
Keywords
Why this is hard to get on your own
When obtaining protein diffraction data and corresponding neural network models, common challenges include inconsistent data formats, large numbers of model files, and missing metadata.
