DeepSData
Curated datasets · a structured reference card for each

A growing dataset library — fields explained, sources you can check

Each card sets out coverage, time span, granularity, license and access method. Metadata was reviewed against each publishing body's official sources in June 2026. 18 are curated in depth on this page. Prefer to see the card format first? View a sample reference card →
State your criteria — availability is assessed first, then searched. Prepared datasets are also available for direct delivery. Start a search assessment See prepared datasets → Browse the full library →
① Public-source search / availability assessment

We deliver source links, search notes and item-by-item verdicts. We do not download third-party raw data on your behalf — obtain it yourself under each source's license.

② Prepared datasets of our own

Collected and prepared by us, with samples provided — these can be delivered directly. See prepared datasets →

③ Restricted / third-party raw data

Governed by the source license and the scope both sides confirm in writing.

Full scope of service and notes on the data

The service covers open datasets that are public and lawful to obtain. The dataset metadata on this page was reviewed against each publishing body's official sources in June 2026 and reflects the official information. The data is provided by third-party bodies; we do not resell third-party datasets as our own.

ECONOMICS & MACRO · 6 IN THIS COLUMN
Economics · Social science · Global development
International open datasets widely used in research, yet often hard to assemble, align or evaluate quickly.
Economics & macroDS-01

Penn World Table 11.0 · cross-country macro panelPWT · Penn World Table

A widely used reference for cross-country economic comparison: it puts each economy's GDP, output, capital and productivity on a comparable basis for macro, growth and development research.

SourceGGDC, University of Groningen
Coverage185 economies
Period1950–2023
ScalePanel of dozens of variables
LicenseCC BY 4.0
AccessFree, no registration
View fields and details

Contents and fields

A national-accounts library built on purchasing power parity (PPP) for cross-country and cross-period comparison: expenditure-side real GDP (CGDPe/RGDPe), output-side real GDP (CGDPo/RGDPo), real GDP at national-accounts definitions (RGDPNA), population and employment, human-capital index, capital stock, total factor productivity (TFP), price levels and PPPs, and more.

Suitable research

A commonly used baseline for cross-country growth, productivity, convergence and development research.

Version PWT 11.0 (2025-10) · Format Excel / Stata / online tool · ID DOI:10.34894/FABVLR

Metadata follows the official source (reviewed June 2026).

Why it's hard to get on your ownPPP benchmarks, chained versus current definitions, and output-side versus expenditure-side GDP are easy to mix up across versions. Reconciling them by hand often takes days and still may not add up. Leave the version mapping and definition alignment to us, and take the data ready to use.
Economics & macroDS-02

World Inequality DatabaseWID · World Inequality Database

A long-run database of global income and wealth distribution that combines national accounts, household surveys and tax data, correcting traditional surveys' undercount at the top.

SourceWorld Inequality Lab
CoverageAbout 216 countries and territories
PeriodModern from 1980, some earlier
ScaleMulti-dimensional distribution series
LicenseOpen access (per official terms)
AccessFree · Stata/R tools
View fields and details

Contents and fields

Income distribution (pre-tax/post-tax, shares and average levels by percentile and decile), wealth distribution and top wealth shares, national income and national-accounts macro variables; recent additions include foreign income, foreign wealth, public income and public spending. Each series carries dimensions such as region code, year, indicator code, population/percentile range, currency and unit.

Suitable research

Distribution-share statistics for distributional structure, redistribution, and macro and development economics.

Updates Rolling (incl. 2024 update) · Format Web export / Stata package / R tools

License type follows the official terms (to be verified); metadata reviewed June 2026.

Why it's hard to get on your ownIndicator codes, currency units, pre- and post-tax definitions and percentile ranges are numerous, and annual updates add new series. Aligning them against the official dictionary, across years and regions, is highly error-prone. Let us organize it into consistent, searchable, well-formed series.
Economics & macroDS-03

BACI global bilateral trade databaseCEPII BACI

A cleaned, reconciled product-level bilateral trade panel at HS 6-digit level, covering the world's major economies — a commonly used benchmark for empirical trade research.

SourceCEPII, France
CoverageMajor economies · about 5,000 products
PeriodAnnual, to 2024
ScaleMulti-year bilateral product flows
LicenseEtalab Open 2.0
AccessFree, no registration
View fields and details

Contents and fields

FieldMeaning
tYear
kProduct category (HS 6-digit, leading zeros kept)
i / jExporter / importer (ISO numeric codes)
vTrade value (thousand USD, current prices)
qQuantity (metric tonnes)

Methodology: harmonizes CIF/FOB definitions on UN Comtrade raw declarations and reconciles mirror data weighted by reporter reliability.

Version 202601 (2026-01, updated each January) · Format CSV (distributed as ZIP)

Units of analysis are expressed as ISO numeric codes; metadata reviewed June 2026.

Why it's hard to get on your ownConflicting mirror data, inconsistent CIF/FOB definitions, HS codes that lose their leading zeros when read as numbers, versions that don't line up across years — the routine work that quietly consumes research time. We have already resolved it into a consistent long table that goes straight into your model.
Economics & macroDS-04

CEPII Gravity databaseCEPII Gravity

A ready set of bilateral variables for gravity-equation research: trade flows, distance, agreements, common language and macro indicators — a commonly used base dataset for empirical international-trade research.

SourceCEPII, France
CoverageSquare bilateral · about 252 economies
Period1948–2020
ScaleBilateral-pair panel
LicenseEtalab Open 2.0
AccessFree, no registration
View fields and details

Contents and fields

① Bilateral trade flows (three sources: IMF DOTS / UN Comtrade / BACI); ② geographic distance measures (weighted distances, contiguity, landlocked, island, latitude/longitude); ③ institutional and trade-facilitation variables (GATT/WTO membership, regional/bilateral trade agreements); ④ proxies for historical and institutional ties (common language, religion, legal-system origin, historical links); ⑤ macro indicators (GDP, population).

Suitable research

Gravity models, bilateral trade, and global-value-chain empirics.

Version 202211 (2022-11) · Format CSV / R / Stata · Citation Conte, Cotterlaz & Mayer (2022)

Regional divisions follow the data source's international statistical classification — a statistical convention only; China's official standards govern. Metadata reviewed June 2026.

Why it's hard to get on your ownThe three trade-flow sources use different definitions, the distance and institutional variables span different years, and reporter codes must be aligned year by year. Assembling a regression-ready bilateral panel by hand often takes weeks. What we deliver is a complete dataset with unified definitions and aligned pairs.
Economics & macroDS-05

Maddison Project historical GDP databaseMaddison Project · MPD

A widely cited benchmark for long-run economic history, with cross-country comparable per-capita GDP and population estimates from AD 1.

SourceGGDC, University of Groningen
Coverage169 countries and territories
PeriodAD 1 – 2022
ScaleLong time series
LicenseCC BY 4.0
AccessFree, no registration
View fields and details

Contents and fields

FieldMeaning
countrycode / yearRegion code / year
cgdppcPer-capita GDP for level comparison (2011 international $)
rgdpnapcReal per-capita GDP (cross-time growth comparison)
popPopulation (thousands)
i_cig / i_bmEstimate source / benchmark-estimate flag

Suitable research

Historical economics, comparative development, and long-run growth and income-level differences.

Version MPD 2023 · Format Excel / Stata · ID DOI:10.34894/INZBF2

Field naming and classification follow the official codebook; metadata reviewed June 2026.

Why it's hard to get on your ownDefinitions are revised across release years, the current-price and real per-capita GDP series each suit different uses, and the flags for benchmark estimates versus interpolation need checking field by field. We deliver data that is mapped, aligned and clearly labelled by definition, ready to use.
Economics & macroDS-06

Barro-Lee educational attainment databaseBarro-Lee Educational Attainment

A widely used reference for measuring human-capital stock, with cross-country educational-attainment estimates by sex and age, cited by the World Bank and a large growth literature.

SourceBarro (Harvard) + Lee (Korea Univ.)
Coverage146 economies
Period1950–2015 (5-year intervals)
ScaleCross-tabs by sex/age
LicenseFree, citation required
AccessFree, no registration
View fields and details

Contents and fields

Educational attainment by sex (total/male/female) and age group: the share of population at each education level (none/primary/secondary/tertiary, incomplete and complete), average years of schooling (yr_sch and by level), plus the enrolment rates, dropout rates and population structure used in estimation; includes Lee-Lee long-run historical data and extension modules such as education quality.

Suitable research

Economics of education, human capital and empirical economic-growth research.

Version 2021-09 (BLv3) · Format Excel / CSV / Stata

The license is stated two ways (GitHub MIT / rights reserved on the official site); verify the authors' authorization before commercial use. Metadata reviewed June 2026.

Why it's hard to get on your ownDefinitions change across release batches, the education-level breakdown is intricate, and stitching together the historical backcast series is demanding — it is easy to end up realigning the data and unsure which version to use. We have completed the version mapping and variable calibration.
GEOSPATIAL & REMOTE SENSING · 5 IN THIS COLUMN
Remote sensing · Population · Land
Geospatial layers that turn what happens on the ground into computable data — commonly used proxies for regional, urban and economic analysis.
Geospatial & remote sensingDS-07

VIIRS nighttime lights (annual composite)VIIRS Nighttime Lights · VNL

Satellite-derived nighttime light intensity used to proxy economic activity, urban expansion and energy use — a commonly used remote-sensing layer for regional economic analysis.

SourceEOG, Colorado School of Mines
CoverageGlobal (75°N–65°S)
Period2012 to present
ResolutionAbout 500m (15 arc-seconds), annual
LicenseOpen, commercial use allowed
AccessFree, registration required (also via GEE)
View fields and details

Contents and fields

BandMeaning
average / average_maskedAverage radiance / masked average radiance
median / maximum / minimumMedian / maximum / minimum radiance
cf_cvg / cvgCloud-free observation count / total observation count

Radiance in nW/cm²/sr, processed to remove clouds, moonlight and fire pixels.

Suitable research

Regional economics, GDP proxies, electricity access and urban expansion.

Version Annual VNL V2.2 · Format GeoTIFF

Official license wording varies (public domain or CC BY; both allow commercial use; attribution to EOG is recommended). This site does not render maps. Metadata reviewed June 2026.

Why it's hard to get on your ownGlobal night-light data is spread across several distribution platforms. Registration hurdles and multiple version formats mean simply obtaining a usable copy takes most of the effort, and year-to-year sensor and algorithm consistency still has to be judged. Leave this tedious part to us.
Geospatial & remote sensingDS-08

Global Human Settlement LayerGHSL · Global Human Settlement Layer

It turns "where people are, how many, and how urbanized" into globally consistent raster layers — hard-to-replace base data for urban and population-exposure research.

SourceEuropean Commission JRC
CoverageGlobal raster
Period1975–2030 (5-year)
Resolution100m / 1km (some 10m)
LicenseEC reuse (cite the source)
AccessFree, no registration
View fields and details

Contents and fields

ProductMeaning
GHS-BUILT-S / V / HBuilt-up area / volume / building height
GHS-POPPopulation distribution grid (people per cell)
GHS-SMOD / DUCDegree of urbanization / urbanization class of administrative units

Suitable research

Urbanization, population distribution, regional economics and disaster exposure.

Version R2023A · Format GeoTIFF

Administrative-unit divisions follow the data source's original conventions and are a technical processing result only; China's official standard maps govern. This site does not render maps. Metadata reviewed June 2026.

Why it's hard to get on your ownVersions and years are interleaved, built-up area and population-grid definitions differ, and aligning with census and UN figures is tedious. Downloading, comparing and unifying coordinates and resolution by hand often takes days. Leave the version checks and definition alignment to us.
Geospatial & remote sensingDS-09

ESA global land cover (10m)ESA WorldCover

A 10-metre global land-cover map based on Sentinel satellites, with independently validated overall accuracy of about 76.7%, clear classes and ready to use as a research base layer.

SourceEuropean Space Agency (ESA), led by VITO
CoverageGlobal
Period2020 / 2021 (two versions)
Resolution10m
LicenseCC BY 4.0
AccessFree (official site / AWS / GEE)
View fields and details

Contents and fields

A single band (Map) records 11 land-cover classes: tree cover, shrubland, grassland, cropland, built-up, bare/sparse vegetation, snow and ice, permanent water bodies, herbaceous wetland, mangroves, moss and lichen. A product user manual and validation report are included.

Suitable research

Land use, agriculture, ecology, urban expansion and environmental research.

Version v200 (2021, released 2022-10) · Format Cloud-Optimized GeoTIFF

The two versions use different algorithms, so use care for change detection; this site does not render maps. Metadata reviewed June 2026.

Why it's hard to get on your ownThe v100 and v200 definition differences, aligning class definitions, comparing accuracy reports and obtaining the raw rasters all cost time and invite errors when done by hand. We have completed the version checks and documentation, so you can take a usable base layer as needed.
Geospatial & remote sensingDS-10

WorldPop global population distributionWorldPop

Gridded population estimates at about 100-metre resolution, downscaled from official censuses — a well-documented public source for spatial demography and regional planning.

SourceUniversity of Southampton and others
CoverageGlobal (customized per country)
PeriodAbout 2000–2021 (projections also)
ResolutionAbout 100m (also 1km)
LicenseCC BY 4.0 (commercial use allowed)
AccessFree and open
View fields and details

Contents and fields

Population counts (estimated residents per grid cell), population density, age/sex population structure, development indicators (poverty / birth rate, etc.), population mobility, and more. Methodology: built on censuses, using random-forest dasymetric reallocation with geographic covariates to downscale to about 100m grids.

Suitable research

Spatial demography, urban studies, disaster assessment, accessibility and public-service planning.

Version Versioned by DOI · Format GeoTIFF / REST API

Model-based estimates by a third-party body; administrative divisions and boundaries follow China's official standards. This site does not render maps. Metadata reviewed June 2026.

Why it's hard to get on your ownFor the same area, population-raster definitions and coordinate systems often need repeated checking across years and versions, and aligning the age and sex layers with covariates is time-consuming. We have completed the version mapping and field unification — ready for direct use after the search.
Geospatial & remote sensingDS-11

WorldClim global climate dataWorldClim

A high-resolution global climate raster baseline, from monthly temperature and precipitation to 19 bioclimatic variables — a common reference for species-distribution and climate-impact assessment.

SourceHijmans et al. (UC Davis/Berkeley)
CoverageGlobal land
PeriodBaseline 1970–2000
Resolution30 arc-seconds – 10 arc-minutes
LicenseFree academic / non-commercial
AccessFree, no registration, direct links
View fields and details

Contents and fields

Monthly minimum/average/maximum temperature (°C), precipitation (mm), solar radiation, wind speed and water-vapour pressure; 19 standardized bioclimatic variables (bio1–bio19, e.g. annual mean temperature, temperature seasonality, annual precipitation); SRTM elevation included. Generated from global weather-station records via thin-plate-spline interpolation.

Suitable research

Species distribution and niche modelling, climate-change impacts, and agricultural and urban climate analysis.

Version 2.1 (2020-01) · Format GeoTIFF (grouped by element/resolution)

The official terms state non-commercial use; redistribution or commercial use is not permitted without authorization. Metadata reviewed June 2026.

Why it's hard to get on your ownComparing variable definitions and coordinate alignment across versions, one by one, often takes days. We have completed the organizing and checks, so you can take it straight into analysis.
HEALTH & MEDICINE · 4 IN THIS COLUMN
Population · Health · Epidemiology
Well-documented, cross-country comparable health and population data with unified definitions — the kind that often requires registration, version checks and careful alignment.
Health & medicineDS-12

Human Mortality DatabaseHMD · Human Mortality Database

A widely used source for mortality and life-table research, with unified calculation methods for actuarial work, life-insurance pricing and demographic analysis.

SourceUC Berkeley & Max Planck Inst. for Demographic Research
CoverageAbout 41 countries and territories
PeriodFrom as early as 1751 (annual)
ScaleAbout 48 population series
LicenseOutputs CC BY 4.0
AccessFree, registration required
View fields and details

Contents and fields

Mortality rates, life tables, death counts, birth counts and exposure-to-risk population by age/sex; includes both period and cohort data, plus the raw inputs used to build the life tables. Companion sub-series: the short-term weekly death series (STMF, for monitoring mortality fluctuations).

Suitable research

Demography, ageing, actuarial science and public health.

Updates Rolling by country · Format CSV/TXT / Excel / R interface

Input data is bound by the original licenses of each country's statistical agency; metadata reviewed June 2026.

Why it's hard to get on your ownNational raw definitions differ, life-table construction is intricate, and version updates and exposure-to-risk alignment are error-prone — and you must register and accept the agreement first. Leave it to us for a standardized, traceable deliverable with unified definitions.
Health & medicineDS-13

Human Fertility DatabaseHFD · Human Fertility Database

A widely used source for high-quality fertility data in developed countries, broken down by mother's age and birth order to the fertility-table level — a comparable baseline for low-fertility research.

SourceMax Planck Inst. for Demographic Research & Vienna Inst. of Demography
CoverageAbout 37 countries and territories
PeriodLongest series per country, recent to 2024
ScalePeriod + cohort fertility data
LicenseOutputs CC BY 4.0
AccessFree, registration required (no-registration lite exists)
View fields and details

Contents and fields

Four data blocks: ① summary indicators (births, crude birth rate, total fertility rate TFR, tempo-adjusted TFR, mean age at childbearing, cohort cumulative fertility); ② detail by age/birth order; ③ period and cohort fertility tables (incl. PATFR); ④ raw inputs. Standardized methods throughout (Lexis format, population denominators, fertility-table computation).

Suitable research

Demography, fertility and family dynamics, and public-policy benchmarking.

Updates Rolling · Format Tab-separated text / Excel (lite)

Input data is bound by the original licenses of each country's statistical agency; metadata reviewed June 2026.

Why it's hard to get on your ownNational birth and population records differ in definition, and the age and birth-order dimensions are uneven. Aligning to the Lexis format, unifying denominators and recomputing fertility tables by hand is time-consuming and error-prone. Here you get standardized, cross-comparable data as a complete set.
Health & medicineDS-14

Demographic and Health SurveysDHS Program · Demographic and Health Surveys

Nationally representative household micro-data covering developing countries, with unified questionnaire definitions — a hard-to-replace primary source for global health and development research.

SourceImplemented by ICF (transitional funding from the Gates Foundation)
Coverage90+ countries · 400+ surveys
Period1984 to present
ScaleNationally representative household micro-data
LicenseFree distribution agreement after registration
AccessFree, application required (24–48h review)
View fields and details

Contents and fields

Modules on fertility and TFR, family planning and contraception, maternal and child health (immunization, illness and survival), nutrition, HIV and malaria, biomarkers and more; organized by recode files (women/children/household/men/HIV, etc.). Survey types include standard DHS, Malaria Indicator Surveys, AIDS Indicator Surveys and Service Provision Assessments.

Suitable research

Global health, and population and development economics.

Format Stata / SPSS / SAS / ASCII; summary indicators via STATcompiler / API · Updates Rolling by country + round

Micro-data must be applied for per the project statement and restricted to the stated use; metadata reviewed June 2026.

Why it's hard to get on your ownMultiple survey rounds, multiple survey types and a field structure split by recode file mean a lot of preparation before the data is analysis-ready: checking definitions version by version, aligning codes and linking across files. We have completed the version mapping and field alignment.
Health & medicineDS-15

Global Burden of Disease Study GBD 2021Global Burden of Disease

A widely cited benchmark for estimating global disease burden, covering 371 diseases and injuries and 88 risk factors for epidemiology and health-policy research.

SourceIHME, University of Washington
Coverage204 countries and territories + subnational
Period1990–2021
Scale371 conditions · 88 risk factors
LicenseFree non-commercial user agreement
AccessFree, registration required
View fields and details

Contents and fields

DimensionValues
MeasuresDeaths, DALYs, YLLs, YLDs, prevalence, incidence, HALE
StrataLocation, year, age, sex, cause, risk factor
UnitNumber / rate / percent

Suitable research

Disease burden, epidemiology and health-policy evaluation.

Version GBD 2021 · Format Web tables / visualizations / CSV (GBD Results tool)

Geographic granularity is described at country/region/subnational levels; place-name conventions follow China's official standards. Metadata reviewed June 2026.

Why it's hard to get on your ownCause classification, risk-factor attribution and aligning definitions across years are intricate. Small differences in indicator definitions or unit conversions across versions can change your conclusions. We have completed the version mapping and definition cross-checks.
MACHINE LEARNING & CORPORA · 3 IN THIS COLUMN
Images · Text · Multilingual
Very large open corpora for computer vision and NLP — substantial in volume, with a learning curve to get started.
Machine learning & corporaDS-16

ImageNet large-scale image databaseImageNet

A foundational benchmark for computer-vision research: over ten million hand-annotated images and more than twenty thousand categories, used by academia and industry as a common yardstick since 2009.

SourceStanford Vision Lab + Princeton
CoverageAbout 14.2 million images · 21,841 classes
PeriodFrom 2009
ScaleSubset ILSVRC-1K: about 1.2M training images
LicenseNon-commercial research and education only
AccessFree, registration required
View fields and details

Contents and fields

Natural images collected from the web, a category label per image (mapped to a WordNet synset), WordNet noun-hierarchy relations, and bounding boxes for object localization in some subsets. Organized by the WordNet noun tree, targeting about 1,000 images per synset.

Suitable research

Image classification, object localization, transfer learning and model-evaluation benchmarks.

Version ImageNet-21K / milestone subset ILSVRC2012 · Format Images (JPEG) + annotations

Since 2019 the official site closed the full 21K download and keeps only the ILSVRC subset; non-commercial academic research only. Metadata reviewed June 2026.

Why it's hard to get on your ownA model can perform well on your own samples yet rank unexpectedly on a recognized benchmark. What's missing isn't compute — it's a yardstick the field agrees on. Leave benchmark access and subset alignment to us.
Machine learning & corporaDS-17

OPUS open parallel corpusOPUS · Open Parallel Corpus

A large-scale open collection of multilingual parallel corpora — over a thousand languages and thousands of language pairs — used for machine translation and multilingual NLP.

SourceHelsinki-NLP, University of Helsinki
CoverageAbout 1,005 languages · 1,214 corpora
PeriodSub-corpora updated on a rolling basis
ScaleAbout 102.9 billion sentence pairs
LicensePer each sub-corpus
AccessFree, no registration
View fields and details

Contents and fields

Bilingual/multilingual sentence-aligned text (bitext): source sentence, target sentence, language-pair identifier, sub-corpus source identifier, sentence-alignment information (XCES stand-off); some processed versions include tokenization, lemmatization and part-of-speech tagging.

Suitable research

Machine translation, cross-lingual models and multilingual NLP.

Format XML+alignment / TMX / Moses plain text · Tools OpusTools / API

Licenses vary by sub-corpus and must be checked before use; metadata reviewed June 2026.

Why it's hard to get on your ownSources are scattered, versions and alignment definitions differ, and sub-corpus formats and licenses each vary. Sorting and integrating them one by one is laborious. We have completed the mapping and unified delivery, so you can take it by language pair, ready to use.
Machine learning & corporaDS-18

Common Crawl web corpusCommon Crawl

A petabyte-scale, standardized web archive covering public web pages worldwide and updated monthly — a base corpus for large-model pretraining and large-scale text research.

SourceCommon Crawl Foundation
CoveragePublic web pages worldwide (over 300 billion pages)
PeriodFrom 2008, monthly updates
ScaleAbout 2.1 billion pages per month / petabyte-scale cumulative
LicenseCommon Crawl Terms
AccessFree, no registration (AWS S3 / HTTPS / HF)
View fields and details

Contents and fields

FormatMeaning
WARCRaw HTTP requests/responses (incl. HTML)
WATExtracted metadata (links, titles, etc. as JSON)
WETExtracted plain-text body only

Also includes URL indexes (CDXJ/columnar) and a hyperlink graph. Main fields: URL, crawl timestamp, HTTP status, MIME type, content and plain text.

Suitable research

Large-scale corpus research, natural language processing and large-model pretraining.

Version CC-MAIN-2026-21 · Format WARC/WAT/WET (gzip) + columnar index

Web-page content copyright belongs to the original sites; users are responsible for compliance cleaning and filtering, and for following each source's license terms and the laws that apply to their use case. Metadata reviewed June 2026.

Why it's hard to get on your ownCrawling the whole web yourself, deduplicating, aligning formats and maintaining definitions across monthly versions tends to cost large amounts of compute and engineering time, and is hard to reproduce. We have mapped out the formats, fields and index structure, so you can take a research-ready corpus directly.
DIDN'T FIND WHAT YOU NEED?
Precise search for hard-to-find open datasets
Tell us the data your research needs and the criteria it must meet. We run real searches across official, institutional and well-documented public sources, and verify matches and gaps against each required criterion.
01
State your criteria
Topic, variables, time and geographic range, format and definition requirements.
02
Availability assessment
We first judge whether the public data can be obtained and where the compliance limits lie, to avoid wasted effort.
03
Real search, item-by-item verdicts
A real multi-source search, judging matches and gaps against each required criterion.
04
Availability report delivered
On a match: dataset notes and source links. When nothing matches, the search directions and near-equivalent sources are still presented.
Open the dataset search assistant See a sample reference card →
DATA-FINDING GUIDES
First, understand how to find data
Companion guides on where to find public data, how to choose a platform, whether a license allows commercial use, and how to look up official data.
Where to find research data Comparing open dataset platforms Where to find ML datasets A primer on panel data Can a dataset be used commercially How to look up official statistics
Need data that isn't among these eighteen?
Tell us your research-data needs. We assess availability first, then run a real search and organize the results.
Start a dataset request See prepared datasets
Search for data