DeepSData
Dataset guide · Machine learning & corpora

High_Quality_Dataset_for_Multi_Source_Heterogeneous_Data_Fusion_and_Health_Diagnosis_of_Industrial_Equipment

A high-quality dataset for multi-source heterogeneous data fusion and health diagnosis of industrial equipment, built by Zhejiang Detasent, with a total size of 6.7TB, 83.74 billion structured time-series records, covering 17 categories, 600 industrial equipment, and 1955 fault-equipment-condition combinations.

AccessFree to access
LicenceLicense statement is still being checked
Producer浙江德塔森特

← Back to dataset library · 中文版

Machine learning & corporaFree to access

Key facts

InstitutionZhejiang Detasent
Coverage17 categories, 600 industrial equipment, 1955 fault-equipment-condition combinations
Time spanSee official page
Scale6.7TB, 83.74 billion structured time-series records

Contents & fields

The dataset consists of three layers: self-collected data from data centers, compliantly integrated IEEE open-source industrial datasets, and standardized fused labeled products. It covers multi-modal data including time-series, images, audio, video, and engineering features, with 26 standardized industrial fault modes, covering extreme complex conditions such as load fluctuations, high/low temperatures, and intermittent start-stop. A standardized data governance system was established, including noise reduction, time-series alignment, and double-person double-label triple review, achieving data integrity of 96.5% and labeling accuracy of 98.5%. No field-level description is disclosed on the source page.

Research uses

Suitable for research and model development in industrial equipment fault diagnosis, multi-source heterogeneous data fusion, health management, and predictive maintenance.

Information comes from the source page. Please check that page for current details and terms.

Keywords

industrial equipmentmulti-source heterogeneous datahealth diagnosisfault diagnosismulti-modaltime-series data

Access and licence

Licence: License statement is still being checked | Free to access

Why this is hard to get on your own

Scarcity of industrial fault samples, difficulty in fusing multi-modal data, and lack of compliant data rights confirmation.

Related datasets

Same domain

Need this data retrieved and prepared?

State the conditions that must be met. We first assess availability, then search in earnest; if it cannot be obtained, we say so.