Key facts
Contents & fields
The dataset includes development and test speech, audio, and image data, along with answer keys, enrollment, and trial files. Data is drawn from the WeCanTalk corpus collected by LDC, containing approximately 355 hours of CTS audio (8kHz A-law sphere format), 53 hours of AfV segments (16kHz FLAC-compressed MS-WAV), 39 hours of video clips (mp4), and 202 selfie images (JPG). No field-level description is publicly available on the source page.
Research uses
Suitable for text-independent speaker recognition technology research, evaluation system development, and performance benchmarking.
Information comes from the source page. Please check that page for current details and terms.
Keywords
Access & license
License: Terms: check the source page | Access: check the source page
Why this is hard to get on your own
Access is restricted and requires request; license terms are custom.
