DeepSData
Dataset guide · 机器学习与语料

Heinrich von Kleist's prose texts corpus analysis using LancsBox X

This corpus contains Heinrich von Kleist's eight Novellen, two essays, and political writings, sourced from Project Gutenberg, analyzed with LancsBox X for word and cluster frequencies.

← Back to dataset library · 中文版

机器学习与语料Free to access

Key facts

InstitutionHarvard Dataverse
CoverageHeinrich von Kleist's prose works (Novellen, essays, political writings)
Time span1789–1811
Scale2 files
LicenseLicense pending verification
AccessHarvard Dataverse

Contents & fields

The corpus includes word and cluster frequencies (one-word to ten-word units, min. frequency 2), relative frequencies, Average Reduced Frequency, range, range percentage, coefficient of variation, Juilland's D, and deviation of proportions. A second file records derivatives of the word "Recht" with cluster counts. No full field-level description is provided on the source page.

Research uses

Suitable for linguistic analysis of Kleist's writing style, literary studies, and exploration of links between his philosophical and political ideas.

This card was drafted from the source page; institution, coverage, time span, scale, fields and license are subject to the official page (pending human review).

Keywords

KleistcorpusLancsBox Xword frequencycluster analysisGerman literature

Access & license

License: License pending verification | Free to access

Why this is hard to get on your own

Access detailed word frequency and cluster data for Kleist's prose, but field-level documentation is missing.

Related datasets

Same domain

Need this data retrieved and prepared?

Tell us your hard requirements. We first assess availability, then retrieve for real — and if it truly cannot be obtained, we say so plainly.