DeepSData
Dataset guide · Machine learning & corpora

Heinrich von Kleist's prose texts corpus analysis using LancsBox X

This corpus contains Heinrich von Kleist's eight Novellen, two essays, and political writings, sourced from Project Gutenberg, analyzed with LancsBox X for word and cluster frequencies.

← Back to dataset library · 中文版

Machine learning & corporaFree to access

Key facts

InstitutionHarvard Dataverse
CoverageHeinrich von Kleist's prose works (Novellen, essays, political writings)
Time span1789–1811
Scale2 files
LicenseLicense statement is still being checked
AccessHarvard Dataverse

Contents & fields

The corpus includes word and cluster frequencies (one-word to ten-word units, min. frequency 2), relative frequencies, Average Reduced Frequency, range, range percentage, coefficient of variation, Juilland's D, and deviation of proportions. A second file records derivatives of the word "Recht" with cluster counts. No full field-level description is provided on the source page.

Research uses

Suitable for linguistic analysis of Kleist's writing style, literary studies, and exploration of links between his philosophical and political ideas.

Information comes from the source page. Please check that page for current details and terms.

Keywords

KleistcorpusLancsBox Xword frequencycluster analysisGerman literature

Access & license

License: License statement is still being checked | Free to access

Why this is hard to get on your own

Access detailed word frequency and cluster data for Kleist's prose, but field-level documentation is missing.

Related datasets

Same domain

Need this data retrieved and prepared?

Tell us your hard requirements. We first assess availability, then retrieve for real — and if it truly cannot be obtained, we say so plainly.