Stylometry with R: A Package for Computational Text Analysis
Maciej Eder, Jan Rybicki and Mike Kestemont
, The R Journal (2016) 8:1, pages 107-121.
Abstract This software paper describes ‘Stylometry with R’ (stylo), a flexible R package for the high level analysis of writing style in stylometry. Stylometry (computational stylistics) is concerned with the quantitative study of writing style, e.g. authorship verification, an application which has considerable potential in forensic contexts, as well as historical research. In this paper we introduce the possibilities of stylo for computational text analysis, via a number of dummy case studies from English and French literature. We demonstrate how the package is particularly useful in the exploratory statistical analysis of texts, e.g. with respect to authorial writing style. Because stylo provides an attractive graphical user interface for high-level exploratory analyses, it is especially suited for an audience of novices, without programming skills (e.g. from the Digital Humanities). More experienced users can benefit from our implementation of a series of standard pipelines for text processing, as well as a number of similarity metrics.
Received: 2015-04-07; online 2015-12-22@article{RJ-2016-007, author = {Maciej Eder and Jan Rybicki and Mike Kestemont}, title = {{Stylometry with R: A Package for Computational Text Analysis}}, year = {2016}, journal = {{The R Journal}}, doi = {10.32614/RJ-2016-007}, url = {https://doi.org/10.32614/RJ-2016-007}, pages = {107--121}, volume = {8}, number = {1} }