DraCor

DraCor
Type of site
Science
Available inEnglish
OwnerFreie Universität Berlin, University of Potsdam
URLhttps://dracor.org/
CommercialNo
Launched2017; 9 years ago (2017)
Current statusOnline

DraCor (Drama Corpora) is an open digital infrastructure developed for the computational study of European drama from Greco-Roman antiquity to the 20th century. The platform hosts plays encoded in the TEI format across various languages, supporting comparative and computational methods in drama studies. As of 2025, the collection comprised over 4,000 texts in more than 20 languages. Data provided by DraCor has seen widespread use in digital humanities research.[1] The project received the Rahtz Prize for TEI Ingenuity by the TEI Consortium in 2022.[2]

Overview

DraCor aims to create reliable, expandable, and interoperable corpora of dramatic literature. The project emphasises the concept of Programmable Corpora,[3] where the data is not only accessible but also designed for computational analysis through APIs and integration with other tools. The platform strives to adhere to FAIR data principles (Findability, Accessibility, Interoperability, Reusability).

Key features

  • Multilingual corpora: Contains drama corpora in more than 20 languages, primarily European.
  • TEI encoding: Texts are encoded according to the TEI guidelines to maintain structural and semantic consistency.
  • API access: Provides a documented Application Programming Interface for programmatic access to texts and metadata.
  • Network visualisations: Generates network graphs representing character co-occurrences within plays.
  • Data download: Offers options to download subsets of texts, such as speeches or stage directions, as well as network data.
  • Open access: Data is openly available for research and related purposes.
  • Programmable Corpora: Supports integration with external analytical tools and programming languages, with API wrappers available for Python (pydracor[4]) and R (rdracor[5]).

Corpora

DraCor's collection of corpora is continuously growing and covers plays in Dutch, English, French, German, Ancient Greek, Hungarian, Italian, Latin, Polish, Russian, Spanish, Swedish, Ukrainian, and other languages. Each corpus is curated by individual scholars or teams[6] and provides rich metadata alongside TEI-encoded texts, supporting analyses of dramatic structures, character interactions, and related topics.

Tools and usage

The DraCor platform includes basic visualisation tools, particularly for network analysis. It also supports programmatic access to the corpora, enabling integration into computational research workflows. This facilitates various types of analyses, including:

Community, development, impact

DraCor is developed through collaboration among researchers from multiple institutions. The DraCor platform is jointly run at the Freie Universität Berlin and the University of Potsdam.[7] As an open-source project, it actively encourages community contributions and feedback. The DraCor community presented its corpora and associated research projects at the DraCor Summit, a five-day event in Berlin in September 2025.[8]

References

  1. ^ "DraCor Research". dracor.org. Retrieved 15 September 2025.
  2. ^ "Rahtz Prize for TEI Ingenuity". tei-c.org. Retrieved 15 September 2025.
  3. ^ Fischer, Frank; Börner, Ingo; et al. (2019). Programmable Corpora: Introducing DraCor, an Infrastructure for the Research on European Drama. DH2019: “Complexities”. Utrecht University. doi:10.5281/zenodo.4284002.
  4. ^ "pydracor". Python Package Index. Retrieved 21 May 2025.
  5. ^ "rdracor". Comprehensive R Archive Network. 26 September 2024. Retrieved 21 May 2025.
  6. ^ "DraCor Corpus Registry". dracor.org. Retrieved 15 September 2025.
  7. ^ "DraCor Credits". dracor.org. Retrieved 15 September 2025.
  8. ^ "DraCor Summit". summit.dracor.org. Retrieved 15 September 2025.

Content Disclaimer

Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.

  1. The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
  2. There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
  3. It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
  4. Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
  5. Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.