Analyzing historical legal textcorpora : German VET and CVET regulations
item.page.bibb.id
783368
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
item.page.bibb.publisherplace
item.page.bibb.participation
BIBB-Mitarbeiter
item.page.bibb.citedinbibb
item.page.bibb.researchfocus
item.page.bibb.reviewof
item.page.bibb.citationdr
Abstract
„The digitization of historical documents has gained particular interest in recent years. The majority of research endeavors aim at digitizing historical documents by extracting text from scanned images. A pipeline that transcribes scanned documents into fully structured texts was utilized to digitize over 900 German VET and CVET regulations. As a preliminary investigation, a basic corpus analysis was conducted to assess the usability of the digitized documents and the necessity for document digitization methods that can generate transcripts that maintain the logical text structure and hierarchy. This paper focuses on the processing of the transcripts created from German VET and CVET regulation images to demonstrate the advantages of fully structured text over plain OCR results and to illustrate that even simple analyses require more information for more comprehensive document understanding.“ (authors‘ abstract; BIBB-Doku)
Description
Keywords
Citation
item.page.bibb.voevzlink
item.page.bibb.additionallink
Endorsement
Review
Supplemented By
Referenced By
Creative Commons license
Except where otherwise noted, this item's license is described as Namensnennung 4.0 International
