Analyzing historical legal textcorpora : German VET and CVET regulations

item.page.bibb.id

783368
Loading...
Thumbnail Image

Date

Journal Title

Journal ISSN

Volume Title

Publisher

item.page.bibb.publisherplace

item.page.bibb.participation

BIBB-Mitarbeiter

item.page.bibb.citedinbibb

item.page.bibb.researchfocus

item.page.bibb.reviewof

item.page.bibb.citationdr

Abstract

„The digitization of historical documents has gained particular interest in recent years. The majority of research endeavors aim at digitizing historical documents by extracting text from scanned images. A pipeline that transcribes scanned documents into fully structured texts was utilized to digitize over 900 German VET and CVET regulations. As a preliminary investigation, a basic corpus analysis was conducted to assess the usability of the digitized documents and the necessity for document digitization methods that can generate transcripts that maintain the logical text structure and hierarchy. This paper focuses on the processing of the transcripts created from German VET and CVET regulation images to demonstrate the advantages of fully structured text over plain OCR results and to illustrate that even simple analyses require more information for more comprehensive document understanding.“ (authors‘ abstract; BIBB-Doku)

Description

Keywords

Citation

item.page.bibb.voevzlink

item.page.bibb.additionallink

Endorsement

Review

Supplemented By

Referenced By

Creative Commons license

Except where otherwise noted, this item's license is described as Namensnennung 4.0 International