Converting German historical legal documents to TEI XML including challenges with table extraction

item.page.bibb.id

784192
Loading...
Thumbnail Image

Date

Journal Title

Journal ISSN

Volume Title

Publisher

item.page.bibb.publisherplace

item.page.bibb.participation

BIBB-Mitarbeiter

item.page.bibb.citedinbibb

item.page.bibb.researchfocus

item.page.bibb.reviewof

item.page.bibb.citationdr

Abstract

"The job archive at the Federal Institute for Vocational Education and Training contains thousands of historical German VET and CVET regulations from the last 100 years. However, these are hardly accessible because they are currently only available in their original paper form. We present a workflow that transcribes images of these regulations into the TEI XML format which preserves the logical document structure and stores metadata. It is widely used for digital archives and represents an important step towards a fully digitalized archive. This paper addresses issues caused by poor page segmentation of the applied OCR methods and presents rules that can reconstruct a large part of the documents' hierarchy. A straightforward table recognition method for tables with borders is presented, as well as a metadata extraction procedure for the selected data set. While our approach is generic and functional, further research is necessary to develop a fully automated workflow." (Authors‘ abstract; BIBB-Doku)

Description

Keywords

Citation

item.page.bibb.voevzlink

item.page.bibb.additionallink

Endorsement

Review

Supplemented By

Referenced By