Converting German historical legal documents to TEI XML including challenges with table extraction
item.page.bibb.id
784192
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
item.page.bibb.publisherplace
item.page.bibb.participation
BIBB-Mitarbeiter
item.page.bibb.citedinbibb
item.page.bibb.researchfocus
item.page.bibb.reviewof
item.page.bibb.citationdr
Abstract
"The job archive at the Federal Institute for Vocational Education and Training contains thousands of historical German VET and CVET regulations from the last 100 years. However, these are hardly accessible because they are currently only available in their original paper form. We present a workflow that transcribes images of these regulations into the TEI XML format which preserves the logical document structure and stores metadata. It is widely used for digital archives and represents an important step towards a fully digitalized archive. This paper addresses issues caused by poor page segmentation of the applied OCR methods and presents rules that can reconstruct a large part of the documents' hierarchy. A straightforward table recognition method for tables with borders is presented, as well as a metadata extraction procedure for the selected data set. While our approach is generic and functional, further research is necessary to develop a fully automated workflow." (Authors‘ abstract; BIBB-Doku)
Description
Keywords
Citation
item.page.bibb.voevzlink
item.page.bibb.additionallink
Collections
Annals of Computer Science and Information Systems, (2024), Vol. 41: Communication Papers of the 19th Conference on Computer Science and Intelligence Systems (FedCSIS), September 8–11, 2024. Belgrade, Serbia / M. Bolanowski [Hrsg.] ; M. Ganzha [Hrsg.] ; L. Maciaszek [Hrsg.] ; M. Paprzycki [Hrsg.] ; D. Ślęzak [Hrsg.]
aktuell und lesenswert Ausgabe 3/2025- SBB und Zeitschriftenartikel
Foko 1/26 BIBB-VÖ
aktuell und lesenswert Ausgabe 3/2025- SBB und Zeitschriftenartikel
Foko 1/26 BIBB-VÖ