Enhancing text recognition of damaged documents through synergistic OCR and large language models

bibb.id784197
bibb.participationBIBB-Mitarbeiterde
dc.contributor.authorAsselborn, Thomas [Verfasser]de
dc.contributor.authorDörpinghaus, Jens [Verfasser]de
dc.contributor.authorKausar, Faraz [Verfasser]de
dc.contributor.authorMöller, Ralf [Verfasser]de
dc.contributor.authorMelzer, Sylvia [Verfasser]de
dc.date.accessioned2025-12-02T11:18:46Z
dc.date.available2025-12-02T11:18:46Z
dc.date.issued2024
dc.description.abstract"Optical Character Recognition (OCR) remains a highly relevant area of research in pattern recognition. Its applications span various domains, including supporting reading for the visually impaired, interpreting Morse codes, capturing postal addresses, evaluating emails, scanning price tags and passports, and extracting text from digitised documents. As the volume of digitised data continues to grow, challenges arise in capturing the semantic structure of documents through logical structure analysis and providing data suitable for information retrieval to answer specific research questions. While classic OCR processes like Tesseract and OCRopus work well for contemporary digitised documents, there is room for improvement in text and word recognition of historical documents that are severely damaged. Large Language Models (LLMs) like GPT-4 can be effectively used for text recognition tasks, utilising their advanced natural language processing capabilities to interpret and reconstruct unclear or damaged text, offering potential for improving the overall text recognition process. However, challenges arise additionally when documents contain e.g. a mixture of single-column and double-column text, images and text, or words not known or blocked by the agents." (Authors‘ abstract; BIBB-Doku)de
dc.description.statementofresponsibilityThomas Asselborn, Jens Dörpinghaus, Faraz Kausar, Ralf Möller, Sylvia Melzerde
dc.description.versionreferiertde
dc.format.extentSeite 29-36de
dc.format.illustrationIllustrationende
dc.format.illustrationFotografiende
dc.format.mediumElektronische Ressourcede
dc.format.mediumSammelbandbeitragde
dc.identifier.uriDOI:10.15439/2024F7400
dc.identifier.urihttps://bibb-dspace.bibb.de/jspui/handle/BIBB/784197
dc.language.isoende
dc.rdacarrier.codecrde
dc.rdacarrier.sourcerdacontentde
dc.rdacarrier.termOnline-Ressourcede
dc.rdacontent.codetxtde
dc.rdacontent.sourcerdacontentde
dc.rdacontent.termTextde
dc.rdamedia.codecde
dc.rdamedia.sourcerdacontentde
dc.rdamedia.termComputermediende
dc.relation.ispartofAnnals of Computer Science and Information Systems, (2024), Vol. 41: Communication Papers of the 19th Conference on Computer Science and Intelligence Systems (FedCSIS), September 8–11, 2024. Belgrade, Serbia / M. Bolanowski [Hrsg.] ; M. Ganzha [Hrsg.] ; L. Maciaszek [Hrsg.] ; M. Paprzycki [Hrsg.] ; D. Ślęzak [Hrsg.]
dc.subject.classificationG 2.2.1 Technologisierungde
dc.subject.ddc370de
dc.subject.ddc600de
dc.subject.otherInformationstechnikde
dc.subject.otherTexterkennungde
dc.subject.otherDigitalisierungde
dc.subject.otherLarge Language Modelde
dc.titleEnhancing text recognition of damaged documents through synergistic OCR and large language modelsde

Files