An open-source pipeline developed by the Library of Congress for reprocessing the OCR of National Digital Newspaper Program (NDNP) data. The application efficiently creates new ALTO XML and PDF files at scale, can be deployed locally or in common cloud environments, uses Tesseract Open-Source OCR engine and custom post-processing steps. - View it on GitHub
Star
7
Rank
2061112