MétaCan
Menu
Back to cohort
Record W7027498689

Conception d'un système de reconnaissance optique de caractères imprimés et manuscrits sur formulaires fédéraux et documents historiques

2024· other· fr· W7027498689 on OpenAlexaboutno aff

Bibliographic record

VenueSémaphore (Université du Québec à Rimouski) · 2024
Typeother
Languagefr
Field
Topic
Canadian institutionsnot available
Fundersnot available
KeywordsFilter (signal processing)Context (archaeology)Subject (documents)Paraphernalia
DOInot available

Abstract

fetched live from OpenAlex

« La préservation, la retranscription et l'accessibilité des documents manuscrits représentent un enjeu coûteux et complexe pour le gouvernement du Canada. De plus, les données pouvant être de nature confidentielle, l'utilisation de solutions privées représente un risque pour l'intégrité de l'information. Ce projet de recherche vise ainsi à développer un système de reconnaissance optique des caractères (OCR) sur des formulaires structurés puis de quantifier le taux d'exactitude et le temps de traitement requis afin d'établir une référence pour de futurs développements. Le taux d'exactitude visé doit être supérieur à 85% en comparaison à un opérateur humain et doit traiter un document de quinze cases en moins d'une minute pour être économiquement viable. Le projet doit aussi permettre de déterminer des pistes de solutions en vue de développer une solution plus complexe pour des documents non-structurés de nature historique. Afin d'atteindre ces objectifs, les principales étapes d'un système de reconnaissance de caractères sont étudiées et décomposées : le prétraitement, la segmentation, l'extraction des caractéristiques, la classification d'images et le post-traitement. Des solutions telles que la méthode ORB, la détection de contours, la transformée en cosinus discrète (DCT), la transformée en ondelettes discrète (DWT), la transformée de Hough et 190 réseaux de neurones différents sont notamment utilisés afin de détecter la position du texte dans l'image, d'extraire les caractéristiques et réaliser la classification des caractères. Sur un formulaire structuré avec une écriture typographique, le taux d'exactitude moyen après post-traitement est de 91,99% et il faut en moyenne 4,02 s pour traiter une case. Pour une écriture manuscrite, le taux d'exactitude après post-traitement est de 94,27% et il faut en moyenne 5,90 s pour traiter une case atteignant ainsi les objectifs fixés en termes de taux d'exactitude. Des améliorations restent cependant à apporter, notamment au niveau de l'écriture typographique, du seuillage ainsi que la segmentation de caractères collés pour améliorer l'exactitude. Au niveau du temps de traitement, la méthode proposée à l'aide de fenêtrage ainsi que le dictionnaire pour les prénoms sont des avenues moins prometteuses. Au niveau des documents historiques non-structurés, ceux-ci ont seulement été abordés. Toutefois, les résultats obtenus pour les formulaires structurés permettent d'établir que les principaux défis se trouvent au niveau de la détection du texte, la correction de l'alignement et de l'inclinaison ainsi que dans la segmentation de l'écriture cursive. Une solution à base de réseau de propositions régionales (RPN), de réseau de neurones convolutifs (CNN) et de réseaux de neurones récurrents (RNN) est suggérée afin de pallier certaines faiblesses observées. -- Mot(s) clé(s) en français : Intelligence artificielle, reconnaissance de caractères manuscrits, HCR, documents historiques, réseaux de neurones, OCR, reconnaissance optique de caractères. »-- « The preservation, transcription, and accessibility of handwritten documents represent a costly and complex challenge for the government of Canada. Additionally, as the data may be of a confidential nature, the use of a private solution poses a risk to the integrity of the information. Therefore, this research project aims to develop an optical character recognition (OCR) system for structured forms and to quantify the accuracy rate and processing time required to establish a benchmark for future development. The targeted accuracy rate should exceed 85% compared to a human operator and should process a fifteen-box document in less than a minute to be economically viable. The project should also identify potential solutions for developing a more complex solution for unstructured documents of a historical nature. To achieve these objectives, the main steps of a character recognition system are studied and broken down: preprocessing, segmentation, feature extraction, image classification, and post-processing. Solutions such as the ORB method, contour detection, discrete cosine transform (DCT), discrete wavelet transform (DWT), Hough transform, and 190 different neural networks are notably used to detect the position of text in the image, extract features, and perform character classification. On a structured form with typewritten text, the average accuracy rate after post-processing is 91.99%, and it takes an average of 4.02 seconds to process a box. For handwritten text, the accuracy rate after post-processing is 94.27%, and it takes an average of 5.90 seconds to process a box, thus meeting the set accuracy objectives. However, improvements are still needed, particularly in terms of typewritten text, thresholding, and segmenting stuck characters to enhance accuracy. Regarding processing time, the proposed method using windowing and a dictionary for first names are less promising avenues. Regarding unstructured historical documents, they have only been addressed. However, the results obtained for structured forms indicate that the main challenges lie in text detection, alignment, and tilt correction, as well as in the segmentation of cursive writing. A solution based on regional proposal networks (RPN), convolutional neural networks (CNN), and recurrent neural networks (RNN) is suggested to address some of the observed weaknesses. -- Mot(s) clé(s) en anglais : Artificial intelligence, Handwritten character recognition, HCR, Historical documents, Neural network, OCR, Optical character recognition. »--

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.006
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.012
Threshold uncertainty score0.037

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.006
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0010.001
Scholarly communication0.0050.004
Open science0.0020.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0110.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.012
GPT teacher head0.231
Teacher spread0.219 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueSémaphore (Université du Québec à Rimouski)French-language works237,207