MétaCan
Menu
Retour à la cohorte
Enregistrement W764208678

Which image processing algorithms best describe the minimal amount of visual information required for immage recognition

2003· article· en· W764208678 sur OpenAlexaboutno aff
Cédric Laloyaux, Christian Schmitt

Notice bibliographique

RevueeScholarship (California Digital Library) · 2003
Typearticle
Langueen
DomaineNeuroscience
ThématiqueNeuroscience and Neural Engineering
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésComputer scienceImage processingArtificial intelligenceObject (grammar)AlgorithmPreprocessorComputer visionCognitive neuroscience of visual object recognitionInformation processingImage (mathematics)Psychology
DOInon disponible

Résumé

récupéré en direct d'OpenAlex

Which image processing algorithms best describe the minimal amount of visual information required for image recognition? Cedric Laloyaux (claloyau@ulb.ac.be) Cognitive Science Research Unit, Universite Libre de Bruxelles CP 122, Av. F.-D. Roosevelt, 50, 1050 Bruxelles, Belgium Cedric Schmitt (cedric.schmitt@umontreal.ca) LBUM, CHUM, Hopital Notre-Dame, Pavillon J.A de Seve (Y-1619) 2099 Alexandre de Seve, Montreal, Quebec, H2L 2W5, Canada Introduction In the framework of visual prosthesis, most methods use a video-camera in order to record visual information (Margalit et al., 2002). Determining the relevant level of information (e.g. contours, junctions) to be send to the retina, the optic nerve or a sensory substitution device for object recognition purpose, is a crucial aspect of raw images preprocessing (Boyle, 2002). For example, in Delbeke et al. (2002), a subject, with a four-contact optic nerve electrode, was able to discriminate simple patterns like “L” or “+” but nothing is said about the image processing algorithm to use in real- world conditions. In this research, we aim to investigate the minimal amount of information that different image processing algorithms provide, for object and scene recognition tasks. Experiment The goal of the study was to compare different image processing algorithms to determine which algorithms provide the best descriptions of the minimal amount of perceptual information required for image recognition. Two categories of images were compared. The following image processing algorithms were applied to 15 images of objects (object category) and 15 pictures of indoor and outdoor scenes likely to be encountered by an observer in motion (mobility category): (1) Canny edge detector based on broadband frequencies (Canny BF), (2) Canny edge detector based only on Low frequencies (Canny LF), (3) The center on-off Marr’s model, (4) a simple threshold method and (5) the Sobel (based on the directional gradient approximation of smooth image) (see Mallot, 2000). Fifteen subjects performed 450 trials: 30 images X 3 image resolutions (32x32, 64x64 with a reliable, but slow, sub-sampling method and 64x64 with a poor, but fast, sub- sampling method) were processed by the five different algorithms. Subjects viewed a series of images: first a fixation cross for 1000 ms followed by a color image consisting of 480x480 pixels for 300 ms, and lastly two processed images: one being the original image processed (the target) and the other being any one of the other pictures from the same category (a distractor). In other words, an object target was presented with an object distracter. The target and the distracter were processed using the same image algorithm. The subjects’ task was to press a key corresponding to the position of the target image on the screen. They were instructed to answer as quickly and accurately as possible. Results and Discussion ANOVA were performed on the error rates and the RT’s. There was no trade-off effect. The RT’s were significantly lower for the two 64x64 resolutions compared to the 32x32 [F(2,28)=42.976 ; p<0.001]. There was a main effect of photo type [F(1,14)=50.247 ; p<0.001], as RT were faster for objects than for scenes. This result supports the hypothesis that navigation based on artificial devices requires an adapted image processing step. There was also a main effect of image processing algorithms [F(4,56)=6.6670; p<0.001]. A planned comparison between the Canny and the Canny LF showed that RT’s were significantly lower for the Canny LF (p<0.05). The thresholding method was also very efficient, and had the same level of efficiency as the Canny LF. However, the number of pixels to send after image processing was 10 times greater for the thresholding method than for the Canny LF. All together, the results showed that human performance differs greatly depending on the image processing algorithms used, and that these algorithms do not require the same amount of pixels in order for images to be minimally recognizable. References Boyle, J.R., Maeder, A.J., Boles, W.W. (2002). Image enhancement for electronic visual prostheses. Australas Phys Eng Sci Med, 25,(2), 81-6. Delbeke, J., Wanet-Defalque, M.C., Gerard, B., Troosters, M., Michaux, G., Veraart, C. (2002). The microsystems based visual prosthesis for optic nerve stimulation. Artificial Organs, 26,(3), 232-234. Margalit, E., Maia, M., Weiland, J.D., Greenberg, R.J., Fujii, G.Y., Torres, G., Piyathaisere, D.V., O'Hearn, T.M., Liu, W., Lazzi, G., Dagnelie, G., Scribner, D.A., de Juan. E., Humayun, M.S. (2002). Retinal prosthesis for the blind. Survey of Ophatlmology, 47,(4), 335-356. Mallot, H.A. (2000). Computational Vision, Cambridge, MA: MIT press.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,003
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesCommunication savante
Catégories consensuellesCommunication savante
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Expérimental (laboratoire) · Signal consensuel: Expérimental (laboratoire)
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,079
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,003
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,001
Études des sciences et des technologies0,0000,000
Communication savante0,0010,015
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,033
Tête enseignante GPT0,252
Écart entre enseignants0,220 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.

Devis d'étudeExpérimental (laboratoire)
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2003
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueeScholarship (California Digital Library)Même sujetNeuroscience and Neural EngineeringTravaux en français237 207