MétaCan
Menu
Retour à la cohorte
Enregistrement W2103797923 · doi:10.1109/icdm.2009.71

Active Learning with Generalized Queries

2009· article· en· W2103797923 sur OpenAlexaff
Jun Du, Charles X. Ling

Notice bibliographique

Revuenon disponible
Typearticle
Langueen
DomaineComputer Science
ThématiqueMachine Learning and Algorithms
Établissements canadiensWestern University
Organismes subventionnairesnon disponible
Mots-clésOracleComputer scienceProbabilistic logicAsk priceConstruct (python library)Artificial intelligenceActive learning (machine learning)Machine learningInformation retrieval

Résumé

récupéré en direct d'OpenAlex

We study active learning with generalized queries in the thesis. In contrast to supervised learning, active learning can usually achieve the same predictive accuracy with much fewer labeled training examples, thus significantly reducing the labeling cost. However, previous studies of active learning mostly assume that the learner can only ask specific queries (i.e., require labels for specific examples by providing all feature values). For instance, if the task is to predict osteoarthritis based on a patient data set with 30 features, the previous active learners could only ask the specific queries as: does this patient have osteoarthritis, if ID is 32765, name is Jane, age is 35, gender is female, weight is 85 kg, blood pressure is 160/90, temperature is 98F, no pain in knees, no history of diabetes, and so on (for all 30 features). However, amongst all these 30 features, many of them may be irrelevant to osteoarthritis diagnosis (such as, ID, name, history of diabetes, etc.). More importantly, for such specific queries, the answers provided by the oracle are also specific. That is, each responded label is only applicable to one specific query (i.e., one specific example). In real-world situations, the oracles (usually human experts) are often more ready to answer generalized queries, such as “are people over age 50 with knee pain likely to have osteoarthritis?” Here only two relevant features (age and type of pain) are mentioned, and the other 28 are considered as don’t-care. Real-world human oracles usually regard such queries as more intuitive and easy to comprehend. More importantly, as one such generalized query can represent a set of specific ones, the corresponding answer provided by the oracle is also applicable to this whole set of specific queries. For instance, in our previous example, the answer for the proposed query is applicable for all people over age 50 with knee pain. Therefore, the active learner can obtain more information from each generalized query (together with the corresponding answer), and furthermore improve the learning more effectively and efficiently. In this thesis, we assume that the oracle is capable of answering such generalized queries, and develop different algorithms to implement such active learning with generalized queries, according to different real-world scenarios (i.e., under different assumptions). As far as we know, no previous work on active learning can deal with such generalized queries. More specifically, we study active learning with generalized queries from the following four perspectives: We theoretically study why and when such generalized queries can help in active learning, and demonstrate the superiority of generalized queries over specific ones through toy examples and learning theories. (See Chapter 2 for details.) We assume that the oracle can answer generalized queries as easily as specific ones (i.e., with the same effort or cost). Thus we develop two novel active learning algorithms to ask as general as possible queries, and simultaneously keep the answers from the oracle as certain as possible. (See Chapter 3 for details.) We make a more realistic assumption that, the more general a query is, the higher cost (effort) it causes to request the label. We therefore study the generalized queries in a cost-sensitive framework, and discuss two scenarios to, either balance the trade-off of the predictive accuracy and the query cost, or minimize the total cost of misclassification and query. (See Chapter 4 for details.) We consider a more relaxed scenario that the oracle could only provide ambiguous answers to generalized queries. That is, the oracle would only respond with either “positive” (“yes”) or “negative” (“no”), where “positive” indicates that at least one of the examples represented by the generalized query can be labeled positive, and “negative” indicates that all such examples would be labeled negative. We then develop another new algorithm to implement active learning with generalized queries under this condition. (See Chapter 5 for details.) Our study in this thesis has thoroughly addressed the advantages and difficulties of active learning with generalized queries. The theoretical study has proved that the query complexity of active learning with generalized queries is significantly lower than active learning with specific ones. The empirical study for a variety scenarios has also demonstrated that, to achieve certain predictive accuracy, active learning with generalized queries requires us to ask significantly fewer queries (or requires us to spend significantly lower labeling cost), compared with active learning with specific ones.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,008
score de la tête « metaresearch » (Gemma)0,030
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,008
Score d'incertitude au seuil0,042

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0080,030
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0030,001
Bibliométrie0,0010,002
Études des sciences et des technologies0,0010,002
Communication savante0,0030,007
Science ouverte0,0050,004
Intégrité de la recherche0,0040,004
Charge utile insuffisante (le modèle a refusé de juger)0,0040,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,005
Tête enseignante GPT0,226
Écart entre enseignants0,221 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations13
Publié2009
Routes d'admission1
Résumé présentoui

Explorer davantage

Même sujetMachine Learning and AlgorithmsTravaux en français237 207