Notice bibliographique
Résumé
Mining data of complex data types (e.g. spatial, multimedial, etc.) is deemed an important research frontier in data mining. These data types are often currently modelled according to the object-relational data model. In this paper we face the problem of de ning general framework for mining object-relational databases instead of concentrating on speci c forms of data mining tailored to speci c data types (e.g. spatial data mining, multimedia data mining, etc.). Such framework allows us to formulate data mining tasks in application domains characterized by objects, properties of objects, relations between objects and concept hierarchies. The hybrid language AL-log and an ILP approach are the building blocks of the framework. Frequent pattern discovery at multiple levels of description granularity is taken as showcase of ORDM. 1 Background and motivation Data models play relevant role in data mining [8]. Yet data mining techniques often make assumptions on the representation of input data which mismatch the data model adopted by the database to be mined. Indeed most techniques, here collectively referred to as Classical Data Mining (CDM), have been developed for data in the single-table form traditionally used in statistics. However, real-world data is seldom of this form. Rather, relational databases are widely used. Only recently large body of research, named Relational Data Mining (RDM) and aimed at overcoming the limits of CDM in dealing with relational databases, has emerged [6]. We would like to emphasize that RDM is not simply data mining in relational databases. This de nition would not be suAEcient to distinguish it from CDM, which has been nonetheless extensively applied to relational databases. We de ne research on RDM as the study of methods and techniques for discovering relational patterns in relational data to emphasize that the relational data model is an invariant of the discovery process. RDM techniques have been mainly developed within the eld of Inductive Logic Programming (ILP), research area at the intersection of machine learning and logic programming [12]. Data in ILP is expected to be represented in Horn clausal logic. Therefore there is natural t between relational databases and ILP techniques as regards the data model. ILP was initially concerned with the synthesis of logic programs from examples and background knowledge. The recent developments, however, have broadened the range of learning problems of ILP from the traditional predictive tasks (classi cation rules) of machine learning to the descriptive ones more peculiar to data mining. E.g., WARMR [4] is an ILP system for frequent pattern discovery where data and patterns are represented in Datalog [3]. Mining data of complex data types (e.g. spatial, multimedial, etc.) is deemed an important research frontier in data mining [8]. These data types are often currently modelled according to the object-relational data model. In this paper we de ne Object-Relational Data Mining as the study of methods and techniques for discovering object-relational patterns in object-relational data and face the problem of de ning general framework for ORDM instead of concentrating on speci c forms of data mining tailored to speci c data types (e.g. spatial data mining, multimedia data mining, etc.). Such framework allows us to formulate data mining tasks in application domains characterized by objects, properties of objects, relations between objects, and concept hierarchies (or taxonomies). An open question in ORDM research is: what approach? ILP seems good candidate, except for the pure relational data model it adopts. Complex data types require appropriate representation and reasoning means. Description Logics (DLs) are particularly interesting because they have been invented for representing and reasoning with structural knowledge and concept hierarchies [1]. Unfortunately in exchange for the ability to model and reason about value restrictions in domains with rich hierarchical structure, DLs o er weaker than usual query language. This makes also pure DLs inadequate as knowledge representation and reasoning means in ORDM problems. Hybrid languages such as AL-log [5] appear more promising because they combine description logics and functionfree Horn clausal logic. In this paper we show that AL-log can be the starting point for the de nition of general framework for ORDM obtained by upgrading ILP solutions for RDM to ILP solutions for ORDM. Also we extend the work on spatial data mining presented in [10]. The paper is organized as follows. Section 2 recalls the link between description logics and databases. Section 3 de nes the data mining task chosen as ORDM showcase. Section 4 presents the application of the AL-log framework to the ORDM showcase. Section 5 concludes the paper with nal remarks and directions for future work. 2 Description logics and databases Description Logics (DLs) are fragments of rst-order logic with equality [1]. From the beginning DLs have been considered general-purpose languages for knowledge representation and reasoning. They were considered especially e ective for those domains where the knowledge could be easily organized along hierarchical structure, based on the is-a relationship. This motivated the use of DLs as modeling language in the design and maintenance of large, hierarchically structured bodies of knowledge. E.g. the description logic ALC [13] allows for the speci cation of structural knowledge in terms of concepts, roles, and individuals. Individuals represent objects in the domain of interest. Concepts represent classes of these objects, while roles represent binary relations between concepts. Complex concepts can be de ned by means of constructs, such as u and t. An ALC knowledge base is two-component system. The intensional component T consists of concept hierarchies spanned by is-a relations between concepts, namely inclusion statements of the form v D (read C is included in where and D are two arbitrary concepts. The extensional component M speci es instance-of relations, e.g. concept assertions of the form : (read a belongs to C) where is an individual and is concept. In ALC an interpretation I = ( I ; ) consists of set I (the domain of I) and function I (the interpretation function of I). E.g., it maps concepts to subsets of I and individuals to elements of I such that 6= b if 6= b (unique names assumption). We say that I is model for v D if D , and for : if 2 . The main reasoning services for ALC knowledge bases are checking whether logically implies an inclusion (i.e. j= v D) or membership assertion (i.e. j= o : C). The former inference is called subsumption check, the latter instance check. Both checks boil down to the more general problem of checking the satis ability of an ALC knowledge base . The relationship between DLs and databases is also rather strong [1]. Several investigations have been carried out on the usage of DLs to formalize objectoriented data models and semantic data models. In these proposals concept descriptions in DLs are used to present the schema of database. On the other hand, since concept description provides necessary and suAEcient conditions for objects to satisfy it, it is natural to treat it as query, thus leading to uni cation of two traditionally distinct languages: the data de nition and data manipulation languages. Unfortunately in exchange for more expressive description of the schema, DLs pay the price of weaker than usual query language. Queries can only return subsets of existing concepts, rather than creating new concepts (as in standard SQL databases). Furthermore, the selection conditions are rather limited. Given the expressive limitations of DL concepts alone as queries, it has been reasonable to consider extending Datalog queries with DLs. In one approach, exempli ed by the AL-log language [5], ALC concept assertions are used essentially as type constraints on variables appearing in function-free Horn clauses. E.g., q(X) item(X,Y), item(X,Z) & X:Order, Y:DairyProduct, Z:GrainsCereals give avor of what unary conjunctive queries look like in AL-log. Here the concept assertions Y:DairyProduct and Z:GrainsCereals restrict the range of the variables Y and Z to individuals of the concepts DairyProduct and GrainsCereals respectively. We claim that AL-log can be used as formal language for objectrelational data models. 3 A case study for ORDM A good representative of the class of data mining tasks which our framework can elegantly deal with is frequent pattern discovery at multiple levels of description granularity. This task and related issues have been already discussed in [10]. For the sake of brevity we only recall the formal problem statement. North American Customer European Customer Customer South American Customer Canadian Customer USA Customer Argentinian Customer Venezuelan Customer Austrian Customer UK Customer 1
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,000 | 0,005 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».