MétaCan
Menu
Retour à la cohorte
Enregistrement W1777383459

Crunching big data in the cloud with Hadoop and BigInsights

2011· article· en· W1777383459 sur OpenAlexaff
Leons Petrazickis, Bradley Steinfeld

Notice bibliographique

RevueConference of the Centre for Advanced Studies on Collaborative Research · 2011
Typearticle
Langueen
DomaineHealth Professions
ThématiqueArtificial Intelligence in Healthcare
Établissements canadiensIBM (Canada)
Organismes subventionnairesnon disponible
Mots-clésComputer scienceBig dataCloud computingPetabyteThe InternetData scienceUnstructured dataWeb trafficVolume (thermodynamics)Data miningWorld Wide Web
DOInon disponible

Résumé

récupéré en direct d'OpenAlex

There is an ongoing information explosion in every field of human endeavour. Enormous, unstructured, immensely valuable data sets are being accumulated. Every device logs numbers and audio and video and text, and then this data is aggregated and stored somewhere. Unfortunately, traditional techniques cannot analyze these data sets. There's too much data to query -- volume! And it's all different -- variety! And it's arriving too fast -- velocity! Finance firms need to analyze transactions to detect fraud and model risk. Energy firms need to analyze old rig performance and wind speeds. IT needs to analyze logs of every type. Service providers need to analyze the prices of various services worldwide. Healthcare providers need to analyze patient data and measurements. New techniques are needed to deal with this Big Data. Fortunately, cloud computing is allowing the emergence of technologies that rely on clusters of commodity hardware to crunch data. Google is one example of a company that had to face and solve a Big Data problem before it could revolutionize internet search and consign countless other early search engines to the dust heap of history. Its page-rank algorithm that it uses to rank results is based on something called Map-Reduce. In the Map phase, the data set (all of the internet) is split into itsy-bitsy chunks, which are transformed from unstructured data (information about a web page) to useful data (the value of web page). In the Reduce phase, the transformed itsy-bitsy chunks are reassembled back together into a Google results page. Because the chunks are itsy-bitsy, the mapping could run on many off-the-shelf computers rather than one big server. This allowed Google to use cheap commodity hardware and put competitors that relied on expensive servers out of business. The hardware side of this approach created Cloud Computing, which is a way of organizing vast amounts of cheap hardware on demand. The software side of this created a lot of useful data analytics applications. Apache Hadoop is one useful data analytics application that's native to the cloud. Hadoop is an open source project led by Yahoo. It makes it straightforward to apply the idea of Map-Reduce to any data set. Many vendors such as Cloudera, HortonWorks, and IBM have their own distribution of Hadoop. The Hadoop ecosystem includes many other open source technologies. The Java libraries are enhanced by the Pig high level language, the HBase database, the Hive data warehouse system, and the Flume log aggregation service. Each of these makes Hadoop more powerful at dealing with larger volumes of data, greater varieties of data, and quicker velocities of data. IBM InfoSphere BigInsights is a distribution of Hadoop. It integrates an IBM-created open source query language called JAQL (Jackal) with the usual components such as Hive, HBase, and Pig. JAQL allows the user to query through large sets of data in JSON (JavaScript Object Notation) form, which is the native data format of Hadoop. The Basic edition of BigInsights is available for download at no charge. It can also be easily deployed on Amazon Elastic Compute Cloud or IBM SmartCloud Enterprise.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,003
score de la tête « metaresearch » (Gemma)0,006
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesÉtudes des sciences et des technologies
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Qualitatif · Signal consensuel: Qualitatif
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,304
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0030,006
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,001
Études des sciences et des technologies0,0020,001
Communication savante0,0000,000
Science ouverte0,0010,001
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,728
Tête enseignante GPT0,576
Écart entre enseignants0,152 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeQualitatif
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2011
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueConference of the Centre for Advanced Studies on Collaborative ResearchMême sujetArtificial Intelligence in HealthcareTravaux en français237 207