Crunching big data in the cloud with Hadoop and BigInsights
Bibliographic record
Abstract
There is an ongoing information explosion in every field of human endeavour. Enormous, unstructured, immensely valuable data sets are being accumulated. Every device logs numbers and audio and video and text, and then this data is aggregated and stored somewhere. Unfortunately, traditional techniques cannot analyze these data sets. There's too much data to query -- volume! And it's all different -- variety! And it's arriving too fast -- velocity! Finance firms need to analyze transactions to detect fraud and model risk. Energy firms need to analyze old rig performance and wind speeds. IT needs to analyze logs of every type. Service providers need to analyze the prices of various services worldwide. Healthcare providers need to analyze patient data and measurements. New techniques are needed to deal with this Big Data. Fortunately, cloud computing is allowing the emergence of technologies that rely on clusters of commodity hardware to crunch data. Google is one example of a company that had to face and solve a Big Data problem before it could revolutionize internet search and consign countless other early search engines to the dust heap of history. Its page-rank algorithm that it uses to rank results is based on something called Map-Reduce. In the Map phase, the data set (all of the internet) is split into itsy-bitsy chunks, which are transformed from unstructured data (information about a web page) to useful data (the value of web page). In the Reduce phase, the transformed itsy-bitsy chunks are reassembled back together into a Google results page. Because the chunks are itsy-bitsy, the mapping could run on many off-the-shelf computers rather than one big server. This allowed Google to use cheap commodity hardware and put competitors that relied on expensive servers out of business. The hardware side of this approach created Cloud Computing, which is a way of organizing vast amounts of cheap hardware on demand. The software side of this created a lot of useful data analytics applications. Apache Hadoop is one useful data analytics application that's native to the cloud. Hadoop is an open source project led by Yahoo. It makes it straightforward to apply the idea of Map-Reduce to any data set. Many vendors such as Cloudera, HortonWorks, and IBM have their own distribution of Hadoop. The Hadoop ecosystem includes many other open source technologies. The Java libraries are enhanced by the Pig high level language, the HBase database, the Hive data warehouse system, and the Flume log aggregation service. Each of these makes Hadoop more powerful at dealing with larger volumes of data, greater varieties of data, and quicker velocities of data. IBM InfoSphere BigInsights is a distribution of Hadoop. It integrates an IBM-created open source query language called JAQL (Jackal) with the usual components such as Hive, HBase, and Pig. JAQL allows the user to query through large sets of data in JSON (JavaScript Object Notation) form, which is the native data format of Hadoop. The Basic edition of BigInsights is available for download at no charge. It can also be easily deployed on Amazon Elastic Compute Cloud or IBM SmartCloud Enterprise.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.020 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.003 | 0.007 |
| Science and technology studies | 0.004 | 0.004 |
| Scholarly communication | 0.008 | 0.015 |
| Open science | 0.006 | 0.012 |
| Research integrity | 0.002 | 0.009 |
| Insufficient payload (model declined to judge) | 0.007 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".