Extending BIRT with geospatial data visualization capabilities by integrating the SOLAPLayers mapping \ncomponent
Bibliographic record
Abstract
Today, Business Intelligence (BI) is an important part for every modern company. BI \nallows analyzing a large amount of data. The overview over the information simplifies to \ntake decisions, which influence the strategies of the company. The different parts of BI \nare supported by a lot of proprietary and open source applications (OSBI). \nSpatialytics from Québec, Canada is a small start‐up company, which established in \n2009. The main goal of Spatialytics is to provide the use of geo‐spatial data for different \nOSBI‐tools. Spatialytics has three software‐tools which covering different aspects of the \nBI‐process. One of these solutions is SOLAPLayers. SOLAPLayers is a reporting tool, \nwhich displays data from different data sources on an interactive, web based dashboard. \nThe main feature is the possibility to retrieve geo‐spatial data and display it on a map \ncomponent. \nThis bachelor thesis forced to restructure the existing SOLAPLayers 2.0 version, to \nprovide a more flexible, extendable and dynamical software component, which can be \nintegrated into other considerable reporting tools. The new version allows providing \nnew data‐source drivers and output formats in an easy way to the framework. In a \nsecond step, a driver for relational databases has been added to the application. This \ndriver was necessary to extend the range of potential users of SOLAPLayers, since not \nevery company owns a data warehouse. The resulting software is not a final version. \nMore data sources and other features will be added to SOLAPLayers before providing \nthe software to the public. \nThe second goal of this project was to create a map component for the popular open \nsource reporting tool BIRT. BIRT is on of the main projects of the Eclipse foundation. \nThe founder and most active collaborator is the company Actuate. Based on different \nfacts, the decision to integrate SOLAPLayers into BIRT was done. Finally, SOLAPLayers is \nused as data source and report item of the new BIRT‐plugin. \nThis document contains the analysis and the project documentation of both parts of this \nbachelor thesis. Additional an excursion on the topic “state of the art of reporting tools” \ncan be found in the appendix of this document. The whole thesis was produced by \nChristoph Süess at Spatialytics in Québec CA during a three month long internship and \nsupervised by Prof. Stefan Keller at the Hochschule für Technik at Rapperswil CH.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.008 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.007 |
| Open science | 0.003 | 0.008 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.028 | 0.015 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".