Knowledge Representation and Artificial Intelligence for Management of Socio-Technical Risks in Megaprojects
Bibliographic record
Abstract
Megaprojects are probabilistically dependent systems prone to progressive failures that are undertaken in significantly incentivized economic and political domains. Current processes in definition, estimation, and financing of these large projects often exclude relevant sources of socio-technical risks with overarching effects on the project. Applications of the artificial intelligence methods in forecasting project performance and outcomes considering socio-technical sources of risks have been mainly ad-hoc solutions catered to specific needs, project types, and data availability. This research proposes a high-level framework for dealing with socio-technical risks by utilizing available sources of data, expert knowledge, and the most applicable analytics methodologies. The framework consists of representation, quantification, and inference connected through a loop of dynamic learning. The representation part defines various sources of risk in an expandable data format and with universal semantics, considering the nature of risk factors as to when and how they should be collected, in conjunction with project performance measures. An ontology is developed for such representation using the linked data and the semantic web format. The quantification aims to measure such sources of risk, which could vary from machine learning algorithms' applications over remote sensing data to pure qualitative judgments in discrete scales. The quantification was exemplified by creating a remoteness risk index, called Nighttime Remoteness Index (NIRI) for risk and resilience assessment of remote projects. The remoteness index takes nighttime satellite imagery as input and produces a remoteness index using machine learning algorithms. The index was validated based on two census bases remoteness indices of Australia and Canada. The inference part aims to incorporate the effects and dependencies of socio-technical risk factors on project outcome variables by defining an expandable object-oriented Bayesian Network (OOBN). The models can be trained based on the collected data from knowledge representation or based on expert knowledge, and different extents of their combinations. The employed methodologies allow for evergreen process of data collection and process optimization in the previous three sections. The framework creates a systematic approach to capturing and modeling project socio-technical risks, enables quality controls across the process, and can be applied to various risk sources such as the Environmental, Social, and Governance (ESG) factors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.004 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.006 | 0.009 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".