A Comprehensive Study and Implementation of Agentic AI via MCP Servers
Bibliographic record
Abstract
Agentic and Multi-Agent Artificial Intelligence (AI) systems are revolutionizing business operations and collaborative decision-making. The convergence of Agentic Artificial Intelligence (Agentic AI) and Multi-Agent Systems (MAS) into enterprise ecosystems marks a paradigm shift to autonomous, adaptive, and context-aware decision-making. A structured approach is taken, classifying the literature into five themes: agentic architectures, multi-agent coordination, ethical governance, tool ecosystems, and enterprise case studies. Comparative analysis picks up improvements like layered agentic frameworks and real-time anomaly detection, as well as revealing substantial deficits in scalability, explainability, persistent learning, and multi-agent orchestration. Standardized evaluation benchmarks, interoperable toolchains, ethical governance models, and hybrid human-agent collaboration frameworks need to be developed through future research. The research concludes that the shift from conceptual innovation to enterprise-level deployment is a multidisciplinary exercise that synergizes AI research, systems engineering, and organizational change management. This review has been designed to steer researchers and practitioners alike to convert Agentic AI into an actionable force for sustainable digital transformation in enterprise settings. Additionally, we introduce a multi-location MCP(Model Context Protocol) Server implementation comprising three servers: (1) a Web/API server, (2) a data source server hosting Text-to-SQL and RAG tools, and (3) a secondary data source server with identical tools, all operating in distributed locations. The MCP servers are critical for aggregating and managing data from heterogeneous sources, enabling seamless integration with Large Language Models (LLMs) for real-time query resolution and reasoning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".