Clasificación de ofertas de vivienda en Madrid mediante modelos de aprendizaje automático y su visualización interactiva
Bibliographic record
Abstract
En un mercado inmobiliario caracterizado por un incremento sostenido de los precios y la urgencia en la toma de decisiones que ello provoca, surge la necesidad de adoptar un enfoque basado en datos que permita a los distintos actores identificar directamente las oportunidades de compraventa relevantes. Este proyecto presenta una solución tecnológica basada en aprendizaje automático para la valoración de las ofertas de vivienda ubicadas en el municipio de Madrid y su posterior visualización a través de un cuadro de mando interactivo. Para ello, se parte de un conjunto de datos publicado por Idealista sobre viviendas en venta durante el año 2018, al que se aplican procesos de limpieza de datos, tratamiento de valores atípicos y normalización de variables. A continuación, se entrenan y estudian modelos de regresión para la predicción del precio de los inmuebles en función de sus características y se utiliza su desviación respecto al precio real para clasificar la calidad de las ofertas. Finalmente, se desarrolla un panel interactivo que permite explorar dinámicamente las ofertas, filtrando por características específicas y visualizando en un mapa los inmuebles disponibles en una zona determinada. Se incluyen adicionalmente indicadores y gráficos para el análisis de tendencias a lo largo del año estudiado. Este cuadro de mando puede ser utilizado tanto por usuarios particulares interesados en comprar una vivienda como por profesionales del sector que deseen analizar orientaciones del mercado. Los resultados obtenidos muestran una notable capacidad predictiva de los modelos seleccionados, permitiendo clasificar con precisión distintas ofertas inmobiliarias en un amplio rango de precios, especialmente para precios inferiores al millón de euros. Además, la visualización interactiva facilita un análisis completo y directo, que sirve de apoyo a la toma de decisiones de los agentes del sector. A modo de conclusión, el estudio realizado refleja patrones en la concentración geográfica de las ofertas y una variabilidad estacional a lo largo del año. Se observa una escasa presencia de ofertas calificadas como malas en las zonas del sur de la localidad, así como una mayor afluencia de viviendas a la venta durante el último trimestre del año, coincidiendo con una reducción proporcional de las ofertas identificadas como de buena calidad. Abstract: In a real estate market characterized by a sustained increase in prices and the urgency in decision-making that it brings, the need to adopt a data-driven approach that enables various stakeholders to directly identify relevant buying and selling opportunities arises. This project presents a technological solution based on machine learning for the rating of housing offers located in the municipality of Madrid, followed by their visualization through an interactive dashboard. To achieve this, the project starts with a dataset published by Idealista on homes for sale during the year 2018, to which processes of data cleaning, outlier treatment, and variable normalization are applied. Regression models are then trained and analyzed to predict property prices based on their characteristics, and the deviation from the actual price is used to classify the quality of each offer. Lastly, an interactive panel is developed to dynamically explore the offers by filtering specific features and visualizing on a map the available properties of a selected area. Indicators and graphs are also included for trend analysis over the year under study. This dashboard can be used by both private individuals interested in purchasing a house and real estate professionals who may want to analyze market trends. The obtained results show a remarkable predictive capacity of the selected models, allowing to accurately classify various real estate offers across a wide price range, especially for houses priced below one million euros. In addition, the interactive visualization facilitates a complete and direct analysis, which supports decision-making for market participants. In conclusion, the study reveals patterns in the geographic concentration of offers and seasonal variability throughout the year. A low presence of poorly rated offers is observed in the southern areas of the municipality, along with a higher number of properties for sale during the last quarter of the year, together with a proportional decrease in offers identified as high quality.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".