DATA SCIENCE FOR DEMOGRAPHIC PROCESSES
Anno accademico 2026/2027 - Docente: ANGELO MAZZARisultati di apprendimento attesi
- Knowledge and understanding
Students will acquire knowledge of the main demographic processes and data sources, with particular reference to population structure, mortality, fertility and migration. They will understand the main stages of a demographic data-science workflow, including data acquisition, cleaning, transformation, visualisation and analysis. Students will also acquire basic knowledge of Large Language Models (LLMs), APIs and AI-assisted coding. - Applying knowledge and understanding
Students will be able to acquire, process, analyse and visualise demographic data using R. They will use official and open data sources, APIs, spatial data and unstructured textual information. Students will be able to use LLMs to support coding, debugging, data transformation and information extraction, while independently validating the resulting code and outputs. - Making judgements
Students will be able to assess the quality, representativeness and limitations of demographic data from different sources. They will critically evaluate AI-generated code and results and identify possible errors, biases and reproducibility issues. - Communication skills
Students will be able to communicate demographic evidence through tables, visualisations, maps and reproducible reports. They will be able to document methodological choices and transparently report the use of AI tools in data analysis. - Learning skills
Students will develop the ability to independently use new demographic datasets, R packages, APIs and AI-based tools, adapting data-science workflows to new research and policy questions.
Modalità di svolgimento dell'insegnamento
The course combines lectures, live coding, guided exercises and laboratory activities.
R will be used throughout the course. Students will progressively develop data-analysis workflows using demographic datasets, APIs, spatial and textual data.
LLMs will be used from the early stages of the course as coding assistants. Students will learn an iterative workflow based on prompting, code generation, execution, debugging, validation and documentation.
Particular attention will be paid to the ability to understand and validate AI-generated code rather than to memorising programming syntax.
All teaching activities and course materials will be in English.
Prerequisiti richiesti
No formal prerequisites are required.
Basic knowledge of statistics and familiarity with quantitative reasoning are useful. Previous experience with R or demographic analysis is not required.
The course is designed for students with heterogeneous academic and technical backgrounds. The essential R skills required for the course will be introduced at the beginning and progressively developed through practical applications and LLM-assisted coding.
Frequenza lezioni
Contenuti del corso
The course introduces data-science methods for the analysis of demographic processes, including population structure, mortality, fertility and migration.
Students will learn to acquire, transform, visualise and analyse demographic data using R and reproducible workflows. Official demographic databases, open data, APIs, spatial data and non-traditional data sources will be considered.
Particular attention will be devoted to LLM-assisted coding (vibe coding) as a data-science workflow. Large Language Models will be used to support R programming, debugging, data cleaning, information extraction and text analysis, with emphasis on validation and reproducibility.
The course also introduces the basic principles of LLMs and their use through APIs, structured outputs and retrieval-based methods. Applications include text-as-data, spatial demographic analysis, survey data and synthetic respondents.
Testi di riferimento
· · IUSSP, Population Analysis for
Policies & Programmes (PAPP).
https://papp.iussp.org/
· Wickham H., Çetinkaya-Rundel M.,
Grolemund G., R for Data Science, 2nd edition.
https://r4ds.hadley.nz/
· United Nations, World Population
Prospects 2024 and related methodological documentation.
https://population.un.org/wpp/
· Lovelace R., Nowosad J., Muenchow
J., Geocomputation with R, 2nd edition.
https://r.geocompx.org/
· Kwon B., Park T., Perez-Cruz F.,
Rungcharoenkitkul P., Large Language Models: A Primer for Economists, BIS
Quarterly Review, 2024.
https://www.bis.org/publications/qr-202412/large-language-models-primer-economists
· Selected papers and technical materials provided by the instructor.
Programmazione del corso
| Argomenti | Riferimenti testi | |
|---|---|---|
| 1 | Demographic processes and data science. Demographic questions, data sources and the data-science workflow. | |
| 2 | R survival kit. Objects, vectors, data frames, functions, packages and basic data manipulation. | |
| 3 | LLM-assisted coding (vibe coding) in R. Prompting for code, debugging, code explanation, validation and iterative development. | |
| 4 | Demographic rates and population structure. Stocks, flows, exposure, rates, ratios, age-sex structure and population pyramids. | |
| 5 | Mortality data. Mortality rates, survival, life tables and life expectancy. | |
| 6 | Fertility, migration and population change. Fertility indicators, migration measures and demographic balance. | |
| 7 | Data wrangling in R. Tidy data, filtering, grouping, joins and reshaping, with LLM assistance. | |
| 8 | Data visualisation and demographic storytelling. ggplot2, demographic graphics and effective communication of population data. | |
| 9 | Official demographic data sources. World Population Prospects and other international demographic databases. | |
| 10 | APIs, JSON and programmatic data access. Retrieving remote data from R and building reproducible data-acquisition workflows. | |
| 11 | Reproducible demographic analysis. Scripts, Quarto, documentation and reproducible reports. | |
| 12 | Spatial demographic data I. Spatial objects, coordinates, boundaries and spatial joins in R. | |
| 13 | Spatial demographic data II. Mapping, geocoding and analysis of spatial demographic inequalities. | |
| 14 | Non-traditional population data. Administrative data, open data and digital traces; coverage, selection and representativeness. | |
| 15 | Inside a Large Language Model. Tokens, embeddings, Transformers, attention, training and open versus proprietary models. | |
| 16 | LLM APIs from R. Programmatic access, prompts, parameters, batch processing and structured outputs. | |
| 17 | Text as data with LLMs. Classification, information extraction and transformation of unstructured text into structured data. | |
| 18 | LLM-assisted data cleaning and harmonisation. Variable matching, metadata interpretation and integration of heterogeneous data sources. | |
| 19 | Survey data and synthetic respondents. Automated coding of open-ended responses, synthetic respondents and empirical validation. | |
| 20 | Validation, bias and reproducibility in AI-assisted research. Hallucinations, benchmarking, representativeness and responsible AI use. | |
| 21 | Integrated demographic data-science project. Data acquisition, R analysis, visualisation and AI-assisted workflow applied to a demographic or policy problem. |
Verifica dell'apprendimento
Modalità di verifica dell'apprendimento
Assessment will be based on an applied data-science project and an oral discussion.
The project will require students to address a demographic or population-related question using real data and R. The analysis may include data acquisition through APIs, data wrangling, visualisation, spatial or textual data analysis and the use of LLM-based tools.
The use of AI tools is allowed and encouraged, but students must document their use and independently validate generated code, data transformations and results.
The oral discussion will assess the student's understanding of the demographic problem, the data-science workflow, the methodological choices adopted and the limitations of the analysis.
Assessment will consider methodological correctness, reproducibility, interpretation of results, appropriate use and validation of AI tools, and clarity of presentation.
Esempi di domande e/o esercizi frequenti
Demographic data
- What is the difference between a demographic rate, a proportion and a ratio?
- How can age structure be represented and compared across populations?
- What data are required to analyse mortality, fertility and migration?
R and demographic data science
- What is tidy data and why is it useful?
- How can demographic data be grouped, transformed and visualised in R?
- What makes a data-analysis workflow reproducible?
Data sources and APIs
- What is an API and how can it be used from R?
- What is the difference between CSV, JSON and data obtained through an API?
- What problems may arise when combining demographic data from different sources?
Large Language Models and vibe coding
- What is a Large Language Model and how does it generate text?
- How can an LLM assist the development and debugging of R code?
- Why should AI-generated code be independently validated?
- What is a structured output and why is it useful in data analysis?
Advanced applications
- How can an LLM transform unstructured text into structured data?
- What are synthetic respondents?
- What are the main problems of bias and representativeness in AI-generated data?