The accumulation of data, including relational ones, has created new opportunities and challenges for the social and computational sciences. Data abundance does not automatically translate into a deeper understanding of phenomena, but requires a set of tools to study their dynamics. In the science-of-science domain in particular, the structures connecting actors, institutions, and scholarly products, as well as the relations among the actors that constitute them, are never purely dyadic or static. Scientific collaborations, grant funding, citation dynamics, and surgical teamwork all involve group interactions of varying size that evolve over time. Reducing such dynamics to simple collections of dyadic ties loses the higher-order information that gives them their deeper meaning. This thesis addresses this challenge through a series of interconnected contributions, moving progressively from foundational questions about data and interpretability to relational systems of increasing complexity. The opening contribution concerns the epistemological and methodological foundations of data-driven research. Taking data and their treatment under the European General Data Protection Regulation (GDPR) as a paradigmatic case, the thesis reflects on the intrinsic nature of data. It then turns to the interpretability problem empirically, applying tree-based ensemble classifiers combined with Recursive Feature Elimination and SHapley Additive exPlanations (SHAP) to identify the determinants of voting intention in Italy. These initial contributions highlight the need for progressively more complex statistical tools, capable of moving beyond individual-level analysis to interactions involving multiple entities. The second part of the thesis develops the theoretical framework of Relational Hyperevent Models through a series of progressive extensions. The work first characterizes the structure and dynamics of the Italian academic statisticians community, modeling the mechanisms driving collaboration within it, such as group persistence, triadic closure, and repetition, while addressing the boundary problem posed by co-authors external to the reference population. It then introduces the Geometrically Weighted Subset Repetition statistic, a new hyperedge measure that bounds the combinatorial growth of subset-repetition counts over long observation windows; this statistic is applied to model the co-evolution of co-authorship and research funding networks across three disciplinary communities. The framework is further extended to tripartite hyperevent structures that connect authors, cited papers, and keywords as a single system of co-evolving relational processes: this is the first application within science of science to jointly model three distinct entity types within a unified statistical framework. Finally, the Relational Hyperevent Outcome Model is applied for the first time outside an academic setting, including non-linear effects. These contributions advance the methodological frontier of dynamic network analysis, moving beyond the dyadic assumption that constrains most existing models, applying Relational Hyperevent Models across different contexts and progressively expanding the number of entities involved, while proposing concrete solutions to open problems in the specification, estimation, and interpretation of statistical models for polyadic interaction data.

L’accumulo di dati, inclusi quelli relazionali, ha generato nuove opportunità e sfide per le scienze sociali e computazionali. Questa disponibilità di dati non si traduce automaticamente in una comprensione profonda dei fenomeni, ma richiede un insieme di strumenti per studiarne le dinamiche. Inoltre, nel dominio della science of science, le strutture che connettono attori, istituzioni e prodotti accademici, nonché le relazioni tra i diversi attori che le costituiscono, non sono mai di natura diadica o statica. Le collaborazioni scientifiche, l’assegnazione di finanziamenti, le dinamiche di citazione e il lavoro d'équipe in ambito chirurgico comportano interazioni di gruppo di dimensioni variabili che evolvono nel tempo. Ridurre tali dinamiche a semplici collezioni di legami diadici comporta la perdita delle informazioni di ordine superiore che ne definiscono il significato più profondo. La presente tesi affronta questa sfida attraverso una serie di contributi interconnessi, muovendosi progressivamente dalle questioni fondative relative ai dati e all’interpretabilità fino a sistemi relazionali di crescente complessità. Il contributo iniziale riguarda i presupposti epistemologici e metodologici della ricerca data-driven. Prendendo i dati e il loro trattamento ai sensi del Regolamento Generale sulla Protezione dei Dati (GDPR) dell'Unione Europea come caso paradigmatico, la tesi riflette sulla natura intrinseca del dato. Successivamente, il lavoro indaga il problema dell'interpretabilità sul piano empirico, applicando classificatori tree-based ensemble combinati con tecniche di Recursive Feature Elimination e SHapley Additive exPlanations (SHAP) per identificare i determinanti dell'intenzione di voto in Italia. Questi contributi iniziali evidenziano la necessità di strumenti statistici progressivamente più complessi, capaci di andare oltre l'analisi individuale, passando a interazioni che coinvolgono più entità. La seconda parte della tesi sviluppa il quadro teorico dei Relational Hyperevent Models attraverso una serie di estensioni progressive. Il lavoro caratterizza in primo luogo la struttura e le dinamiche della comunità accademica degli statistici italiani, modellando i meccanismi che guidano la collaborazione, come la persistenza di gruppo, la chiusura triadica e la ripetizione, e affrontando al contempo il boundary problem (problema dei confini di una rete) posto dai coautori esterni alla popolazione di riferimento. Viene quindi introdotta la statistica Geometrically Weighted Subset Repetition, una nuova misura per gli iperedge (hyperedge) capace di limitare la crescita combinatoria dei conteggi di ripetizione di sottoinsiemi su ampie finestre temporali; tale metrica è stata applicata per modellare la co-evoluzione delle reti di co-autorialità e di finanziamento della ricerca in tre diverse comunità disciplinari. Il quadro di riferimento è stato ulteriormente esteso a strutture di ipereventi tripartiti, che connettono autori, articoli citati e keyword in un unico sistema di processi relazionali che coevolvono: si tratta della prima applicazione, nell'ambito della science of science, che modella congiuntamente tre tipologie di entità distinte all'interno di un unico impianto statistico. Infine, il Relational Hyperevent Outcome Model viene applicato per la prima volta al di fuori del contesto accademico, includendo effetti non lineari. Tali contributi avanzano la frontiera metodologica dell'analisi delle reti dinamiche, superando l'assunzione diadica che vincola gran parte dei modelli esistenti, applicando i Relational Hyperevent Model

Metodi statistici per la dinamica di ordine superiore delle reti. Estensioni e applicazioni dei Relational Hyperevent Models nella collaborazione in team e nei big data accademici / Fabbrucci Barbagli, A.G.. - (2026 Sep 16).

Metodi statistici per la dinamica di ordine superiore delle reti. Estensioni e applicazioni dei Relational Hyperevent Models nella collaborazione in team e nei big data accademici

FABBRUCCI BARBAGLI, AMIN GINO
2026-09-16

Abstract

The accumulation of data, including relational ones, has created new opportunities and challenges for the social and computational sciences. Data abundance does not automatically translate into a deeper understanding of phenomena, but requires a set of tools to study their dynamics. In the science-of-science domain in particular, the structures connecting actors, institutions, and scholarly products, as well as the relations among the actors that constitute them, are never purely dyadic or static. Scientific collaborations, grant funding, citation dynamics, and surgical teamwork all involve group interactions of varying size that evolve over time. Reducing such dynamics to simple collections of dyadic ties loses the higher-order information that gives them their deeper meaning. This thesis addresses this challenge through a series of interconnected contributions, moving progressively from foundational questions about data and interpretability to relational systems of increasing complexity. The opening contribution concerns the epistemological and methodological foundations of data-driven research. Taking data and their treatment under the European General Data Protection Regulation (GDPR) as a paradigmatic case, the thesis reflects on the intrinsic nature of data. It then turns to the interpretability problem empirically, applying tree-based ensemble classifiers combined with Recursive Feature Elimination and SHapley Additive exPlanations (SHAP) to identify the determinants of voting intention in Italy. These initial contributions highlight the need for progressively more complex statistical tools, capable of moving beyond individual-level analysis to interactions involving multiple entities. The second part of the thesis develops the theoretical framework of Relational Hyperevent Models through a series of progressive extensions. The work first characterizes the structure and dynamics of the Italian academic statisticians community, modeling the mechanisms driving collaboration within it, such as group persistence, triadic closure, and repetition, while addressing the boundary problem posed by co-authors external to the reference population. It then introduces the Geometrically Weighted Subset Repetition statistic, a new hyperedge measure that bounds the combinatorial growth of subset-repetition counts over long observation windows; this statistic is applied to model the co-evolution of co-authorship and research funding networks across three disciplinary communities. The framework is further extended to tripartite hyperevent structures that connect authors, cited papers, and keywords as a single system of co-evolving relational processes: this is the first application within science of science to jointly model three distinct entity types within a unified statistical framework. Finally, the Relational Hyperevent Outcome Model is applied for the first time outside an academic setting, including non-linear effects. These contributions advance the methodological frontier of dynamic network analysis, moving beyond the dyadic assumption that constrains most existing models, applying Relational Hyperevent Models across different contexts and progressively expanding the number of entities involved, while proposing concrete solutions to open problems in the specification, estimation, and interpretation of statistical models for polyadic interaction data.
16-set-2026
DE STEFANO, DOMENICO
38
2024/2025
Settore STAT-03/B - Statistica sociale
Università degli Studi di Trieste
File in questo prodotto:
File Dimensione Formato  
Statistical Methods for Higher-Order Network Dynamics.pdf

embargo fino al 16/09/2027

Descrizione: Statistical Methods for Higher-Order Network Dynamics
Tipologia: Tesi di dottorato
Dimensione 10.11 MB
Formato Adobe PDF
10.11 MB Adobe PDF   Visualizza/Apri   Richiedi una copia
Statistical Methods for Higher-Order Network Dynamics_1.pdf

embargo fino al 16/09/2027

Descrizione: Statistical Methods for Higher-Order Network Dynamics
Tipologia: Tesi di dottorato
Dimensione 10.11 MB
Formato Adobe PDF
10.11 MB Adobe PDF   Visualizza/Apri   Richiedi una copia
Statistical Methods for Higher-Order Network Dynamics_2.pdf

embargo fino al 16/09/2027

Descrizione: Statistical Methods for Higher-Order Network Dynamics
Tipologia: Tesi di dottorato
Dimensione 10.11 MB
Formato Adobe PDF
10.11 MB Adobe PDF   Visualizza/Apri   Richiedi una copia
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11368/3147099
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact