Modern data-driven organizations rely on pipelines that transform data from infrastructure, applications, and users, where correctness and performance are critical for reliability and business value. These pipelines are often developed and optimized by teams separate from the data owners defining semantics, so meaningful testing and benchmarking require sharing representative data, even though real data is often sensitive and restricted by legal and commercial constraints. To enable privacy-preserving data sharing while preserving utility, our framework, named SHIELD, automatically constructs stream processing (SP) pipelines to transform live sensitive data into shareable data while preserving the characteristics required for downstream processing and optimization. Internally, SHIELD leverages evolutionary computation to synthesize executable SP queries under predefined privacy and utility requirements. Using real-world use cases, we show SHIELD can synthesize privacy-preserving pipelines that retain analytical value and scale to realistic workloads.
SHIELD: Evolutionary Synthesis of Privacy-Preserving Pipelines for Live Stream Data Sharing / Perelli, S., Medvet, E., Gulisano, V.. - (2026), pp. 82-94. (ACM International Conference on Distributed and Event-based Systems Lisbona, Portogallo June 2026) [10.1145/3809481.3812614].
SHIELD: Evolutionary Synthesis of Privacy-Preserving Pipelines for Live Stream Data Sharing
Eric Medvet;Vincenzo Gulisano
2026-01-01
Abstract
Modern data-driven organizations rely on pipelines that transform data from infrastructure, applications, and users, where correctness and performance are critical for reliability and business value. These pipelines are often developed and optimized by teams separate from the data owners defining semantics, so meaningful testing and benchmarking require sharing representative data, even though real data is often sensitive and restricted by legal and commercial constraints. To enable privacy-preserving data sharing while preserving utility, our framework, named SHIELD, automatically constructs stream processing (SP) pipelines to transform live sensitive data into shareable data while preserving the characteristics required for downstream processing and optimization. Internally, SHIELD leverages evolutionary computation to synthesize executable SP queries under predefined privacy and utility requirements. Using real-world use cases, we show SHIELD can synthesize privacy-preserving pipelines that retain analytical value and scale to realistic workloads.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


