Recording provenance of workflow runs with RO-Crate
Authors:
Simone Leo,
Michael R. Crusoe,
Laura Rodríguez-Navas,
Raül Sirvent,
Alexander Kanitz,
Paul De Geest,
Rudolf Wittner,
Luca Pireddu,
Daniel Garijo,
José M. Fernández,
Iacopo Colonnelli,
Matej Gallo,
Tazro Ohta,
Hirotaka Suetake,
Salvador Capella-Gutierrez,
Renske de Wit,
Bruno P. Kinoshita,
Stian Soiland-Reyes
Abstract:
Recording the provenance of scientific computation results is key to the support of traceability, reproducibility and quality assessment of data products. Several data models have been explored to address this need, providing representations of workflow plans and their executions as well as means of packaging the resulting information for archiving and sharing. However, existing approaches tend to…
▽ More
Recording the provenance of scientific computation results is key to the support of traceability, reproducibility and quality assessment of data products. Several data models have been explored to address this need, providing representations of workflow plans and their executions as well as means of packaging the resulting information for archiving and sharing. However, existing approaches tend to lack interoperable adoption across workflow management systems. In this work we present Workflow Run RO-Crate, an extension of RO-Crate (Research Object Crate) and Schema.org to capture the provenance of the execution of computational workflows at different levels of granularity and bundle together all their associated objects (inputs, outputs, code, etc.). The model is supported by a diverse, open community that runs regular meetings, discussing development, maintenance and adoption aspects. Workflow Run RO-Crate is already implemented by several workflow management systems, allowing interoperable comparisons between workflow runs from heterogeneous systems. We describe the model, its alignment to standards such as W3C PROV, and its implementation in six workflow systems. Finally, we illustrate the application of Workflow Run RO-Crate in two use cases of machine learning in the digital image analysis domain.
A corresponding RO-Crate for this article is at https://w3id.org/ro/doi/10.5281/zenodo.10368989
△ Less
Submitted 16 July, 2024; v1 submitted 12 December, 2023;
originally announced December 2023.
Decomposition of quantitative Gaifman graphs as a data analysis tool
Authors:
José Luis Balcázar,
Marie Ely Piceno,
Laura Rodríguez-Navas
Abstract:
We argue the usefulness of Gaifman graphs of first-order relational structures as an exploratory data analysis tool. We illustrate our approach with cases where the modular decompositions of these graphs reveal interesting facts about the data. Then, we introduce generalized notions of Gaifman graphs, enhanced with quantitative information, to which we can apply more general, existing decompositio…
▽ More
We argue the usefulness of Gaifman graphs of first-order relational structures as an exploratory data analysis tool. We illustrate our approach with cases where the modular decompositions of these graphs reveal interesting facts about the data. Then, we introduce generalized notions of Gaifman graphs, enhanced with quantitative information, to which we can apply more general, existing decomposition notions via 2-structures; thus enlarging the analytical capabilities of the scheme. The very essence of Gaifman graphs makes this approach immediately appropriate for the multirelational data framework.
△ Less
Submitted 11 August, 2018; v1 submitted 14 May, 2018;
originally announced May 2018.