Toward Total Recall: Enhancing FAIRness through AI-Driven Metadata Standardization

13 February 2025

Sowmya S. Sundaram

Rafael S Gonçalves

ArXiv (abs)PDF HTML

Main:17 Pages

5 Figures

3 Tables

Abstract

Current metadata often suffer from incompleteness, inconsistency, and incorrect formatting, hindering effective data reuse and discovery. Using GPT-4 and a metadata knowledge base (CEDAR), we devised a method that standardizes metadata in scientific data sets, ensuring the adherence to community standards. The standardization process involves correcting and refining metadata entries to conform to established guidelines, significantly improving search performance and recall metrics. The investigation uses BioSample and GEO repositories to demonstrate the impact of these enhancements, showcasing how standardized metadata lead to better retrieval outcomes. The average recall improves significantly, rising from 17.65\% with the baseline raw datasets of BioSample and GEO to 62.87\% with our proposed metadata standardization pipeline. This finding highlights the transformative impact of integrating advanced AI models with structured metadata curation tools in achieving more effective and reliable data retrieval.

View on arXiv

Comments on this paper