Provenance Requirements for AI & ML

A big thank you to @yoanapopova and @kate_sam for another excellent CKAN Monthly Live this month, and to @steven_decosta for joining us to introduce the Objective Observer Initiative and the Provenance Sandbox.

Steven’s presentation explored the use of CKAN to manage machine learning datasets alongside information about their provenance, intent, and use. After offering a quick walk-through of some of the ways he’s extended CKAN to support this kind of work, the provided a lengthier and really interesting case for for the importance of provenance as open data is increasingly being reused in machine learning and AI applications.

In the days since, I’ve found myself thinking a lot about his argument. Given that my work sits adjacent to many of you whose work engages you more directly with these kinds of questions and challenges, I thought I might ask you: What are your thoughts about provenance requirements and the ways in which they may change or be forced to change as open data continue to be reused for ML & AI applications?

Has your organization changed the provenance information it captures or preserves in response to these uses? Are there provenance requirements that have become more important or more difficult to meet?

I’d be super interested in hearing how others are approaching and thinking about this! (I will also work with Yoana to make sure that one of us shares the official recap once its available!)

3 Likes

Sorry to have missed this CML episode! (on a sidenote, we should index all the CML episodes in the Ecosystem Catalog)

This was a big topic of discussion during UN Open Source Week and there are several experiments underway.

On our end, we’re currently integrating with TypedStandards.org - as we find that we generative AI claims are best done by a third-party, using an open source standard - as you simply cannot self-certify.

There are also some initiatives like https://ethicstoolkit.ai/ that were created by practitioners - and we were fortunate enough to go through a workshop with them to validate our Gen AI experiments.

IMHO, “Open Data is Great Again” :wink: because of AI - as it turns out that FAIR Data is AI-Ready Data - as its the Perfect Context for AI as borne out by our Data Schematics experiment - where for instance, we can comprehensively describe 16 years of NYC 311 data - all ~27 gigabytes, 42 million rows of it - with just 59 kilobytes of statistical FAIR metadata.

Though the example below is just for 10,000 rows, the statistical FAIR metadata should be about the same size if we compile it for all 42 million rows.

https://dathere.com/2026/08/data-schematic-a-neuro-symbolic-visual-story-telling-data-explorer-dictionary/