A big thank you to @yoanapopova and @kate_sam for another excellent CKAN Monthly Live this month, and to @steven_decosta for joining us to introduce the Objective Observer Initiative and the Provenance Sandbox.
Steven’s presentation explored the use of CKAN to manage machine learning datasets alongside information about their provenance, intent, and use. After offering a quick walk-through of some of the ways he’s extended CKAN to support this kind of work, the provided a lengthier and really interesting case for for the importance of provenance as open data is increasingly being reused in machine learning and AI applications.
In the days since, I’ve found myself thinking a lot about his argument. Given that my work sits adjacent to many of you whose work engages you more directly with these kinds of questions and challenges, I thought I might ask you: What are your thoughts about provenance requirements and the ways in which they may change or be forced to change as open data continue to be reused for ML & AI applications?
Has your organization changed the provenance information it captures or preserves in response to these uses? Are there provenance requirements that have become more important or more difficult to meet?
I’d be super interested in hearing how others are approaching and thinking about this! (I will also work with Yoana to make sure that one of us shares the official recap once its available!)