Technical blog post by Sharif Islam
Demystifying Data Spaces for biodiversity projects
This article was originally written and published by Sharif Islam from Naturalis Biodiversity Center.
"Data space" is one of those terms we're now seeing more and more in project proposals and at conferences. Everyone nods along as if they get the idea, but it's often hard to pin down.
So to put it in context, here's how Data Space Support Center defines it:
"A distributed system defined by a governance framework that enables secure and trustworthy data transactions between participants while supporting trust and data sovereignty. A data space is implemented by one or more infrastructures and enables one or more use cases."
International Data Spaces Association (IDSA) gives a similar definition, with a bit more detail:
"A data space is a governance framework and a set of supporting services that enable organisations to share data in a trusted, sovereign, and interoperable way. It provides an agreed set of policies, semantic models, protocols, and processes that allow participants to retain control over their data while making it discoverable and usable by others under clearly defined conditions."
Both are accurate. But still, hard to visualise when the language is governance and infrastructure. And none of it tells you who is actually doing what to which dataset. That's usually where the confusion sets in: people picture a platform, when what's actually being described is a set of relationships and rules between people and organisations.
The official portal for European data has a good picnic analogy: a data portal is where one host brings the table and cutlery for everyone to share, and a data space is where everyone brings their own table and shares directly with each other, under their own terms. Slower to coordinate, but more resilient.

Data space as Picnic. Source: https://data.europa.eu/en/publications/datastories/when-open-data-meets-data-spaces
In the Biodiversity Meets Data (BMD) project, we're looking into various "data space" concepts and implementations. And we've found it more useful to start with roles than technology: who provides data, who consumes it, who owns the right to decide how it's shared, and who keeps the whole arrangement usable.
We're also learning a lot from projects that came before us, like the Green Deal Data Space (greendealdata.eu) and AD4GD: All Data for Green Deal (github.com/AD4GD), who've already worked through a lot of this in practice.
It gets interesting once you attach the concepts to a real use case. Take a species record tagged with a particular habitat type under EU directive and a conservation status. See example below. The status label and the species name might sit in different repositories as open, citable facts. But more detail is often linked in other databases, or not linked or restricted. For example, location data might be blurred, and the full shapefile or map behind it can sit behind a login with its own redistribution terms. We're all aware of this, and there are manual workarounds that sort of work.

Source: https://biodiversity.europa.eu/species/8709 https://eunis.eea.europa.eu/species/8709#threat_status. Image generated by ChatGPT.
These kinds of links and graphs are common in biodiversity projects, but as data scale and complexity keep increasing, we can't keep relying on manual effort.
Sources used for the example: https://biodiversity.europa.eu/species/8709 and https://eunis.eea.europa.eu/species/8709#threat_status. More info at IUCN (https://www.iucnredlist.org/species/22696495/204182473).
What often goes unsaid is that none of this is glamorous "AI" work.
It's tracking revisions, catching a new or changed metadata field before it breaks something three systems downstream, being able to audit who changed what and reproduce the same result twice. But that's the actual difference between a data space and a folder full of CSV files.
Strip away the governance language, and trust and sovereignty in a data space mostly come down to three questions: do you have the data I need, can I combine it with something else, and can I use it with my tool? That's it. If we can show people the concept at that level, it gets a lot easier to explain why the rest of it matters.
Find out more in BMD GitHub on our ongoing work. And watch this space for more updates as we do more.


