Skip navigation
ChatGPT Image 6 лип. 2026 р., 21_54_01

Why EOSC begins not with a repository, but with metadata

Baumann K., Cecconi B., Dietrich M., Doerr M., Jouneau T., Tewatia P., Kokkinaki A., Med M., Marek J. et al. Landscape overview for the EOSC Federation Semantic Interoperability in 2025. EOSC Association Task Force: Technical and Semantic Interoperability, 2026. https://zenodo.org/records/19134873

The European Open Science Cloud may appear to be a large file storage system or another catalogue of links to repositories. At present, however, its real purpose is to create an environment in which research data, software, publications, training materials, services and research infrastructures can be found, understood, linked to one another and reused. To achieve this, it is not enough simply to publish a file in open access. The file must be described in a way that is understandable not only to humans, but also to machines, catalogues, search services, repositories and future artificial intelligence tools.

This is why, in addition to the availability of a repository, the core of practical integration with EOSC is the quality of metadata and the consistency of vocabularies. Metadata answer basic questions: what kind of resource this is, who created it, when it was created, which research it belongs to, how it should be cited, under which licence it can be used, and with which publications, projects or datasets it is associated. Vocabularies, thesauri and ontologies answer another question: what exactly do the terms used in the description mean?

For example, the field “method” in a dataset may contain different values: DFT, molecular dynamics, finite element modelling, microscopy, spectroscopy, survey, simulation. If these values are entered freely, in different languages or with different spellings, the system may treat them as unrelated words. If, however, they are linked to a controlled vocabulary, grouping, search, filtering, comparison with other datasets and the development of interdisciplinary services become possible.

In this sense, the FAIR principles are not reduced to the slogan “make data open”. FAIR means that data should be findable, accessible, interoperable and reusable. The components “findable”, “interoperable” and “reusable” depend most strongly on metadata, identifiers, licences, description standards, vocabularies and semantic links. Without these elements, an open file remains an isolated object. With them, it becomes part of a broader space of research data and services.

Metadata as an infrastructure layer

In ordinary research practice, metadata are often perceived as an additional description: title, authors, a short abstract and keywords. For EOSC, such a minimum is not sufficient. Metadata must perform an infrastructural function. They should support search, citation, automatic exchange between systems, and links with publications, projects, institutions, funding, data versions, software and workflows.

Therefore, metadata in EOSC should be considered at several levels.

The first level is the general description of a research object. It includes the title, authors, organisations, description, date, resource type, identifier, licence, access conditions, language, keywords and related resources. This level is necessary for almost any object: dataset, software, publication, training material, service or workflow.

The second level is the description in a catalogue. It is needed so that resources can be represented in catalogues, aggregators, open science portals and search services. At this level, both the properties of an individual dataset and the description of a repository, service, organisation, collection, thematic space or research infrastructure become important.

The third level is the domain-specific description. It depends on the particular field of science. In materials science, this may include chemical composition, crystal structure, calculation method, type of experiment, software package, modelling parameters, sample type, measured or calculated property. In bioinformatics, the fields will be different; in ecology, social sciences or astronomy, they will differ again. This is why EOSC cannot operate through a single universal standard for all disciplines. Compatibility between general standards and disciplinary profiles is required.

The fourth level is the semantic level. It defines how the values of fields are linked to controlled vocabularies, thesauri, ontologies and external identifiers. Without this, the system sees only text. With it, the system begins to see concepts, relationships, classes of resources and types of research objects.

Why vocabularies are not an add-on, but a necessary part of FAIR

Vocabularies are often perceived as a secondary detail. In fact, they are one of the main instruments of semantic interoperability. If different repositories, institutes or research groups use different names for the same concepts, data integration becomes difficult or impossible. If they use controlled terms linked to stable identifiers, the data can be compared, aggregated and reused in broader contexts.

In the context of EOSC, a dictionary service should be understood as a managed service for maintaining terms. It may include controlled value lists, multilingual labels, definitions of concepts, links between broader and narrower concepts, equivalents in other vocabularies, versions, URIs and machine-readable formats. Such a service is needed both for user convenience and for automated metadata completion, quality control of descriptions and the construction of links between systems.

For example, if the Competence Center supports a metadata template for DFT calculations, it may contain fields such as “software”, “method”, “material system”, “calculated property” and “workflow stage”. For each of these fields, it is preferable to have not free text, but a controlled list or a link to a vocabulary. This helps avoid situations where one group writes “Quantum Espresso”, another writes “QE”, a third writes “quantum-espresso”, and a fourth writes “PWscf”, although in some cases these refer to related software components of the same environment.

For a local Competence Center, this means that work with FAIR data is not limited to consultations on the repository. It is necessary to create and maintain a practical set of metadata profiles, README templates, JSON examples, mapping tables, controlled vocabularies and mappings between local terms and international standards.

From local description to EOSC compatibility

A practical mistake of many open data initiatives is that they begin with local convenience: how to describe a dataset quickly, how to fill in a form, how to add a file to a repository. This is a necessary first step, but it is not sufficient. If the local description cannot be mapped to DataCite, Dublin Core, OpenAIRE, DCAT or disciplinary standards, the resource will remain weakly visible beyond the local portal.

EOSC compatibility means that each local metadata profile must be linked to a broader context. Title, authors, organisations, identifier, date, licence, resource type and related publications must be represented in a way that external services can understand. Disciplinary fields must be described through clear schemas. Local terms must be documented and, where possible, linked to international vocabularies or ontologies.

This is where the role of a metadata schema and crosswalk registry emerges. A crosswalk is a table or formal description of correspondences between different metadata schemas. For example, a local field such as “dataset lead” may correspond to the contributor role in DataCite; a field such as “related publication” may correspond to relatedIdentifier; a field such as “access type” may correspond to access rights; a field such as “workflow stage” may correspond to a domain vocabulary or a separate application profile. Without such correspondences, integration remains manual. With them, automation becomes possible.

The role of the Competence Center

For the NASU Competence Center for Data Management, this topic has practical significance. The Center can become both a place for consultations on DataverseUA or the FAIR principles and a focal point for developing and maintaining minimum metadata profiles for Ukrainian research groups. The aim is not to create “our own standard instead of international ones”, but to build a clear bridge between the local practices of institutes and EOSC standards.

Such a bridge should include several elements. First, basic metadata profiles for dataset, software, workflow, training material and service. Second, README and DMP templates linked to these profiles. Third, controlled vocabularies for resource types, disciplines, methods, software, licences, access modes and workflow stages. Fourth, examples in formats suitable for machine processing: JSON, CSV, Markdown and, over time, SKOS/RDF. Fifth, mappings and crosswalks to DataCite, OpenAIRE, DCAT and other standards.

In this model, the website of the Competence Center ceases to be merely an informational resource. It becomes a working access point to guidelines, templates, vocabularies, training materials and examples that help institutes prepare data for publication, harvesting, indexing and further integration with EOSC.

The transition from a declaration of open science to EOSC begins with the concrete quality of the description of research objects. If data have stable identifiers, clear metadata, licences, links to publications, workflow descriptions and controlled terms, they can become part of the European space of FAIR data. If these elements are absent, even an open file remains almost invisible.

This is why core standards for metadata and dictionary services should be regarded as one of the fundamental competencies of a modern Research Data Management Center.