Repository and catalogue metadata profiles
Metadata requirements and mappings used to expose repository records to catalogues, aggregators, discovery services and research infrastructures.
This resource explains how locally managed metadata can be transformed into interoperable records for DOI services, OAI-PMH harvesting, OpenAIRE, data catalogues and web-based discovery.
It distinguishes repository metadata, exchange profiles, transfer protocols and catalogue descriptions and shows how they work together.
Resource information
Resource type: Metadata profile overview
Profile level: Repository, exchange and catalogue level
Primary audience: Repository managers, metadata specialists, data stewards, software developers and research infrastructure providers
Primary uses: Metadata export, harvesting, mapping, catalogue integration, validation and discovery
Main specifications: DataCite Metadata Schema, DCMI Metadata Terms, OAI-PMH, OpenAIRE Guidelines, DCAT, DCAT-AP and Schema.org
Current catalogue baseline: DCAT 3 and DCAT-AP 3.0.1
Last reviewed: July 2026
Purpose and scope
A repository metadata profile defines how a repository records and manages information about deposited research resources.
A catalogue metadata profile defines how selected information is represented for discovery and exchange across organisations, platforms and research infrastructures.
The two profiles may use different classes, element names, obligation levels and controlled vocabularies. A mapping is therefore required between the local repository record and each destination profile.
Important: a repository record is normally richer than any single exported catalogue record. Export should preserve the information required by the destination without treating the external profile as a complete replacement for the local metadata model.
Repository and catalogue profiles
Repository profile
Defines the fields, validation rules and deposit workflow used to create and manage a local repository record.
Exchange profile
Selects and constrains metadata elements for transfer to a particular aggregator, service or infrastructure.
Transfer mechanism
Moves records between systems through OAI-PMH, an API, a metadata feed or another documented interface.
Catalogue profile
Represents datasets, distributions, data services and catalogues for cross-system discovery.
Four layers of metadata exchange
1. Local record
Metadata is created and maintained in the repository according to local deposit, curation and publication requirements.
2. Mapping
Local fields and values are mapped to the properties and controlled vocabularies of the destination profile.
3. Exposure
The mapped record is exposed through OAI-PMH, an API, RDF, JSON-LD or another supported representation.
4. Catalogue indexing
The receiving service validates, harvests, normalises and indexes the record for discovery and reuse.
Main standards and profiles
| Specification or profile | Primary function | Typical use |
|---|---|---|
| Local repository schema | Records the full metadata required for deposit, management, publication and curation. | Repository interface, internal database and deposit workflow. |
| DataCite Metadata Schema | Identifies, cites and connects research resources. | DOI registration and exchange of citation-oriented metadata. |
| DCMI Metadata Terms | Provides general cross-domain metadata properties and classes. | Generic repository records, mappings and RDF descriptions. |
| OAI Dublin Core — oai_dc | Provides a simple, broadly interoperable XML metadata format. | Basic OAI-PMH harvesting where richer profiles are unavailable. |
| OpenAIRE Guidelines | Constrain repository metadata for integration into the OpenAIRE information space. | OpenAIRE harvesting, validation and inclusion in the OpenAIRE Graph. |
| DCAT 3 | Provides an RDF vocabulary for interoperable data catalogues. | Description of catalogues, datasets, distributions and data services. |
| DCAT-AP 3.0.1 | Constrains DCAT for exchange between European data catalogues. | European catalogue portals and national or institutional profiles. |
| Schema.org | Embeds structured metadata in web pages. | Search-engine and web-based dataset discovery through JSON-LD, RDFa or Microdata. |
| OAI-PMH 2.0 | Transfers metadata records from a repository to harvesters. | Incremental harvesting using metadata formats identified by metadata prefixes. |
Protocol, format and profile are not the same
Protocol: defines how records are requested and transferred. OAI-PMH is a harvesting protocol.
Syntax: defines how the record is technically encoded, for example XML, JSON or RDF.
Vocabulary or schema: defines properties and classes, for example DCMI Terms, DataCite or DCAT.
Application profile: selects properties, obligation levels, cardinalities and controlled values for a particular use case.
Selecting the destination profile
| Destination | Primary profile or format | Additional requirement |
|---|---|---|
| DOI registration service | DataCite Metadata Schema | DOI, landing-page URL and valid DataCite metadata. |
| Generic OAI-PMH harvester | oai_dc or another declared metadata format | Valid OAI-PMH endpoint, identifiers, datestamps and metadataPrefix. |
| OpenAIRE | Relevant OpenAIRE Guidelines | Registration, validation and harvesting through OpenAIRE PROVIDE. |
| European data catalogue | DCAT-AP 3.0.1 or a compatible national profile | RDF representation and conformity with profile cardinalities and controlled vocabularies. |
| General web discovery | Schema.org Dataset | Structured data embedded in the landing page, preferably as JSON-LD. |
| Scientific domain catalogue | Core catalogue profile plus domain extension | Domain entities, vocabularies, variables and specialised relationships. |
Minimum metadata for repository exchange
| Element | Purpose in exchange | Requirement |
|---|---|---|
| Record identifier | Uniquely identifies the metadata record within the repository or harvesting interface. | Required |
| Resource identifier | Identifies the described dataset or other research resource, preferably through a DOI or another persistent URI. | Required |
| Landing-page URL | Directs users and machines to the authoritative repository record. | Required |
| Title | Provides the principal name of the resource. | Required |
| Creators | Supports attribution, citation and identity linking. | Required |
| Description | Allows users and discovery services to evaluate the resource. | Required |
| Resource type | Distinguishes datasets, software, publications and other resources. | Required |
| Publisher and publication date | Supports citation and identifies responsibility for publication. | Required at publication |
| Subjects | Supports thematic indexing, classification and search. | Recommended |
| Access rights | Indicates whether the resource is open, embargoed, restricted or closed. | Required |
| Licence | Identifies the legal conditions governing reuse. | Required where applicable |
| Distribution or access location | Identifies downloadable representations, access pages or data services. | Recommended |
| Format and media type | Helps systems and users select a usable distribution. | Recommended |
| Related resources | Connects publications, software, projects, versions and derived datasets. | Recommended |
| Modified date | Supports incremental harvesting and record synchronisation. | Required for harvesting |
Example metadata crosswalk
A crosswalk records how the same concept is represented in different metadata models. The mapping may require transformation rather than a direct replacement of one field name with another.
| Concept | DataCite | DCMI | DCAT / DCAT-AP | Schema.org |
|---|---|---|---|---|
| Title | titles.title |
dct:title |
dct:title |
name |
| Identifier | doi |
dct:identifier |
dct:identifier |
identifier |
| Creator | creators |
dct:creator |
dct:creator |
creator |
| Description | descriptions |
dct:description |
dct:description |
description |
| Subject | subjects |
dct:subject |
dcat:theme or dct:subject |
keywords or about |
| Publisher | publisher |
dct:publisher |
dct:publisher |
publisher |
| Licence | rightsList |
dct:license |
dct:license |
license |
| Distribution | Partly represented through Format, Size and the landing page | dct:hasFormat |
dcat:distribution |
distribution |
Why mappings are not always one-to-one
Different scopes: DataCite describes a citable research resource, while DCAT also models catalogues, distributions and data services.
Different cardinalities: one profile may permit several values where another expects only one.
Different controlled vocabularies: resource types, access rights and relationships may use different value sets.
Different structures: a structured creator with ORCID and affiliation may become an unstructured text value in a simpler export.
Different obligation levels: an optional local field may be mandatory for a destination catalogue.
OAI-PMH repository exposure
An OAI-PMH endpoint enables external services to retrieve metadata records from a repository without downloading or transferring the research data files themselves.
Base URL: stable address of the OAI-PMH endpoint.
Metadata prefix: identifies the requested metadata
format, such as oai_dc.
OAI identifier: uniquely identifies the metadata item within the endpoint.
Datestamp: records creation, modification or deletion for incremental harvesting.
Set: optionally groups records by collection, institution, type or another repository-defined category.
Deletion status: allows a harvester to recognise records removed from the repository.
Resumption token: supports harvesting large result sets in multiple responses.
OpenAIRE compatibility
OpenAIRE compatibility requires more than enabling an OAI-PMH endpoint. The exported records must follow the relevant OpenAIRE application profile and pass the applicable validation and registration process.
Register the source
Register the repository or data source through OpenAIRE PROVIDE.
Expose records
Provide a stable harvesting endpoint and the required metadata format.
Validate metadata
Check mandatory fields, controlled values, identifiers and relationships against the selected guideline.
Monitor harvesting
Review indexing status, validation reports and changes in source requirements.
DCAT and DCAT-AP catalogue model
| Class | Purpose |
|---|---|
| Catalogue | A curated collection of metadata descriptions for datasets, data services or related resources. |
| Catalogue record | The catalogue's own administrative record describing an entry and its update history. |
| Dataset | A collection of data published or curated by a single agent and available for access or download. |
| Dataset series | A collection of related datasets organised as a sequence or coherent group. |
| Distribution | An accessible representation of a dataset in a particular format, language or access arrangement. |
| Data service | A service providing access to or processing of datasets through a defined interface. |
Dataset and distribution
Dataset
The conceptual or intellectual data resource described by title, creators, subject, scope and provenance.
Distribution
A particular downloadable or accessible representation of that dataset.
Access URL
A page or endpoint through which access to the distribution can be obtained.
Download URL
A direct location from which a particular distribution can be downloaded.
Schema.org metadata for repository landing pages
Schema.org metadata can be embedded in a dataset landing page to make its content understandable to general web discovery services.
Dataset: describes the dataset and its principal descriptive properties.
DataCatalog: describes the repository or catalogue containing the dataset.
DataDownload: describes a downloadable representation, its format and content URL.
JSON-LD: is commonly used to embed the structured description in the HTML landing page.
Example of a catalogue-oriented record
The following simplified JSON-LD example illustrates the separation between a dataset and one of its distributions. It is explanatory and does not replace a complete DCAT-AP record.
{
"@context": {
"dcat": "http://www.w3.org/ns/dcat#",
"dct": "http://purl.org/dc/terms/"
},
"@id": "https://doi.org/10.xxxx/example-dataset",
"@type": "dcat:Dataset",
"dct:title": {
"@value": "Example research dataset",
"@language": "en"
},
"dct:description": {
"@value": "Data generated during the example study.",
"@language": "en"
},
"dct:identifier": "https://doi.org/10.xxxx/example-dataset",
"dct:creator": {
"@id": "https://orcid.org/0000-0000-0000-0000"
},
"dct:publisher": {
"@id": "https://ror.org/example"
},
"dct:license": {
"@id": "https://creativecommons.org/licenses/by/4.0/"
},
"dcat:landingPage": {
"@id": "https://repository.example.org/dataset/123"
},
"dcat:distribution": {
"@type": "dcat:Distribution",
"dcat:accessURL": {
"@id": "https://repository.example.org/dataset/123"
},
"dcat:downloadURL": {
"@id": "https://repository.example.org/files/data.csv"
},
"dct:format": "text/csv"
}
}
Validating exchanged metadata
Syntax validation
Checks whether XML, JSON, JSON-LD or RDF is technically well formed.
Profile validation
Checks mandatory properties, cardinalities, data types and controlled values.
Identifier validation
Checks that DOI, ORCID, ROR, licence and vocabulary URIs are valid and resolvable.
Harvesting validation
Tests endpoint behaviour, metadata prefixes, datestamps, sets, pagination and deleted records.
Common implementation errors
Treating OAI-PMH as a metadata schema: the endpoint works, but the exported records do not follow a defined application profile.
Assuming every mapping is one-to-one: structured metadata is flattened or important qualifiers are lost.
Using the file URL as the dataset identifier: the intellectual resource, landing page and individual distributions are not distinguished.
Combining licence and access status: legal reuse conditions and technical availability become ambiguous.
Omitting modification datestamps: harvesters cannot update records incrementally.
Exporting names as unstructured strings: ORCID, affiliation and contributor roles are lost.
Not describing distributions: catalogues can find the dataset but cannot determine its formats or access locations.
Using outdated destination requirements: validation fails even though the local metadata record is complete.
How to implement a repository-to-catalogue profile
1. Inventory local fields
Document local field names, structures, cardinalities, vocabularies and validation rules.
2. Select destinations
Identify the DOI service, aggregator, catalogue and web-discovery profiles that must be supported.
3. Create mappings
Define transformations, obligation levels, controlled-value mappings and rules for information that cannot be transferred.
4. Validate and monitor
Test exports and harvesting, document errors and monitor changes in external specifications.
Recommended implementation documentation
Local metadata dictionary: definitions of all repository fields.
Application profile: obligation levels, cardinalities and value rules.
Crosswalk table: mapping from local fields to every destination profile.
Controlled-value mapping: correspondence between local and external vocabularies.
Export specification: supported formats, endpoints and metadata prefixes.
Validation rules: syntax, profile and identifier checks.
Change log: documented updates to mappings and supported profile versions.
Explanatory status
This page provides explanatory and implementation guidance for mapping repository metadata to external catalogue and harvesting profiles.
It does not replace the official specifications, validation rules, onboarding procedures or technical requirements maintained by DataCite, DCMI, OpenAIRE, W3C, SEMIC or individual catalogue providers.
The requirements of a destination service may be stricter than the general metadata standard on which its profile is based.
Profile versions, controlled vocabularies and validation rules should be reviewed before implementing or updating a production metadata export.