Skip navigation

Core dataset metadata profile

A cross-disciplinary metadata profile defining the core elements required to identify, describe, cite, access, understand and reuse a research dataset.

The profile provides a common baseline for dataset preparation, repository deposit, DOI registration, catalogue exposure and FAIR-oriented metadata review.

It can be used for datasets from different research disciplines and extended with domain-specific elements describing methods, instruments, materials, software, models, parameters and scientific workflows.

ChatGPT Image 27 лип. 2026 р., 22_26_51

Resource information

Resource type: Metadata application profile

Profile level: Dataset-level, cross-disciplinary

Primary audience: Researchers, data stewards, repository managers, metadata specialists and research support staff

Primary uses: Dataset preparation, repository deposit, DOI registration, metadata review, catalogue exposure and data reuse

Alignment: DataCite Metadata Schema, DCMI Metadata Terms, Dataverse citation metadata and OpenAIRE interoperability requirements

Profile version: 1.0

Last reviewed: July 2026

Purpose and scope

The Core dataset metadata profile defines a common set of elements for describing a dataset independently of its scientific discipline, repository platform or file format.

It supports four related tasks: identifying and citing the dataset, explaining its content and context, defining access and reuse conditions, and recording the technical and provenance information required for interpretation.

The profile describes the dataset as a whole. Detailed scientific parameters, instrument settings, simulation configurations and file-specific variables should be added through domain-specific or file-level metadata.

Important: the profile defines the minimum common description of a dataset. It does not replace disciplinary metadata standards, repository requirements or the documentation contained in README, manifest and provenance files.

What the core profile supports

Identification

Identifies the dataset, its creators, version, publisher, persistent identifier and recommended citation.

Discovery

Describes the subject, content, methods, dates and coverage so that the dataset can be found and evaluated.

Access and reuse

Records access conditions, licence, contact information and relationships with publications, software and other datasets.

Interpretation

Provides formats, size, provenance, technical requirements and links to the documentation needed to understand the data.

Requirement levels

Each element is assigned a requirement level indicating when it should be included in a metadata record.

Required

Must be recorded for every dataset described using this profile.

Required at publication

May be absent during preparation but must be completed before the dataset is formally published.

Recommended

Should be included when the information is available because it improves discovery, interpretation or reuse.

Conditional

Must be included when the described condition applies to the dataset.

Cardinality

1: exactly one value is required.

1–n: one or more values are required.

0–1: the element is optional but may occur only once.

0–n: the element may be repeated when several values apply.

Core metadata elements

Identification and citation

Content, responsibility and discovery

Access, rights and relationships

Technical information and provenance

Minimum metadata for dataset publication

Before publication, the metadata record should contain enough information to identify the dataset, attribute responsibility, explain its content and establish the conditions under which it can be accessed and reused.

Required identification: title, creator and resource type.

Required description: description, subject or keywords and contact point.

Required publication information: publisher, publication year, persistent identifier and version where applicable.

Required rights information: access status and reuse licence.

Required relationships: identifiers of directly related publications, software, projects or dataset versions where these exist.

Metadata quality rules

Use identifiers

Record DOI, ORCID, ROR and other persistent identifiers as complete, resolvable URIs where the receiving system supports them.

Use controlled values

Apply recognised vocabularies for resource types, contributor roles, languages, licences, access rights and relationship types.

Separate concepts

Record creators, contributors, contacts, access rights, licences and related objects in separate structured elements.

Describe relationships

Do not provide an identifier alone. State whether the related object documents, cites, supplements, is a version of or is derived from the dataset.

Core and domain-specific metadata

The core profile should be applied to every dataset. Domain-specific metadata should then be added when scientific interpretation depends on specialised entities, methods, instruments, parameters or vocabularies.

Core profile: title, creators, description, subjects, identifiers, dates, access, rights, relationships and basic technical information.

Domain extension: materials, samples, instruments, experimental conditions, simulation parameters, models, variables, quality indicators and domain-specific workflows.

File-level metadata: filename, format, size, checksum, variable names, units and other properties of individual files.

Example of a core metadata record

The following simplified example shows how the principal elements may be represented in a structured record. It is illustrative and is not a replacement for a repository-specific or DataCite API format.

{
  "title": "Example research dataset",
  "resourceType": "Dataset",
  "creators": [
    {
      "name": "Researcher, Example",
      "orcid": "https://orcid.org/0000-0000-0000-0000",
      "affiliation": {
        "name": "Example Research Institution",
        "ror": "https://ror.org/example"
      }
    }
  ],
  "description": "Data generated and processed during the example study.",
  "subjects": [
    "research data",
    "example discipline"
  ],
  "publisher": "Example Data Repository",
  "publicationYear": "2026",
  "identifier": "https://doi.org/10.xxxx/example",
  "version": "1.0",
  "language": "en",
  "accessRights": "open",
  "licence": {
    "name": "Creative Commons Attribution 4.0 International",
    "uri": "https://creativecommons.org/licenses/by/4.0/"
  },
  "relatedIdentifiers": [
    {
      "identifier": "https://doi.org/10.xxxx/example-publication",
      "relationType": "IsSupplementTo"
    }
  ],
  "formats": [
    "text/csv",
    "application/json"
  ],
  "documentation": [
    "README.md",
    "manifest.csv",
    "provenance.json"
  ]
}

How to apply the profile

1. Describe the dataset

Complete the required elements using the dataset content, project documentation and information supplied by the creators.

2. Add identifiers

Verify ORCID, ROR, grant identifiers and persistent identifiers for related research outputs.

3. Extend the profile

Add domain-specific, workflow and file-level metadata required to interpret the particular dataset.

4. Validate the record

Check completeness, controlled values, relationships, access conditions and the metadata exported by the repository.

Explanatory status

This profile provides a common implementation baseline for the preparation and review of dataset metadata. It does not replace the official specifications maintained by DataCite, DCMI, repository providers or research infrastructures.

Individual repositories and scientific communities may require additional elements, controlled vocabularies or validation rules. The requirements of the destination repository and any applicable domain-specific profile should therefore be checked before publication.

The profile may be updated when underlying standards, repository requirements or interoperability guidelines change.