Core dataset metadata profile
A cross-disciplinary metadata profile defining the core elements required to identify, describe, cite, access, understand and reuse a research dataset.
The profile provides a common baseline for dataset preparation, repository deposit, DOI registration, catalogue exposure and FAIR-oriented metadata review.
It can be used for datasets from different research disciplines and extended with domain-specific elements describing methods, instruments, materials, software, models, parameters and scientific workflows.
Resource information
Resource type: Metadata application profile
Profile level: Dataset-level, cross-disciplinary
Primary audience: Researchers, data stewards, repository managers, metadata specialists and research support staff
Primary uses: Dataset preparation, repository deposit, DOI registration, metadata review, catalogue exposure and data reuse
Alignment: DataCite Metadata Schema, DCMI Metadata Terms, Dataverse citation metadata and OpenAIRE interoperability requirements
Profile version: 1.0
Last reviewed: July 2026
Purpose and scope
The Core dataset metadata profile defines a common set of elements for describing a dataset independently of its scientific discipline, repository platform or file format.
It supports four related tasks: identifying and citing the dataset, explaining its content and context, defining access and reuse conditions, and recording the technical and provenance information required for interpretation.
The profile describes the dataset as a whole. Detailed scientific parameters, instrument settings, simulation configurations and file-specific variables should be added through domain-specific or file-level metadata.
Important: the profile defines the minimum common description of a dataset. It does not replace disciplinary metadata standards, repository requirements or the documentation contained in README, manifest and provenance files.
What the core profile supports
Identification
Identifies the dataset, its creators, version, publisher, persistent identifier and recommended citation.
Discovery
Describes the subject, content, methods, dates and coverage so that the dataset can be found and evaluated.
Access and reuse
Records access conditions, licence, contact information and relationships with publications, software and other datasets.
Interpretation
Provides formats, size, provenance, technical requirements and links to the documentation needed to understand the data.
Requirement levels
Each element is assigned a requirement level indicating when it should be included in a metadata record.
Required
Must be recorded for every dataset described using this profile.
Required at publication
May be absent during preparation but must be completed before the dataset is formally published.
Recommended
Should be included when the information is available because it improves discovery, interpretation or reuse.
Conditional
Must be included when the described condition applies to the dataset.
Cardinality
1: exactly one value is required.
1–n: one or more values are required.
0–1: the element is optional but may occur only once.
0–n: the element may be repeated when several values apply.
Core metadata elements
Identification and citation
| Element | Requirement | Cardinality | Content and value guidance | Main alignment |
|---|---|---|---|---|
| Title | Required | 1–n | A specific and informative name for the dataset. Alternative or translated titles may be recorded separately. |
DataCite Title; DCMI title |
| Creator | Required | 1–n | Person or organisation primarily responsible for creating the dataset. Record family and given names separately where supported. |
DataCite Creator; DCMI creator |
| Creator identifier | Recommended | 0–n | Persistent identifier for a creator. Use an ORCID iD for an individual researcher where available. | DataCite nameIdentifier |
| Affiliation | Recommended | 0–n | Organisation associated with the creator when the dataset was produced. Include a ROR identifier where available. | DataCite affiliation |
| Publisher | Required at publication | 0–1 / 1 | Organisation responsible for making the dataset available. This is commonly the repository or hosting institution. |
DataCite Publisher; DCMI publisher |
| Publication year | Required at publication | 0–1 / 1 | Year in which the dataset was or will be made publicly available. |
DataCite PublicationYear; DCMI issued |
| Resource type | Required | 1 | General and, where appropriate, specific type of the resource. The general value for this profile is Dataset. |
DataCite ResourceType; DCMI type |
| Persistent identifier | Required at publication | 0–1 / 1 | Globally resolvable persistent identifier assigned to the published dataset, normally a DOI. |
DataCite Identifier; DCMI identifier |
| Version | Recommended | 0–1 | Version designation of the dataset. Use a consistent institutional or semantic versioning convention. | DataCite Version |
Content, responsibility and discovery
| Element | Requirement | Cardinality | Content and value guidance | Main alignment |
|---|---|---|---|---|
| Description | Required | 1–n | A concise explanation of the dataset content, purpose, scope, principal variables or objects and potential uses. |
DataCite Description; DCMI description |
| Subject and keywords | Required | 1–n | Topics, disciplines, objects or methods represented in the dataset. Use controlled vocabularies where suitable. |
DataCite Subject; DCMI subject |
| Language | Recommended | 0–1 | Main language of textual content or documentation. Use an established language code. |
DataCite Language; DCMI language |
| Contributor | Conditional | 0–n | Person or organisation contributing to collection, processing, management, validation or distribution without being a primary creator. |
DataCite Contributor; DCMI contributor |
| Contact point | Required | 1–n | Person, role-based address or organisational unit that can respond to questions about the dataset. |
DataCite Contributor:
ContactPerson; Dataverse Dataset Contact |
| Methods | Recommended | 0–n | Summary of the methods used to collect, generate, process or analyse the data, with links to fuller documentation where available. | DataCite Description: Methods |
| Relevant dates | Recommended | 0–n | Dates associated with creation, collection, updating, availability, acceptance or withdrawal. Record the date type explicitly. |
DataCite Date; DCMI date |
| Spatial coverage | Conditional | 0–n | Geographic area, named place, coordinates or bounding box covered by the dataset. |
DataCite GeoLocation; DCMI spatial |
| Temporal coverage | Conditional | 0–n | Time period represented by the data, distinct from the date when the dataset was created or published. |
DataCite Date:
Coverage; DCMI temporal |
| Funding information | Conditional | 0–n | Funder name, funder identifier, award number and award title for research that supported creation of the dataset. | DataCite FundingReference |
Access, rights and relationships
| Element | Requirement | Cardinality | Content and value guidance | Main alignment |
|---|---|---|---|---|
| Licence | Required at publication | 0–1 / 1 | Standard licence or rights statement defining permitted reuse. Record both the licence name and persistent URI. |
DataCite Rights; DCMI license |
| Access rights | Required | 1 | Access category such as open, embargoed, restricted or closed, recorded separately from the reuse licence. |
DataCite Rights; DCMI accessRights; OpenAIRE access terms |
| Access conditions | Conditional | 0–1 | Procedure, eligibility requirements, approval process or technical conditions that apply when access is not fully open. | Repository terms of access |
| Embargo end date | Conditional | 0–1 | Date on which embargoed data becomes openly available. | DataCite Date: Available |
| Data sensitivity statement | Conditional | 0–1 | Statement indicating whether the dataset contains personal, confidential, security-sensitive or otherwise controlled information. |
Local profile; repository access policy |
| Related identifiers | Recommended | 0–n | Persistent identifiers for related publications, software, projects, instruments, samples, datasets or earlier and later versions. Always specify the relationship type. |
DataCite RelatedIdentifier; DCMI relation |
Technical information and provenance
| Element | Requirement | Cardinality | Content and value guidance | Main alignment |
|---|---|---|---|---|
| Format | Recommended | 0–n | File formats or media types included in the dataset. Use recognised format names or media types where possible. |
DataCite Format; DCMI format |
| Size | Recommended | 0–n | Total data volume, number of files, number of records or another measure relevant to the dataset. |
DataCite Size; DCMI extent |
| Technical requirements | Conditional | 0–n | Software, operating environment, libraries, codecs, hardware or other requirements needed to open or process the data. | DataCite Description: TechnicalInfo |
| Provenance | Recommended | 0–n | Information about data origin, processing history, transformations, quality-control actions and relationships between inputs and outputs. |
DCMI provenance; related workflow documentation |
| File inventory or manifest | Recommended | 0–1 | Link to or identification of a structured file inventory containing filenames, paths, roles, formats, sizes and checksums. | Local package profile |
| Checksums | Conditional | 0–n | Cryptographic checksum for individual files or the package, used to verify integrity. | Repository or file-level metadata |
Minimum metadata for dataset publication
Before publication, the metadata record should contain enough information to identify the dataset, attribute responsibility, explain its content and establish the conditions under which it can be accessed and reused.
Required identification: title, creator and resource type.
Required description: description, subject or keywords and contact point.
Required publication information: publisher, publication year, persistent identifier and version where applicable.
Required rights information: access status and reuse licence.
Required relationships: identifiers of directly related publications, software, projects or dataset versions where these exist.
Metadata quality rules
Use identifiers
Record DOI, ORCID, ROR and other persistent identifiers as complete, resolvable URIs where the receiving system supports them.
Use controlled values
Apply recognised vocabularies for resource types, contributor roles, languages, licences, access rights and relationship types.
Separate concepts
Record creators, contributors, contacts, access rights, licences and related objects in separate structured elements.
Describe relationships
Do not provide an identifier alone. State whether the related object documents, cites, supplements, is a version of or is derived from the dataset.
Core and domain-specific metadata
The core profile should be applied to every dataset. Domain-specific metadata should then be added when scientific interpretation depends on specialised entities, methods, instruments, parameters or vocabularies.
Core profile: title, creators, description, subjects, identifiers, dates, access, rights, relationships and basic technical information.
Domain extension: materials, samples, instruments, experimental conditions, simulation parameters, models, variables, quality indicators and domain-specific workflows.
File-level metadata: filename, format, size, checksum, variable names, units and other properties of individual files.
Example of a core metadata record
The following simplified example shows how the principal elements may be represented in a structured record. It is illustrative and is not a replacement for a repository-specific or DataCite API format.
{
"title": "Example research dataset",
"resourceType": "Dataset",
"creators": [
{
"name": "Researcher, Example",
"orcid": "https://orcid.org/0000-0000-0000-0000",
"affiliation": {
"name": "Example Research Institution",
"ror": "https://ror.org/example"
}
}
],
"description": "Data generated and processed during the example study.",
"subjects": [
"research data",
"example discipline"
],
"publisher": "Example Data Repository",
"publicationYear": "2026",
"identifier": "https://doi.org/10.xxxx/example",
"version": "1.0",
"language": "en",
"accessRights": "open",
"licence": {
"name": "Creative Commons Attribution 4.0 International",
"uri": "https://creativecommons.org/licenses/by/4.0/"
},
"relatedIdentifiers": [
{
"identifier": "https://doi.org/10.xxxx/example-publication",
"relationType": "IsSupplementTo"
}
],
"formats": [
"text/csv",
"application/json"
],
"documentation": [
"README.md",
"manifest.csv",
"provenance.json"
]
}
How to apply the profile
1. Describe the dataset
Complete the required elements using the dataset content, project documentation and information supplied by the creators.
2. Add identifiers
Verify ORCID, ROR, grant identifiers and persistent identifiers for related research outputs.
3. Extend the profile
Add domain-specific, workflow and file-level metadata required to interpret the particular dataset.
4. Validate the record
Check completeness, controlled values, relationships, access conditions and the metadata exported by the repository.
Explanatory status
This profile provides a common implementation baseline for the preparation and review of dataset metadata. It does not replace the official specifications maintained by DataCite, DCMI, repository providers or research infrastructures.
Individual repositories and scientific communities may require additional elements, controlled vocabularies or validation rules. The requirements of the destination repository and any applicable domain-specific profile should therefore be checked before publication.
The profile may be updated when underlying standards, repository requirements or interoperability guidelines change.