Skip navigation

Domain-specific metadata profiles

Metadata profiles extending a common dataset description with the scientific entities, methods, instruments, parameters, vocabularies and relationships required by a particular research domain.

Domain-specific profiles help research communities describe data at the level needed for scientific interpretation, comparison, validation and reuse.

They complement general repository and citation metadata rather than replacing it and should preserve mappings to cross-disciplinary standards and discovery services.

ChatGPT Image 27 лип. 2026 р., 22_26_38 (3)

Resource information

Resource type: Metadata profile guidance

Profile level: Scientific domain, method, instrument, workflow and data-type level

Primary audience: Researchers, data stewards, domain experts, repository managers, metadata specialists and research software engineers

Primary uses: Scientific documentation, repository deposit, machine-actionable data exchange, workflow reproducibility, domain discovery and reuse

Profile basis: Core dataset metadata, community standards, controlled vocabularies, domain ontologies and provenance models

Last reviewed: July 2026

Purpose and scope

A domain-specific metadata profile defines the information required to understand and reuse data produced within a particular scientific domain, method, instrument or workflow.

It extends general dataset metadata with descriptions of scientific objects, samples, variables, instruments, experimental conditions, computational parameters, software, models, quality indicators and processing history.

The profile should be based on standards and terminology accepted by the relevant research community and should remain mappable to general repository, DOI and catalogue metadata.

Important: a domain-specific profile should add the scientific detail required for interpretation without duplicating core information such as title, creators, publisher, DOI, licence and access rights.

Why domain-specific metadata is needed

Scientific interpretation

Records the entities, variables, methods and conditions needed to understand what was measured, calculated or observed.

Comparison

Uses consistent properties, units and vocabularies so that data from different studies can be compared.

Reproducibility

Captures instruments, software, parameters, configurations and processing steps required to repeat an experiment or computation.

Machine actionability

Represents domain concepts and relationships in structured forms that can be interpreted by software and research infrastructures.

Layers of a complete metadata description

Core dataset metadata

Identifies and cites the dataset and records its creators, description, access conditions, licence and related outputs.

Domain metadata

Describes scientific entities, variables, methods, instruments, conditions and community terminology.

Workflow and provenance

Connects inputs, processing steps, software, parameters, agents and outputs.

File-level metadata

Records filenames, formats, sizes, checksums, variables, units and the role of each file within the package.

Types of domain-specific metadata profiles

What a domain profile must define

Extending the core dataset profile

A domain profile should explicitly state which metadata elements are inherited from the Core dataset metadata profile and which elements are introduced or constrained by the domain extension.

Inherit: title, creators, DOI, description, publisher, dates, access rights, licence, funding and general relationships.

Extend: scientific entities, methods, instruments, variables, parameters, conditions, units and quality indicators.

Constrain: specify narrower resource types, controlled values, permitted units and required relationships.

Map: document how domain properties are represented in repository, catalogue and exchange metadata.

Examples of domain standards and profiles

Selecting an appropriate domain standard

1. Identify the data

Determine the scientific objects, methods, data structures and workflows that must be represented.

2. Review community practice

Check standards used by established domain repositories, research infrastructures, journals and international projects.

3. Assess implementation

Verify that documentation, schemas, controlled values, examples, validators and software support are available.

4. Record the decision

Document the selected version, local constraints, mappings, extensions and maintenance responsibility.

Selection criteria

Community adoption: the standard is recognised and used by the relevant research community.

Scientific coverage: it represents the entities, properties and processes needed for the intended data.

Machine actionability: schemas, persistent terms and structured serialisations are available.

Validation: conformance can be tested using documented or machine-actionable rules.

Governance: responsibility for maintenance and version management is identifiable.

Interoperability: mappings to core, repository and cross-domain metadata can be established.

Implementation support: examples, libraries, parsers, converters or reference implementations are available.

Controlled vocabularies and ontologies

A domain profile should identify the terminology used for scientific entities and values instead of relying on unrestricted text wherever a recognised vocabulary is available.

Term label: a human-readable value displayed to users.

Persistent URI: a machine-resolvable identifier for the concept.

Vocabulary source: the ontology, thesaurus or controlled list from which the term is selected.

Vocabulary version: the release used when the metadata was created.

Mapping: documented equivalence or relationship with terminology used by another profile.

Units and quantitative values

Separate value and unit: store the numerical value independently from its unit.

Use standard units: prefer a recognised unit system and machine-readable unit identifiers.

Record uncertainty: include uncertainty, tolerance or confidence information where scientifically relevant.

Define ranges: distinguish a single value, interval, minimum, maximum and detection limit.

Preserve original values: document transformations when values are normalised or converted.

Domain metadata for computational workflows

Example: DFT → MLIP → MD metadata extension

A materials-science workflow may extend the core dataset description with separate metadata groups for electronic-structure calculations, machine-learned interatomic potentials and molecular-dynamics simulation.

Domain metadata for experimental data

Provenance in a domain profile

Provenance should show how a scientific object or dataset was created, transformed and connected to the agents, activities and resources involved.

Entity: sample, input file, dataset, model, software object or output.

Activity: measurement, simulation, conversion, training, validation or analysis.

Agent: researcher, organisation, software service or instrument responsible for an action.

Relationship: generated by, used, derived from, attributed to or associated with.

Execution detail: date, version, configuration, status and identifier of a particular run.

Machine-actionable representation

Stable identifiers

Use persistent identifiers or URIs for properties, entities, vocabularies, agents and related objects.

Structured values

Represent quantities, units, names, dates, parameters and relationships as separate machine-readable components.

Explicit conformance

Record the profile name, version and persistent reference to which the metadata conforms.

Validation rules

Publish schemas or constraint definitions that allow automated conformance checking.

Domain and cross-domain interoperability

Domain metadata should preserve scientific precision while also exposing a cross-disciplinary layer that enables discovery and reuse outside the original community.

Domain layer: detailed scientific entities, methods, variables, conditions and relationships.

Cross-domain layer: dataset identity, agents, description, access, rights, provenance and general relationships.

Mapping layer: documented correspondences between domain terms and shared metadata models.

Discovery layer: metadata exposed through repository, DataCite, DCAT, Schema.org or another catalogue profile.

Example of a domain profile record

The following simplified example illustrates how general dataset metadata can be combined with a materials-science workflow extension. It is explanatory and does not replace a formal schema.

{
  "profile": {
    "name": "Computational materials workflow profile",
    "version": "1.0",
    "conformsTo": "https://example.org/profiles/materials-workflow/1.0"
  },
  "dataset": {
    "identifier": "https://doi.org/10.xxxx/example",
    "title": "Example DFT to MLIP to MD dataset",
    "resourceType": "Dataset"
  },
  "materialSystem": {
    "chemicalFormula": "SiC",
    "structureType": "crystalline",
    "periodicity": "three-dimensional"
  },
  "workflow": {
    "stages": [
      {
        "id": "dft-01",
        "type": "electronic-structure-calculation",
        "software": {
          "name": "Example DFT software",
          "version": "1.0"
        },
        "parameters": {
          "exchangeCorrelationFunctional": "example-value",
          "energyCutoff": {
            "value": 500,
            "unit": "eV"
          }
        },
        "outputs": [
          "reference-structures",
          "energies",
          "forces"
        ]
      },
      {
        "id": "mlip-01",
        "type": "model-training",
        "used": [
          "dft-01"
        ],
        "outputs": [
          "trained-model",
          "evaluation-metrics"
        ]
      },
      {
        "id": "md-01",
        "type": "molecular-dynamics-simulation",
        "used": [
          "mlip-01"
        ],
        "parameters": {
          "ensemble": "NVT",
          "temperature": {
            "value": 300,
            "unit": "K"
          }
        },
        "outputs": [
          "trajectory",
          "thermodynamic-results"
        ]
      }
    ]
  }
}

Validating a domain-specific profile

Structural validation

Checks required sections, element names, data types, cardinalities and nesting.

Vocabulary validation

Checks whether terms, identifiers, units and relationship types come from permitted sources.

Scientific validation

Checks domain rules, dependencies, ranges and combinations that cannot be tested through syntax alone.

Package validation

Checks that metadata references correspond to available files, identifiers, checksums, software and outputs.

Common implementation errors

Creating a local profile before reviewing community standards: existing terminology and implementation experience are unnecessarily duplicated.

Adding fields without definitions: different researchers interpret the same field differently.

Using unrestricted text for everything: values cannot be compared or processed consistently.

Embedding units in text: quantitative values become difficult to validate and convert.

Mixing dataset and file metadata: properties of the entire dataset are confused with properties of individual files.

Recording a workflow as one narrative paragraph: inputs, outputs, parameters and dependencies cannot be interpreted automatically.

Omitting profile versions: later users cannot determine which rules governed the metadata record.

Breaking cross-domain discovery: the domain record does not retain mappings to title, creators, identifiers, rights and general relationships.

How to develop a domain-specific metadata profile

1. Analyse the workflow

Identify scientific objects, data flows, methods, decisions, parameters and outputs.

2. Reuse standards

Select existing schemas, vocabularies, identifiers and community reporting requirements.

3. Define the profile

Specify elements, obligation levels, cardinalities, value constraints, mappings and examples.

4. Test and maintain

Validate the profile with real datasets, publish its version and establish a change-management process.

Recommended profile documentation

Scope statement: domains, methods and object types covered.

Conceptual model: principal entities and relationships.

Element specification: definitions, rules and examples.

Controlled vocabularies: permitted terminology and units.

Mappings: core, repository, catalogue and external standards.

Serialisations: supported JSON, JSON-LD, XML, RDF or tabular formats.

Validation resources: schemas, shapes, tests and error messages.

Examples: complete representative records and packages.

Version history: releases, changes and migration guidance.

Governance: maintainers, approval process and issue reporting.

Explanatory status

This page provides explanatory and implementation guidance for selecting, adapting and developing domain-specific metadata profiles.

It does not replace the official specifications, schemas, controlled vocabularies, reporting guidelines or validation rules maintained by scientific communities and standards organisations.

A local profile should reuse established standards wherever possible, clearly identify all extensions and preserve mappings to general repository and discovery metadata.

Standards, ontologies and community practices may change. Their current versions and governance status should therefore be reviewed before a profile is implemented or revised.