Domain-specific metadata profiles
Metadata profiles extending a common dataset description with the scientific entities, methods, instruments, parameters, vocabularies and relationships required by a particular research domain.
Domain-specific profiles help research communities describe data at the level needed for scientific interpretation, comparison, validation and reuse.
They complement general repository and citation metadata rather than replacing it and should preserve mappings to cross-disciplinary standards and discovery services.
Resource information
Resource type: Metadata profile guidance
Profile level: Scientific domain, method, instrument, workflow and data-type level
Primary audience: Researchers, data stewards, domain experts, repository managers, metadata specialists and research software engineers
Primary uses: Scientific documentation, repository deposit, machine-actionable data exchange, workflow reproducibility, domain discovery and reuse
Profile basis: Core dataset metadata, community standards, controlled vocabularies, domain ontologies and provenance models
Last reviewed: July 2026
Purpose and scope
A domain-specific metadata profile defines the information required to understand and reuse data produced within a particular scientific domain, method, instrument or workflow.
It extends general dataset metadata with descriptions of scientific objects, samples, variables, instruments, experimental conditions, computational parameters, software, models, quality indicators and processing history.
The profile should be based on standards and terminology accepted by the relevant research community and should remain mappable to general repository, DOI and catalogue metadata.
Important: a domain-specific profile should add the scientific detail required for interpretation without duplicating core information such as title, creators, publisher, DOI, licence and access rights.
Why domain-specific metadata is needed
Scientific interpretation
Records the entities, variables, methods and conditions needed to understand what was measured, calculated or observed.
Comparison
Uses consistent properties, units and vocabularies so that data from different studies can be compared.
Reproducibility
Captures instruments, software, parameters, configurations and processing steps required to repeat an experiment or computation.
Machine actionability
Represents domain concepts and relationships in structured forms that can be interpreted by software and research infrastructures.
Layers of a complete metadata description
Core dataset metadata
Identifies and cites the dataset and records its creators, description, access conditions, licence and related outputs.
Domain metadata
Describes scientific entities, variables, methods, instruments, conditions and community terminology.
Workflow and provenance
Connects inputs, processing steps, software, parameters, agents and outputs.
File-level metadata
Records filenames, formats, sizes, checksums, variables, units and the role of each file within the package.
Types of domain-specific metadata profiles
| Profile type | Scope | Examples of additional metadata |
|---|---|---|
| Disciplinary profile | A broad scientific field or research community. | Domain entities, classifications, variables, methods and standard terminology. |
| Method-specific profile | A particular experimental or computational method. | Method configuration, calibration, parameters, uncertainty and quality controls. |
| Instrument-specific profile | Data generated using a particular type of scientific instrument. | Instrument model, detector, acquisition settings, calibration and sample environment. |
| Workflow profile | A sequence of experimental, computational or data-processing steps. | Inputs, outputs, software, dependencies, parameters and execution relationships. |
| Data-type profile | A recurring scientific data structure or digital object. | Dimensions, variables, coordinate systems, units, encoding and quality properties. |
| Repository application profile | Implementation of domain metadata within a repository. | Selected elements, obligation levels, cardinalities, vocabularies and local mappings. |
What a domain profile must define
| Profile component | Required documentation |
|---|---|
| Element identifier | Stable technical name or persistent URI identifying the metadata property. |
| Label | Human-readable name displayed in forms and documentation. |
| Definition | Precise explanation of the concept represented by the element. |
| Requirement level | Required, recommended, optional or conditional. |
| Cardinality | Number of permitted or required occurrences. |
| Data type | Text, number, date, Boolean value, identifier, controlled term, structured object or relationship. |
| Value constraint | Permitted format, range, pattern or controlled vocabulary. |
| Units | Required unit or unit vocabulary for quantitative values. |
| Vocabulary or ontology | Authoritative terminology used to represent entities and values. |
| Relationships | Typed links to samples, methods, instruments, software, files, agents and outputs. |
| Mapping | Correspondence with core, repository, DataCite, catalogue or other domain standards. |
| Example | Valid representative value or structured record. |
| Validation rule | Machine-actionable condition used to test profile conformance. |
Extending the core dataset profile
A domain profile should explicitly state which metadata elements are inherited from the Core dataset metadata profile and which elements are introduced or constrained by the domain extension.
Inherit: title, creators, DOI, description, publisher, dates, access rights, licence, funding and general relationships.
Extend: scientific entities, methods, instruments, variables, parameters, conditions, units and quality indicators.
Constrain: specify narrower resource types, controlled values, permitted units and required relationships.
Map: document how domain properties are represented in repository, catalogue and exchange metadata.
Examples of domain standards and profiles
| Domain or use case | Example | What it supports |
|---|---|---|
| Life sciences | Bioschemas Dataset | A Schema.org-based profile for consistent web descriptions of life sciences datasets. |
| Computational workflows | Bioschemas ComputationalWorkflow | Description of executable or repeatable workflows, their inputs, outputs, software requirements and related resources. |
| Biodiversity | Darwin Core | Standard concepts and terms for sharing information about biological diversity, occurrences, organisms, taxa and events. |
| Social, economic and health sciences | DDI-Lifecycle and DDI-CDI | Documentation of studies, variables, data structures, processing and data integration. |
| Bioimaging | OME Data Model and OME-XML | Image dimensions, pixel types, acquisition metadata, annotations, instruments and regions of interest. |
| Neutron, X-ray and related experiments | NeXus application definitions | Domain-specific structures for experimental data, samples, instruments, processing and method-specific metadata. |
| Materials science | NOMAD Metainfo | Hierarchical and extensible schemas for computational and experimental materials data. |
| Materials databases | OPTIMADE | A common API specification and data structures for interoperable retrieval from materials databases. |
Selecting an appropriate domain standard
1. Identify the data
Determine the scientific objects, methods, data structures and workflows that must be represented.
2. Review community practice
Check standards used by established domain repositories, research infrastructures, journals and international projects.
3. Assess implementation
Verify that documentation, schemas, controlled values, examples, validators and software support are available.
4. Record the decision
Document the selected version, local constraints, mappings, extensions and maintenance responsibility.
Selection criteria
Community adoption: the standard is recognised and used by the relevant research community.
Scientific coverage: it represents the entities, properties and processes needed for the intended data.
Machine actionability: schemas, persistent terms and structured serialisations are available.
Validation: conformance can be tested using documented or machine-actionable rules.
Governance: responsibility for maintenance and version management is identifiable.
Interoperability: mappings to core, repository and cross-domain metadata can be established.
Implementation support: examples, libraries, parsers, converters or reference implementations are available.
Controlled vocabularies and ontologies
A domain profile should identify the terminology used for scientific entities and values instead of relying on unrestricted text wherever a recognised vocabulary is available.
Term label: a human-readable value displayed to users.
Persistent URI: a machine-resolvable identifier for the concept.
Vocabulary source: the ontology, thesaurus or controlled list from which the term is selected.
Vocabulary version: the release used when the metadata was created.
Mapping: documented equivalence or relationship with terminology used by another profile.
Units and quantitative values
Separate value and unit: store the numerical value independently from its unit.
Use standard units: prefer a recognised unit system and machine-readable unit identifiers.
Record uncertainty: include uncertainty, tolerance or confidence information where scientifically relevant.
Define ranges: distinguish a single value, interval, minimum, maximum and detection limit.
Preserve original values: document transformations when values are normalised or converted.
Domain metadata for computational workflows
| Metadata group | Information to record |
|---|---|
| Workflow identity | Workflow name, version, identifier, purpose and responsible creators. |
| Workflow stages | Ordered or dependency-based description of computational steps. |
| Inputs | Input datasets, configurations, structures, models and parameter files. |
| Outputs | Generated datasets, trained models, trajectories, metrics, reports and visualisations. |
| Software | Program names, versions, source identifiers, dependencies and execution environment. |
| Parameters | Method parameters, numerical settings, thresholds, random seeds and convergence criteria. |
| Computing environment | Operating system, libraries, container, hardware, accelerators and scheduler configuration. |
| Execution | Start and end time, status, resource use, logs and execution identifier. |
| Provenance | Typed links connecting agents, activities, inputs, transformations and outputs. |
Example: DFT → MLIP → MD metadata extension
A materials-science workflow may extend the core dataset description with separate metadata groups for electronic-structure calculations, machine-learned interatomic potentials and molecular-dynamics simulation.
| Workflow stage | Examples of domain metadata |
|---|---|
| Material or system | Composition, structure, phase, dimensionality, defects, supercell, boundary conditions and structure identifier. |
| DFT calculation | Code and version, exchange-correlation functional, pseudopotentials, basis or cutoff, k-point sampling, spin settings and convergence criteria. |
| Reference data | Structures, energies, forces, stresses, sampling method, data split and filtering procedures. |
| MLIP training | Model architecture, software, hyperparameters, random seed, training and validation sets, loss function and checkpoints. |
| Model evaluation | Metrics, test sets, error definitions, uncertainty indicators, benchmark results and applicability limits. |
| MD simulation | Engine, model identifier, ensemble, temperature, pressure, timestep, duration, thermostat, barostat and initial configuration. |
| Outputs | Trajectories, thermodynamic quantities, structures, derived properties, plots and analysis results. |
| Provenance chain | Links from structures and DFT calculations to training records, model versions, MD runs and final datasets. |
Domain metadata for experimental data
| Metadata group | Information to record |
|---|---|
| Sample | Identity, composition, preparation, geometry, state and sample history. |
| Instrument | Instrument type, manufacturer, model, facility, identifier and component configuration. |
| Acquisition | Acquisition mode, date, duration, scan or exposure settings and measurement sequence. |
| Environment | Temperature, pressure, atmosphere, field, humidity and other experimental conditions. |
| Calibration | Standards, reference materials, calibration date, method and correction procedures. |
| Processing | Background subtraction, filtering, reconstruction, normalisation, fitting and derived-data generation. |
| Quality | Measurement uncertainty, detection limits, artefacts, exclusions and quality-control outcomes. |
Provenance in a domain profile
Provenance should show how a scientific object or dataset was created, transformed and connected to the agents, activities and resources involved.
Entity: sample, input file, dataset, model, software object or output.
Activity: measurement, simulation, conversion, training, validation or analysis.
Agent: researcher, organisation, software service or instrument responsible for an action.
Relationship: generated by, used, derived from, attributed to or associated with.
Execution detail: date, version, configuration, status and identifier of a particular run.
Machine-actionable representation
Stable identifiers
Use persistent identifiers or URIs for properties, entities, vocabularies, agents and related objects.
Structured values
Represent quantities, units, names, dates, parameters and relationships as separate machine-readable components.
Explicit conformance
Record the profile name, version and persistent reference to which the metadata conforms.
Validation rules
Publish schemas or constraint definitions that allow automated conformance checking.
Domain and cross-domain interoperability
Domain metadata should preserve scientific precision while also exposing a cross-disciplinary layer that enables discovery and reuse outside the original community.
Domain layer: detailed scientific entities, methods, variables, conditions and relationships.
Cross-domain layer: dataset identity, agents, description, access, rights, provenance and general relationships.
Mapping layer: documented correspondences between domain terms and shared metadata models.
Discovery layer: metadata exposed through repository, DataCite, DCAT, Schema.org or another catalogue profile.
Example of a domain profile record
The following simplified example illustrates how general dataset metadata can be combined with a materials-science workflow extension. It is explanatory and does not replace a formal schema.
{
"profile": {
"name": "Computational materials workflow profile",
"version": "1.0",
"conformsTo": "https://example.org/profiles/materials-workflow/1.0"
},
"dataset": {
"identifier": "https://doi.org/10.xxxx/example",
"title": "Example DFT to MLIP to MD dataset",
"resourceType": "Dataset"
},
"materialSystem": {
"chemicalFormula": "SiC",
"structureType": "crystalline",
"periodicity": "three-dimensional"
},
"workflow": {
"stages": [
{
"id": "dft-01",
"type": "electronic-structure-calculation",
"software": {
"name": "Example DFT software",
"version": "1.0"
},
"parameters": {
"exchangeCorrelationFunctional": "example-value",
"energyCutoff": {
"value": 500,
"unit": "eV"
}
},
"outputs": [
"reference-structures",
"energies",
"forces"
]
},
{
"id": "mlip-01",
"type": "model-training",
"used": [
"dft-01"
],
"outputs": [
"trained-model",
"evaluation-metrics"
]
},
{
"id": "md-01",
"type": "molecular-dynamics-simulation",
"used": [
"mlip-01"
],
"parameters": {
"ensemble": "NVT",
"temperature": {
"value": 300,
"unit": "K"
}
},
"outputs": [
"trajectory",
"thermodynamic-results"
]
}
]
}
}
Validating a domain-specific profile
Structural validation
Checks required sections, element names, data types, cardinalities and nesting.
Vocabulary validation
Checks whether terms, identifiers, units and relationship types come from permitted sources.
Scientific validation
Checks domain rules, dependencies, ranges and combinations that cannot be tested through syntax alone.
Package validation
Checks that metadata references correspond to available files, identifiers, checksums, software and outputs.
Common implementation errors
Creating a local profile before reviewing community standards: existing terminology and implementation experience are unnecessarily duplicated.
Adding fields without definitions: different researchers interpret the same field differently.
Using unrestricted text for everything: values cannot be compared or processed consistently.
Embedding units in text: quantitative values become difficult to validate and convert.
Mixing dataset and file metadata: properties of the entire dataset are confused with properties of individual files.
Recording a workflow as one narrative paragraph: inputs, outputs, parameters and dependencies cannot be interpreted automatically.
Omitting profile versions: later users cannot determine which rules governed the metadata record.
Breaking cross-domain discovery: the domain record does not retain mappings to title, creators, identifiers, rights and general relationships.
How to develop a domain-specific metadata profile
1. Analyse the workflow
Identify scientific objects, data flows, methods, decisions, parameters and outputs.
2. Reuse standards
Select existing schemas, vocabularies, identifiers and community reporting requirements.
3. Define the profile
Specify elements, obligation levels, cardinalities, value constraints, mappings and examples.
4. Test and maintain
Validate the profile with real datasets, publish its version and establish a change-management process.
Recommended profile documentation
Scope statement: domains, methods and object types covered.
Conceptual model: principal entities and relationships.
Element specification: definitions, rules and examples.
Controlled vocabularies: permitted terminology and units.
Mappings: core, repository, catalogue and external standards.
Serialisations: supported JSON, JSON-LD, XML, RDF or tabular formats.
Validation resources: schemas, shapes, tests and error messages.
Examples: complete representative records and packages.
Version history: releases, changes and migration guidance.
Governance: maintainers, approval process and issue reporting.
Explanatory status
This page provides explanatory and implementation guidance for selecting, adapting and developing domain-specific metadata profiles.
It does not replace the official specifications, schemas, controlled vocabularies, reporting guidelines or validation rules maintained by scientific communities and standards organisations.
A local profile should reuse established standards wherever possible, clearly identify all extensions and preserve mappings to general repository and discovery metadata.
Standards, ontologies and community practices may change. Their current versions and governance status should therefore be reviewed before a profile is implemented or revised.