Skip navigation

Controlled vocabularies

Standardised terms, codes and concept identifiers used to create consistent, interoperable and machine-readable research metadata.

This section helps researchers, data stewards and repository staff select controlled values for resource types, contributor roles, relationships, access conditions, licences, scientific subjects and units of measure.

Controlled vocabularies reduce ambiguity and make it possible to compare, validate, exchange and search metadata across repositories, catalogues, research infrastructures and scientific domains.

Resource information

Resource type: Controlled vocabulary guidance

Primary audience: Researchers, data stewards, metadata specialists, repository managers, software developers and research infrastructure providers

Coverage: Resource types, roles, relationships, lifecycle states, access rights, licences, subject vocabularies, ontologies and units of measure

Primary uses: Metadata creation, repository deposit, profile implementation, validation, catalogue exchange and semantic interoperability

Representation models: Code lists, taxonomies, thesauri, SKOS concept schemes, ontologies and unit vocabularies

Last reviewed: July 2026

Browse controlled vocabulary resources

Select the resource that corresponds to the type of metadata value you need to record: general descriptive values, relationships and lifecycle states, access and licensing conditions, or specialised scientific terminology.

Semantic resources for FAIR description of computational workflows

Practical use of domain ontologies, controlled vocabularies, provenance models and standardised units, illustrated by the DFT → MLIP → MD workflow in materials science.

What is a controlled vocabulary?

A controlled vocabulary is a maintained set of authorised terms or codes used to represent concepts consistently within metadata, databases and information systems.

Each term should have a clear meaning and, where possible, a stable identifier. More advanced vocabularies may also define multilingual labels, hierarchical relationships, synonyms and mappings to terms in other vocabularies.

Free text: allows users to enter any wording but may produce spelling variants, synonyms and ambiguous values.

Controlled label: limits users to an approved list of human-readable values.

Controlled identifier: records a stable URI or code representing the concept independently of its displayed label.

Types of semantic resources

Code list

A finite set of permitted codes or values without a complex conceptual hierarchy.

Taxonomy or thesaurus

Organises concepts through broader, narrower, related and alternative terms.

Ontology

Formally represents classes, properties, relationships and logical constraints within a domain.

Unit vocabulary

Provides standard codes and identifiers for quantities, dimensions and units of measure.

Selecting an appropriate vocabulary

1. Identify the field

Determine the concept that must be controlled: resource type, role, relationship, access status, subject or unit.

2. Check the profile

Review the metadata schema, repository or catalogue requirements for prescribed value vocabularies.

3. Evaluate the vocabulary

Check governance, scope, identifiers, versioning, documentation, machine-readable access and community adoption.

4. Record the value

Store the concept identifier, preferred label, vocabulary name and version where the receiving system supports them.

Vocabulary selection criteria

Relevance: the vocabulary represents the intended concept accurately and at the required level of detail.

Authority: the maintainer and governance process are clearly identified.

Community use: the vocabulary is recognised by the relevant repository, infrastructure or scientific community.

Persistent identifiers: concepts have stable URIs or codes that do not depend only on their labels.

Machine-readable access: the vocabulary is available in a structured format or through an API.

Versioning: changes, releases and deprecated concepts are documented.

Multilingual support: labels in several languages are provided where needed.

Mapping: relationships to equivalent or related vocabularies are available or can be documented.

Licence: reuse conditions permit implementation in the intended metadata system.

How to record a controlled value

A metadata record should preserve both a machine-readable identifier and a human-readable label where the format supports structured values.

{
  "value": "open access",
  "conceptUri": "http://purl.org/coar/access_right/c_abf2",
  "vocabulary": "COAR Access Rights",
  "vocabularyVersion": "1.1",
  "language": "en"
}

Labels and identifiers

Preferred label: the authorised name of the concept in a particular language.

Alternative label: a synonym, abbreviation or variant used to support search and user interfaces.

Definition: an explanation that distinguishes the concept from related terms.

Concept URI: a stable identifier used in machine-readable metadata.

Scheme URI: an identifier for the vocabulary or concept scheme containing the term.

Language tag: identifies the language of the displayed label without changing the concept itself.

Mapping between vocabularies

Different repositories and infrastructures may use different vocabularies for the same metadata concept. A mapping should state the strength and direction of the relationship rather than assuming that similarly named terms are identical.

Exact match: the two concepts can normally be used interchangeably.

Close match: the concepts are sufficiently similar for many applications but are not fully equivalent.

Broad match: the target concept is more general.

Narrow match: the target concept is more specific.

Related match: the concepts are associated but do not represent the same meaning.

Common implementation errors

Storing only a label: changes in language or wording make it difficult to recognise the underlying concept.

Mixing vocabularies without recording the source: identical codes may have different meanings.

Using access status as a licence: availability and legal permission for reuse are different metadata properties.

Creating local terms too early: established community vocabularies and mappings are overlooked.

Embedding units in free text: numerical values cannot be validated or converted reliably.

Ignoring vocabulary versions: later users cannot determine which definitions and values were applied.

Using deprecated concepts: metadata is created with values that are no longer recommended by the maintainer.

Assuming label similarity means equivalence: mappings introduce semantic errors.

Explanatory status

The materials in this section provide explanatory and implementation guidance for selecting and applying controlled vocabularies in research metadata.

They do not replace the official vocabularies, ontology specifications, licence texts, repository requirements or validation rules maintained by the relevant organisations.

Vocabulary terms, mappings and versions may change over time. The current authoritative source should therefore be checked before implementing or revising a production metadata profile.

Local values should be introduced only when no appropriate maintained vocabulary exists. Such extensions should be documented, assigned stable identifiers and mapped to established concepts where possible.