How to prepare data for cataloguing
Cataloguing makes a research dataset visible as a structured and searchable resource in a repository, institutional catalogue, disciplinary portal or federated discovery service.
This guideline explains how to define the catalogue object, select the appropriate destination, prepare the required description, identifiers, access information and relationships, and validate the final catalogue record.
Resource information
Resource type
Step-by-step data cataloguing guideline
Intended users
Researchers, data curators, repository staff and catalogue managers
Recommended use
Before repository submission, catalogue registration, indexing or aggregation
Related output
A complete, consistent and catalogue-ready dataset record
What does cataloguing mean?
Cataloguing is the process of representing a dataset as a structured record that can be searched, filtered, indexed, linked and displayed in a catalogue or repository.
The catalogue record normally describes the dataset and points to its landing page or access location. It does not necessarily contain all dataset files.
A catalogue record is not the dataset itself
The dataset consists of the research files and supporting documentation. The catalogue record provides a structured description, identifiers, access information and links through which users can discover and obtain the dataset.
Where can data be catalogued?
The destination determines which metadata fields, controlled values and technical requirements must be applied.
Repository
A repository record describes and provides access to a deposited dataset and its files.
Institutional catalogue
An institutional catalogue aggregates datasets and other resources produced by an organisation.
Disciplinary portal
A domain catalogue may require specialist terminology, variables, methods or disciplinary classifications.
Federated discovery service
An aggregator indexes metadata supplied by repositories or other registered data sources.
Dataset-level and service-level descriptions
Before completing the record, determine which type of resource is being catalogued.
Dataset record
Describes a specific research dataset, its creators, content, methods, files, access conditions, licence and related outputs.
Resource or service record
Describes a repository, database, software service, research platform, instrument or another reusable research resource.
This guideline focuses primarily on dataset records. Use the Resource description template for cataloguing when describing a repository, service, platform or infrastructure resource.
Cataloguing preparation workflow
Move from defining the catalogue object to validating the final record and its public landing page.
1. Define
Define the resource type, scope, version and responsible organisation.
2. Select
Select the repository, catalogue or discovery service.
3. Map
Map the available metadata to the destination fields and controlled values.
4. Connect
Add identifiers and relationships to people, organisations, projects and outputs.
5. Validate
Check completeness, syntax, consistency, links and display.
Step 1 — Define the catalogue object
A catalogue record should represent one identifiable resource. Avoid creating records for undefined groups of files or entire projects unless the catalogue explicitly supports collections.
Questions to resolve
- Is the object a dataset, collection, database, service or software resource?
- What is the scientific and technical scope of the object?
- Does the record describe one version or a continuing resource?
- Where are the data or service accessed?
- Who is responsible for the resource?
- Who will maintain and update the catalogue record?
Step 2 — Select the catalogue destination
Review the destination before preparing the final record. A local metadata file may contain more information than the catalogue can display or exchange.
| Destination characteristic | What to check |
|---|---|
| Resource scope | Which resource types and disciplines are accepted? |
| Required fields | Which metadata elements must be completed? |
| Controlled values | Which vocabularies, roles, resource types and access terms are supported? |
| Identifiers | Which identifier schemes are accepted for resources, people and organisations? |
| Relationships | Which related objects and relation types can be represented? |
| Technical interface | Can metadata be imported, exported, harvested or accessed through an API? |
| Review process | Is the record moderated, validated or approved before publication? |
Step 3 — Prepare the minimum catalogue record
A useful catalogue record must contain enough information to identify, discover, understand and access the resource.
| Metadata element | Expected content |
|---|---|
| Title | Clear and specific name of the dataset or resource |
| Resource type | Dataset, collection, database, software, service or another defined type |
| Description | Content, purpose, scope, methods and intended use |
| Creators or providers | People and organisations responsible for the resource |
| Identifier or landing page | Persistent identifier or stable URL resolving to the resource record |
| Subjects and keywords | Scientific discipline, research object, method and data type |
| Access conditions | Open, restricted, embargoed or closed access and access procedure |
| Licence or rights | Permitted reuse and applicable rights statement |
| Dates and version | Publication, update or availability dates and current version |
| Contact | Current contact point for questions or access requests |
Step 4 — Write a catalogue-oriented title and description
Catalogue users often see only the title, resource type, creators and a short description in search results. These elements must be understandable without opening the complete dataset package.
Avoid
Data from the project
Results of calculations
Supplementary materials
Research database
Prefer
DFT calculation results for structural and electronic properties of cubic silicon carbide
Genomic variant dataset for the studied clinical cohort
Time series of radionuclide concentrations in surface waters
Recommended description structure
- State what the resource contains.
- Explain its scientific purpose.
- Identify the method or source.
- Define the temporal, spatial, material or disciplinary scope.
- Identify the principal data types or outputs.
- State the access conditions and important limitations.
Step 5 — Select subjects and controlled terms
Subject information determines how a record is classified, filtered and retrieved. Use controlled vocabularies where an appropriate disciplinary or general scheme exists.
Recommended subject coverage
- scientific discipline;
- research object, material, population or location;
- method, instrument or computational approach;
- data type or processing level;
- principal measured or calculated properties;
- relevant temporal or spatial coverage.
When a controlled term is used, preserve the term, vocabulary name and vocabulary URI where the catalogue supports them.
Step 6 — Add persistent identifiers
Persistent identifiers reduce ambiguity and enable reliable linking between datasets, people, organisations, projects and publications.
Research object
DOI, Handle, accession number or another identifier assigned by an authoritative system.
People
ORCID or another supported persistent person identifier.
Organisations
ROR or another authoritative organisation identifier.
Funding and projects
Grant, award, programme or project identifiers where available.
Do not invent identifiers
Record only identifiers that have been assigned by the relevant system and verify that each identifier resolves to the correct object.
Step 7 — Record access and reuse conditions
A catalogue record must explain whether the data can be accessed and what users are permitted to do with them.
Access
- open access;
- embargoed access;
- restricted access;
- closed access;
- access-request procedure;
- embargo end date.
Reuse
- licence or rights statement;
- rights holder;
- required attribution;
- third-party limitations;
- recommended citation;
- conditions of permitted reuse.
Step 8 — Link related research objects
Relationships increase discovery and allow catalogues to represent the research context around the dataset.
| Related object | Example relationship |
|---|---|
| Publication | The dataset supplements or is documented by an article |
| Source dataset | The dataset was derived from another dataset |
| New version | The record is a version of an earlier dataset |
| Software | The dataset was generated or processed using identified software |
| Workflow | The dataset was generated by a documented computational or experimental workflow |
| Project | The dataset was produced within a research project or grant |
| Collection | The dataset is part of a larger collection |
Where supported, record the related identifier, identifier type and relation type as separate structured values.
Step 9 — Map local metadata to catalogue fields
Do not copy values mechanically from a local spreadsheet into a catalogue form. Confirm the meaning, repeatability, controlled values and expected structure of every destination field.
| Local information | Catalogue mapping decision |
|---|---|
| Dataset author | Map to creator; use contributor only for secondary roles |
| Institution | Map to affiliation, provider or publisher according to its actual role |
| Project description | Extract only information relevant to the catalogued dataset |
| Open access | Map to access status, not automatically to a licence |
| Related article DOI | Map to a related identifier with the appropriate relationship |
| Method keywords | Map to subjects or methods without duplicating inconsistent variants |
| Local file path | Do not expose internal computer or server paths as public access URLs |
Step 10 — Prepare the landing page
The catalogue record should resolve to a stable landing page that provides users with sufficient information about the dataset and explains how to obtain it.
The landing page should display
- dataset title and resource type;
- creators and organisations;
- description and subjects;
- persistent identifier;
- version and relevant dates;
- access status and access instructions;
- licence or rights statement;
- related publications and research objects;
- contact information;
- available files or a link to the files.
Step 11 — Validate the catalogue record
Validate the record after it has been transferred to the catalogue or repository, not only in the local working template.
| Validation area | Questions |
|---|---|
| Identity | Does the record describe the intended dataset or resource? |
| Completeness | Are all required and relevant fields completed? |
| Consistency | Do title, creators, version, licence and access information match the dataset package? |
| Identifiers | Do all persistent identifiers resolve to the correct objects? |
| Controlled values | Are resource types, access terms, roles and relationships encoded correctly? |
| Links | Do landing-page, download and related-resource links work? |
| Display | Is the record understandable in search results and on the landing page? |
| Export | Does exported metadata preserve identifiers, repeated fields and relationships? |
| Sensitivity | Does the public record avoid protected or confidential information? |
When is a record catalogue-ready?
A record is catalogue-ready when the resource is clearly defined, the required fields are complete, terminology and identifiers are valid, access and reuse conditions are explicit, relationships are structured and the public landing page resolves correctly.
Catalogue readiness does not necessarily mean that the dataset is openly accessible or that all disciplinary metadata requirements have been satisfied.
Common cataloguing problems
Wrong object type
A dataset is described as a project, publication, service or general organisational activity.
Generic title
The title does not identify the dataset, subject, method or scope.
Project-level description
The description explains the project but not the content of the catalogued dataset.
Broken links
The record links to temporary pages, internal paths or unavailable resources.
Unverified identifiers
DOI, ORCID, organisation or project identifiers resolve to the wrong objects.
Access–licence confusion
Open access, file availability and permitted reuse are represented as one undifferentiated value.
Free-text relationships
Related publications and datasets are mentioned only in the description without structured identifiers.
Catalogue mismatch
The published record differs from the approved metadata file or dataset package.
Recommended working method
Prepare and approve the dataset metadata in a structured local file before entering the information into a catalogue. Create a mapping between the local fields and the catalogue fields and document any transformations or omitted values.
After publication, compare the public landing page and exported metadata with the approved local record. Retain the mapping and validation result for future updates.
Important notes
- Define whether you are cataloguing a dataset, collection, service or another type of resource.
- Describe the dataset itself rather than the entire research project.
- Complete the destination catalogue’s required fields in addition to the local core metadata profile.
- Use only verified persistent identifiers.
- Preserve controlled vocabulary terms and their identifiers where supported.
- Keep access status and licence information separate.
- Record relationships using identifiers and relation types rather than only free-text statements.
- Do not expose internal server paths, credentials or confidential technical information.
- Validate the public landing page and exported metadata after publication.
- Update the catalogue record when the dataset version, access, licence, creators or location changes.