Skip navigation

How to prepare a FAIR data package

A FAIR data package brings together research data, metadata, documentation and supporting files in a coherent structure that can be understood, checked, preserved, published and reused.

This guideline explains how to define the package scope, organise the files, prepare the README and manifest, record metadata and provenance, document access and reuse conditions, and validate the completed package before repository deposit.

ChatGPT Image 24 лип. 2026 р., 22_34_09 (2)

Resource information

Resource type

Step-by-step FAIR data package preparation guideline

Intended users

Researchers, dataset authors, data curators and data stewards

Recommended use

Before curation, repository deposit, dataset publication or preservation

Related output

An organised, documented and internally consistent research data package

What is a FAIR data package?

A FAIR data package is an organised collection of digital objects prepared so that the dataset can be identified, discovered, accessed under defined conditions, interpreted by people and machines, and reused beyond the original research project.

The package normally combines the research files with structured metadata, human-readable documentation, file-level descriptions, technical information, provenance and access and reuse statements.

FAIR does not mean that every file must be openly accessible

Restricted or embargoed files can form part of a FAIR-oriented package when the metadata remain discoverable and the conditions and procedure for obtaining access are clearly documented.

The package is more than a folder of files

Files

Research data, scripts, configuration files, models, visualisations and supporting outputs.

Documentation

README, methods, data dictionaries, protocols and instructions for understanding and using the files.

Structured records

Dataset metadata, manifest, provenance, workflow and software-environment information.

Conditions

Access status, licence, rights holder, restrictions, citation and contact information.

Minimum and extended package

Package composition depends on the data type, research method and level of reproducibility required.

Minimum package

  • principal research data files;
  • README.md or equivalent documentation;
  • file manifest or inventory;
  • dataset metadata record;
  • licence or rights statement;
  • access and contact information.

Extended package

  • input and intermediate data;
  • scripts and configuration files;
  • software-environment record;
  • workflow and provenance metadata;
  • quality-control or validation evidence;
  • data dictionaries and disciplinary metadata.

FAIR data package preparation workflow

Build the package in a controlled sequence from scope definition to final validation.

1. Define

Define the dataset, package scope, version and responsible persons.

2. Organise

Organise folders and files using clear names and roles.

3. Document

Prepare the README, manifest, metadata and technical records.

4. Connect

Record provenance and relationships between files and research outputs.

5. Validate

Check completeness, integrity, consistency and publication readiness.

Step 1 — Define the package scope

Define which dataset or research result the package represents before copying files into the final directory. A package should correspond to one clearly identifiable digital object or a deliberately defined collection.

Questions to resolve

  • What scientific result or dataset will be published?
  • Which files belong to this package?
  • Which version is being prepared?
  • Are raw, processed and derived data included?
  • Which files are essential for interpretation or reproduction?
  • Which files will be open, restricted or excluded?
  • Who approves the final package?

Step 2 — Inventory the available files

Create an initial inventory before reorganising the files. The inventory helps identify duplicates, obsolete versions, undocumented formats and missing components.

Inventory field What to record
Current path Location of the file before package preparation
File name Current name and proposed final name
File role Input, raw, processed, derived, output, documentation or software
Format File format, extension and required software
Version File or model version where relevant
Access status Open, restricted, embargoed or excluded
Action Include, rename, convert, document, replace or remove

Step 3 — Design the folder structure

Use a folder structure that reflects the scientific and technical roles of the files. Avoid structures that reproduce personal computer paths or temporary project organisation.

Example package structure

fair-data-package/
│
├── README.md
├── manifest.csv
├── metadata.json
│
├── data/
│   ├── raw/
│   ├── processed/
│   └── derived/
│
├── documentation/
│   ├── methods.md
│   └── data_dictionary.csv
│
├── scripts/
│   └── process_data.py
│
├── configuration/
│   └── parameters.yml
│
├── environment/
│   └── software_environment.yml
│
├── provenance/
│   └── provenance.json
│
└── quality/
    └── validation_report.pdf

Clear roles

Separate data, documentation, software, configuration, provenance and quality-control files.

Limited depth

Avoid unnecessarily deep folder hierarchies that make paths hard to understand and maintain.

Stable paths

Use relative paths and keep them consistent across the README, manifest and structured records.

Repository awareness

Check whether the target repository preserves folders or presents all uploaded files as a single file list.

Step 4 — Apply consistent file naming

File names should remain understandable outside the original working environment and should distinguish versions, samples, dates or processing stages where required.

Avoid

final.csv

final_new.csv

results2_fixed.csv

data_from_PC_old.zip

Prefer

sample01_raw_2026-07-15.csv

sample01_cleaned_v1.1.csv

sic_dft_total_energy_v1.csv

survey_responses_anonymised_v2.csv

  • Use short but meaningful names.
  • Use one naming pattern throughout the package.
  • Avoid spaces and unstable punctuation where possible.
  • Use ISO-style dates such as YYYY-MM-DD.
  • Represent versions consistently.
  • Do not use words such as new, latest or final-final.
  • Do not rename files without updating the manifest and documentation.

Step 5 — Prepare the README

Place the README at the package root. It should provide the main human-readable explanation of the dataset and guide users through the package.

README section Expected content
Dataset overview Title, purpose, scope and scientific context
Creators and contact Responsible people, institutions and contact information
Package contents Folder structure, file groups and principal files
Methods How the data were collected, calculated or processed
Technical requirements Formats, software, instruments and dependencies
Quality and limitations Validation, uncertainty, exclusions and known limitations
Access and reuse Access conditions, licence, citation and restrictions
Related outputs Publications, software, projects and related datasets

Step 6 — Create the manifest

The manifest is a structured inventory of package files. It enables file-level checking and supports automated processing, integrity verification and curation.

Recommended manifest fields

  • file_path
  • file_name
  • file_role
  • description
  • format or media_type
  • size_bytes
  • checksum
  • checksum_algorithm
  • access_status
  • related_step or another relationship field

Step 7 — Prepare the dataset metadata

Prepare one structured metadata record describing the package as a whole. The record should correspond to the README, the manifest and the planned repository entry.

Identification

Title, resource type, version, dates and identifier

Responsibility

Creators, contributors, contacts and organisations

Description

Content, methods, subject, coverage and technical characteristics

Conditions and links

Access, licence, funding, publications, software and related datasets

Step 8 — Document software and environment

Include the information required to open, process or reproduce the files. The level of detail should correspond to the technical complexity of the dataset.

Possible environment records

  • requirements.txt for Python dependencies;
  • environment.yml for a Conda environment;
  • container image or definition reference;
  • software and plugin version table;
  • operating system and execution-platform information;
  • instrument model and acquisition software version;
  • configuration and parameter files used for processing.

Step 9 — Record provenance and workflow

Document how the principal package objects were produced, transformed or derived. This is especially important for computational, simulation, imaging and multi-stage processing workflows.

Provenance element What to document
Entity Input, intermediate, output, dataset, model or software object
Activity Experiment, processing step, calculation, conversion or validation
Agent Person, organisation, instrument or software responsible
Used Input objects used by an activity
Generated Objects produced by an activity
Derived from Source object from which another object was derived
Execution record Date, software version, parameters, status and validation result

Step 10 — Define access and reuse conditions

Review the package at both dataset and file level. Different files within one package may require different access conditions.

Access information

  • open, restricted, embargoed or closed;
  • embargo end date;
  • reason for restriction;
  • procedure for requesting access;
  • responsible contact.

Reuse information

  • rights holder;
  • licence or rights statement;
  • required attribution;
  • citation recommendation;
  • third-party restrictions.

Step 11 — Validate the completed package

Validate the final package version rather than the earlier working directory.

Validation area Questions
Completeness Are all declared and required files present?
File integrity Can principal files be opened and do checksums correspond?
Paths Do README, manifest, metadata and workflow paths match the package?
Consistency Do title, creators, version, licence and access status agree across records?
Syntax Are CSV, JSON, XML, YAML and other structured files valid?
Documentation Can a researcher outside the original team understand the package?
Security Have credentials, personal data and unintended confidential files been removed?
Repository readiness Does the package satisfy the target repository requirements?

Freeze the package before generating final checksums

Assign a package version and generate final checksums only after files have been renamed, corrected and approved. Any subsequent file change requires checksum regeneration and another consistency review.

Common package preparation problems

Unclear scope

The package combines unrelated project files without defining one dataset or collection.

Archive-only deposit

All files are hidden in one archive although individual files could be described and accessed more effectively.

Missing documentation

Files are present, but their roles, formats, variables and methods are not explained.

Path mismatch

Paths in the manifest, README or workflow record do not match the actual package.

Uncontrolled versions

Several files are labelled final, new or corrected without a defined versioning scheme.

Missing provenance

Outputs are included without documenting the inputs, software or processing steps that produced them.

Licence inconsistency

The repository, README and metadata record state different reuse conditions.

Sensitive content

Credentials, personal data or confidential technical information remain in files or metadata.

Recommended working method

Create the FAIR data package in a separate controlled directory rather than reorganising the only existing copy of the research files. Preserve the original source materials until the package has been checked and approved.

Use the README, manifest and metadata templates together. Update all three whenever files are added, removed, renamed or replaced.

Important notes

  • Prepare the package for one clearly defined dataset or collection.
  • Retain the original source files until the curated package has been approved.
  • Use relative paths consistently across all package records.
  • Do not generate final checksums before the files are frozen.
  • Do not include passwords, tokens, private keys or confidential configuration values.
  • Review licence and access conditions separately.
  • Explain proprietary or specialist formats and provide open exports where scientifically appropriate.
  • Include only the files needed to understand, verify or reuse the dataset.
  • Re-run the pre-check after every change to the final package.
  • Apply repository-specific requirements in addition to this general guideline.