MATSCI-NASU
MATSCI-NASU is a preparatory project developing the governance, operational readiness, research resources and pilot use cases required for a future NASU thematic EOSC candidate Node in materials science.
The project connects computational materials-science workflows, FAIR data curation, repository publication, digital research services and an institutional pathway towards participation in the EOSC Federation.
Its central demonstrator follows a multi-stage computational workflow: Density Functional Theory, Machine Learning Interatomic Potentials and Molecular Dynamics. The demonstrator is used to test how research artefacts can be transformed into documented, connected and reusable digital research objects.
Project information
Full project title: Designing the Governance, Core Readiness, and Pilot Use Case for the NASU Materials Science Thematic Node
Acronym: MATSCI-NASU
Project type: EOSC Gravity Preparatory Grant
Funding context: EOSC Gravity, Horizon Europe Grant Agreement No. 101188045
Lead organisation: Bogolyubov Institute for Theoretical Physics of the National Academy of Sciences of Ukraine
Project partners: BITP, IPMS and Kyiv Academic University
Scientific domain: Computational and experimental materials science
Implementation period: May–October 2026
Current status: Preparatory project and pilot demonstrator
Last reviewed: August 2026
Project status
MATSCI-NASU is not yet an enrolled or production EOSC Node.
It is a preparatory project that develops and tests the organisational, operational, metadata and scientific components required for a future thematic Node application.
The project uses real materials-science artefacts and workflows to validate its proposed approach, but it does not claim that the complete EOSC Federation onboarding process has already been completed.
Its outputs should therefore be understood as a readiness package, implementation baseline and validated demonstrator for subsequent Node development.
Purpose of MATSCI-NASU
MATSCI-NASU addresses the gap between computational environments, where materials-science data are generated, and repository environments, where research outputs must be documented, preserved, discovered and reused.
Develop governance
Define institutional roles, responsibilities, decision-making and cooperation for a future thematic Node.
Prepare core readiness
Establish an initial pathway for identity, catalogues, support, monitoring and service management.
Validate scientific use cases
Test the proposed approach using real computational materials-science artefacts and workflows.
Build community capacity
Produce reusable training, metadata, curation and onboarding resources for researchers and institutions.
Why materials science?
Materials science combines experimental measurements, atomistic simulations, machine learning, large datasets, software and computational infrastructure.
Scientific results often depend on a chain of interconnected objects: material structures, simulation parameters, labelled configurations, trained models, molecular-dynamics trajectories, analysis scripts and derived properties.
Data-intensive research
Simulations and experiments generate large and technically complex collections of research artefacts.
Connected workflows
Scientific conclusions depend on traceable links between structures, calculations, models and results.
Computing dependence
Research requires cloud, high-performance computing, storage and specialised software environments.
Reuse potential
Well-documented data and models can support new simulations, validation, comparison and AI applications.
The problem addressed
Computational outputs are commonly stored in local working directories without consistent metadata, licences, provenance or repository publication.
The relationship between the computing environment and the final published research object is frequently undocumented.
Large files, trained models, software environments and derived results require different storage and publication strategies.
Researchers need coordinated support covering computing access, metadata, FAIR packaging, repository deposit and future EOSC integration.
Ukrainian materials-science organisations also need a realistic institutional pathway towards participation in the EOSC Federation.
Project partners
BITP
Leads project coordination, governance, operational readiness, computing access and the pathway to Federation-compatible services.
IPMS
Provides materials-science expertise and owns the pilot computational workflow, scientific artefacts and domain validation.
Kyiv Academic University
Provides FAIR data competence, metadata and curation support, training design and reusable learning resources.
Adopter community
Research institutes and teams review templates, contribute use cases and provide feedback on future services.
Role of the NASU FAIR Data Competence Center
The NASU FAIR Data Competence Center provides the internal FAIR curation and research-data-support layer of the MATSCI-NASU pilot.
Research-object inventory: identifying the data, software, models, workflows and supporting documentation produced by the pilot.
Metadata profiling: defining general and workflow-specific metadata required to describe computational materials-science objects.
FAIR packaging: organising files, README documents, manifests, provenance and repository metadata.
Curatorial review: checking structure files, metadata sources, package completeness, licences and reuse conditions.
Training: translating the pilot into reusable examples and learning materials for researchers, data stewards and curators.
Repository preparation: preparing stable research objects for future publication through DataverseUA or another suitable repository.
Project organisation
Steering Group
One representative from each partner coordinates milestones, priorities, risks and institutional decisions.
Technical Working Group
BITP and IPMS coordinate infrastructure assumptions, metadata, packaging and provenance requirements.
Training Working Group
KAU and project partners develop the training matrix, guidance, templates and reusable examples.
Use-case team
Researchers and curators jointly prepare, review and improve the pilot research object.
From pilot use case to Node readiness
1. Scientific use case
Start with a real DFT → MLIP → MD workflow and its scientific artefacts.
2. FAIR packaging
Add metadata, documentation, provenance, licences, manifests and repository-facing structure.
3. Service readiness
Define access, catalogue, helpdesk, monitoring and operational support processes.
4. Federation pathway
Convert tested gaps and requirements into a prioritised Node onboarding roadmap.
Core readiness areas
Governance
Institutional mandate, partner responsibilities, decision-making and sustainability.
Access and identity
Institutional identities, community roles, authorisation and a path towards federated access.
Services and operations
Catalogue, user support, monitoring, incident handling and service ownership.
Research resources
Curated datasets, software, workflows, models and reusable scientific packages.
Initial federating capabilities
MATSCI-NASU is designing a minimal set of capabilities that can be tested through the pilot before wider institutional deployment.
| Capability | Initial MATSCI-NASU approach |
|---|---|
| Federated AAI | Community AAI proxy, institutional identity pathway and future connection to EOSC-compatible identity services |
| Resource and service catalogue | Common descriptions of datasets, workflows, computing resources, training and support services |
| Helpdesk | One entry point for users with specialised support retained by the responsible partner |
| Service monitoring | Basic availability and performance monitoring for pilot services and endpoints |
| Service management | Lightweight FitSM-oriented procedures for ownership, incidents, access requests, changes and reporting |
Authentication and access pathway
Institutional identity: researchers should use recognised institutional identities rather than separate local accounts for every service.
Community proxy: MATSCI-NASU should manage project affiliations, groups, roles and entitlements through an AAI proxy.
Federation connection: pilot services may be connected to EOSC-compatible identity services, including EGI Check-in where appropriate.
Authorisation: access to computing, storage and restricted research objects remains governed by provider and project policies.
Initial service portfolio
Computational access
Access to cloud or computing environments supporting materials modelling and workflow execution.
FAIR curation
Metadata profiling, package review, provenance documentation and repository preparation.
Repository publication
Publication of stable research objects through DataverseUA with DOI or other persistent identifiers.
Training and support
Practical guidance, clinics, templates and support for researchers and future adopters.
Research resources planned for MATSCI-NASU
| Resource category | Examples |
|---|---|
| Computational datasets | Structures, input parameters, labelled configurations, raw and processed outputs and validation data |
| Scientific workflows | DFT, MLIP and molecular-dynamics workflows and associated post-processing |
| Research software | Scripts, notebooks, configurations, models and executable environments |
| Reproducibility resources | Model cards, evaluation summaries, manifests, provenance and workflow documentation |
| Metadata resources | Profiles, controlled terms, templates, README files and DMP guidance |
| Training resources | Courses, practical examples, checklists and onboarding materials |
Anchor pilot use case
Material system: silicon carbide
Research-object type: multi-stage computational materials-science workflow
Workflow: DFT-labelled dataset → MLIP / NequIP model → Molecular Dynamics → thermomechanical properties
DFT software: Quantum ESPRESSO
MLIP software: NequIP and PyTorch
Atomistic processing and MD layer: ASE
FAIR objective: represent files, metadata, software, models and results as one connected and reusable research object
SiC computational workflow
1. DFT calculations
Generate labelled structures containing energies, forces and stress using Quantum ESPRESSO.
2. MLIP development
Train and evaluate a NequIP machine-learning interatomic potential using the labelled configurations.
3. Molecular Dynamics
Use the trained potential in atomistic simulations and calculate material properties.
4. FAIR publication
Connect the dataset, model, workflow, results and documentation in a repository-ready package.
Research-object structure
| Workflow stage | Principal digital objects |
|---|---|
| Material definition | Structures, chemical composition, phases or polytypes and source information |
| DFT | Input files, pseudopotentials, calculation parameters, outputs and labelled structures |
| MLIP | Training configuration, dataset splits, model checkpoint, deployable model, logs and metrics |
| Molecular Dynamics | MD configurations, input structures, trajectories, scripts and calculated properties |
| Documentation | README, file manifest, workflow description, provenance, licences and citation information |
| Curation | Inventory, file mapping, metadata sources, FAIR assessment and curator notes |
Provenance across the workflow
Initial structure → DFT calculation: records which material structure and calculation settings generated a labelled configuration.
DFT dataset → trained model: records which configurations and train-validation-test split were used to develop the MLIP.
Model → MD run: records which model version, configuration and environment produced a trajectory.
MD run → property result: records which trajectory, script and parameters generated a reported material property.
Research object → publication: links the package with the associated dataset record, DOI, article, authors, project and licence.
FAIR curation workflow
1. Inventory
Identify files, formats, workflow stages, roles, size and publication status.
2. Mapping
Connect received files with the expected research objects and metadata fields.
3. Technical checks
Inspect atomistic structures, JSON, YAML and CSV files and verify machine readability.
4. FAIR assessment
Record strengths, gaps, required author responses and repository preparation actions.
FAIR package structure
dataset/
├── README.md
├── dataset.json
├── material.json
├── provenance.json
├── workflow.md
├── manifest.csv
├── LICENSE.txt
├── CITATION.cff
├── checksums.csv
│
├── dft/
│ ├── dft.json
│ ├── config/
│ ├── structures/
│ ├── pseudopotentials/
│ ├── convergence/
│ ├── input/
│ ├── output/
│ └── scripts/
│
├── mlip/
│ ├── mlip.json
│ ├── config/
│ ├── dataset/
│ ├── model/
│ ├── metrics/
│ ├── logs/
│ └── scripts/
│
├── md/
│ ├── md.json
│ ├── config/
│ ├── inputs/
│ ├── trajectories/
│ ├── results/
│ └── scripts/
│
└── curation/
├── README_curated.md
├── file_inventory.csv
├── file_mapping.csv
├── metadata_sources.csv
├── fair_assessment.csv
└── request_to_authors.csv
Current pilot evidence
The received package has a clear DFT → MLIP → MD directory structure.
README and MANIFEST documents are available.
The principal labelled DFT dataset is readable programmatically with ASE.
The inspected dataset contains structures with energy, force and stress labels.
NequIP training configuration, a model, evaluation metrics and an MD configuration are represented in the package.
Elastic-constant and thermal-expansion results are provided in machine-readable formats.
Artefacts are distinguished as real, sample, pointer or synthesised, supporting transparent interpretation of the pilot.
Current FAIR status
Current designation: FAIR v0.1 candidate
The package is sufficiently structured for an initial FAIR publication review.
This status does not mean that the complete computational workflow has been independently reproduced.
It confirms that the package has undergone an initial inventory, metadata, machine-readability and documentation review.
A publication-ready version requires additional author confirmation, complete provenance, licensing and final repository metadata.
SiC metadata profile
The pilot metadata profile describes the complete workflow rather than only the final files.
| Metadata layer | Examples |
|---|---|
| Research object | Title, creators, organisation, project, description, rights and related publication |
| Material | Chemical composition, structure, phase, polytype and source |
| DFT | Software, pseudopotentials, functional, cut-offs, k-points, convergence and labelled quantities |
| MLIP | Training configuration, splits, hyperparameters, model version and evaluation metrics |
| MD | Model used, ensembles, temperature, pressure, timestep, trajectory and property calculation |
| Provenance | Typed links connecting structures, datasets, models, runs and property results |
| Publication | Repository, DOI, licences, file locations, checksums and access conditions |
Remaining pilot gaps
Material definition: confirm the final scope and SiC structures or polytypes represented by the publication package.
Dataset splits: provide or confirm the complete training, validation and test split.
Model identification: identify the final deployable model and its authoritative version.
Evaluation: confirm the final test metrics, limitations and domain of applicability.
Environment: document complete software versions, dependencies and execution environment.
Provenance: strengthen calculation-level and cross-stage provenance records.
Licensing: establish compatible licences for data, scripts, models and pseudopotentials.
Storage and publication: determine which artefacts are deposited in DataverseUA and which require external large-file storage.
Next curation outputs
FAIR v0.2 review
Update the assessment after receiving the outstanding information from the authors.
Model card
Document the trained potential, data, evaluation, limitations and intended use.
Provenance record
Connect the DFT dataset, trained model, MD runs and derived results.
Repository package
Add licences, split manifest, checksums and final DataverseUA metadata.
Repository publication pathway
1. Freeze the release
Identify the stable data, model, software and result versions covered by the publication.
2. Complete the package
Add metadata, manifests, provenance, checksums, licences and documentation.
3. Deposit research objects
Publish suitable components through DataverseUA or another trusted repository.
4. Connect identifiers
Link the dataset, software, model, workflow, article, project and responsible organisations.
Training and capacity building
MATSCI-NASU uses the pilot as a training environment in which researchers and data-support specialists learn by producing real research objects and operational artefacts.
FAIR packaging
README files, metadata profiles, manifests, licences and packaging checklists.
Model documentation
Model cards, training-data descriptions, evaluation summaries and limitations.
MD reproducibility
Input configurations, representative trajectories, property tables and analysis scripts.
Node onboarding
Access, licensing, discovery, support and readiness-gap assessment.
Community engagement
Open webinars introduce the project approach and collect feedback from researchers and support professionals.
Clinics allow research teams to test metadata, packaging and access assumptions on concrete artefacts.
An adopter cohort can review templates and propose additional computational or experimental use cases.
Reusable materials are shared with NASU institutes and future thematic infrastructure initiatives.
Project results are consolidated into a showcase contribution and a Node onboarding roadmap.
Principal project outputs
| Output | Purpose |
|---|---|
| Project Plan | Defines scope, activities, outputs, indicators and risks |
| Governance framework | Defines partner roles, decision-making and future Node organisation |
| Project Charter | Provides the foundational document for a future Node application |
| Pilot use-case report | Documents the compute-to-repository journey and identified readiness gaps |
| Metadata profile | Defines the minimum description of a DFT → MLIP → MD research object |
| Packaging templates | Standardise README, manifests, model cards, provenance and reproducibility resources |
| EOSC Academy support matrix | Maps training topics, audiences, formats and reusable materials |
| Node onboarding roadmap | Prioritises future organisational, operational and technical actions |
2026 implementation plan
| Period | Planned stage |
|---|---|
| May 2026 | Project baseline, kick-off, research-object inventory and metadata profile draft |
| June 2026 | Governance and operational baseline, roles and support assumptions |
| July 2026 | Metadata and packaging validation, pilot clinic and first curation review |
| August 2026 | Project Charter, pilot plan and consolidation of core project documents |
| September 2026 | Pilot onboarding exercise and compute-to-repository pathway test |
| October 2026 | Use-case report, training matrix, showcase contribution and onboarding roadmap |
Expected benefits
For researchers
Better documentation, repository publication, reproducibility and access to reusable materials-science resources.
For institutions
A reusable governance, metadata and operational baseline for materials-science services.
For EOSC
A pathway for Ukrainian materials-science resources and communities to participate in federated services.
For future projects
Templates, training materials and implementation lessons reusable beyond the initial pilot.
Scaling beyond the pilot
The SiC demonstrator provides a reusable pattern for additional computational materials-science workflows.
The same packaging logic can be adapted to experimental datasets, instruments and characterisation workflows.
Additional NASU institutes may contribute datasets, software, computing, training or domain services.
Shared metadata, support and service-description templates reduce the cost of onboarding new use cases.
The project’s roadmap can also inform other thematic Node initiatives within NASU.
What MATSCI-NASU is not
Not yet an enrolled EOSC Node: the project prepares and validates a future candidate pathway.
Not only a repository: it connects scientific workflows, computing, FAIR support, publication and user services.
Not only an IT platform: governance, metadata, curation, training and community coordination are essential components.
Not a fully reproduced SiC workflow: the current pilot validates package structure and readiness rather than independently repeating all calculations.
Not a final metadata standard: the current profile is a working baseline that requires domain review and further iteration.
Not limited to SiC: silicon carbide is the anchor demonstrator for a broader materials-science service model.
Readiness checklist
Institutional mandate and coordinating organisation
Agreed governance and partner responsibilities
Defined scientific and user communities
Production-relevant research resources and services
Complete service and resource descriptions
AAI and authorisation pathway
Catalogue and metadata exchange approach
User-support and helpdesk procedures
Monitoring and service-management baseline
Security, legal and access policies
Validated FAIR research-object package
Multi-Node or Federation-oriented scientific use case
Training and adopter-community programme
Sustainable staffing and resource commitments
Prioritised Node onboarding roadmap
How to use this page
For researchers
Review the pilot workflow and FAIR package as an example for documenting computational research.
For data stewards
Use the metadata, curation and gap sections to design support for similar scientific workflows.
For institutions
Use the governance and readiness sections when considering participation in a thematic Node.
For EOSC projects
Use the pilot as evidence of a staged transition from scientific workflow to Federation-oriented service readiness.
Explanatory status
This page presents the objectives, implementation model and current pilot evidence of MATSCI-NASU.
MATSCI-NASU is a preparatory project and should not be interpreted as an already enrolled or production EOSC Node.
Planned project outputs are distinguished from results that have already been produced and reviewed.
The SiC package currently has the status of a FAIR v0.1 candidate. Its assessment confirms initial package and metadata readiness but does not constitute independent reproduction of the complete computational workflow.
Project status, deliverables and technical arrangements will be updated as the Project Charter, pilot report and onboarding roadmap are completed.