OSCARS demonstrator: making experimental data findable across facilities through shared semantics

We are happy to share a demonstrator from the OSCARS project Findable Big Data from Various Material Characterisation Techniques.

The demonstrator focuses on a fairly practical problem: **how can we find and process experimental data from different facilities and platforms in a consistent way, when those facilities describe and store their data differently?**

Here, we use X-ray absorption spectroscopy (XAS) data from the synchrotron facilities at ESRF and BESSY II as an example. The two facilities have their own data infrastructures and discovery mechanisms, but the demonstrator connects them through a common semantic layer and a shared processing workflow.

A key part of the work in this OSCARS project has been connecting the Photon and Neutron Experimental Techniques (PaNET) ontology with the NeXus Ontology, as well as with facility-specific ontologies such as the ESRF Experimental Techniques (ESRFET) ontology. This connection provides a common way of describing experimental techniques across facilities. The resulting ontology integration has been incorporated into NOMAD, the research data management platform developed by the FAIRmat project.

As part of the OSCARS project, NOMAD’s pynxtools plugin was extended with an ontology service that injects technique terms into the metadata of parsed datasets, making them searchable by common ontology terms. For ESRF specifically, the nomad-semantic-web-service plugin adds a PaNET->ESRFET mapping and ICAT+ catalogue discovery that is used in this demonstrator to find datasets there in the first place.

The XAS demonstrator puts this into practice end to end:

  • It discovers XAS datasets at ESRF and BESSY II using the respective facility-specific discovery mechanisms.
  • The datasets are brought into a common NeXus-based representation (via pynxtools-xas for ESRF’s raw files; BESSY II’s data already arrives this way).
  • The standardized data are passed to a common EXAFS processing workflow.
  • The processed results are written back to NOMAD as new entries.

This means that the raw data, standardized data, and processed results can be kept together and remain searchable in NOMAD. The screenshot below shows an example of a processed result — the EXAFS Fourier transform of a XAS spectrum— directly in NOMAD after the workflow has completed.

Where to find it: NOMAD upload

This demonstrator shows that combining shared ontologies, standardized data representations, and the NOMAD platform makes it possible to connect and process experimental data across facilities in practice.

We would be interested in hearing from others working on similar cross-facility or cross-platform data discovery and processing challenges.

Further information can be found here:

On behalf of the team at FAIRmat, ESRF, and Helmholtz-Zentrum Berlin (HZB),
Lukas Pielsticker

1 Like

It was great working with FAIRmat and ESRF and demonstrating (re)use of standardised data (NeXus :heart:) and (re)use of processing workflow (EWOKS :heart:)… We, BESSY II, are looking forward to having more such use case and bringing best practices into action! :handshake:

1 Like

Thanks @lukas.pielsticker for this post. It is a good example of how data can be found and shared between photon sources. Do you have an idea of what the scientific application could be? I presume it applies to most use cases which need XAS data. I was wondering if you had a specific application in mind while doing this or was it more a demonstration of how data can be found and combined across facilities? Just curious!

I think first and foremost it is a demonstrator for now.

For future endeavours, it may however be useful for any task that wants to compare or aggregrate XAS data from different sources. Examples could be round robins to see how different data from the various facilities/beam line is or even machine learning approaches were you actually want to have data with a larger variety.

1 Like