Which enrichments are included?

COMET’s first pilot project matched preprints from the arXiv repository (registered with DataCite DOIs) to their corresponding published articles (registered with Crossref DOIs) in various journals. The resulting dataset—over 850,000 new preprint-to-article connections—fills a gap in the scholarly record, enabling a better understanding of the research timeline and its impact.

Following the release of ROR’s updated matching strategy, COMET and DataCite embarked on a project to match author affiliations to ROR IDs in DataCite. The enriched dataset includes over 20 million matches across 5,812,774 unique DOIs in the affiliation field of the creator property (as of April 2026).

How does the enrichment service work?

The diagram below shows DataCite’s current enrichment pipeline. DataCite ingests an enriched data file from COMET that contains a record of each enrichment. It includes provenance metadata connecting each enrichment to both its contributors and the resources that the enrichment was derived from, such as project documentation and related datasets.

Enrichments are stored in a separate datastore so that their application is additive, allowing granular comparisons to the metadata submitted by DOI record owners.

Four-lane workflow diagram showing how a COMET enrichment record flows through the DataCite enrichments pipeline into the metadata store, then out via the DataCite REST API.

For a walkthrough of the pipeline, view the 22 April COMET Community Meeting recording. To find out more about DataCite’s metadata enrichments service and how to use the API endpoint, visit the support documentation on metadata enrichments.

How can I give feedback?

This is a new service, so what the community flags now will directly inform its next round of features. Share your feedback through DataCite’s survey or join the discussion on the open request for comments.