Michael Markie
Michael Markie, Head of Innovation Initiatives at eLife

Q1: Why does better quality scholarly metadata matter for your organisation and the communities you work with?

At eLife Pathways, we believe that building a modern, open science ecosystem is impossible without high-quality, open, and machine-readable metadata. If we want to unlock the full potential of our current knowledge base, the underlying metadata needs to be complete.

Better quality scholarly metadata is a major consideration in the work we are doing. In collaboration with others, we are developing a modular science schema that treats the basic building blocks of a research paper—questions, claims, and evidence, data, code, etc.—as individual components. Each of these components will be explicitly labelled and linked via a unified metadata layer, enabling researchers and systems to recombine and reuse scientific information in entirely new ways, creating a universally connected, natively interoperable knowledge base.

A key attribute of modular science is its ability to connect with multiple outputs, enabling it to be recomposed in new ways. With this in mind, we are heavily invested in improving existing scholarly metadata, as well as ensuring the new schema provides a framework for metadata to be fully accurate and easily improved if needed.

Q2: What have you observed or experienced that makes you think the current model for maintaining scholarly metadata needs to change?

We recently conducted some discovery work with a funder to link preprints to their published articles and related research outputs, including data, software, peer reviews, and other assets, to create a knowledge graph of linked research outputs.

During this work, we became aware of the variance in the quality of information provided across depositors and the gaps in the metadata that has been made available.
Looking at the nodes in this knowledge graph across a spectrum of preprints and articles gave us a zoomed-out view of a landscape with metadata potholes, and we saw that the effort required to fill them is not trivial. You can always make changes as original depositors, but beyond that, when making corrections is out of your control, it’s not possible to make changes systematically. So we really need collaborative ways to start filling these potholes.

Q3: Where do you see the most significant barriers—technical, institutional, or cultural—to making collaborative metadata enrichment work in practice?

There is already a wealth of curated scholarly metadata available, but it is not well connected. Fortunately, the technology and expertise needed to make these connections exist, as demonstrated by the successful COMET pilot projects that have enhanced millions of records. The main challenge lies in creating an automated and trusted pipeline that integrates these assertions into legacy publishing platforms while maintaining the integrity of the original sources.

To achieve this, collaboration among third-party developers, publishers, and repositories is essential. Adopting shared standards for assertions and building export pipelines are not purely technical challenges; rather, they are cultural ones. Overcoming this barrier will require genuine commitment and cooperation from all stakeholders involved.

Q4: How does the COMET Model apply—or potentially apply—in your specific context? What would putting it into practice actually look like for your organisation and/or community?

As a collaborative partner with COMET, eLife Pathways is well-positioned to actively support the COMET Model. We will be directly involved in testing and improving parsing models and evaluation benchmarks for preprints, as well as helping to establish community standards for metadata curation and enrichment. Additionally, we will explore ways to create a practical publishing pathway for enriching data back to DOIs through our open-source publishing platform, Kotahi.

Another key focus of our work is to enhance COMET’s efforts in improving funding metadata. The eLife journal already provides funding metadata for reviewed preprints, and we aim to leverage this experience to develop a model for extracting and parsing funding statements from the broader biomedical preprints corpus.

Our collaboration with COMET will enable us to adopt new methodologies that we can share across the ecosystem. We will work with preprint servers and the eLife journal to streamline the entry and flow of high-quality metadata throughout the system.

Q5: What conditions or changes would need to be in place for that to work well?

The biggest challenge (as with most things) would be dedicating the time to step back from current processes and consider how to change them to incorporate a new way of doing things.
If an organisation’s processes are efficient and closely linked to a specific workflow, the reluctance to change becomes greater, especially if the change does not provide an immediate benefit. When existing workflows are deeply embedded and do not yield immediate internal returns on investment, the resistance to change is high. At eLife Pathways, our objective is to communicate the advantages of our approach and assist our partners in providing improved metadata for the entire ecosystem. When collaborating with others, we will focus on demonstrating how collaborative curation can enhance their existing systems without disrupting them.

Q6: What would you want other practitioners working on similar problems to understand—either about the challenge itself, or about what you’ve learnt so far?

I’ll revisit my analogy about potholes. It’s a common problem: we are aware of the gaps in our road infrastructure, and we often find ourselves waiting (impatiently) for a central authority to address them. However, in some European countries, a more collaborative approach brings together local authorities, utility companies, and technology firms to share real-time road data and coordinate resurfacing schedules, resulting in much faster, more efficient road repairs.

In scholarly research, missing or incomplete metadata fields are like potholes. They disrupt the journey for everyone downstream. What we’ve learned so far is that collaboration is the most effective way to address these issues. Rather than assigning blame, it’s more productive to identify areas for improvement and support those depositing the data. Metadata experts from publishers, repositories, and institutions each bring valuable skills to the table. By combining these strengths through the COMET Model, we can resolve metadata issues more efficiently, creating a smoother experience for everyone as we work together to build better infrastructure.