Onderstaande printversie van het indicatorenboek werd door uw browser gegenereerd, en zal niet steeds optimaal ogen. Via de ingebouwde printfunctie op de website van het Indicatorenboek (ronde knop rechts bovenaan) kan u een printvriendelijke PDF genereren met mooi ogende lay-out.
8.2.2Criteria for building and using metrics in a responsible way
The use of bibliometric and other quantitative indicators in research evaluation has become increasingly complex. As illustrated in the Multidimensional Research Assessment Matrix (AUBR, 2010) and further developed by Moed (2017), there exists a broad array of indicators and methods intended to assess both scholarly and non-academic research impacts. While Moed (2017) offers concrete recommendations and evaluations of specific metrics, the AUBR matrix provides a more general overview of methods and their potential applications.
To integrate quantitative and qualitative approaches meaningfully, diverse data types must be harmonised. Daraio and Glänzel (2016) proposed a standardised data integration model to support this process.
For bibliometric indicators to be meaningful and robust, they must meet several foundational conditions: i) data quality is crucial; ii) metrics must ensure comparability and allow benchmarking (commensurability), and iii) results should be reproducible (validatability). Bookstein (1997) further warned that measurement efforts are often undermined by the “demons” randomness, ambiguity, and conceptual fuzziness. These challenges affect both metric design and interpretation.
To be considered fit for purpose in research assessment, indicators must meet a core set of criteria. They must be
- valid – measuring what they are designed and intended to measure;
- meaningful – significance of results;
- reliable – statistically and based on data of sufficient coverage and availability;
- robust – not sensitive to changes or fluctuations in the system;
- normalisable - although some widely used indicators, such as the h-index, fall short in this regard
- standardisable – to enable comparability and reproducibility.
Meeting these criteria ensures that indicators are suitable for comparative assessments and benchmarking exercises.
Most importantly, metrics must be selected based on their “fitness for purpose” – that is, their ability to align with the specific aims of the evaluation. Users should be aware of the acceptable margins of error (cf. Moed, 2017) and interpret the results in light of any possible methodological limitations, uncertainties or flaws.
Although both ex-ante and ex-post assessments are valuable, they require different types of data and interpretative approaches. A thoughtful balance between qualitative and quantitative perspectives is therefore essential. Qualitative aspects and dimensions – as academic recognition, diversity, and societal engagement – must not be neglected.
Finally, caution is warranted when using composite indicators, which often suffer from non-transparency, arbitrary weighting, and interdependencies among components. Their tendency to reduce the multidimensional realities to a single numerical value may obscure more than it reveals (Gumpenberger et al, 2016).