Finer-grained topic matching design
The study pairs OA and NOA papers using TF-IDF cosine similarity on titles and abstracts, addressing topic at a finer level than journal- or discipline-level controls.
↳ Introduction §1; Methods §3.1.3
Assembling the evidence…
Research topics vary in their citation potential. In a metric-wise scientific milieu, it would be probable that authors tend to select citation-attractive topics especially when choosing open access (OA) outlets that are more likely to attract citations. Applying a matched-pairs study design, this research aims to examine the role of research topics in the citation advantage of OA papers. Using a comparative citation analysis method, it investigates a sample of papers published in 47 Elsevier article processing charges (APC)-funded journals in different access models including non-open access (NOA), APC, Green and mixed Green-APC. The contents of the papers are analysed using natural language processing techniques at the title and abstract level and served as a basis to match the NOA papers to their peers in the OA models. The publication years and journals are controlled for in order to avoid their impacts on the citation numbers. According to the results, the OA citation advantage that is observed in the whole sample still holds even for the highly similar OA and NOA papers. This implies that the OA citation surplus is not an artefact of the OA and NOA papers’ differences in their topics and, therefore, in their citation potential. This leads to the conclusion that OA authors’ self-selectivity, if it exists at all, is not responsible for the OA citation advantage, at least as far as selection of topics with probably higher citation potentials is concerned.
the OA citation advantage that is observed in the whole sample still holds even for the highly similar OA and NOA papers
repeated-document dependence may make the matched-pair significance anti-conservative
the OA citation surplus is not an artefact of the OA and NOA papers’ differences in their topics
exclusionary inference rests on an observational lexical topic proxy with dependent matches
OA authors’ self-selectivity, if it exists at all, is not responsible for the OA citation advantage, at least as far as selection of topics with probably higher citation potentials is concerned
topic matching cannot securely exclude the proposed selection mechanism under the unresolved design limitations
the OA papers significantly outperform their highly similar NOA peers no matter if they are published in the same or different years, journals or document types
Table 6 reports consistent stratified directions, though repeated matches temper inferential certainty
Derived from the full evaluation — not a separate score.
Strengths
The study pairs OA and NOA papers using TF-IDF cosine similarity on titles and abstracts, addressing topic at a finer level than journal- or discipline-level controls.
↳ Introduction §1; Methods §3.1.3
The paper describes the KNIME preprocessing and matching nodes, citation normalisation, similarity clustering, and Wilcoxon comparisons sufficiently to expose the major analytical choices.
↳ Methods §§3.1.3–3.2
The introduction discusses visibility, early access, quality advantage, quality bias, journal prestige, collaboration, and contrary findings rather than presenting OACA causation as settled.
↳ Introduction §1
Limitations
Methods §3.1.3 allows papers to appear in multiple matches, but Tables 4–6 do not adjust the Wilcoxon inference for repeated-document dependence. Reported significance may therefore be anti-conservative.
↳ Methods §3.1.3; Tables 4–6
TF-IDF cosine similarity over titles and abstracts is treated as equalising topic and citation potential without comparison to alternative semantic measures or quantified measurement error.
↳ Methods §§3.1.3–3.2
The abstract and conclusion say topic self-selection is not responsible, although the analysis uses an observational lexical-matching design and leaves material inferential limitations unresolved.
↳ Abstract; Conclusion §6
The article makes a distinct contribution by matching OA and NOA papers at the title-and-abstract level and then examining citation differences across similarity strata. Its methods are reported in useful detail, and Table 6 extends the analysis across year, journal, and document-type groupings. The principal constraint is that §3.1.3 permits documents to enter multiple matches while Tables 4–6 do not adjust for the resulting dependence. The abstract and conclusion also convert this observational result into an exclusionary claim about topic-based self-selection, while the limitations section does not address the dependence problem or fully examine the validity of the topic proxy.
Nabu’s assessment, alongside the field’s view.
Are you an author of this paper?
Limited2.9
Confidence mediumThe study provides a genuine but incremental advance by matching OA and NOA papers at the title-and-abstract topic level rather than relying only on journal or discipline categories. Its contribution is bounded by one publisher, 47 hybrid journals, and a narrow lexical operationalisation of topic.
papers’ topics are more specific and therefore more prone to reflect authors’ interests and preferences
The workflow reports sampling, TF-IDF processing, cosine matching, normalisation, Wilcoxon tests, and stratification by publication factors. Methodological Rigour is limited by repeated documents across matches without dependence adjustment, an unvalidated lexical topic proxy, unexplained seed selection, and causal-exclusion language.
some NOA papers matched more than one document in the OA models and vice versa
The research aims, workflow, and progression from overall to similarity-stratified results are readily followable. The abstract and conclusion nevertheless express the central observational result as exclusion of topic-based self-selection, creating a material mismatch between wording and evidence.
the OACA does not result from OA papers’ differences in topics
The introduction engages competing OACA explanations and both confirming and disconfirming studies, while the conclusion acknowledges publisher scope and subgroup-size limits. It does not acknowledge repeated-match dependence or adequately trace limitations of the topic proxy through the exclusionary conclusion.
there is not yet any certainty about the causation of the phenomenon
Lower confidence on Methodological Rigour — domain match limited.
Concerns4 of 4 checks
The reported counts and rank statistics are arithmetically plausible, but the inferential analysis does not account for dependence created when documents participate in multiple matched pairs. That issue could affect the statistical support for the central results.
Conflicts and funding are disclosed, and human-participant ethics approval is not expected for this bibliometric study. The absence of shared data, matching outputs, and code limits independent inspection but does not constitute a suspicious availability claim.
Flags: 1 declared / 5 total
77 of 77 checkable references verified
86 references in manuscript 9 have no canonical index record — counted, but not index-checkable 8 references confirmed by manual review
No retraction notice found in Retraction Watch.
Sources: Retraction Watch ✓
Where this paper’s evidence sits on the path from initial observation to real-world use.
The output is an observational bibliometric finding that narrows one proposed explanation rather than tested policy guidance or an intervention. The KNIME workflow is described but not released as a reusable artefact.
Further studies are required to dig deep into topics’ distributions
AI-generated, human-governed. Something look off? Contact us to request a review.