Abstract Generative AI (GenAI) tools, and large language models (LLMs) in particular, are being rapidly adopted across the scientific research process, including in funding evaluation and peer review. While their efficiency gains are well documented, their systematic biases remain comparatively understudied. A central concern is whether LLMs can recognise and value scientific novelty, an essential property for the long-run productivity of science. This paper examines 139 monodisciplinary research proposals submitted in 2021 to the Austrian Science Fund (FWF) in mathematics, physics and astronomy, chemistry, and biology. Each proposal carries expert-assigned scores for novelty, scientific quality, team capacity, and feasibility, along with the actual binary funding decision. Using GPT-4o, we generate AI-based funding decisions through within-field pairwise comparisons aggregated into field-level rankings, simulating a competitive selection process. AI-based funding decisions matched human funding decisions in 60% of cases. Conditional logistic regression with field fixed effects shows that, among the four review dimensions, novelty is the only one significantly associated with disagreement. Conditional and multinomial logit analyses further show that novelty systematically predicts the direction of disagreement. Proposals rated as more novel by human reviewers are substantially more likely to be funded by humans and not by the LLM. The observed pattern is consistent with a conservative tendency in LLM-based evaluation, potentially reflecting the statistical logic of next-token prediction trained on past scientific outputs. While the design does not allow us to identify the underlying mechanism, the results raise concerns that LLM-assisted evaluation may under-select proposals that human reviewers identify as highly novel.