Five concrete theorem-level advances
Table 1 identifies five results against specific prior states of the art, including improved regret, competition-complexity, revenue, and price-of-anarchy bounds.
↳ Results §3, Table 1
Crunching the numbers. Responsibly.
We present Cogentic, a multi-agent harness for automated proof discovery on open research problems. While frontier language models can generate strong mathematical ideas in a single shot, single-shot generation is often insufficient for open problems that require exploring multiple competing conjectures, overcoming subtle technical obstructions, and retaining intermediate progress over a long horizon. Cogentic addresses these challenges through an iterative prove--verify loop in which an orchestrator allocates a population of independent provers across distinct proof directions, subjects their output to adversarial verification by several specialized components, and promotes confirmed intermediate results into a persistent verified ledger that later rounds build on. The harness is designed to be able to solve research-level math and theoretical computer science problems. Using Gemini as the base model, Cogentic produced novel results on five open problems across online learning, auction theory, and mechanism design. Each result was independently verified by domain experts and is developed in full in companion papers. We list these results, and new ones as they are verified, at https://sites.google.com/view/cogentic .
There is a deterministic, proper, anytime algorithm for online inverse linear optimization with cumulative shortfall R_T = O(d), uniform in the horizon T, using O(d²) arithmetic and one linear optimization per round.
a coherent proof sketch supports the bound, while the complete proof is deferred to a companion paper
There is an algorithm for prediction with expert advice which uses no knowledge of the horizon and whose regret over n experts satisfies R_t ≤ (1 + O(√(ln ln n/ln n)))√(t ln n/2) simultaneously for all t≥1.
the stated construction directly addresses the anytime bound and is supported by a detailed proof sketch
For any m≥n≥1, if the buyer distribution F_B first-order stochastically dominates the seller distribution F_S, then adding exactly two sellers suffices for Seller Trade Reduction to achieve at least the first-best gains from trade.
the theorem is precisely scoped and linked to a companion proof, but was enumerated by only one source review
For n=2 bidders, the standard Proportional First-Price Auction achieves a Price of Anarchy of at most 1.5; moreover this is tight.
the claim is precisely qualified, though the full companion paper is listed as forthcoming
Derived from the full evaluation — not a separate score.
Strengths
Table 1 identifies five results against specific prior states of the art, including improved regret, competition-complexity, revenue, and price-of-anarchy bounds.
↳ Results §3, Table 1
Figure 1 and Sections 2–2.5 explain how provers, adversarial verifiers, ledgers, process adaptation, and manuscript auditing interact across rounds.
↳ Sections 2–2.5, Figure 1
The case studies identify prior gaps and include contextual qualifications, such as the subsequent tight regret result and the authors’ role in checking and extending outputs.
↳ Results §3.1; Discussion §4
Limitations
The Introduction and Results describe operation from a problem statement without expert hints or human mathematical intervention, while Appendix A supplies target theorems, references, mechanisms, and desired bounds.
↳ Introduction §1; Results §3; Appendix A.2–A.5
The paper does not compare Cogentic with single-agent Gemini, repeated sampling, verifier-free operation, or another harness, and it does not isolate any architectural component.
↳ Sections 2–2.5; Results §3
Only successful cases and approximate call counts are reported, without attempted-problem counts, failed runs, cost distributions, uncertainty, or executable configurations.
↳ Introduction §1; Results §3, Table 1
Table 1 and Sections 3.1–3.5 support a meaningful contribution score by pairing five theorem statements with concrete prior bounds and open questions. Methodological Rigour is constrained by the absence of baselines, ablations, failed-run denominators, and reproducible harness details, leaving the orchestration advantage untested. The architectural exposition and theorem summaries are clear, but Appendix A materially qualifies the no-hints and no-human-intervention framing in the Introduction and Results. Prior-work positioning is strongest at the individual-result level and less complete for the nearest agentic mathematics systems.
Nabu’s assessment, alongside the field’s view.
Are you an author of this paper?
Sound3.0
Confidence highThe five theorem-level results address explicitly identified open bounds and conjectures, with prior states and claimed improvements presented side by side. The harness contribution is less clearly separated from prompting, base-model capability, and expert follow-up.
Open problems resolved or improved by Cogentic.
The multi-agent workflow is described conceptually, but no baseline, component ablation, success denominator, run distribution, or reproducible configuration tests whether orchestration caused the reported performance. The autonomy inconsistency in Results §3 and Appendix A is separately reflected in the Red reliability flag.
Rounds continue until a draft clears verification or until the budget runs out.
The architecture, workflow, and theorem statements are readily followable. Reporting is reduced because the no-hints and no-human-intervention language gives a materially stronger account of autonomy than the prompts and Discussion support.
A key feature of the system is that it only needs a problem statement, without any expert hints.
Each mathematical result is positioned against a concrete prior bound or open question, and the paper acknowledges later work and human refinement. Positioning remains incomplete because it declines comparison with the closest agentic systems and does not fully trace the stated human role into the autonomy and efficiency claims.
We do not attempt to give a comprehensive survey or compare these efforts.
Lower confidence on Contribution, Positioning — domain match limited.
Concerns4 of 4 checks
The paper’s description of autonomous operation is materially inconsistent with the task guidance and subsequent human involvement documented elsewhere in the supplied text. This affects attribution to the harness, not the correctness of the stated theorem results.
No human-subject or comparable ethics requirement is apparent, and the paper makes no suspicious code or data availability claim. The absence of companion proofs and executable harness materials is an assessability limitation rather than a conduct concern.
Flags: 0 declared / 5 total
3 of 3 checked DOIs point to the cited work
59 references in manuscript 47 of 47 checkable references found in an index 12 have no canonical index record — counted, but not index-checkable 5 references confirmed by manual review
No retraction notice found in Retraction Watch.
Sources: Retraction Watch ✓
Where this paper’s evidence sits on the path from initial observation to real-world use.
The reported outputs constitute a proof of concept for expert-assisted mathematical research, but operation still depends on human verification. No released harness, deployment protocol, or independently measured operating characteristics are reported.
Cogentic produces natural language proofs which are verified by experts.
AI-generated, human-governed. Something look off? Contact us to request a review.