Researcher Knowledge Workflows
How I Test AI Research Tools
Every review on this site makes claims — this tool finds relevant papers, that one hallucinates citations, this one is worth $20 a month. Claims are cheap. This page is the methodology those claims rest on, published in full so you can hold every future review to it.
If a Noted Scholar review ever deviates from what follows, that’s a defect. Report it.
The problem with AI tool reviews
Search for a review of almost any AI research tool and you’ll find three kinds of pages. Vendor blogs reviewing their own competitors, which is exactly as reliable as it sounds. Affiliate listicles assembled from screenshots of pricing pages, written by people who plainly never ran a real search in the tool. And the occasional honest academic blog post — usually excellent, usually two years out of date.
The gap is testing that is current, documented, and done by someone who actually reads papers for a living. That gap is the reason this site exists, and this methodology is how we try to fill it without becoming the third kind of page ourselves.
Rule 1: Real research tasks only
Every tool is tested against genuine research work, not demos designed to flatter it.
That means, concretely:
- Literature searches with known answers. When we test a research assistant like Elicit or Consensus, we run questions from fields we know well — questions where we already know the key papers a competent search must surface. A tool that misses the landmark paper in its top twenty results gets that miss documented, with the paper named.
- A real reference library. Reference managers and sync workflows get tested against a working Zotero library of several thousand items with a decade of accumulated mess — duplicate entries, broken PDFs, inconsistent tags. Any tool works on a clean library of forty items. Almost nothing survives a real one, and the review says so.
- Manuscripts in progress. Writing and editing tools are run on actual drafts headed for actual submission — with AI-use policies checked first, and disclosure handled as the venue requires. Not sample paragraphs written to be easy.
A demo shows you the tool at its best. A workflow shows you the tool on a Tuesday.
Rule 2: Two weeks minimum
No verdict is published on less than two weeks of regular use.
First impressions in this category are consistently wrong, in both directions. Tools with beautiful onboarding fall apart in week two when you hit the limits of the free tier, the context window, or the citation database. Tools with ugly interfaces — Zotero has never won a beauty contest — reveal their depth slowly. The two-week floor exists because both errors are expensive for you: adopting a tool that collapses under real use costs you a migration, and skipping a genuinely good tool because its first hour is rough costs you the compounding benefit.
For head-to-head comparisons, the same tasks run through every tool in the comparison, in the same weeks. Testing tool A in January and tool B in June, in a category where models change monthly, is not a comparison — it’s two unrelated anecdotes.
Rule 3: Failures are the review
Every review documents where the tool broke. Not as a token “cons” list — as the substance of the review.
For AI research tools the failures that matter are specific and recurring, so we check for each of them every time:
| Failure mode | How we check |
|---|---|
| Hallucinated citations | Every reference the tool produces gets verified against the actual source. Fabrication rates are reported as counts, not vibes. |
| Missed landmark papers | Search results checked against fields where we know the canon. |
| Summary distortion | Tool summaries compared against papers we’ve read closely — especially methods and limitations sections, which summarizers love to skip. |
| Sync and data integrity | Libraries and notes checked after round-trips. A reference manager that silently drops annotations is disqualifying, whatever else it does. |
| Pricing traps | Free-tier limits, auto-renewal behavior, and cancellation flow all tested with our own card. |
A tool that fails gracefully — that tells you when it’s unsure — routinely outranks a more capable tool that fails silently. Silent failure in research tooling produces retraction-grade mistakes, and the scoring reflects that.
Rule 4: We pay, and we disclose
Every subscription tested here is bought at the public price with our own money. No review copies, no “creator accounts,” no extended trials arranged through a marketing contact. Vendors do not know a review is happening until it’s published.
Money still enters the picture in one place: affiliate links. Some reviews link to tools through programs that pay this site a commission if you subscribe. Three commitments keep that honest:
- Any article containing affiliate links says so at the top, before the first link — per FTC guidance, and per basic respect.
- Verdicts are drafted from testing notes before affiliate status is considered. Free tools beat paid tools here regularly, which is easy to verify: our most-recommended tools — Zotero, Obsidian, NotebookLM — pay us nothing.
- The full policy is at the affiliate disclosure page, permanently.
The scoring rubric
Reviews score tools on five axes, weighted for research work rather than general productivity:
- Retrieval quality — does it find the right papers, and does it know what it missed?
- Trustworthiness — hallucination rate, source transparency, graceful failure.
- Workflow fit — does it connect to the tools researchers actually use (Zotero, Obsidian, Word, LaTeX), or does it demand you live inside it?
- Honest pricing — what the useful tier actually costs, not the teaser tier.
- Durability — export paths and data ownership, so that when the tool dies or pivots — and in this market, assume it will — your work survives.
Numbers are a summary, not a substitute. The verdict paragraph, and especially the “who should not use this” section, matters more than the score.
What gets re-tested
AI tools do not stay reviewed. A model swap can change a tool’s hallucination rate overnight, in either direction. So: every review of a paid tool is re-verified quarterly — pricing, headline features, and a spot-check of the original test tasks. Material changes get a visible “Updated” date and a changelog note; a review that no longer holds gets rewritten or pulled, not quietly left to rank.
If you catch a stale claim before we do, email hello@notedscholar.com — corrections get fixed and credited.
Holding us to it
This methodology is a standing promise, and it’s falsifiable on purpose. Every review names the tasks it ran, the failures it found, and the date it was last verified. If you run the same tasks and get a different result, one of us has learned something — write in either way.
That’s the whole system. It’s slower than screenshotting a pricing page. It’s the only version of this work worth your reading time.