Buildability report · AI Search

Can Scite be vibe coded?

Smart Citations that show whether a paper's citations support, contrast or merely mention its claims

Keep itWeak replacementNot faithfully

The AI classification of a citation as supporting, contrasting or mentioning is a solvable NLP task, but the product's real value is the licensed full-text corpus across 40+ publishers plus preprint servers, kept current and cross-referenced at scale. No individual or small team can replicate that data-access moat; a DIY build only works on open-access papers, which is a small fraction of the literature that matters for a lot of fields.

Jump to the build brief ↓
Buildability18/100

Legacy-calibrated assessment

Current price$20/mo

Checked Aug 2026

Current annual cost$240

What you pay today, before any DIY hosting

ConsequenceOperational risk

high editorial confidence

Full report reviewNot dated

Tracked separately from the pricing check

The score by layer

Buildability by layer

Scoring method ↗
Interface28

Screens, forms, and focused interactions

Core workflow18

The repeatable job the product performs

Data access5

Availability and legality of required data

Operations5

Uptime, queues, support, and maintenance

Trust & safety18

Security, compliance, and user confidence

What an LLM can build

The achievable core

  • Pull open-access full text (arXiv, PubMed Central, bioRxiv) for a paper's cited works, then use an LLM to classify each citing sentence as supporting, contrasting or neutral.
  • Ground answers in sources you select and cite them visibly.
  • A responsive interface with real empty, loading, success, and error states.
Where the clone breaks

The parts a prompt cannot buy

  • coverage of paywalled publishers (Wiley, Cambridge, Wolters Kluwer, etc.)
  • pre-built citation database spanning hundreds of millions of citation statements
  • retraction/correction flags sourced from Crossref and PubMed
  • browser extension overlay on Google Scholar and journal pages
  • The useful dataset is owned, accumulated, or expensive to reproduce.
  • Reliability at the vendor's scale is an operations problem, not a prompt.
Choose the sensible path

Build, switch, or keep paying

Build the focused core

Narrower, with trade-offs

Pull open-access full text (arXiv, PubMed Central, bioRxiv) for a paper's cited works, then use an LLM to classify each citing sentence as supporting, contrasting or neutral.

Use the build brief ↓
Use an existing alternative

2 checked options

  • Semantic ScholarFree academic search engine with citation graphs and TLDR summaries, though without support/contrast classification.
  • COREFree aggregator of over 200 million open-access research papers with full-text search.
Defensibility

Why people still pay

Researchers pay for reliable, broad coverage of paywalled literature and for a maintained, cross-checked citation graph rather than re-scraping and re-classifying papers themselves every time.

proprietary data

The useful dataset is owned, accumulated, or expensive to reproduce.

scale infra

Reliability at the vendor's scale is an operations problem, not a prompt.

Production build brief

The brief

Context, requirements, acceptance criteria, non-goals, and the full production standard — as Markdown, ready for any coding agent.

Raw URL ↗

Build brief — a focused alternative to Scite

Verdict: Not faithfully · Buildability: 18/100 · Category: AI Search

Source: https://www.canitbevibecoded.com/scite-ai

Independent editorial assessment from Can It Be Vibe Coded? Not affiliated with, endorsed by, or derived from Scite. Verify current pricing and capabilities before acting.

Context

Scite — Smart Citations that show whether a paper's citations support, contrast or merely mention its claims. It currently costs $20/mo.

The AI classification of a citation as supporting, contrasting or mentioning is a solvable NLP task, but the product's real value is the licensed full-text corpus across 40+ publishers plus preprint servers, kept current and cross-referenced at scale. No individual or small team can replicate that data-access moat; a DIY build only works on open-access papers, which is a small fraction of the literature that matters for a lot of fields.

This brief describes a focused, single-operator replacement for the part of Scite that is genuinely reproducible. It is deliberately narrower than the product it replaces, and it says so in writing. Build the useful core; do not pretend to have rebuilt the rest.

What you are building

Pull open-access full text (arXiv, PubMed Central, bioRxiv) for a paper's cited works, then use an LLM to classify each citing sentence as supporting, contrasting or neutral.

Ground answers in sources you select and cite them visibly.

A responsive interface with real empty, loading, success, and error states.

Requirements

Functional

Access to open-access full-text sources (arXiv, PMC, bioRxiv).

Data and integrations

OpenAI/Anthropic API key.

Each of these needs a real account, credential, or quota. Set them up before writing feature code.

Non-functional

Accessibility: semantic markup, labelled controls, visible focus, and reduced-motion support.

Security: server-side secrets, validated input, and no credentials in the client bundle.

Reliability: retries with backoff on external calls, and a clear failure state when a provider is down.

Portability: the operator can export their data and leave without losing it.

Implementation brief

Build me an open-access citation-context checker as a Next.js app.

Stack: Next.js + TypeScript, SQLite via better-sqlite3, Tailwind. No auth, single user, localhost only.

Core loop:

1. A search box where I paste a DOI or arXiv ID; fetch its metadata and reference list via the free Semantic Scholar API (no key needed).

2. For each citing paper available as open-access full text (via arXiv or PubMed Central APIs), pull the sentence(s) around the citation.

3. Send each citation sentence to the Anthropic API (key from .env) with a fixed prompt asking it to classify the citation as supporting, contrasting, or mentioning, with a one-line reason.

4. Show results in a table: citing paper, classification, and the quoted sentence, filterable by classification.

Out of scope: paywalled-publisher content, browser extension, manuscript reference-check upload, accounts, alerts.

Include a README noting this only works for open-access papers and citing papers, unlike a licensed full-text database.

Delivery standard

Inspect the repository first, then write a short implementation plan before writing code.

Deliver the smallest complete end-to-end workflow first; every primary control must work against persisted data.

Use real validation and storage; never substitute fake dashboards, decorative controls, hard-coded success states, or mock integrations.

Include responsive layouts plus genuine empty, loading, success, validation, and failure states.

Keep secrets server-side in environment variables, provide .env.example, and never commit credentials or user data.

Add structured logs around every external call and return actionable errors without leaking sensitive details.

Write unit tests for the core logic and one automated test of the main user journey.

Finish with a README covering setup, architecture, data location, backups, tests, deployment, and known limitations.

Acceptance criteria

A clean install starts the app using only the README and .env.example.

The primary journey works from first visit through saved result, reload, edit, export, and deletion where applicable.

Invalid input, missing configuration, provider failure, and an empty database each have a usable state.

The interface works at 390px and 1440px, is keyboard navigable, and shows visible focus on every control.

Tests, type checking, linting, and a production build all pass with no ignored failures.

No part of the interface implies a live integration, security guarantee, or scale capability that was not actually built and verified.

Non-goals

Do not build these, and do not claim to have replaced them:

Coverage of paywalled publishers (Wiley, Cambridge, Wolters Kluwer, etc.).

Pre-built citation database spanning hundreds of millions of citation statements.

Retraction/correction flags sourced from Crossref and PubMed.

Browser extension overlay on Google Scholar and journal pages.

Reference-check tool for uploaded manuscripts.

What you still own after launch

Secure credentials, rotate secrets, and handle provider rate limits.

Run migrations, backups, restores, and dependency updates.

Test the critical journey after every model, API, or hosting change.

Monitor failures and fix the edge cases a first prompt will miss.

Risk

Operational risk. The code is achievable; dependable data, integrations, and ongoing operations are the real cost.

Editorial confidence in this assessment: high. No reviewed project implementation is linked yet.

Existing alternatives

Before building, compare these checked options:

Semantic Scholar — Free academic search engine with citation graphs and TLDR summaries, though without support/contrast classification

CORE — Free aggregator of over 200 million open-access research papers with full-text search

Prior art

Working open-source software you can read, fork, or borrow from before starting:

Semantic Scholar API — free API with citation graphs and some citation-intent classification for open-access papers


Generated by Can It Be Vibe Coded? · Full report: https://www.canitbevibecoded.com/scite-ai

After the agent stops

You still own the product

  • Secure credentials, rotate secrets, and handle provider rate limits.
  • Run migrations, backups, restores, and dependency updates.
  • Test the critical journey after every model, API, or hosting change.
  • Monitor failures and fix the edge cases a first prompt will miss.
Evidence, not screenshots

Projects built from this idea

No reviewed implementation has been linked for Scite yet. A submission is evidence for review, not automatic proof that the whole product was replaced.

Built a version of Scite?Submit the project as evidence for this report.

Submissions are private until reviewed. Approval adds a link; reproduced verification requires a separate acceptance check.

Start from working software

Open-source prior art

Practical questions

Before you start

Can Scite be vibe coded?

Not faithfully. The AI classification of a citation as supporting, contrasting or mentioning is a solvable NLP task, but the product's real value is the licensed full-text corpus across 40+ publishers plus preprint servers, kept current and cross-referenced at scale. No individual or small team can replicate that data-access moat; a DIY build only works on open-access papers, which is a small fraction of the literature that matters for a lot of fields.

What can an AI coding agent reproduce from Scite?

Pull open-access full text (arXiv, PubMed Central, bioRxiv) for a paper's cited works, then use an LLM to classify each citing sentence as supporting, contrasting or neutral. Ground answers in sources you select and cite them visibly. A responsive interface with real empty, loading, success, and error states.

What will a DIY Scite replacement still be missing?

coverage of paywalled publishers (Wiley, Cambridge, Wolters Kluwer, etc.); pre-built citation database spanning hundreds of millions of citation statements; retraction/correction flags sourced from Crossref and PubMed; browser extension overlay on Google Scholar and journal pages; The useful dataset is owned, accumulated, or expensive to reproduce.; Reliability at the vendor's scale is an operations problem, not a prompt.

What do I still own after building a Scite alternative?

Secure credentials, rotate secrets, and handle provider rate limits. Run migrations, backups, restores, and dependency updates. Test the critical journey after every model, API, or hosting change. Monitor failures and fix the edge cases a first prompt will miss.