Can Scite be vibe coded?
Smart Citations that show whether a paper's citations support, contrast or merely mention its claims
The AI classification of a citation as supporting, contrasting or mentioning is a solvable NLP task, but the product's real value is the licensed full-text corpus across 40+ publishers plus preprint servers, kept current and cross-referenced at scale. No individual or small team can replicate that data-access moat; a DIY build only works on open-access papers, which is a small fraction of the literature that matters for a lot of fields.
Jump to the build brief ↓Legacy-calibrated assessment
Checked Aug 2026
What you pay today, before any DIY hosting
high editorial confidence
Tracked separately from the pricing check
Buildability by layer
Screens, forms, and focused interactions
The repeatable job the product performs
Availability and legality of required data
Uptime, queues, support, and maintenance
Security, compliance, and user confidence
The achievable core
- Pull open-access full text (arXiv, PubMed Central, bioRxiv) for a paper's cited works, then use an LLM to classify each citing sentence as supporting, contrasting or neutral.
- Ground answers in sources you select and cite them visibly.
- A responsive interface with real empty, loading, success, and error states.
The parts a prompt cannot buy
- coverage of paywalled publishers (Wiley, Cambridge, Wolters Kluwer, etc.)
- pre-built citation database spanning hundreds of millions of citation statements
- retraction/correction flags sourced from Crossref and PubMed
- browser extension overlay on Google Scholar and journal pages
- The useful dataset is owned, accumulated, or expensive to reproduce.
- Reliability at the vendor's scale is an operations problem, not a prompt.
Build, switch, or keep paying
Narrower, with trade-offs
Pull open-access full text (arXiv, PubMed Central, bioRxiv) for a paper's cited works, then use an LLM to classify each citing sentence as supporting, contrasting or neutral.
Use the build brief ↓2 checked options
- Semantic Scholar ↗Free academic search engine with citation graphs and TLDR summaries, though without support/contrast classification.
- CORE ↗Free aggregator of over 200 million open-access research papers with full-text search.
Recommended
Researchers pay for reliable, broad coverage of paywalled literature and for a maintained, cross-checked citation graph rather than re-scraping and re-classifying papers themselves every time.
Visit Scite ↗Why people still pay
Researchers pay for reliable, broad coverage of paywalled literature and for a maintained, cross-checked citation graph rather than re-scraping and re-classifying papers themselves every time.
The useful dataset is owned, accumulated, or expensive to reproduce.
Reliability at the vendor's scale is an operations problem, not a prompt.
The brief
Context, requirements, acceptance criteria, non-goals, and the full production standard — as Markdown, ready for any coding agent.
Build brief — a focused alternative to Scite
Context
Scite — Smart Citations that show whether a paper's citations support, contrast or merely mention its claims. It currently costs $20/mo.
The AI classification of a citation as supporting, contrasting or mentioning is a solvable NLP task, but the product's real value is the licensed full-text corpus across 40+ publishers plus preprint servers, kept current and cross-referenced at scale. No individual or small team can replicate that data-access moat; a DIY build only works on open-access papers, which is a small fraction of the literature that matters for a lot of fields.
This brief describes a focused, single-operator replacement for the part of Scite that is genuinely reproducible. It is deliberately narrower than the product it replaces, and it says so in writing. Build the useful core; do not pretend to have rebuilt the rest.
What you are building
Pull open-access full text (arXiv, PubMed Central, bioRxiv) for a paper's cited works, then use an LLM to classify each citing sentence as supporting, contrasting or neutral.
Ground answers in sources you select and cite them visibly.
A responsive interface with real empty, loading, success, and error states.
Requirements
Functional
Access to open-access full-text sources (arXiv, PMC, bioRxiv).
Data and integrations
OpenAI/Anthropic API key.
Each of these needs a real account, credential, or quota. Set them up before writing feature code.
Non-functional
Accessibility: semantic markup, labelled controls, visible focus, and reduced-motion support.
Security: server-side secrets, validated input, and no credentials in the client bundle.
Reliability: retries with backoff on external calls, and a clear failure state when a provider is down.
Portability: the operator can export their data and leave without losing it.
Implementation brief
Build me an open-access citation-context checker as a Next.js app.
Stack: Next.js + TypeScript, SQLite via better-sqlite3, Tailwind. No auth, single user, localhost only.
Core loop:
1. A search box where I paste a DOI or arXiv ID; fetch its metadata and reference list via the free Semantic Scholar API (no key needed).
2. For each citing paper available as open-access full text (via arXiv or PubMed Central APIs), pull the sentence(s) around the citation.
3. Send each citation sentence to the Anthropic API (key from .env) with a fixed prompt asking it to classify the citation as supporting, contrasting, or mentioning, with a one-line reason.
4. Show results in a table: citing paper, classification, and the quoted sentence, filterable by classification.
Out of scope: paywalled-publisher content, browser extension, manuscript reference-check upload, accounts, alerts.
Include a README noting this only works for open-access papers and citing papers, unlike a licensed full-text database.
Delivery standard
Inspect the repository first, then write a short implementation plan before writing code.
Deliver the smallest complete end-to-end workflow first; every primary control must work against persisted data.
Use real validation and storage; never substitute fake dashboards, decorative controls, hard-coded success states, or mock integrations.
Include responsive layouts plus genuine empty, loading, success, validation, and failure states.
Keep secrets server-side in environment variables, provide .env.example, and never commit credentials or user data.
Add structured logs around every external call and return actionable errors without leaking sensitive details.
Write unit tests for the core logic and one automated test of the main user journey.
Finish with a README covering setup, architecture, data location, backups, tests, deployment, and known limitations.
Acceptance criteria
A clean install starts the app using only the README and .env.example.
The primary journey works from first visit through saved result, reload, edit, export, and deletion where applicable.
Invalid input, missing configuration, provider failure, and an empty database each have a usable state.
The interface works at 390px and 1440px, is keyboard navigable, and shows visible focus on every control.
Tests, type checking, linting, and a production build all pass with no ignored failures.
No part of the interface implies a live integration, security guarantee, or scale capability that was not actually built and verified.
Non-goals
Do not build these, and do not claim to have replaced them:
Coverage of paywalled publishers (Wiley, Cambridge, Wolters Kluwer, etc.).
Pre-built citation database spanning hundreds of millions of citation statements.
Retraction/correction flags sourced from Crossref and PubMed.
Browser extension overlay on Google Scholar and journal pages.
Reference-check tool for uploaded manuscripts.
What you still own after launch
Secure credentials, rotate secrets, and handle provider rate limits.
Run migrations, backups, restores, and dependency updates.
Test the critical journey after every model, API, or hosting change.
Monitor failures and fix the edge cases a first prompt will miss.
Risk
Operational risk. The code is achievable; dependable data, integrations, and ongoing operations are the real cost.
Editorial confidence in this assessment: high. No reviewed project implementation is linked yet.
Existing alternatives
Before building, compare these checked options:
Semantic Scholar — Free academic search engine with citation graphs and TLDR summaries, though without support/contrast classification
CORE — Free aggregator of over 200 million open-access research papers with full-text search
Prior art
Working open-source software you can read, fork, or borrow from before starting:
Semantic Scholar API — free API with citation graphs and some citation-intent classification for open-access papers
Generated by Can It Be Vibe Coded? · Full report: https://www.canitbevibecoded.com/scite-ai
You still own the product
- Secure credentials, rotate secrets, and handle provider rate limits.
- Run migrations, backups, restores, and dependency updates.
- Test the critical journey after every model, API, or hosting change.
- Monitor failures and fix the edge cases a first prompt will miss.
Projects built from this idea
No reviewed implementation has been linked for Scite yet. A submission is evidence for review, not automatic proof that the whole product was replaced.
Built a version of Scite?Submit the project as evidence for this report.
Open-source prior art
Before you start
Can Scite be vibe coded?
Not faithfully. The AI classification of a citation as supporting, contrasting or mentioning is a solvable NLP task, but the product's real value is the licensed full-text corpus across 40+ publishers plus preprint servers, kept current and cross-referenced at scale. No individual or small team can replicate that data-access moat; a DIY build only works on open-access papers, which is a small fraction of the literature that matters for a lot of fields.
What can an AI coding agent reproduce from Scite?
Pull open-access full text (arXiv, PubMed Central, bioRxiv) for a paper's cited works, then use an LLM to classify each citing sentence as supporting, contrasting or neutral. Ground answers in sources you select and cite them visibly. A responsive interface with real empty, loading, success, and error states.
What will a DIY Scite replacement still be missing?
coverage of paywalled publishers (Wiley, Cambridge, Wolters Kluwer, etc.); pre-built citation database spanning hundreds of millions of citation statements; retraction/correction flags sourced from Crossref and PubMed; browser extension overlay on Google Scholar and journal pages; The useful dataset is owned, accumulated, or expensive to reproduce.; Reliability at the vendor's scale is an operations problem, not a prompt.
What do I still own after building a Scite alternative?
Secure credentials, rotate secrets, and handle provider rate limits. Run migrations, backups, restores, and dependency updates. Test the critical journey after every model, API, or hosting change. Monitor failures and fix the edge cases a first prompt will miss.