Can 100 Questions be vibe coded?
Evidence-linked brand visibility benchmarks across OpenAI, Claude, Gemini, and Grok
A personal CLI that asks the same questions across four model APIs and compares the answers is achievable solo, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work.
Jump to the build brief ↓Checked Jul 2026
What you pay today, before any DIY hosting
high editorial confidence
Buildability by layer
Screens, forms, and focused interactions
The repeatable job the product performs
Availability and legality of required data
Uptime, queues, support, and maintenance
Security, compliance, and user confidence
The achievable core
- Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report.
- Automate a bounded research or reporting workflow using permitted data sources.
- A responsive interface with real empty, loading, success, and error states.
The parts a prompt cannot buy
- reliable orchestration and retries across four providers
- normalized citations and evidence-linked metrics
- competitor and missed-question extraction
- stored point-in-time reports and comparisons
- The last 20 percent is sync, migration fidelity, speed, and edge cases.
- Connectors, OAuth flows, and vendor API changes require constant upkeep.
Why people still pay
They pay for a repeatable, frozen benchmark with provider failures handled, citations normalized, every metric tied to evidence, and a report that is ready to act on.
The last 20 percent is sync, migration fidelity, speed, and edge cases.
Connectors, OAuth flows, and vendor API changes require constant upkeep.
The brief
Context, requirements, acceptance criteria, non-goals, and the full production standard — as Markdown, ready for any coding agent.
Build brief — a focused alternative to 100 Questions
Context
**100 Questions** — Evidence-linked brand visibility benchmarks across OpenAI, Claude, Gemini, and Grok. It currently costs $9/mo.
A personal CLI that asks the same questions across four model APIs and compares the answers is achievable solo, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work.
This brief describes a focused, single-operator replacement for the part of 100 Questions that is genuinely reproducible. It is deliberately narrower than the product it replaces, and it says so in writing. Build the useful core; do not pretend to have rebuilt the rest.
What you are building
Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report.
Automate a bounded research or reporting workflow using permitted data sources.
A responsive interface with real empty, loading, success, and error states.
Requirements
Functional
Provider-specific web search or grounding tools.
Durable run storage.
URL and citation normalization.
Report generation.
Data and integrations
OpenAI, Anthropic, Google Gemini, and xAI API keys.
Each of these needs a real account, credential, or quota. Set them up before writing feature code.
Non-functional
Accessibility: semantic markup, labelled controls, visible focus, and reduced-motion support.
Security: server-side secrets, validated input, and no credentials in the client bundle.
Reliability: retries with backoff on external calls, and a clear failure state when a provider is down.
Portability: the operator can export their data and leave without losing it.
Implementation brief
Build me a local AI visibility benchmark for one brand. Requirements:
Use Node 22, TypeScript, official provider SDKs, SQLite, and a CLI.
`benchmark --domain example.com --description "..."` creates one immutable run.
Generate 25 buyer questions from the domain and description, or accept a JSON question file.
Ask the exact same questions through OpenAI, Anthropic, Gemini, and xAI.
Use each provider's supported web-search or grounding tool; keys live only in `.env`.
Limit concurrency per provider, retry transient failures, and preserve failed cells in the report.
Store prompts, raw answers, citations, timestamps, model ids, and errors in SQLite.
Detect exact and case-insensitive brand mentions; allow aliases in a config file.
Extract named competitors with one structured LLM pass after all answers are stored.
Normalize citation URLs by hostname, canonical URL, and stripped tracking parameters.
Compute visibility by provider, answer coverage, owned-domain citation rate, and top sources.
Show missed questions where competitors appear but the target brand does not.
Render a self-contained static HTML report with filters and expandable raw evidence.
Export questions, answer metrics, competitors, and citations as CSV files.
Every aggregate metric must link back to the answer rows used to calculate it.
Out of scope: accounts, billing, teams, scheduled monitoring, and recommendation generation.
Include fixture-based tests for mention detection, URL normalization, and metric calculations.
README: setup, provider-specific grounding caveats, estimated API cost, and exact run commands.
Delivery standard
Inspect the repository first, then write a short implementation plan before writing code.
Deliver the smallest complete end-to-end workflow first; every primary control must work against persisted data.
Use real validation and storage; never substitute fake dashboards, decorative controls, hard-coded success states, or mock integrations.
Include responsive layouts plus genuine empty, loading, success, validation, and failure states.
Keep secrets server-side in environment variables, provide .env.example, and never commit credentials or user data.
Add structured logs around every external call and return actionable errors without leaking sensitive details.
Write unit tests for the core logic and one automated test of the main user journey.
Finish with a README covering setup, architecture, data location, backups, tests, deployment, and known limitations.
Acceptance criteria
A clean install starts the app using only the README and .env.example.
The primary journey works from first visit through saved result, reload, edit, export, and deletion where applicable.
Invalid input, missing configuration, provider failure, and an empty database each have a usable state.
The interface works at 390px and 1440px, is keyboard navigable, and shows visible focus on every control.
Tests, type checking, linting, and a production build all pass with no ignored failures.
No part of the interface implies a live integration, security guarantee, or scale capability that was not actually built and verified.
Non-goals
Do not build these, and do not claim to have replaced them:
Reliable orchestration and retries across four providers.
Normalized citations and evidence-linked metrics.
Competitor and missed-question extraction.
Stored point-in-time reports and comparisons.
The last 20 percent is sync, migration fidelity, speed, and edge cases.
Connectors, OAuth flows, and vendor API changes require constant upkeep.
What you still own after launch
Secure credentials, rotate secrets, and handle provider rate limits.
Run migrations, backups, restores, and dependency updates.
Test the critical journey after every model, API, or hosting change.
Monitor failures and fix the edge cases a first prompt will miss.
Maintain every third-party integration as APIs and OAuth rules change.
Risk
**Operational risk.** The code is achievable; dependable data, integrations, and ongoing operations are the real cost.
Editorial confidence in this assessment: high. No independent one-shot implementation is linked yet.
Generated by [Can It Be Vibe Coded?](https://www.canitbevibecoded.com) · Full report: https://www.canitbevibecoded.com/100-questions
You still own the product
- Secure credentials, rotate secrets, and handle provider rate limits.
- Run migrations, backups, restores, and dependency updates.
- Test the critical journey after every model, API, or hosting change.
- Monitor failures and fix the edge cases a first prompt will miss.
- Maintain every third-party integration as APIs and OAuth rules change.
Before you start
Can 100 Questions be vibe coded?
Partly, if you narrow it. A personal CLI that asks the same questions across four model APIs and compares the answers is achievable solo, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work.
What can an AI coding agent reproduce from 100 Questions?
Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report. Automate a bounded research or reporting workflow using permitted data sources. A responsive interface with real empty, loading, success, and error states.
What will a DIY 100 Questions replacement still be missing?
reliable orchestration and retries across four providers; normalized citations and evidence-linked metrics; competitor and missed-question extraction; stored point-in-time reports and comparisons; The last 20 percent is sync, migration fidelity, speed, and edge cases.; Connectors, OAuth flows, and vendor API changes require constant upkeep.
What do I still own after building a 100 Questions alternative?
Secure credentials, rotate secrets, and handle provider rate limits. Run migrations, backups, restores, and dependency updates. Test the critical journey after every model, API, or hosting change. Monitor failures and fix the edge cases a first prompt will miss. Maintain every third-party integration as APIs and OAuth rules change.