100 Questions
Evidence-linked brand visibility benchmarks across OpenAI, Claude, Gemini, and Grok
A personal CLI that asks the same questions across four model APIs and compares the answers is weekend-buildable, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work.
Build verification: not recorded. How we judge buildability
What you give up
- reliable orchestration and retries across four providers
- normalized citations and evidence-linked metrics
- competitor and missed-question extraction
- stored point-in-time reports and comparisons
- polished exports and action recommendations
Why people still pay
They pay for a repeatable, frozen benchmark with provider failures handled, citations normalized, every metric tied to evidence, and a report that is ready to act on.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Runtime and tools: Node, TypeScript, Commander, better-sqlite3 and the configured providers' documented SDKs.
- Before starting: A Node runtime, provider keys, reviewed question JSON and editable model/pricing configuration; use a five-question pilot before larger runs.
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule — domain: Store QuestionSet, Run, ProviderAttempt, Citation and CostReservation; a provider denominator includes completed grounded answers only, while failed, uncited and unknown outcomes stay visible.
Project rule — scope and recovery: Preserve identical question text and model/tool settings per comparison. Estimated reservations are not guaranteed spend caps, and API answers do not represent consumer search-interface rankings.
Project rule — acceptance: Run three questions where one provider times out and another returns no citations; neither outcome may inflate grounded coverage or silently trigger a second charged request.
Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.
Recommended skill: sharp-edges — review configuration and API defaults against the app-specific invariants and recovery boundaries above; this is not a security certification. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Recommended skill: web-design-guidelines — review keyboard access, focus, validation, error recovery and the readable work/review interface or HTML report. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Implementation plan
Phase 1
Pin the working slice and create its example input: Run a reviewed five-question pilot across configured search-capable model providers, then compare brand mentions and owned-domain citations in an evidence-linked report. Confirm setup: A Node runtime, provider keys, reviewed question JSON and editable model/pricing configuration; use a five-question pilot before larger runs.
Phase 2
Implement durable result files and command invariants before formatting terminal output: Store QuestionSet, Run, ProviderAttempt, Citation and CostReservation; a provider denominator includes completed grounded answers only, while failed, uncited and unknown outcomes stay visible.
Phase 3
Connect CLI commands to real saved state and explicit result/exit statuses. Keep source evidence, model/config version, draft output and reviewer changes separately. Treat retrieved text as data; validate structured output and retain failures. Never silently send private material to a fallback provider.
Phase 4
Expose the app-specific limits and recovery path in context: Preserve identical question text and model/tool settings per comparison. Estimated reservations are not guaranteed spend caps, and API answers do not represent consumer search-interface rankings.
Phase 5
Walk through this concrete acceptance case and preserve its exported evidence: Run three questions where one provider times out and another returns no citations; neither outcome may inflate grounded coverage or silently trigger a second charged request. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.
Build me a local four-provider AI visibility benchmark CLI as a limited substitute for 100 Questions. Requirements: - Use Node, TypeScript, Commander, and better-sqlite3. Use the official openai, @anthropic-ai/sdk, and @google/genai packages, plus xAI's documented HTTPS API through fetch; keep model IDs and search-tool options in config rather than hardcoding dated versions. - Generate 25 buyer questions from a supplied brand description, or import a reviewed JSON question list. Show the list before sending anything and default to a five-question pilot; a full run needs an explicit flag. - Send identical question text to each configured provider with its documented web-search capability. If search or citation metadata is unavailable, record that outcome rather than treating an uncited answer as web-grounded. - Before each call, reserve a configurable estimated maximum from remaining run budget using the user's price table, max output tokens, and search allowance. Stop dispatch when it will not fit; show that actual provider charges may differ and are not a guaranteed hard cap. - Store request, provider/model/tool settings, response, citation metadata, usage, status, and attempt time in SQLite. Retry only definite rate-limit or transient failures; mark ambiguous timeouts 'unknown' for manual retry to avoid duplicate charges. - Compute brand and competitor mentions and owned-domain citations from completed grounded answers only. Report per-provider completed, failed, and unknown counts as denominators; retain raw evidence beside every metric. - Export static escaped HTML and CSV. A SHA-256 digest may detect accidental file changes but must not be called tamper-proof. Keep API keys in .env and out of reports; no telemetry. README: provider setup, current pricing checks, cost uncertainty, and partial-run recovery. EDITORIAL IMPLEMENTATION CONTRACT Working slice: Run a reviewed five-question pilot across configured search-capable model providers, then compare brand mentions and owned-domain citations in an evidence-linked report. Data and invariants: Store QuestionSet, Run, ProviderAttempt, Citation and CostReservation; a provider denominator includes completed grounded answers only, while failed, uncited and unknown outcomes stay visible. Boundary and recovery: Preserve identical question text and model/tool settings per comparison. Estimated reservations are not guaranteed spend caps, and API answers do not represent consumer search-interface rankings. Acceptance walkthrough: Run three questions where one provider times out and another returns no citations; neither outcome may inflate grounded coverage or silently trigger a second charged request. Record actual dependency versions, permissions and provider access in setup instructions. Preserve originals, expose partial failures and document backup/restore. These are acceptance requirements, not a claim of a completed or production-certified build. Add the domain, recovery and acceptance rules to AGENTS.md so future edits preserve them.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
Alternatives to building your own
no votes, no pay-to-list · just what's real
Questions about 100 Questions
Can you build your own 100 Questions with AI?
Partly. A personal CLI that asks the same questions across four model APIs and compares the answers is weekend-buildable, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work.
What does the 100 Questions build prompt cover?
The prompt starts with this scope: Run a reviewed five-question pilot across configured search-capable model providers, then compare brand mentions and owned-domain citations in an evidence-linked report. Full-product capabilities excluded from the comparison include: reliable orchestration and retries across four providers; normalized citations and evidence-linked metrics; competitor and missed-question extraction. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the 100 Questions prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this 100 Questions project take?
The catalogue estimate is multi-day for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing 100 Questions?
reliable orchestration and retries across four providers; normalized citations and evidence-linked metrics; competitor and missed-question extraction; stored point-in-time reports and comparisons; polished exports and action recommendations. They pay for a repeatable, frozen benchmark with provider failures handled, citations normalized, every metric tied to evidence, and a report that is ready to act on.
What price is this guide comparing against?
The recorded First benchmark plan is $9 one-time (one-time per benchmark), checked 2026-07-31. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.
What can I use instead of building 100 Questions?
Elmo: Prompt-by-prompt AI visibility with citations and competitors across the major engines; the queries still need model or scraper credentials. Check each option's license, hosting needs and feature limits.