Pangram
A classifier that scores whether a piece of text was written by an AI model.
The interface is a text box and a percentage, which is exactly the kind of thing that tricks you into thinking it is a weekend project. The product is not the text box: it is a classifier trained on a very large, continuously refreshed corpus of human writing paired with output from every model release, tuned hard against false positives because accusing a real person of cheating is the failure mode that ends the company. You can absolutely build a local detector from perplexity and burstiness features in an afternoon, and it will be confidently wrong often enough to be useless for any decision that matters. Nobody outside your own head will accept your homemade score, and the calibration drifts every time a new frontier model ships. Build it to understand the problem, not to rely on it.
Build verification: not recorded. How we judge buildability
What you give up
- Calibration: a real false positive rate you can quote, instead of a vibe
- Coverage of new models, which changes every few weeks whether you update or not
- Sentence-level and mixed-authorship detection rather than one blunt document score
- Any external credibility, since a self-built score persuades exactly zero teachers, editors or clients
- Throughput, batch uploads, API access and document parsing
Why people still pay
Because the number has to be defensible to someone else. Institutions and publishers are not buying a classifier, they are buying a third party willing to stand behind a false positive rate, plus continuous retraining against whatever model came out last month. A local heuristic detector gives you a plausible-sounding percentage with no error bars, which is worse than nothing when the outcome is an accusation.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Local Python with torch and transformers
- A couple of GB of disk for small model weights, CPU works but is slow
- Your own labelled samples of human and AI text if you want any idea of accuracy
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule, data: Model source documents, editable pages or notes, internal links, revisions, and exports; retain source IDs and timestamps.
Project rule, behavior: Features to compute per submission:.
Project rule, recovery: On invalid input or interrupted processing, keep the original record, show the failed step, and permit a safe retry.
Implementation plan
Phase 1, architecture and data
Stack, no substitutions: Python 3.11, FastAPI + uvicorn, transformers + torch, a single server-rendered HTML page with vanilla JS. No build step, no database, no Docker. Model source documents, editable pages or notes, internal links, revisions, and exports; retain source IDs and timestamps.
Phase 2, implement
Features to compute per submission:.
Phase 3, implement
Burstiness: standard deviation of per-sentence mean log-probability.
Phase 4, review and output
Rank-based signal: fraction of tokens that were in the model's top-10 predictions. Scoring runs entirely locally with gpt2 (small) loaded once at startup via transformers. Download on first run and cache it.
Phase 5, recovery and acceptance
On invalid input or interrupted processing, keep the original record, show the failed step, and permit a safe retry. Verify this invariant with a saved fixture: A classifier score cannot be presented as proof of authorship; a failed detector run remains unknown and never overwrites source text. State the practical limit: Calibration: a real false positive rate you can quote, instead of a vibe.
Build a local AI-text-detection playground so I can see how weak naive detection actually is. No cloud calls, no accounts, no telemetry. Stack, no substitutions: Python 3.11, FastAPI + uvicorn, transformers + torch, a single server-rendered HTML page with vanilla JS. No build step, no database, no Docker. Core loop: 1. One page with a large textarea and an Analyze button. 2. POST /analyze takes the text and returns a JSON report. 3. Scoring runs entirely locally with gpt2 (small) loaded once at startup via transformers. Download on first run and cache it. Features to compute per submission: - Mean token log-probability and perplexity under gpt2. - Burstiness: standard deviation of per-sentence mean log-probability. - Rank-based signal: fraction of tokens that were in the model's top-10 predictions. - Style stats: sentence length mean and variance, type-token ratio, punctuation counts, count of common LLM filler phrases from a hardcoded list. Output: a 0-100 score from a simple weighted formula over those features with the weights in a single WEIGHTS dict at the top of the file, per-sentence heat colouring in the page, and a full feature table so I can see what drove the score. Calibration, and be honest about it: - Add a CLI command: python calibrate.py --human dir --ai dir that runs the features over two folders of .txt files, prints mean and spread per class, and prints the ROC AUC of the current weights. - The web page must show a fixed banner: this is an uncalibrated heuristic, not evidence, false positives are common. Explicitly out of scope: user accounts, saved history, PDF or DOCX parsing, any hosted API, model fine-tuning, mixed-authorship segmentation, batch uploads. Deliver: main.py, calibrate.py, features.py, templates/index.html, requirements.txt, and a README with run instructions and one paragraph stating plainly why this cannot match a trained commercial detector. No .env needed since there are no secrets; say so in the README.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy AGENTS.md · generated from this app's build plan
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
No prior-art project is listed yet. Compare the scoped build with the paid product before choosing.
Questions about Pangram
Can you build your own Pangram with AI?
A full replacement is not the recommended project. The interface is a text box and a percentage, which is exactly the kind of thing that tricks you into thinking it is a weekend project. The product is not the text box: it is a classifier trained on a very large, continuously refreshed corpus of human writing paired with output from every model release, tuned hard against false positives because accusing a real person of cheating is the failure mode that ends the company. You can absolutely build a local detector from perplexity and burstiness features in an afternoon, and it will be confidently wrong often enough to be useless for any decision that matters. Nobody outside your own head will accept your homemade score, and the calibration drifts every time a new frontier model ships. Build it to understand the problem, not to rely on it.
What does the Pangram build prompt cover?
The prompt starts with this scope: Paste text, score it locally with a small language model's token log-probabilities plus a few style statistics, and get a hand-wavy human-or-machine guess with a loud accuracy disclaimer. Full-product capabilities excluded from the comparison include: Calibration: a real false positive rate you can quote, instead of a vibe; Coverage of new models, which changes every few weeks whether you update or not; Sentence-level and mixed-authorship detection rather than one blunt document score. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the Pangram prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this Pangram project take?
The catalogue estimate is one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing Pangram?
Calibration: a real false positive rate you can quote, instead of a vibe; Coverage of new models, which changes every few weeks whether you update or not; Sentence-level and mixed-authorship detection rather than one blunt document score; Any external credibility, since a self-built score persuades exactly zero teachers, editors or clients; Throughput, batch uploads, API access and document parsing. Because the number has to be defensible to someone else. Institutions and publishers are not buying a classifier, they are buying a third party willing to stand behind a false positive rate, plus continuous retraining against whatever model came out last month. A local heuristic detector gives you a plausible-sounding percentage with no error bars, which is worse than nothing when the outcome is an accusation.
What price is this guide comparing against?
The recorded Individual plan is $20/mo (monthly, 300,000 words/month), checked 2026-08-18. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.
What can I use instead of building Pangram?
No alternative is listed in this entry yet. That is a gap in this catalogue, not proof that no suitable product exists. Compare the paid product and the proposed scope before committing to a build.