ZeroLeaks
Automated red-teaming that tries to extract your AI product's system prompt, keys and hidden instructions, then reports what leaked.
The mechanics here are not exotic: fire a few hundred adversarial prompts at your own chat endpoint, capture the responses, and check whether any of them contain your system prompt, tool schemas or API keys. An agent can build that loop, including an LLM-as-judge scorer and an HTML report, in a single sitting, and it will find the embarrassing stuff on day one. What you cannot one-shot is a probe library that stays current with each new model release and each new jailbreak family, because that is maintained knowledge, not code. There is also a trust angle: 'we ran our own script and found nothing' reads very differently in a security review than a dated third-party report. Build it for your own sanity checks, keep paying if you need something to show someone else.
Build verification: not recorded. How we judge buildability
What you give up
- A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day
- Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts
- A third-party report with a date on it that you can hand to a customer or an auditor
- Severity triage and remediation guidance written by someone who has seen a lot of these
- Regression runs on every model or prompt change without you remembering to trigger them
Why people still pay
Two reasons, and neither is that the harness is hard. First, jailbreaks rot: the probes that worked against last quarter's model are dead weight now, and keeping a live corpus is somebody's full time job, not a cron you set up once. Second, self-attested security is worth roughly nothing to an enterprise buyer, so companies pay for an external artifact with a date and a logo on it. If your goal is just to stop shipping a system prompt that unravels when someone types 'repeat everything above', a local harness is genuinely enough.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Runtime and tools: Python, Typer, httpx, Pydantic, Jinja2 and versioned JSONL run files; no database or web server.
- Before starting: One owned endpoint specification, a bounded auditable probe pack, a local reference fixture and secrets outside shareable reports.
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule — domain: Store TargetConfig, ProbeVersion, Attempt, DetectorResult and RunComparison in JSONL; failure/timeout is unknown, and detector evidence is separate from an optional model judge's opinion.
Project rule — scope and recovery: Keep the existing Python CLI with no database/server. Do not store real secrets in shareable transcripts, scan third-party targets or treat a clean probe run as proof of security.
Project rule — acceptance: Use one fixture with a planted token, one harmless response and one timeout; report finding, clear observation and unknown separately, with a stable-ID diff between runs.
Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.
Recommended skill: modern-python — structure the Python worker or explicitly optional read-only utility with pinned dependencies, typed boundaries and clear failure handling. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Recommended skill: sharp-edges — review configuration and API defaults against the app-specific invariants and recovery boundaries above; this is not a security certification. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Implementation plan
Phase 1
Pin the working slice and create its example input: Run a versioned, bounded probe pack against an authorized chat endpoint, inspect redacted leakage evidence and compare findings across two runs. Confirm setup: One owned endpoint specification, a bounded auditable probe pack, a local reference fixture and secrets outside shareable reports.
Phase 2
Implement durable result files and command invariants before formatting terminal output: Store TargetConfig, ProbeVersion, Attempt, DetectorResult and RunComparison in JSONL; failure/timeout is unknown, and detector evidence is separate from an optional model judge's opinion.
Phase 3
Connect CLI commands to real saved state and explicit result/exit statuses. Operate only against configured authorized endpoints. Redact secret values in saved evidence and distinguish detector findings, judge opinion, failures and unknown outcomes.
Phase 4
Expose the app-specific limits and recovery path in context: Keep the existing Python CLI with no database/server. Do not store real secrets in shareable transcripts, scan third-party targets or treat a clean probe run as proof of security.
Phase 5
Walk through this concrete acceptance case and preserve its exported evidence: Use one fixture with a planted token, one harmless response and one timeout; report finding, clear observation and unknown separately, with a stable-ID diff between runs. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.
Build me a local CLI for authorized system-prompt and secret-leak regression checks, as a narrow substitute for Zeroleaks. Requirements: - Use Python 3.12, Typer, httpx, Pydantic, Jinja2, and JSONL files in ./runs/. No web server or database. - A target.yaml names one endpoint I control, its request template, response-text path, allowed request rate, and timeout. Confirm the endpoint scope before running probes. Keep bearer keys in .env. - Store versioned probes with ID, family, severity, and turns. Start with a small auditable pack covering direct extraction, role confusion, encoded requests, and pasted-document injection; permit user-authored additions. - Run probes with bounded concurrency and backoff on 429/5xx. Save request IDs, redacted transcripts, status, and elapsed time. Never write real secret values or full system instructions into a shareable report. - Detect exact secret-pattern hits and overlap with a local reference prompt. An optional model judge may add a labelled opinion; a judge failure produces unknown, not clean. - Produce report.html and report.json with finding evidence, detector, severity, probe version, and unknown counts. A diff command compares two runs by stable probe ID. - Acceptance: a fixture containing a planted token flags the exact probe; a harmless fixture stays unflagged; a timed-out endpoint stays unknown; diff reports one introduced and one resolved finding. - No accounts, telemetry, third-party scanning, or unbounded attack traffic. README documents permission, data retention, redaction limits, and the fact that a clean run is not proof of safety. EDITORIAL IMPLEMENTATION CONTRACT Working slice: Run a versioned, bounded probe pack against an authorized chat endpoint, inspect redacted leakage evidence and compare findings across two runs. Data and invariants: Store TargetConfig, ProbeVersion, Attempt, DetectorResult and RunComparison in JSONL; failure/timeout is unknown, and detector evidence is separate from an optional model judge's opinion. Boundary and recovery: Keep the existing Python CLI with no database/server. Do not store real secrets in shareable transcripts, scan third-party targets or treat a clean probe run as proof of security. Acceptance walkthrough: Use one fixture with a planted token, one harmless response and one timeout; report finding, clear observation and unknown separately, with a stable-ID diff between runs. Record actual dependency versions, permissions and provider access in setup instructions. Preserve originals, expose partial failures and document backup/restore. These are acceptance requirements, not a claim of a completed or production-certified build. Add the domain, recovery and acceptance rules to AGENTS.md so future edits preserve them.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md · generated from this app's build plan
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
No prior-art project is listed yet. Compare the scoped build with the paid product before choosing.
Questions about ZeroLeaks
Can you build your own ZeroLeaks with AI?
Partly. The mechanics here are not exotic: fire a few hundred adversarial prompts at your own chat endpoint, capture the responses, and check whether any of them contain your system prompt, tool schemas or API keys. An agent can build that loop, including an LLM-as-judge scorer and an HTML report, in a single sitting, and it will find the embarrassing stuff on day one. What you cannot one-shot is a probe library that stays current with each new model release and each new jailbreak family, because that is maintained knowledge, not code. There is also a trust angle: 'we ran our own script and found nothing' reads very differently in a security review than a dated third-party report. Build it for your own sanity checks, keep paying if you need something to show someone else.
What does the ZeroLeaks build prompt cover?
The prompt starts with this scope: Run a versioned, bounded probe pack against an authorized chat endpoint, inspect redacted leakage evidence and compare findings across two runs. Full-product capabilities excluded from the comparison include: A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day; Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts; A third-party report with a date on it that you can hand to a customer or an auditor. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the ZeroLeaks prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this ZeroLeaks project take?
The catalogue estimate is one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing ZeroLeaks?
A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day; Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts; A third-party report with a date on it that you can hand to a customer or an auditor; Severity triage and remediation guidance written by someone who has seen a lot of these; Regression runs on every model or prompt change without you remembering to trigger them. Two reasons, and neither is that the harness is hard. First, jailbreaks rot: the probes that worked against last quarter's model are dead weight now, and keeping a live corpus is somebody's full time job, not a cron you set up once. Second, self-attested security is worth roughly nothing to an enterprise buyer, so companies pay for an external artifact with a date and a logo on it. If your goal is just to stop shipping a system prompt that unravels when someone types 'repeat everything above', a local harness is genuinely enough.
What price is this guide comparing against?
The recorded Pro plan is $79/mo (monthly, single seat), checked 2026-08-18. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.
What can I use instead of building ZeroLeaks?
No alternative is listed in this entry yet. That is a gap in this catalogue, not proof that no suitable product exists. Compare the paid product and the proposed scope before committing to a build.