Harvey

An enterprise AI assistant for legal work: research, drafting, review and diligence across a firm's document estate.

NOT REALLY · consider alternatives
price variesestimated build time one sittingreplaced by 0 people

You can absolutely build a local RAG assistant over your own PDFs in an afternoon, and for reading your lease or a vendor contract that is genuinely enough. Harvey is not sold on that loop. It is sold on trained-and-evaluated legal workflows, curated case law and regulatory sources under license, deployment that survives a law firm's security review, and the ability to put a name behind an output that a partner will bill against. The thing you cannot one-shot is the confidence to rely on the answer, which in legal work is the entire product. A personal replacement is fine for personal stakes and dangerous the moment money or a filing depends on it.

Build verification: not recorded. How we judge buildability

What you give up

  • Licensed primary law: case law, statutes, filings and regulatory corpora you cannot legally scrape together
  • Workflow products that have been evaluated by actual lawyers: diligence checklists, redline review, deposition prep
  • Firm-grade deployment: SSO, data residency, retention controls, audit logs, security questionnaires answered
  • Anyone to blame. Your hallucination is your malpractice exposure
  • Integration into the systems legal work actually lives in: DMS, iManage, Word, the review platform

Why people still pay

A law firm is not paying for text generation, it is paying for defensibility. The output has to be traceable to a licensed source, the deployment has to pass a security review that takes months, the workflows have to have been tested against how associates actually do diligence, and there has to be a vendor contract with indemnities when something goes wrong. Individual lawyers also cannot use a homemade tool on client matters without answering awkward questions about where the privileged data went. A local RAG box over your own files solves none of that and does not need to; it solves reading your own documents faster, which is a different and much smaller job.

Your build guide

The stack, security requirements, and agent rules for a focused replacement.

Before you start

  • An LLM API key, or a local model via Ollama if you would rather nothing leaves the machine
  • Your own documents: no licensed case law, no statute databases, no primary sources
  • Python 3.11 and a willingness to read the cited passage rather than trust the summary
01
Build a local document Q&A tool for my own contracts and PDFs. Python 3.11, single project, no web framework, no accounts, no telemetry.
02
Model source documents, editable pages, links or annotations, and exported versions; store source IDs and timestamps for each.
03
Interface for this AI for law firms and in-house legal teams workflow: a focused AI for law firms and in-house legal teams input, review, and export interface.
engineering roadmap

Implementation plan

1

Phase 1, architecture and data

Build a local document Q&A tool for my own contracts and PDFs. Python 3.11, single project, no web framework, no accounts, no telemetry. Model source documents, editable pages, links or annotations, and exported versions; store source IDs and timestamps for each.

2

Phase 2, implement

ingest PATH: walk a folder, extract text per page, chunk to roughly 800 tokens with 100 overlap, store chunk text plus doc name plus page number, embed and index. Skip files already ingested unless --force.

3

Phase 3, implement

docs: list ingested files, page counts, chunk counts.

4

Phase 4, review and output

In scope: local-only storage in ./index.db, deterministic chunking, a --model flag, plain text output. Print a one-line disclaimer after every answer: this is a reading aid over my own files, not legal advice, verify each citation.

5

Phase 5, recovery and acceptance

If An LLM API key, or a local model via Ollama if you would rather nothing leaves the machine is unavailable, keep the source record and show a recoverable error instead of a fabricated result. Verify this invariant with a saved fixture: An invalid input or interrupted operation must retain the source and show a recoverable state; exported records must reload with the same IDs. State the practical limit: Licensed primary law: case law, statutes, filings and regulatory corpora you cannot legally scrape together.

the pro prompt
Build a local document Q&A tool for my own contracts and PDFs. Python 3.11, single project, no web framework, no accounts, no telemetry.

Stack, no alternatives:
- CLI with Typer
- pypdf for text extraction, page numbers preserved
- SQLite with the sqlite-vec extension for vector storage
- OpenAI API for embeddings and answers, key from .env via python-dotenv, plus an --ollama flag that swaps to a local model at http://localhost:11434

Commands:
- ingest PATH: walk a folder, extract text per page, chunk to roughly 800 tokens with 100 overlap, store chunk text plus doc name plus page number, embed and index. Skip files already ingested unless --force.
- ask "QUESTION": retrieve top 12 chunks, then answer with an LLM that is instructed to answer only from the provided chunks and to say "not in these documents" when the answer is absent. Every claim must carry an inline citation like [contract.pdf p.4].
- sources "QUESTION": print the retrieved chunks verbatim with file and page, no LLM, so I can read the raw text.
- docs: list ingested files, page counts, chunk counts.

In scope: local-only storage in ./index.db, deterministic chunking, a --model flag, plain text output.
Out of scope: web UI, multi-user, cloud sync, any bundled case law or statute data, any attempt to cite external legal sources.

Print a one-line disclaimer after every answer: this is a reading aid over my own files, not legal advice, verify each citation.
Include a README with setup, .env.example, and a short section explaining that answers are only as good as the documents I ingested.

$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy AGENTS.md · generated from this app's build plan

prior art · use these instead of building, if you'd rather

No prior-art project is listed yet. Compare the scoped build with the paid product before choosing.

share on X ↗

Questions about Harvey

Can you build your own Harvey with AI?

A full replacement is not the recommended project. You can absolutely build a local RAG assistant over your own PDFs in an afternoon, and for reading your lease or a vendor contract that is genuinely enough. Harvey is not sold on that loop. It is sold on trained-and-evaluated legal workflows, curated case law and regulatory sources under license, deployment that survives a law firm's security review, and the ability to put a name behind an output that a partner will bill against. The thing you cannot one-shot is the confidence to rely on the answer, which in legal work is the entire product. A personal replacement is fine for personal stakes and dangerous the moment money or a filing depends on it.

What does the Harvey build prompt cover?

The prompt starts with this scope: Indexes a folder of your own contracts and PDFs locally, then answers questions about them with quotes and page citations so you can check every claim yourself. Full-product capabilities excluded from the comparison include: Licensed primary law: case law, statutes, filings and regulatory corpora you cannot legally scrape together; Workflow products that have been evaluated by actual lawyers: diligence checklists, redline review, deposition prep; Firm-grade deployment: SSO, data residency, retention controls, audit logs, security questionnaires answered. Follow the implementation plan and its prerequisites before expanding the build.

How do I use the prompt, AGENTS.md and agent skills?

Start with the Harvey prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.

How long will this Harvey project take?

The catalogue estimate is one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.

What would I give up by replacing Harvey?

Licensed primary law: case law, statutes, filings and regulatory corpora you cannot legally scrape together; Workflow products that have been evaluated by actual lawyers: diligence checklists, redline review, deposition prep; Firm-grade deployment: SSO, data residency, retention controls, audit logs, security questionnaires answered; Anyone to blame. Your hallucination is your malpractice exposure; Integration into the systems legal work actually lives in: DMS, iManage, Word, the review platform. A law firm is not paying for text generation, it is paying for defensibility. The output has to be traceable to a licensed source, the deployment has to pass a security review that takes months, the workflows have to have been tested against how associates actually do diligence, and there has to be a vendor contract with indemnities when something goes wrong. Individual lawyers also cannot use a homemade tool on client matters without answering awkward questions about where the privileged data went. A local RAG box over your own files solves none of that and does not need to; it solves reading your own documents faster, which is a different and much smaller job.

What can I use instead of building Harvey?

No alternative is listed in this entry yet. That is a gap in this catalogue, not proof that no suitable product exists. Compare the paid product and the proposed scope before committing to a build.

Every week, more subscriptions die.

New verdicts, new prompts, the week's most-doomed apps.
One email. Unsubscribe in one click.

last week:100 Questions · KINDA1of10 · KINDA1Password · KINDA+1090 more

free forever · no scanner spam · the prompt stays on the site, the deaths come to you

$weekly: what got a verdict, what died.