Hedy
Listens to your meetings, classes and calls live and feeds you suggested questions, context and notes while they are still happening.
The pieces are all commodity now: capture audio, stream it through a local Whisper variant, keep a rolling transcript, and every few seconds ask a model for the next good question. An agent will get you that in one sitting on a laptop, and the transcript-plus-summary half will genuinely be fine. What it will not get you is a phone in your pocket that survives being backgrounded during a two hour lecture, or sub-second suggestions that arrive before the moment passes. Live assistance is a latency and UX problem more than an AI problem, and that is exactly the part that takes weeks of fiddling. Build it if you mostly sit at a desk; keep paying if you mostly do not.
Build verification: not recorded. How we judge buildability
What you give up
- Mobile apps that keep recording reliably when the screen is off, which is where most of these sessions actually happen
- Tuned latency: your DIY suggestions arrive a beat late, and a beat late is useless in conversation
- Speaker diarization and meeting-type presets that shape the assistant for a sales call vs a lecture vs a doctor visit
- Calendar and conferencing integrations that join and label sessions for you
- Hosted, searchable history across every session with no laptop babysitting
Why people still pay
Because real-time is unforgiving. A post-hoc summarizer can be sloppy and still useful, but a live copilot that lags four seconds or drops audio when you switch apps is worse than nothing, and people pay to not think about that. The mobile side compounds it: background audio on iOS is a permissions and lifecycle minefield that nobody wants to solve for themselves. Add the meeting-type presets and calendar hookups and the subscription is mostly buying tuning you would otherwise do by hand for a month.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Runtime and tools: Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor.
- Before starting: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key.
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule — domain: Store Session, AudioWindow, TranscriptRevision and Suggestion; suggestions cite the latest finalized transcript window and never pretend to know unrecorded context or speaker intent.
Project rule — scope and recovery: Require clear participant permission and show recording status. Start with one supported capture route; suggestions are prompts for the user, not hidden real-time coaching guarantees.
Project rule — acceptance: Pause recording, correct a negation and resume; invalidate questions based on the old wording and do not process audio while paused.
Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.
Recommended skill: modern-python — structure the Python worker or explicitly optional read-only utility with pinned dependencies, typed boundaries and clear failure handling. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Recommended skill: web-design-guidelines — review keyboard access, focus, validation, error recovery and the readable work/review interface or HTML report. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Implementation plan
Phase 1
Pin the working slice and create its example input: Transcribe short meeting windows into a private side panel and offer optional, evidence-linked follow-up questions that the user can dismiss. Confirm setup: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key.
Phase 2
Implement persistence and write-time invariants before decorating the UI: Store Session, AudioWindow, TranscriptRevision and Suggestion; suggestions cite the latest finalized transcript window and never pretend to know unrecorded context or speaker intent.
Phase 3
Connect the working view to real saved state. Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words.
Phase 4
Expose the app-specific limits and recovery path in context: Require clear participant permission and show recording status. Start with one supported capture route; suggestions are prompts for the user, not hidden real-time coaching guarantees.
Phase 5
Walk through this concrete acceptance case and preserve its exported evidence: Pause recording, correct a negation and resume; invalidate questions based on the old wording and do not process audio while paused. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.
Build the following focused alternative to Hedy. This is a deliberately limited personal or small-team substitute, not parity with the paid service. WORKING SLICE Transcribe short meeting windows into a private side panel and offer optional, evidence-linked follow-up questions that the user can dismiss. SETUP AND ARCHITECTURE Use Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor. Prerequisites: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key. Before integrating anything, record actual versions and permissions, plus model files or provider limits only where used, in the README; make unavailable dependencies visible rather than simulating success. DOMAIN MODEL AND INVARIANTS Store Session, AudioWindow, TranscriptRevision and Suggestion; suggestions cite the latest finalized transcript window and never pretend to know unrecorded context or speaker intent. IMPLEMENTATION CONTRACT Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words. Provide an input/setup view, the main work view, and a review/export view appropriate to this workflow. Preserve the last saved state if a job or save fails. Include empty, loading, permission-denied, partial and retryable-error states. Log identifiers and error categories without secret values or unnecessary private content. APP-SPECIFIC BOUNDARY AND RECOVERY Require clear participant permission and show recording status. Start with one supported capture route; suggestions are prompts for the user, not hidden real-time coaching guarantees. ACCEPTANCE SCENARIO Pause recording, correct a negation and resume; invalidate questions based on the old wording and do not process audio while paused. Also reopen the app after an interrupted operation, confirm the saved record/export remains inspectable, and document the recovery action. These are implementation acceptance requirements, not a claim that this guide has been tested. DELIVERY Deliver a runnable repository with migrations or project-format versioning, a non-sensitive example, environment/permission setup, the exact manual acceptance steps, and a backup/export-and-restore walkthrough. Implement the working slice before optional integrations; list any deferred paid-product capabilities honestly. Do not add capabilities outside the working slice just to resemble the original product. PROJECT RULES FOR AGENTS.md Keep the domain invariants above executable at the write boundary. Propose scope changes before adding providers or permissions. Never fabricate source evidence, publish results, identity matches or successful delivery. Preserve user originals and require an explicit confirmation for destructive changes or external publication.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
No prior-art project is listed yet. Compare the scoped build with the paid product before choosing.
Questions about Hedy
Can you build your own Hedy with AI?
Partly. The pieces are all commodity now: capture audio, stream it through a local Whisper variant, keep a rolling transcript, and every few seconds ask a model for the next good question. An agent will get you that in one sitting on a laptop, and the transcript-plus-summary half will genuinely be fine. What it will not get you is a phone in your pocket that survives being backgrounded during a two hour lecture, or sub-second suggestions that arrive before the moment passes. Live assistance is a latency and UX problem more than an AI problem, and that is exactly the part that takes weeks of fiddling. Build it if you mostly sit at a desk; keep paying if you mostly do not.
What does the Hedy build prompt cover?
The prompt starts with this scope: Transcribe short meeting windows into a private side panel and offer optional, evidence-linked follow-up questions that the user can dismiss. Full-product capabilities excluded from the comparison include: Mobile apps that keep recording reliably when the screen is off, which is where most of these sessions actually happen; Tuned latency: your DIY suggestions arrive a beat late, and a beat late is useless in conversation; Speaker diarization and meeting-type presets that shape the assistant for a sales call vs a lecture vs a doctor visit. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the Hedy prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this Hedy project take?
The catalogue estimate is a weekend for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing Hedy?
Mobile apps that keep recording reliably when the screen is off, which is where most of these sessions actually happen; Tuned latency: your DIY suggestions arrive a beat late, and a beat late is useless in conversation; Speaker diarization and meeting-type presets that shape the assistant for a sales call vs a lecture vs a doctor visit; Calendar and conferencing integrations that join and label sessions for you; Hosted, searchable history across every session with no laptop babysitting. Because real-time is unforgiving. A post-hoc summarizer can be sloppy and still useful, but a live copilot that lags four seconds or drops audio when you switch apps is worse than nothing, and people pay to not think about that. The mobile side compounds it: background audio on iOS is a permissions and lifecycle minefield that nobody wants to solve for themselves. Add the meeting-type presets and calendar hookups and the subscription is mostly buying tuning you would otherwise do by hand for a month.
What price is this guide comparing against?
The recorded Pro plan is $12.99/mo (monthly per user), checked 2026-08-18. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.
What can I use instead of building Hedy?
No alternative is listed in this entry yet. That is a gap in this catalogue, not proof that no suitable product exists. Compare the paid product and the proposed scope before committing to a build.