Transkriptor
Audio and video transcription, translation, summaries, and collaboration
The visible transcription loop is buildable, but a credible replacement needs more than the first screen. Transkriptor earns its keep through capture, integrations, reliability, so expect a weekend or multi-day build and a narrower personal scope.
Build verification: not recorded. How we judge buildability
What you give up
- meeting-bot auto-join
- live multi-speaker accuracy
- calendar and CRM integrations
- cross-call team analytics
- production codecs, rendering speed, and media templates
Why people still pay
Transkriptor: Customers pay for automatic capture, dependable speaker handling, search across calls, and notes arriving without manual file wrangling.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Runtime and tools: Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor.
- Before starting: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key.
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule — domain: Store MediaSource, TranscriptionJob, Segment, Correction and Export; corrections are layered over machine output and a failed rerun cannot erase the last reviewed transcript.
Project rule — scope and recovery: Start with one language/model configuration and manual speaker labels. Translation and live meeting joining are separate; never promise universal accuracy or supported formats without checking the decoder.
Project rule — acceptance: Import a silent segment and a corrupt file, then edit a valid transcript; show no-speech/decoding failures distinctly and preserve the reviewed export.
Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.
Recommended skill: modern-python — structure the Python worker or explicitly optional read-only utility with pinned dependencies, typed boundaries and clear failure handling. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Recommended skill: web-design-guidelines — review keyboard access, focus, validation, error recovery and the readable work/review interface or HTML report. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Implementation plan
Phase 1
Pin the working slice and create its example input: Import audio/video, extract audio, transcribe locally, correct text/speakers and export ordered text/subtitles with an optional reviewed summary. Confirm setup: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key.
Phase 2
Implement persistence and write-time invariants before decorating the UI: Store MediaSource, TranscriptionJob, Segment, Correction and Export; corrections are layered over machine output and a failed rerun cannot erase the last reviewed transcript.
Phase 3
Connect the working view to real saved state. Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words.
Phase 4
Expose the app-specific limits and recovery path in context: Start with one language/model configuration and manual speaker labels. Translation and live meeting joining are separate; never promise universal accuracy or supported formats without checking the decoder.
Phase 5
Walk through this concrete acceptance case and preserve its exported evidence: Import a silent segment and a corrupt file, then edit a valid transcript; show no-speech/decoding failures distinctly and preserve the reviewed export. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.
Build the following focused alternative to Transkriptor. This is a deliberately limited personal or small-team substitute, not parity with the paid service. WORKING SLICE Import audio/video, extract audio, transcribe locally, correct text/speakers and export ordered text/subtitles with an optional reviewed summary. SETUP AND ARCHITECTURE Use Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor. Prerequisites: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key. Before integrating anything, record actual versions and permissions, plus model files or provider limits only where used, in the README; make unavailable dependencies visible rather than simulating success. DOMAIN MODEL AND INVARIANTS Store MediaSource, TranscriptionJob, Segment, Correction and Export; corrections are layered over machine output and a failed rerun cannot erase the last reviewed transcript. IMPLEMENTATION CONTRACT Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words. Provide an input/setup view, the main work view, and a review/export view appropriate to this workflow. Preserve the last saved state if a job or save fails. Include empty, loading, permission-denied, partial and retryable-error states. Log identifiers and error categories without secret values or unnecessary private content. APP-SPECIFIC BOUNDARY AND RECOVERY Start with one language/model configuration and manual speaker labels. Translation and live meeting joining are separate; never promise universal accuracy or supported formats without checking the decoder. ACCEPTANCE SCENARIO Import a silent segment and a corrupt file, then edit a valid transcript; show no-speech/decoding failures distinctly and preserve the reviewed export. Also reopen the app after an interrupted operation, confirm the saved record/export remains inspectable, and document the recovery action. These are implementation acceptance requirements, not a claim that this guide has been tested. DELIVERY Deliver a runnable repository with migrations or project-format versioning, a non-sensitive example, environment/permission setup, the exact manual acceptance steps, and a backup/export-and-restore walkthrough. Implement the working slice before optional integrations; list any deferred paid-product capabilities honestly. Do not add capabilities outside the working slice just to resemble the original product. PROJECT RULES FOR AGENTS.md Keep the domain invariants above executable at the write boundary. Propose scope changes before adding providers or permissions. Never fabricate source evidence, publish results, identity matches or successful delivery. Preserve user originals and require an explicit confirmation for destructive changes or external publication.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md · generated from this app's build plan
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
Transkriptor pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| free | $0/user | $0/user | 90 transcription minutes per month. |
| lite | $9.99/user | — | 300 transcription minutes per month.A numeric annual price was not exposed on the live public page. |
| pro | $19.99/user | $8.33/user | 2,400 transcription minutes per month.$99.99 billed yearly. |
| team | $30/user | $20/user | 3,000 transcription minutes per seat/month.$240 per seat billed yearly. |
| bulk 100 hours | $60/workspace | $30/workspace | 100 transcription hours in the billing allocation.$360 billed yearly on annual billing. |
| bulk 250 hours | $150/workspace | $75/workspace | 250 transcription hours in the billing allocation.$900 billed yearly on annual billing. |
| bulk 500 hours | $300/workspace | $150/workspace | 500 transcription hours in the billing allocation.$1,800 billed yearly on annual billing. |
| bulk 1,000 hours | $600/workspace | $300/workspace | 1,000 transcription hours in the billing allocation.$3,600 billed yearly on annual billing. |
free tier90 transcription minutes per month.
billingmonthly + annual; subscriptions auto-renew
hidden costsUnused monthly transcription minutes expire after 1 month rather than accumulating. Education pricing is 50% off for eligible users.
pricing sources checked 2026-08-14 · pricing source ↗
Questions about Transkriptor
Can you build your own Transkriptor with AI?
Partly. The visible transcription loop is buildable, but a credible replacement needs more than the first screen. Transkriptor earns its keep through capture, integrations, reliability, so expect a weekend or multi-day build and a narrower personal scope.
What does the Transkriptor build prompt cover?
The prompt starts with this scope: Import audio/video, extract audio, transcribe locally, correct text/speakers and export ordered text/subtitles with an optional reviewed summary. Full-product capabilities excluded from the comparison include: meeting-bot auto-join; live multi-speaker accuracy; calendar and CRM integrations. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the Transkriptor prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this Transkriptor project take?
The catalogue estimate is multi-day for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing Transkriptor?
meeting-bot auto-join; live multi-speaker accuracy; calendar and CRM integrations; cross-call team analytics; production codecs, rendering speed, and media templates. Transkriptor: Customers pay for automatic capture, dependable speaker handling, search across calls, and notes arriving without manual file wrangling.
What can I use instead of building Transkriptor?
The prior-art section lists whisper.cpp as starting points. Review their current scope, license and maintenance before adopting one.