Skip to content
IntroductionHow it works, guides, and what you can transcribe.
QuickstartSubmit a link, wait for the job, export the transcript. Three requests.
AuthenticationAPI keys, scopes, Idempotency-Key and X-Request-Id.
JobsCreate, wait for, list, cancel and retry jobs.
TranscriptsThe transcript object and txt, srt, vtt, json exports.
BatchesUp to 50 videos in one request.
UploadsTranscribe your own audio and video files.
WebhooksGet called when a job or batch finishes; verify the signature.
ErrorsEvery error code, what it means and what to do.
Rate limitsSubmits, jobs in progress and reads per plan; the 429 response.
Pricing and credits1 credit per caption transcript, 2 per minute of AI transcription. Plans and extra credits.
SourcesYouTube, TikTok, your files and direct links: accepted URLs and modes.
MCPFind and transcribe videos from Claude, Cursor, Windsurf or your own agent.
OpenAPI SpecificationThe OpenAPI 3.1 spec for codegen and typed clients.
Claude CodeOne command adds TranscriptDock to Claude Code.
Claude DesktopAdd TranscriptDock as an MCP server in Claude Desktop's configuration file.
CursorAdd TranscriptDock as an MCP server in Cursor.
WindsurfAdd TranscriptDock to Windsurf so Cascade can find and transcribe videos.
OpenClawConnect TranscriptDock to OpenClaw autonomous agents.
19 results
Guides

TranscriptDock compared with the alternatives

An honest comparison with youtube-transcript-api, Supadata and AssemblyAI: what each one does, what it does not do, and when to pick it over TranscriptDock.


Short version: if you only need YouTube captions and can run Python on a residential IP, the open source youtube-transcript-api is free and good. If you need TikTok, a hosted endpoint, word timing, subtitle exports or an MCP server for Claude and Cursor, TranscriptDock or Supadata is the shorter path. If you already have audio files and want the best raw recognition, go straight to AssemblyAI.

Facts below were checked on 2026-09-16 against each project's public documentation or pricing page. Prices and limits change; the linked pages are authoritative.

Side by side#

FeatureTranscriptDockyoutube-transcript-apiSupadataAssemblyAI
SourcesYouTube, TikTok links; audio uploadsYouTube onlyYouTube, TikTok, Instagram, XAudio file or audio URL you supply
Existing captionsCreator and platform captions, word timing when the track has itYes (creator and auto-generated)YesNo, always runs recognition
AI transcriptionYouTube, TikTok, uploads and direct links, 2 credits per minuteNoYes, 2 credits per minuteYes, its core product
Hosted / needs infraHosted API + MCP serverPython library on your machine; cloud IPs are blocked without a proxyHosted APIHosted API
MCP serverRemote endpoint and npx packageNoYes (GitHub)No first-party server listed
ExportsJSON, SRT, VTT, TXTObjects with start and duration; SRT/WebVTT formattersJSONJSON, SRT, VTT
Free tier50 credits, no cardFree (open source)100 credits per monthTrial credits
Entry price$19 per month, 2,000 credits$0 (plus a proxy if you run in the cloud)$5 per month billed annually, 300 credits$0.15 per audio hour (Universal-2)

youtube-transcript-api#

A Python library that fetches the transcript YouTube already has for a video, with no API key and no browser. Its README notes that YouTube "has started blocking most IPs that are known to belong to cloud providers", so production use from AWS, GCP or a VPS needs a rotating residential proxy, which the library supports. It does not cover TikTok and does not do AI transcription.

  • Pick it when: YouTube only, a script or notebook, you control the IP.
  • Pick TranscriptDock when: you need a hosted endpoint, TikTok, batch jobs, webhooks, exports or MCP.

Supadata#

A hosted transcript API covering YouTube, TikTok, Instagram and X, with optional AI transcription for videos that have no captions (2 credits per generated minute per its pricing page) and integrations for Make, Zapier, n8n and an MCP server. It is the closest alternative in shape.

  • Pick it when: you need Instagram or X today, or their no-code integrations.
  • Pick TranscriptDock when: you want typed error codes with a retryable flag, idempotent submits, a versioned OpenAPI contract, per-job credit caps, and a transcript cache so re-exports never cost a second request.

AssemblyAI#

A speech-to-text API: you give it an audio file or URL and it returns a transcript with word timings. It does not ingest YouTube or TikTok links or use existing captions, so for social video you would still fetch and extract the audio yourself. Its listed price is $0.15 per audio hour for Universal-2 and $0.21 for Universal-3.5 Pro.

  • Pick it when: you already have audio and want to own the pipeline.
  • Pick TranscriptDock when: the input is a link. Captions are fetched first (no recognition cost) and recognition is only used when a video has none.