Reference
Sources
YouTube, TikTok, your files and direct links: accepted URLs, what each mode does, and limits.
Whatever the source, you get the same transcript object: text, timed segments, and word timings when AI transcription produced it. Limits for every source: 2 hours and 250 MB.
Sources#
| Source | Accepted links | captions_only | auto / transcribe |
|---|---|---|---|
| YouTube | youtube.com/watch?v=…, youtu.be/…, /shorts/…, /live/… | Creator or automatic captions, 1 credit | Paid plans |
| TikTok | tiktok.com/@user/video/…, vm.tiktok.com/… share links | Creator captions, 1 credit | Paid plans |
| Your file | mp3, wav, m4a, ogg, aac, mp4, webm via POST /v1/uploads | Not available | Paid plans |
| Direct link | Any public https URL that serves an audio or video file | Not available | Paid plans |
GET /v1/capabilities with your key returns this table for your plan, so a client can show only what will work.
Captions#
- Public videos only. Private, members-only and age-gated videos fail with a
SOURCE_*code. - YouTube has two kinds: captions the creator uploaded and YouTube's automatic ones.
caption_preferencepicks; the default takes the creator's when present. caption_languagesasks for specific tracks, e.g.["es", "en"].POST /v1/language-discoverieslists what a YouTube video has.- Timing is per caption cue (
timing_granularity: "segment").
AI transcription#
- Listens to the audio; works with or without captions. Detects the language, or takes a hint in
language. - Returns word timings and confidence (
timing_granularity: "word"). - Takes roughly a quarter of the media length. Use
?wait=25or a webhook rather than tight polling. - 2 credits per started minute, 1 minute minimum; cap a job with
max_credits.