The interactive reference needs JavaScript. The same reference as plain text: /docs?format=md · raw spec: /openapi.yaml.
# Transcribr API — API reference Base URL: https://api.staging.audiotranscribr.com Interactive version (needs a browser): https://api.staging.audiotranscribr.com/docs · Machine-readable spec: https://api.staging.audiotranscribr.com/openapi.yaml Turn audio or video into text. Send a URL or upload a file, poll until the job is `completed`, then download the transcript as **JSON, SRT, VTT, TXT or PDF**. Speaker labels are included on every plan. ## Quickstart 1. **Get a key** — create a free account at [audiotranscribr.com/app](https://audiotranscribr.com/app) (email and a one-time code, no card; your first transcription is free, then 5 minutes a month), open **API & connectors** and create a key. Paid plans are on the [pricing page](https://api.staging.audiotranscribr.com/billing/pricing). 2. **Create a transcription** ```bash curl -s -X POST "$API/v1/transcriptions" -H "x-api-key: $KEY" \ -d '{"url": "https://audiotranscribr.com/samples/clip-22-451-01.mp3", "name": "Sample clip"}' ``` That URL is a one-minute public-domain sample we host (US Supreme Court oral argument; five more at `clip-22-451-02.mp3` … `-06.mp3`), so the command runs as-is. To upload a file instead, omit `url`; the response contains `upload.url` — `PUT` the bytes there: ```bash curl -X PUT --upload-file interview.m4a "<upload.url>" ``` 3. **Poll** `GET /v1/transcriptions/{id}` until `status` is `completed` (usually seconds; roughly 1 minute per hour of audio). The response then includes a presigned download link per format. 4. **Download** — use those links, or `GET /v1/transcriptions/{id}/result?format=pdf`. ## Using it from an AI agent (MCP) The API is also a remote [MCP](https://modelcontextprotocol.io) server (Streamable HTTP) with two doors: `POST /v1/mcp` takes the same `x-api-key` header (Claude Code, Cursor, VS Code, scripts), and `POST /mcp` takes an OAuth sign-in (Claude.ai and ChatGPT connectors: add the URL, sign in with your account email and a one-time code, no key). Tools: `transcribe_url`, `create_upload`, `get_transcription`, `get_transcript_text`, `list_transcriptions`, `get_account`, `delete_transcription`. ```bash claude mcp add --transport http -s user audiotranscribr "$API/v1/mcp" --header "x-api-key: $KEY" ``` Manage keys, usage and billing in the dashboard at https://audiotranscribr.com/app (sign in with your email, no password). A plain-text summary of this API for AI assistants is at [`/llms.txt`](/llms.txt). ## Good to know - **Auth**: send your key in the `x-api-key` header. A missing, wrong or cancelled key returns `403`. - **Statuses**: `awaiting_upload` → `queued` → `transcribing` → `completed` | `failed`. - **Errors** are JSON: `{"error": {"code": "...", "message": "..."}}`. `402 quota_exceeded` means the plan's included minutes are used up (plans with overage never return it). - **Rate limits**: each plan has a requests-per-second limit, a monthly request quota and a cap on transcriptions in progress at once (see [pricing](https://api.staging.audiotranscribr.com/billing/pricing)). Any `429` carries a `Retry-After` header — honour it and back off exponentially. Codes: `rate_limited`, `request_quota_exceeded`, `too_many_active_jobs`, `blocked`. - **Lost your key?** [Recover it by email](https://api.staging.audiotranscribr.com/billing/recover), or rotate it with `POST /v1/billing/rotate-key`. - **URLs** must point straight at the file (YouTube, Spotify, Drive and other page links are rejected up front). - **Languages**: see the section below — 35 languages detected automatically, 60+ with an explicit `language` tag. - **Postman**: a ready-to-run collection in the public workspace at https://www.postman.com/transcribr-5195959/audiotranscribr. - **Limits**: files up to 5 GB by upload or URL, audio or video (large videos have their audio track extracted first, which adds a few minutes). Multi-hour recordings are fine (a 2.7-hour file is in our test set and completes in under a minute); a job still processing after 1 hour is cancelled, marked `failed`, and not billed. Source media is deleted after 7 days, transcripts after 30 days. - **Usage** is metered per second of audio, per billing period — see `GET /v1/account`. Failed jobs are not billed. ## Languages Leave `language` out and the language is detected automatically from the audio. Detection covers 35 languages: Bulgarian, Catalan, Chinese, Czech, Danish, Dutch, English, Estonian, Finnish, Flemish, French, German, Greek, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Latvian, Lithuanian, Malay, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Spanish, Swedish, Thai, Turkish, Ukrainian, Vietnamese. The response's `language` field reports what was detected (`fr`, `pt`, …). Pass `language` (a BCP-47 tag such as `es`, `pt-BR`, `zh-TW`, `ar`, `he`) to skip detection — faster, and the only way to reach the languages detection does not cover. With an explicit tag, 60+ languages are available, including Afrikaans, Arabic, Armenian, Assamese, Belarusian, Bengali, Bosnian, Croatian, Georgian, Gujarati, Hebrew, Kannada, Kazakh, Macedonian, Marathi, Mongolian, Nepali, Pashto, Persian, Punjabi, Serbian, Slovenian, Tagalog, Tamil, Telugu and Urdu, plus regional variants (`en-GB`, `en-IN`, `es-419`, `fr-CA`, `de-CH`, `pt-PT`). Recordings that switch between English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian and Dutch can pass `language: "multi"`. Speaker labels work in every language. ## Errors, and what to do about them Every error is `{"error": {"code": "...", "message": "..."}}`. Treat the HTTP status as the retry signal and the `code` as the reason. | Status | `code` | Meaning | Retry? | |---|---|---|---| | 400 | `invalid_json`, `invalid_url`, `invalid_format`, `invalid_cursor`, `unknown_plan` | Bad request; the message says which field | No — fix the request | | 400 | `not_a_media_url`, `media_unreachable` (the URL answered 4xx/5xx) | The URL is a web page (Spotify, Drive, a YouTube channel or playlist…), not a media file or a single YouTube video | No — use a direct file link, a YouTube video link, or upload | | 400 | `private_url` | The URL points at a private or local address | No | | 400 | `media_too_large` | The URL's file is over 5 GB | No — extract the audio track and send that | | 402 | `quota_exceeded` | The plan's included minutes are used up (plans with overage never return this) | No — upgrade with `POST /v1/billing/change-plan` | | 402 | `video_quota_exceeded`, `youtube_not_included` | The plan's YouTube videos for the period are used up, or the plan has none | No — upgrade with `POST /v1/billing/change-plan` | | 403 | `invalid_api_key` | Missing, unknown, disabled or cancelled key | No — check the `x-api-key` header; recover or rotate the key | | 403 | `account_suspended`, `subscription_inactive` | Account blocked or subscription ended | No — contact us / resubscribe | | 404 | `not_found`, `format_unavailable` | No such transcription, route or rendering | No | | 409 | `not_ready` | Results requested before `status` is `completed` | Yes — poll `GET /v1/transcriptions/{id}` first | | 409 | `in_progress`, `same_plan`, `no_active_subscription`, `payment_method_required` | State conflict; the message says what to do | After fixing the state | | 429 | `rate_limited` | Over the plan's requests-per-second limit | Yes — after `Retry-After` (1 s), with exponential backoff | | 429 | `too_many_active_jobs` | Plan's concurrent-job cap reached | Yes — after `Retry-After` (30 s) or when a job finishes | | 429 | `request_quota_exceeded` | Monthly request quota used | Yes — next month, or upgrade | | 429 | `blocked` | Too many requests from one IP address (firewall) | Yes — after `Retry-After` (5 min) | | 500 | `internal_error` | Our fault | Yes — retry with backoff; if it persists, email us | | 503 | `billing_not_configured`, `prices_not_configured` | Billing routes temporarily unavailable | Yes — later | **Transcription failures** are not HTTP errors: the job's `status` becomes `failed` and `error` names the cause — `ProviderPermanentError` (the media could not be fetched or decoded: unreachable URL, login-protected link, corrupt or silent file, unsupported codec), `ProviderTransientError` (the speech-recognition provider kept failing — resubmit), or `States.Timeout` (processing exceeded 1 hour). Failed jobs are never billed. Uploads that never arrive stay `awaiting_upload` until the upload link expires (2 hours) and are then removed automatically (usually within a day). **Safe to retry:** any 429 (honour `Retry-After`), 500, 503, and jobs that failed with `ProviderTransientError`. Creating a transcription is not idempotent — do not retry a `POST /v1/transcriptions` that returned 201 or you will pay for two jobs. ## Endpoints ### GET /v1/transcriptions List transcriptions _(x-api-key header)_ Newest first. Parameters: - `limit` (query): integer - `cursor` (query): string — `next_cursor` from the previous page. Responses: - `200`: OK - `data`: object[] - `next_cursor`: string, nullable ### POST /v1/transcriptions Create a transcription _(x-api-key header)_ With `url`, transcription starts immediately. Without it, the response contains a presigned `upload.url`; `PUT` the file bytes there and transcription starts when the upload finishes. Request body (JSON): - `url`: string (uri) — Either a **YouTube video link** (youtube.com/watch, youtu.be, Shorts, embed): its caption track is imported in seconds, counted as one of the plan's `included_videos` and never against the minutes; no speaker labels; a video without captions fails with `no_captions`. Or a **direct**, publicly reachable http(s) link to the media file itself (a podcast enclosure URL, a file on your server or storage, a signed link that has not expired). Other pages that *play* media — Spotify, SoundCloud, Apple Podcasts, Drive share links, episode pages — are not media and are rejected with `400 not_a_media_url`; download the file and upload it instead. YouTube channels and playlists are rejected the same way. Files behind a login cannot be fetched. - `name`: string — Label for your own reference; used as the PDF title. - `diarize`: boolean — Identify speakers. - `language`: string — BCP-47 language tag. Omit to auto-detect. Responses: - `201`: Created - `400`: Error - `402`: The plan's included minutes are used up and the plan has no overage. - `error`: object - `code`: string - `message`: string - `429`: Rate limited, monthly request quota used, or too many transcriptions in progress. See `Retry-After`. - `error`: object - `code`: string - `message`: string ### GET /v1/transcriptions/{id} Get a transcription _(x-api-key header)_ When `status` is `completed`, `results` holds a presigned download URL per format. Responses: - `200`: OK - `404`: Error ### DELETE /v1/transcriptions/{id} Delete a transcription _(x-api-key header)_ Removes the source media, every result file and the record. Not allowed while in progress. Responses: - `204`: Deleted - `404`: Error - `409`: Error ### GET /v1/transcriptions/{id}/result Download a result _(x-api-key header)_ Parameters: - `id` (path, required): string - `format` (query): json | srt | vtt | txt | pdf Responses: - `302`: Redirect to a presigned download URL. - `404`: Error - `409`: The transcription is not completed yet. - `error`: object - `code`: string - `message`: string ### GET /v1/account Plan and usage for the current billing period _(x-api-key header)_ Responses: - `200`: OK - `plan`: string - `status`: active | past_due | canceled | provisioning - `period_start`: string - `minutes_used`: number - `minutes_included`: integer, nullable - `minutes_remaining`: integer, nullable - `overage_cents_per_minute`: number, nullable — null = hard cap: new jobs return 402 once the allowance is used. ### POST /v1/billing/portal Create a Stripe customer-portal session _(x-api-key header)_ Returns a short-lived URL where the customer manages cards, invoices and cancellation. Responses: - `201`: Created - `url`: string (uri) - `404`: Error ### POST /v1/billing/change-plan Switch plan _(x-api-key header)_ Prorated immediately. Paid plans need a payment method on file (add one through the portal). Request body (JSON, required): - `plan`: free | payg | starter | pro | scale (required) - `interval`: month | year — year is available on plans with a yearly price (Pro). A yearly subscription keeps the monthly allowance and has no overage (hard cap). Switching month↔year is prorated by Stripe. Responses: - `200`: Plan change submitted. - `409`: Error ### POST /v1/billing/rotate-key Rotate your API key _(x-api-key header)_ Issues a new key and retires the one used for this request. Allow up to a minute for the change to take effect everywhere. Responses: - `201`: Created - `api_key`: string - `note`: string ### GET /billing/plans List plans _(no API key)_ Responses: - `200`: OK - `currency`: string - `plans`: object[] ### POST /mcp MCP server (OAuth sign-in) _(OAuth bearer token)_ Model Context Protocol server, stateless JSON-only Streamable HTTP: one JSON-RPC request in, one JSON response out (`GET` returns 405). Same seven tools, plan, quota and minutes as `POST /v1/mcp`, which takes an `x-api-key` instead of a bearer token. Authenticate with `Authorization: Bearer <access token>` (OAuth 2.1, authorization code + PKCE S256). Clients discover the flow from the `WWW-Authenticate` header of the `401` response and the documents under `/.well-known/`, and register themselves at `POST /oauth/register`. In Claude.ai or ChatGPT just add this URL as a custom connector and sign in with your account email and the one-time code we send. The account must already exist (create a free one at [`/billing/start?plan=free`](https://api.staging.audiotranscribr.com/billing/start?plan=free)). Access tokens last 60 minutes, refresh tokens 30 days. Request body (JSON, required): - `jsonrpc`: 2.0 (required) - `id`: object - `method`: string (required) - `params`: object Responses: - `200`: JSON-RPC response. - `202`: Notification accepted (no body). - `401`: Missing, expired or invalid token (or a token issued for another resource: `invalid_token`; or a sign-in not attached to an account: `unknown_user`). `WWW-Authenticate: Bearer resource_metadata="…/.well-known/oauth-protected-resource/mcp"` says where to start. - `error`: object - `code`: string - `message`: string - `403`: `account_suspended` or `subscription_inactive`. - `error`: object - `code`: string - `message`: string ### POST /oauth/register Register an OAuth client (RFC 7591) _(no API key)_ Dynamic client registration for MCP clients. Redirect URIs must be `https`, loopback `http` (`localhost`, `127.0.0.1`, `[::1]`) or a private-use scheme. Identical public-client registrations return the same `client_id`. 20 registrations per IP per hour (`429 rate_limited`). Request body (JSON, required): - `client_name`: string - `redirect_uris`: string[] (required) - `token_endpoint_auth_method`: none | client_secret_basic | client_secret_post - `grant_types`: authorization_code | refresh_token[] - `scope`: string — Space-separated; unsupported scopes are ignored. Responses: - `201`: Registered. - `client_id`: string - `client_secret`: string — Only when a secret was requested. - `client_id_issued_at`: integer - `client_name`: string - `redirect_uris`: string[] - `grant_types`: string[] - `response_types`: string[] - `token_endpoint_auth_method`: string - `scope`: string - `400`: `invalid_redirect_uri` or `invalid_client_metadata`. - `error`: string - `error_description`: string - `429`: `rate_limited`. ### GET /.well-known/oauth-protected-resource Protected resource metadata (RFC 9728) _(no API key)_ Names the authorization server and scopes for the MCP resource. Same document at `/.well-known/oauth-protected-resource/mcp`. Responses: - `200`: OK ### GET /.well-known/oauth-protected-resource/mcp Protected resource metadata for /mcp (RFC 9728) _(no API key)_ The URL given in the `WWW-Authenticate` header of a `401` from `/mcp`. Responses: - `200`: OK ### GET /.well-known/oauth-authorization-server Authorization server metadata (RFC 8414) _(no API key)_ Authorize, token and revocation endpoints (hosted by Amazon Cognito), our `registration_endpoint`, and S256 as the only PKCE method. Responses: - `200`: OK ### GET /.well-known/openid-configuration OpenID Connect discovery _(no API key)_ Same document as `/.well-known/oauth-authorization-server`. Responses: - `200`: OK ### GET /v1/me Account overview _(OAuth bearer token)_ The dashboard API behind https://audiotranscribr.com/app. It takes an OAuth access token with the `account` scope (not an API key) and is documented for completeness; it is not meant for third-party use. The first call after a sign-in links the user to their customer account. Responses: - `200`: OK - `email`: string, nullable - `customer_id`: string - `status`: string - `plan`: object, nullable - `period_start`: string - `minutes_used`: number - `minutes_included`: integer, nullable - `minutes_remaining`: integer, nullable - `keys`: object[] - `max_keys`: integer - `plans`: object[] - `401`: Error - `403`: Error ### POST /v1/me/keys Create a named API key _(OAuth bearer token)_ Up to 5 keys per account. The key value is returned once. Request body (JSON): - `label`: string Responses: - `201`: Created - `id`: string - `label`: string - `hint`: string - `created_at`: string (date-time) - `api_key`: string — Shown only in this response. - `409`: Error ### DELETE /v1/me/keys/{id} Revoke an API key _(OAuth bearer token)_ The last remaining key cannot be revoked (`409 last_key`). Parameters: - `id` (path, required): string Responses: - `204`: Revoked - `404`: Error - `409`: Error ### GET /v1/me/transcriptions Recent transcriptions with download links _(OAuth bearer token)_ Newest first, 20 per page. Completed jobs carry a download link per format, valid for one hour (`results_expire_in`). Parameters: - `cursor` (query): string — `next_cursor` from the previous page. Responses: - `200`: OK - `data`: object[] - `next_cursor`: string, nullable - `results_expire_in`: integer ### POST /v1/me/transcriptions Create an upload job from the dashboard _(OAuth bearer token)_ Creates an `awaiting_upload` transcription and returns the presigned `upload` URL; the browser PUTs the file straight to it. Same plan, quota and concurrency rules as `POST /v1/transcriptions`. Files up to 5 GB. Request body (JSON): - `name`: string - `language`: string - `diarize`: boolean - `size`: integer — File size in bytes (checked against the 5 GB limit before a job is created) Responses: - `201`: Job created; `upload.url` is where to PUT the file. - `402`: quota_exceeded - `413`: file_too_large - `429`: too_many_active_jobs ### DELETE /v1/me/transcriptions/{id} Delete a transcription _(OAuth bearer token)_ Same as `DELETE /v1/transcriptions/{id}`. Parameters: - `id` (path, required): string Responses: - `204`: Deleted - `404`: Error - `409`: Error ### POST /v1/me/portal Open the Stripe billing portal _(OAuth bearer token)_ Responses: - `201`: Created - `url`: string (uri) ### POST /v1/me/change-plan Switch plan _(OAuth bearer token)_ Paid to paid changes the subscription in place (prorated). When a card must be collected (Free to paid, or no card on file) the response is `201` with a Stripe Checkout `checkout_url` to open instead. Going back to Free is done by cancelling in the billing portal. Request body (JSON, required): - `plan`: free | payg | starter | pro | scale (required) - `interval`: month | year — year is available on plans with a yearly price (Pro). A yearly subscription keeps the monthly allowance and has no overage (hard cap). Switching month↔year is prorated by Stripe. Responses: - `200`: Plan change submitted. - `201`: Open `checkout_url` to add a card. - `checkout_url`: string (uri) - `409`: Error ## Schemas ### Transcription - `id`: string - `name`: string, nullable - `status`: awaiting_upload | queued | transcribing | completed | failed - `source`: upload | url | youtube — youtube = caption import - `diarize`: boolean - `language`: string, nullable — Requested or detected language. - `duration_seconds`: number, nullable - `summary`: string, nullable — A few-sentence summary of the recording - `error`: string, nullable — Human-readable failure message when `status` is `failed`. YouTube imports can fail with no captions (download the audio and upload it instead) - `created_at`: string (date-time) - `updated_at`: string (date-time) ### ErrorBody - `error`: object - `code`: string - `message`: string ### TranscriptJson Shape of the `json` result file. - `provider`: string - `model`: string - `language`: string - `durationSeconds`: number - `text`: string - `segments`: object[] - `words`: object[]