> For the complete documentation index, see [llms.txt](https://developers.oxylabs.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developers.oxylabs.io/api-reference/dedicated-scrapers/media-scrapers.md).

# Media Scrapers

Reference for Web API Media scraper endpoints: metadata, subtitles, downloads, channels, search, autocomplete, and trainability.

<table><thead><tr><th width="110">Platform</th><th>Endpoints</th></tr></thead><tbody><tr><td><a href="#youtube">YouTube</a></td><td><code>/media/youtube/metadata</code>, <code>/media/youtube/subtitles</code>, <code>/media/youtube/download</code>, <code>/media/youtube/channel</code>, <code>/media/youtube/search</code>, <code>/media/youtube/search/max</code>, <code>/media/youtube/autocomplete</code>, <code>/media/youtube/video/trainability</code></td></tr></tbody></table>

### YouTube

#### `POST /v1/scrape/media/youtube/metadata`

Returns a video's title, channel, description, duration, and view and like counts.

**Modes:** sync · async · batch

<table><thead><tr><th width="139.5">Name</th><th width="102">Type</th><th width="92">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>string</td><td>yes</td><td>Video ID, not a URL</td></tr><tr><td><code>output</code></td><td>array of strings</td><td>no</td><td>Formats to return: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code>. Default <code>["html"]</code>. <code>json</code> returns parsed fields. <code>screenshot</code> requires <code>run_js: true</code></td></tr><tr><td><code>json</code></td><td>object</td><td>no</td><td>Custom extraction: <code>{"prompt": string}</code> or <code>{"schema": object}</code>, never both. Overrides the built-in parser</td></tr><tr><td><code>location</code></td><td>string</td><td>no</td><td>Two-letter country code to fetch from, e.g. <code>DE</code></td></tr><tr><td><code>device</code></td><td>string</td><td>no</td><td><code>desktop</code> (default) or <code>mobile</code></td></tr><tr><td><code>run_js</code></td><td>boolean</td><td>no</td><td>Executes page JavaScript before capture. Slower — use a client timeout of 120 seconds or more</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Async only. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response: <code>{"type": string, "url": string}</code>. <code>type</code> is <code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Async only. <code>url</code> is not validated</td></tr></tbody></table>

#### `POST /v1/scrape/media/youtube/subtitles`

Returns a video's subtitle track: the transcript.

**Modes:** sync · async · batch

<table><thead><tr><th width="173">Name</th><th width="100.5">Type</th><th width="106">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>string</td><td>yes</td><td>Video ID, not a URL</td></tr><tr><td><code>language_code</code></td><td>string</td><td>no</td><td>Subtitle language to fetch, e.g. <code>en</code>, <code>lt</code></td></tr><tr><td><code>subtitle_origin</code></td><td>string</td><td>no</td><td><code>uploader_provided</code> or <code>auto_generated</code>. Omit to use all available origins</td></tr><tr><td><code>output</code></td><td>array of strings</td><td>no</td><td>Formats to return: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code>. Default <code>["html"]</code>. <code>json</code> returns parsed fields. <code>screenshot</code> requires <code>run_js: true</code></td></tr><tr><td><code>json</code></td><td>object</td><td>no</td><td>Required with <code>output: ["json"]</code> — this endpoint has no built-in parser. <code>{"prompt": string}</code> or <code>{"schema": object}</code>, never both</td></tr><tr><td><code>location</code></td><td>string</td><td>no</td><td>Two-letter country code to fetch from, e.g. <code>DE</code></td></tr><tr><td><code>device</code></td><td>string</td><td>no</td><td><code>desktop</code> (default) or <code>mobile</code></td></tr><tr><td><code>run_js</code></td><td>boolean</td><td>no</td><td>Executes page JavaScript before capture. Slower — use a client timeout of 120 seconds or more</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Async only. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response: <code>{"type": string, "url": string}</code>. <code>type</code> is <code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Async only. <code>url</code> is not validated</td></tr></tbody></table>

A video with no matching track returns an empty result, not an error.

#### `POST /v1/async/scrape/media/youtube/download`

Returns a video's media, or a clip of it.

**Modes:** async · batch

<table><thead><tr><th width="158">Name</th><th width="102.5">Type</th><th width="87.5">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>string</td><td>yes</td><td>Video ID, not a URL</td></tr><tr><td><code>audio_format</code></td><td>string</td><td>no</td><td>Container for audio downloads</td></tr><tr><td><code>audio_language</code></td><td>string</td><td>no</td><td>Audio track, for videos with dubs</td></tr><tr><td><code>download_type</code></td><td>string</td><td>no</td><td><code>video</code> or <code>audio</code>. Audio is much smaller when only speech is needed</td></tr><tr><td><code>end_at</code></td><td>string</td><td>no</td><td>End of the range to fetch, e.g. <code>00:02:45</code></td></tr><tr><td><code>start_at</code></td><td>string</td><td>no</td><td>Start of the range to fetch, e.g. <code>00:01:30</code>. Omit to fetch from the start</td></tr><tr><td><code>video_format</code></td><td>string</td><td>no</td><td>Container for video downloads</td></tr><tr><td><code>video_quality</code></td><td>string</td><td>no</td><td>Resolution. Higher is bigger, slower, and costlier</td></tr><tr><td><code>output</code></td><td>array of strings</td><td>no</td><td>Formats to return: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code>. Default <code>["html"]</code>. <code>json</code> returns parsed fields. <code>screenshot</code> requires <code>run_js: true</code></td></tr><tr><td><code>json</code></td><td>object</td><td>no</td><td>Custom extraction: <code>{"prompt": string}</code> or <code>{"schema": object}</code>, never both. Overrides the built-in parser</td></tr><tr><td><code>location</code></td><td>string</td><td>no</td><td>Two-letter country code to fetch from, e.g. <code>DE</code></td></tr><tr><td><code>device</code></td><td>string</td><td>no</td><td><code>desktop</code> (default) or <code>mobile</code></td></tr><tr><td><code>run_js</code></td><td>boolean</td><td>no</td><td>Executes page JavaScript before capture. Slower — use a client timeout of 120 seconds or more</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Async only. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response: <code>{"type": string, "url": string}</code>. <code>type</code> is <code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Async only. <code>url</code> is not validated</td></tr></tbody></table>

#### `POST /v1/scrape/media/youtube/channel`

Returns a channel's videos.

**Modes:** sync · async

<table><thead><tr><th width="137.5">Name</th><th width="93.5">Type</th><th width="89">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>string</td><td>yes</td><td>Channel ID</td></tr><tr><td><code>limit</code></td><td>integer</td><td>no</td><td>Number of videos to return. The default returns more than most callers use</td></tr><tr><td><code>output</code></td><td>array of strings</td><td>no</td><td>Formats to return: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code>. Default <code>["html"]</code>. <code>json</code> returns parsed fields. <code>screenshot</code> requires <code>run_js: true</code></td></tr><tr><td><code>json</code></td><td>object</td><td>no</td><td>Custom extraction: <code>{"prompt": string}</code> or <code>{"schema": object}</code>, never both. Overrides the built-in parser</td></tr><tr><td><code>location</code></td><td>string</td><td>no</td><td>Two-letter country code to fetch from, e.g. <code>DE</code></td></tr><tr><td><code>device</code></td><td>string</td><td>no</td><td><code>desktop</code> (default) or <code>mobile</code></td></tr><tr><td><code>run_js</code></td><td>boolean</td><td>no</td><td>Executes page JavaScript before capture. Slower — use a client timeout of 120 seconds or more</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Async only. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response: <code>{"type": string, "url": string}</code>. <code>type</code> is <code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Async only. <code>url</code> is not validated</td></tr></tbody></table>

#### `POST /v1/scrape/media/youtube/search`

Returns YouTube search results.

**Modes:** sync · async

<table><thead><tr><th width="170">Name</th><th width="111.5">Type</th><th width="98.5">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>string</td><td>yes</td><td>Search phrase</td></tr><tr><td><code>360</code></td><td>boolean</td><td>no</td><td>Only 360° videos</td></tr><tr><td><code>3d</code></td><td>boolean</td><td>no</td><td>Only 3D videos</td></tr><tr><td><code>4k</code></td><td>boolean</td><td>no</td><td>Only 4K videos</td></tr><tr><td><code>creative_commons</code></td><td>boolean</td><td>no</td><td>Only Creative Commons-licensed videos</td></tr><tr><td><code>duration</code></td><td>string</td><td>no</td><td><code>&#x3C;4</code> under four minutes, <code>4-20</code> four to twenty minutes, <code>>20</code> over twenty minutes</td></tr><tr><td><code>hd</code></td><td>boolean</td><td>no</td><td>Only HD videos</td></tr><tr><td><code>hdr</code></td><td>boolean</td><td>no</td><td>Only HDR videos</td></tr><tr><td><code>live</code></td><td>boolean</td><td>no</td><td>Only live streams</td></tr><tr><td><code>purchased</code></td><td>boolean</td><td>no</td><td>Only purchased content</td></tr><tr><td><code>sort_by</code></td><td>string</td><td>no</td><td>Relevance, upload date, view count, or rating</td></tr><tr><td><code>subtitles</code></td><td>boolean</td><td>no</td><td>Only videos with a subtitle track</td></tr><tr><td><code>type</code></td><td>string</td><td>no</td><td>Restricts to videos, channels, playlists, or movies</td></tr><tr><td><code>upload_date</code></td><td>string</td><td>no</td><td><code>last_hour</code>, <code>today</code>, <code>this_week</code>, <code>this_month</code>, or <code>this_year</code></td></tr><tr><td><code>vr180</code></td><td>boolean</td><td>no</td><td>Only VR180 videos</td></tr><tr><td><code>output</code></td><td>array of strings</td><td>no</td><td>Formats to return: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code>. Default <code>["html"]</code>. <code>json</code> returns parsed fields. <code>screenshot</code> requires <code>run_js: true</code></td></tr><tr><td><code>json</code></td><td>object</td><td>no</td><td>Custom extraction: <code>{"prompt": string}</code> or <code>{"schema": object}</code>, never both. Overrides the built-in parser</td></tr><tr><td><code>location</code></td><td>string</td><td>no</td><td>Two-letter country code to fetch from, e.g. <code>DE</code></td></tr><tr><td><code>device</code></td><td>string</td><td>no</td><td><code>desktop</code> (default) or <code>mobile</code></td></tr><tr><td><code>run_js</code></td><td>boolean</td><td>no</td><td>Executes page JavaScript before capture. Slower — use a client timeout of 120 seconds or more</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Async only. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response: <code>{"type": string, "url": string}</code>. <code>type</code> is <code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Async only. <code>url</code> is not validated</td></tr></tbody></table>

#### `POST /v1/scrape/media/youtube/search/max`

Returns a deeper result set than `/search` for the same parameters, at higher cost and latency.

**Modes:** sync · async

<table><thead><tr><th width="171">Name</th><th width="108.5">Type</th><th width="94">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>string</td><td>yes</td><td>Search phrase</td></tr><tr><td><code>360</code></td><td>boolean</td><td>no</td><td>Only 360° videos</td></tr><tr><td><code>3d</code></td><td>boolean</td><td>no</td><td>Only 3D videos</td></tr><tr><td><code>4k</code></td><td>boolean</td><td>no</td><td>Only 4K videos</td></tr><tr><td><code>creative_commons</code></td><td>boolean</td><td>no</td><td>Only Creative Commons-licensed videos</td></tr><tr><td><code>duration</code></td><td>string</td><td>no</td><td><code>&#x3C;4</code> under four minutes, <code>4-20</code> four to twenty minutes, <code>>20</code> over twenty minutes</td></tr><tr><td><code>hd</code></td><td>boolean</td><td>no</td><td>Only HD videos</td></tr><tr><td><code>hdr</code></td><td>boolean</td><td>no</td><td>Only HDR videos</td></tr><tr><td><code>live</code></td><td>boolean</td><td>no</td><td>Only live streams</td></tr><tr><td><code>purchased</code></td><td>boolean</td><td>no</td><td>Only purchased content</td></tr><tr><td><code>sort_by</code></td><td>string</td><td>no</td><td>Relevance, upload date, view count, or rating</td></tr><tr><td><code>subtitles</code></td><td>boolean</td><td>no</td><td>Only videos with a subtitle track</td></tr><tr><td><code>type</code></td><td>string</td><td>no</td><td>Restricts to videos, channels, playlists, or movies</td></tr><tr><td><code>upload_date</code></td><td>string</td><td>no</td><td><code>last_hour</code>, <code>today</code>, <code>this_week</code>, <code>this_month</code>, or <code>this_year</code></td></tr><tr><td><code>vr180</code></td><td>boolean</td><td>no</td><td>Only VR180 videos</td></tr><tr><td><code>output</code></td><td>array of strings</td><td>no</td><td>Formats to return: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code>. Default <code>["html"]</code>. <code>json</code> returns parsed fields. <code>screenshot</code> requires <code>run_js: true</code></td></tr><tr><td><code>json</code></td><td>object</td><td>no</td><td>Custom extraction: <code>{"prompt": string}</code> or <code>{"schema": object}</code>, never both. Overrides the built-in parser</td></tr><tr><td><code>location</code></td><td>string</td><td>no</td><td>Two-letter country code to fetch from, e.g. <code>DE</code></td></tr><tr><td><code>device</code></td><td>string</td><td>no</td><td><code>desktop</code> (default) or <code>mobile</code></td></tr><tr><td><code>run_js</code></td><td>boolean</td><td>no</td><td>Executes page JavaScript before capture. Slower — use a client timeout of 120 seconds or more</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Async only. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response: <code>{"type": string, "url": string}</code>. <code>type</code> is <code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Async only. <code>url</code> is not validated</td></tr></tbody></table>

#### `POST /v1/scrape/media/youtube/autocomplete`

Returns YouTube's search suggestions for a partial phrase.

**Modes:** sync · async

<table><thead><tr><th width="136">Name</th><th width="97">Type</th><th width="88">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>string</td><td>yes</td><td>Partial phrase to complete</td></tr><tr><td><code>language</code></td><td>string</td><td>no</td><td>Suggestion language, e.g. <code>en</code></td></tr><tr><td><code>output</code></td><td>array of strings</td><td>no</td><td>Formats to return: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code>. Default <code>["html"]</code>. <code>json</code> returns parsed fields. <code>screenshot</code> requires <code>run_js: true</code></td></tr><tr><td><code>json</code></td><td>object</td><td>no</td><td>Custom extraction: <code>{"prompt": string}</code> or <code>{"schema": object}</code>, never both. Overrides the built-in parser</td></tr><tr><td><code>location</code></td><td>string</td><td>no</td><td>Two-letter country code to fetch from, e.g. <code>DE</code></td></tr><tr><td><code>device</code></td><td>string</td><td>no</td><td><code>desktop</code> (default) or <code>mobile</code></td></tr><tr><td><code>run_js</code></td><td>boolean</td><td>no</td><td>Executes page JavaScript before capture. Slower — use a client timeout of 120 seconds or more</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Async only. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response: <code>{"type": string, "url": string}</code>. <code>type</code> is <code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Async only. <code>url</code> is not validated</td></tr></tbody></table>

#### `POST /v1/scrape/media/youtube/video/trainability`

Returns a video's declared AI-training permission signal.

**Modes:** sync · async

<table><thead><tr><th width="130">Name</th><th width="92.5">Type</th><th width="89.5">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>string</td><td>yes</td><td>Video ID, not a URL</td></tr><tr><td><code>output</code></td><td>array of strings</td><td>no</td><td>Formats to return: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code>. Default <code>["html"]</code>. <code>json</code> returns parsed fields. <code>screenshot</code> requires <code>run_js: true</code></td></tr><tr><td><code>json</code></td><td>object</td><td>no</td><td>Custom extraction: <code>{"prompt": string}</code> or <code>{"schema": object}</code>, never both. Overrides the built-in parser</td></tr><tr><td><code>location</code></td><td>string</td><td>no</td><td>Two-letter country code to fetch from, e.g. <code>DE</code></td></tr><tr><td><code>device</code></td><td>string</td><td>no</td><td><code>desktop</code> (default) or <code>mobile</code></td></tr><tr><td><code>run_js</code></td><td>boolean</td><td>no</td><td>Executes page JavaScript before capture. Slower — use a client timeout of 120 seconds or more</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Async only. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response: <code>{"type": string, "url": string}</code>. <code>type</code> is <code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Async only. <code>url</code> is not validated</td></tr></tbody></table>

The signal is declared by the rights holder. It is not a legal clearance.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://developers.oxylabs.io/api-reference/dedicated-scrapers/media-scrapers.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `build a script that syncs our docs to a CMS` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
