> For the complete documentation index, see [llms.txt](https://developers.oxylabs.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developers.oxylabs.io/api-reference/jobs-and-discovery.md).

# Jobs & Discovery

Reference for submitting, batching, and collecting async scrape jobs, and for listing endpoints and parameters at runtime.

<table><thead><tr><th width="112.5">Section</th><th>Endpoints</th></tr></thead><tbody><tr><td><a href="https://claude.ai/chat/f8236b05-b9bb-4809-af2d-33fe5a159352#jobs">Jobs</a></td><td><code>POST /v1/async/scrape/{path}</code>, <code>POST /v1/async/scrape/batch/{path}</code>, <code>GET /v1/async/scrape/{request_id}</code></td></tr><tr><td><a href="https://claude.ai/chat/f8236b05-b9bb-4809-af2d-33fe5a159352#discovery">Discovery</a></td><td><code>GET /v1/scrapers</code>, <code>OPTIONS /v1/{endpoint}</code></td></tr></tbody></table>

## Jobs

### `POST /v1/async/scrape/{path}`

Submits one scrape job. `{path}` is any scrape path with `/v1/async/scrape/` in place of `/v1/scrape/`.

**Modes:** `async`

<table><thead><tr><th width="151">Name</th><th width="98">Type</th><th width="106.5">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for the finished job. Malformed URLs are rejected with <code>400</code></td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Delivers the result to a bucket instead of the response</td></tr><tr><td><code>storage.type</code></td><td>string</td><td>yes, with <code>storage</code></td><td><code>s3</code>, <code>s3_gzip</code>, <code>s3_compatible</code>, <code>gcs</code>, or <code>tos</code>. Validated</td></tr><tr><td><code>storage.url</code></td><td>string</td><td>yes, with <code>storage</code></td><td>Destination for the result. Not validated: a malformed value is accepted, the job runs and is billed, and delivery fails</td></tr></tbody></table>

All other parameters are those of the endpoint being submitted.

**Response** — `202 Accepted`

| Field                 | Type    | Description                                 |
| --------------------- | ------- | ------------------------------------------- |
| `status`              | string  | `pending`                                   |
| `params`              | object  | The request as executed, defaults filled in |
| `metadata.timestamp`  | integer | Unix timestamp in seconds                   |
| `metadata.request_id` | string  | The job ID. Used to collect the result      |

### `POST /v1/async/scrape/batch/{path}`

Submits one job per value in a list. `{path}` is any scrape path that has a batch mode.

**Modes:** `batch`

<table><thead><tr><th width="143.5">Name</th><th width="153">Type</th><th width="194.5">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>query</code></td><td>array of strings</td><td>yes, if the endpoint takes <code>query</code></td><td>Values, one job per value</td></tr><tr><td><code>url</code></td><td>array of strings</td><td>yes, if the endpoint takes <code>url</code></td><td>URLs, one job per value</td></tr><tr><td><code>callback_url</code></td><td>string</td><td>no</td><td>Webhook URL for each finished job</td></tr><tr><td><code>storage</code></td><td>object</td><td>no</td><td>Bucket delivery for each job. Same fields as above</td></tr></tbody></table>

All other parameters are those of the endpoint being submitted and apply to every job.

**Response** – `202 Accepted`

<table><thead><tr><th width="261">Field</th><th width="141.5">Type</th><th>Description</th></tr></thead><tbody><tr><td><code>data[]</code></td><td>array</td><td>One object per submitted job</td></tr><tr><td><code>data[].status</code></td><td>string</td><td><code>pending</code></td></tr><tr><td><code>data[].params</code></td><td>object</td><td>The job's request as executed</td></tr><tr><td><code>data[].metadata.request_id</code></td><td>string</td><td>The job ID</td></tr></tbody></table>

There is no batch collection route: collect each job by its `request_id`. `/v1/scrape/batch` returns `404`.

### `GET /v1/async/scrape/{request_id}`

Returns the status and result of a submitted job.

| Name         | Type   | Required | Description                                                    |
| ------------ | ------ | -------- | -------------------------------------------------------------- |
| `request_id` | string | yes      | `metadata.request_id` from the submit response. Path parameter |

**Response** – `200 OK`

| Field       | Type   | Description                                                            |
| ----------- | ------ | ---------------------------------------------------------------------- |
| `status`    | string | `pending` while the job is queued or running, `done` when finished     |
| `results[]` | array  | Same shape as the realtime response. Populated when `status` is `done` |

An unknown or expired `request_id` returns `404`.

## Discovery

### `GET /v1/scrapers`

Lists every implemented scrape path: realtime, async, and batch.

**Response** – `200 OK`

<table><thead><tr><th width="122">Field</th><th width="157.5">Type</th><th>Description</th></tr></thead><tbody><tr><td><code>scrapers</code></td><td>array of strings</td><td>Scrape paths, e.g. <code>/v1/scrape/amazon/product</code>, <code>/v1/async/scrape/amazon/product</code></td></tr></tbody></table>

`GET /v1/async/scrape/{request_id}` is not in the list.

### `OPTIONS /v1/{endpoint}`

Returns an example request body for an endpoint: every accepted parameter, with a type in place of each value.

<table><thead><tr><th width="114">Name</th><th width="99">Type</th><th width="109">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>endpoint</code></td><td>string</td><td>yes</td><td>Any path from <code>GET /v1/scrapers</code>. Path parameter</td></tr></tbody></table>

#### **Type values**

<table><thead><tr><th width="239">Value</th><th>Meaning</th></tr></thead><tbody><tr><td><code>"string"</code></td><td>Text</td></tr><tr><td><code>"bool"</code></td><td>Boolean</td></tr><tr><td><code>"int"</code></td><td>Integer</td></tr><tr><td><code>Nested object</code></td><td>Object parameter, e.g. <code>json</code></td></tr><tr><td><code>Single-element array</code></td><td>Array of that type</td></tr></tbody></table>

Names and types only: no descriptions, allowed values, or required flags. Async paths also list `storage` and `callback_url`.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://developers.oxylabs.io/api-reference/jobs-and-discovery.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `build a script that syncs our docs to a CMS` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
