> For the complete documentation index, see [llms.txt](https://developers.oxylabs.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developers.oxylabs.io/products/web-api/scrape.md).

# /scrape

Learn how to use Oxylabs Web API /scrape endpoint to fetch and parse any web page in the format you want.

The `/v1/scrape` endpoint fetches web pages by URL, automatically handling headless rendering, proxy rotation, access challenges, and localization. It returns structured page data in Markdown, HTML, JSON, or as PNG screenshots.

## Request sample

The endpoint accepts a JSON payload defining the target URL, requested output formats, rendering options, and geo-location context.

```bash
curl https://webapi.oxylabs.io/v1/scrape \
  -H "Authorization: Bearer $OXYLABS_WEB_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
      "url": "https://artificialintelligenceact.eu/implementation-timeline/",
      "output": ["markdown"]
  }'
```

### Request parameters

<table><thead><tr><th width="98">Parameter</th><th width="566">Description</th><th width="93">Type</th></tr></thead><tbody><tr><td><mark style="background-color:green;"><code>url</code></mark></td><td>Absolute <code>http</code> / <code>https</code> target URL.</td><td>string</td></tr><tr><td><code>output</code></td><td>Response format: <code>markdown</code>, <code>html</code>, <code>json</code>, <code>screenshot</code> (includes <code>video</code> for <code>/media</code> <a href="/products/web-api/scrape/dedicated-scrapers.md">dedicated scrapers</a>). Supports multiple values. Default: <code>html</code></td><td>array of strings</td></tr><tr><td><code>json</code></td><td>When <code>output: ["json"]</code>, can be used for custom AI parsing through <code>prompt</code> or <code>schema</code>. See <a href="/products/web-api/scrape/ai-parsing.md">AI Parsing</a>.</td><td>object</td></tr><tr><td><code>location</code></td><td>Two-letter ISO 3166-1 alpha-2 country code (e.g., <code>"US"</code>, <code>"DE"</code>, <code>"JP"</code>).</td><td>string</td></tr><tr><td><code>device</code></td><td>User-agent and viewport emulation: <code>"desktop"</code> (default) or <code>"mobile"</code>. </td><td>string</td></tr><tr><td><code>run_js</code></td><td>Enables JavaScript rendering when set to <code>true</code>. Default: <code>false</code>.</td><td>boolean</td></tr></tbody></table>

&#x20;    – mandatory parameter.

<details>

<summary><strong><code>output</code></strong></summary>

The `output` array allows fetching multiple output formats in a single request execution. Below is an example of combined output request:

```json
{
  "url": "https://example.com/product",
  "output": ["markdown", "json"],
  "json": {
    "prompt": "Extract the product name, stock status, and price as a number."
  }
}
```

<table data-header-hidden><thead><tr><th width="119"></th><th></th></tr></thead><tbody><tr><td><code>markdown</code></td><td>Converts the DOM tree into clean, structured Markdown, stripping navigation scripts and styling formatting. Optimized for ingestion into LLM context windows to minimize token usage.</td></tr><tr><td><code>html</code></td><td>The raw page markup. Use it when you need the markup itself (attributes, embedded JSON-LD, inline data). Not recommended for page content collection.</td></tr><tr><td><code>screenshot</code></td><td>Base64-encoded PNG image of the fully rendered page. <strong><code>run_js: true</code> is required</strong>, otherwise the request is rejected.</td></tr><tr><td><code>json</code></td><td>Structured extraction in the same request as the retrieval (no selectors to write or maintain). Describe the fields in a <code>prompt</code>, or pin their exact shape with a <code>schema</code>. See <a href="/products/web-api/scrape/ai-parsing.md">AI parsing</a>.</td></tr></tbody></table>

</details>

<details>

<summary><strong><code>run_js</code></strong></summary>

Executes the page's own JavaScript before the content is captured.

Dropping `run_js` defaults to the optimal execution path for the given source path. Explicitly setting `"run_js": true` uses a headless browser instance to execute client-side scripts, dynamic XHR requests, and Single Page Application (SPA) rendering.

```json
{
  "url": "https://example.com/spa-dashboard",
  "output": ["markdown"],
  "run_js": true
}
```

* **Latency impact:** Headless browser rendering increases request duration significantly. Set client HTTP timeouts to at least `150` seconds. When high volume is required, [async delivery](/products/web-api/scrape/delivery-modes.md) is recommended.
* **Resource optimization:** Always try fetching content without it. For static pages or RSS/atom feeds, pass `"run_js": false` to maximize request throughput and minimize latency. Choose rendering only when necessary.

</details>

<details>

<summary><strong><code>location</code></strong></summary>

Geo-targeting routes requests through localized proxy exit nodes. This ensures localized pricing, language variations, and region-exclusive content are served correctly.

```json
{
  "url": "https://example.com/pricing",
  "output": ["html"],
  "location": "DE",
  "check_empty_geo": true
}
```

{% hint style="info" %}
When using [Dedicated Scrapers](/products/web-api/scrape/dedicated-scrapers.md), some localization parameters may be named differently.
{% endhint %}

</details>

## Output sample

A successful request returns an HTTP `200 OK` status with a standard envelope. Fields inside `results[]` with unrequested output formats return `null`.

```json
{
  "status": 200,
  "results": [
    {
      "html": null,
      "markdown": "# Implementation Timeline\n\nProviders need to be compliant...",
      "screenshot": null,
      "json": null,
      "metadata": {
        "page": 1,
        "run_js": false,
        "url": "https://artificialintelligenceact.eu/implementation-timeline/"
      }
    }
  ],
  "params": {
    "url": "https://artificialintelligenceact.eu/implementation-timeline/",
    "output": ["markdown"],
    "json": { "schema": null, "prompt": null },
    "location": null,
    "device": "desktop",
    "run_js": null
  },
  "metadata": {
    "timestamp": 1787431568,
    "request_id": "7503078851572948993"
  }
}
```

{% hint style="info" %}
**`results[]` field with unrequested output formats return `null`.** \
For example, asking for `["markdown"]` populates `markdown` and leaves `html`, `screenshot` and `json` as null. Any [dedicated scraper](/products/web-api/scrape/dedicated-scrapers.md) puts its structured fields under `json`.
{% endhint %}

### Output dictionary

<table><thead><tr><th width="208">Field</th><th>Description</th><th width="92">Type</th></tr></thead><tbody><tr><td><code>status</code></td><td>HTTP status code, <code>200</code> on success. See <a href="/products/web-api/troubleshooting.md">Error codes</a>.</td><td>integer</td></tr><tr><td><code>results</code></td><td>One entry for the requested URL.</td><td>array</td></tr><tr><td><code>results[].markdown</code></td><td>Page as Markdown, if requested. <code>null</code> otherwise.</td><td>string</td></tr><tr><td><code>results[].html</code></td><td>Raw result markup, if requested. <code>null</code> otherwise.</td><td>string</td></tr><tr><td><code>results[].json</code></td><td>Parsed fields from a <a href="/products/web-api/scrape/dedicated-scrapers.md">dedicated scraper</a> or <a href="/products/web-api/scrape/ai-parsing.md">AI parser</a>, if requested <code>null</code> otherwise.</td><td>object</td></tr><tr><td><code>results[].screenshot</code></td><td>Base64-encoded image string, if requested. <code>null</code> otherwise.</td><td>string</td></tr><tr><td><code>results[].metadata.url</code></td><td>Fetched URL for scraping after redirects.</td><td>string</td></tr><tr><td><code>params</code></td><td>All parameters use for the requests with applied defaults.</td><td>object</td></tr><tr><td><code>metadata.timestamp</code></td><td>Unix timestamp in seconds.</td><td>integer</td></tr><tr><td><code>metadata.request_id</code></td><td>Identifier for this request. <strong>Include it in support tickets.</strong></td><td>string</td></tr></tbody></table>

## Latency and concurrency

A scrape renders a real page with budget in **seconds, not milliseconds**. We recommend a couple of of safe practices to build around that:

<table data-header-hidden><thead><tr><th width="123"></th><th></th></tr></thead><tbody><tr><td><strong>Client timeout</strong></td><td>150 seconds is a safe default, while a 10-second timeout may fail on pages that would have succeeded. For higher volume (e.g. media, large SERP runs, etc.) with a risk of client timeout, run the same scrapers with async delivery instead of blocking the request. See <a href="/products/web-api/scrape/delivery-modes.md">Delivery Modes</a>.</td></tr><tr><td><strong>Concurrency limit</strong></td><td>Scraping ten search results in parallel multiplies cost and latency for pages you won't read. Scrape the two or three that plausibly answer the question and only increase concurrency gradually until <code>429</code> error is reached.</td></tr></tbody></table>

## Dedicated scrapers

Beyond generic URL scraping, target-specific scrapers can be used with `POST /v1/scrape/{target}/{collection}`. They accept target-specific parameters (e.g., ASINs, browse nodes, store zip codes) and return structured fields under `results[].json`:

```bash
curl https://webapi.oxylabs.io/v1/scrape/amazon/product \
  -H "Authorization: Bearer $OXYLABS_WEB_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "query": "B0935DN1BN",
    "output": ["json"],
    "location": "10115",
    "domain": "de",
    "currency": "EUR"
  }'
```

See [Dedicated scrapers](/products/web-api/scrape/dedicated-scrapers.md) for target overviews (Google, Amazon, YouTube, Bing, AI & LLMs, Commerce).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://developers.oxylabs.io/products/web-api/scrape.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
