> For the complete documentation index, see [llms.txt](https://developers.oxylabs.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developers.oxylabs.io/products/web-api/scrape/ai-parsing.md).

# AI Parsing

Extract structured JSON from any page in one Web API request, using a natural-language prompt or a JSON Schema definition.

{% hint style="danger" %}
**Costs extra:** Using AI parsing with your `/scrape` requests uses significantly more account credit.
{% endhint %}

The Web API provides structured JSON extraction in the same HTTP call that fetches the target web page. Combine `output: ["json"]` with the root-level `json` parameter to extract custom fields using either natural language prompts or strict OpenAPI / JSON Schema definitions.

<table><thead><tr><th width="143">Field</th><th>Description</th><th width="85">Type</th></tr></thead><tbody><tr><td><code>json.prompt</code></td><td>Natural-language description of the fields to extract.</td><td>string</td></tr><tr><td><code>json.schema</code></td><td>An <a href="https://spec.openapis.org/oas/v3.1.0#schema-object">OpenAPI Schema Object</a> the result must conform to.</td><td>object</td></tr></tbody></table>

{% hint style="success" %}
Use `prompt` when exploring a page and you don't yet know its shape. Use `schema` when the output feeds code, where a renamed or missing key would break it.
{% endhint %}

## Dedicated parsers vs. Custom AI extraction

The Web API supports two parsing mechanisms, both returning data under the `json` key in the `results[]` array:

<table><thead><tr><th width="101">Type</th><th>Method</th><th>Targets</th><th>Use for</th></tr></thead><tbody><tr><td>Built-in parser</td><td>Set <code>output: ["json"]</code> without providing a root <code>json</code> object.</td><td>Dedicated targets with pre-built parsers (e.g., <code>/scrape/amazon/product</code>, <code>/scrape/google/search</code>).</td><td>Standardized e-commerce and SERP schemas maintained by Oxylabs.</td></tr><tr><td>Custom AI Extraction</td><td>Set <code>output: ["json"]</code> and pass <code>json: { "prompt": ... }</code> or <code>json: { "schema": ... }</code> object.</td><td>Every route, including generic <code>/v1/scrape</code> and targets without dedicated parsers.</td><td>Custom schemas, unique page layouts, or non-standard fields.</td></tr></tbody></table>

{% hint style="warning" %}
Both pre-built parsers and custom AI extractions populate the `results[].json` key in the response payload. If you call a dedicated scraper (e.g. `/v1/scrape/amazon/product`) with `output: ["json"]` and use a custom `json` extraction object, the custom AI extraction will override the built-in parser output. Choose one option per request.
{% endhint %}

## Extraction modes: `prompt` vs. `schema`

When using custom AI extraction, the root `json` object accepts either a `prompt` string or a `schema` object. Do not pass both parameters in the same request.

### Natural language (`prompt`)

Pass a natural language instruction describing the desired fields. The AI model inspects the page DOM and constructs a matching JSON object.

```json
{
  "url": "https://example.com/product/10928",
  "output": ["json"],
  "json": {
    "prompt": "Extract the main product title, current sale price as a number, currency code, and a list of image URLs."
  }
}
```

#### When to use

* Rapid prototyping
* Experimental data extraction
* Simple flat key-value extractions where exact data types are not strictly enforced later

### Schema-defined (`schema`)

Pass a JSON Schema object conforming to the OpenAPI 3.1 / OAS 3.0 Schema specification. The extraction pipeline guarantees the response matches your defined types, nested structures, and required fields.

```json
{
  "url": "https://example.com/product/10928",
  "output": ["json"],
  "json": {
    "schema": {
      "type": "object",
      "properties": {
        "title": { "type": "string" },
        "price": { 
          "type": "number", 
          "description": "The final discounted price shown to the user, excluding tax." 
        },
        "currency": { "type": "string" },
        "in_stock": { "type": "boolean" },
        "categories": {
          "type": "array",
          "items": { "type": "string" }
        }
      },
      "required": ["title", "price", "currency", "in_stock"]
    }
  }
}
```

#### When to use:&#x20;

* Production pipelines
* Automated ingestion databases
* Strongly typed applications where missing or unexpected data types may cause errors

## Common schema patterns

### 1. Repeating rows and lists

Extract search results, tables, or product listings by wrapping a row object in an array:

```json
{
  "url": "https://example.com/search?q=laptop",
  "output": ["json"],
  "json": {
    "schema": {
      "type": "object",
      "properties": {
        "products": {
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "name": { "type": "string" },
              "price": { "type": "number" },
              "rating": { "type": "number", "nullable": true },
              "url": { 
                "type": "string", 
                "description": "Absolute URL to the product detail page" 
              }
            },
            "required": ["name", "price", "url"]
          }
        }
      },
      "required": ["products"]
    }
  }
}
```

### 2. Standardized enum and nested records

Map different text across different websites into fixed internal vocabularies using `enum`:

```json
{
  "url": "https://example.com/product/123",
  "output": ["json"],
  "json": {
    "schema": {
      "type": "object",
      "properties": {
        "title": { "type": "string" },
        "availability": {
          "type": "string",
          "enum": ["in_stock", "out_of_stock", "preorder", "discontinued"],
          "description": "Map the page wording to one of these standardized state strings"
        },
        "seller": {
          "type": "object",
          "properties": {
            "name": { "type": "string" },
            "rating": { "type": "number", "nullable": true }
          },
          "required": ["name"]
        }
      },
      "required": ["title", "availability"]
    }
  }
}
```

For example, `enum` is what makes the field comparable across sites: "In stock", "Available now" and "Ships today" all land on `in_stock`, so downstream code branches on one value instead of large number of string matches.

### 3. Article and blog metadata

Extract publishing metadata for structured database storage:

```json
{
  "url": "https://example.com/blog/article",
  "output": ["json"],
  "json": {
    "schema": {
      "type": "object",
      "properties": {
        "headline": { "type": "string" },
        "author": { "type": "string", "nullable": true },
        "published_at": {
          "type": "string",
          "description": "Publication date formatted as an ISO 8601 string (YYYY-MM-DD)"
        },
        "tags": { "type": "array", "items": { "type": "string" } }
      },
      "required": ["headline", "published_at"]
    }
  }
}
```

## Localized extraction

`location` applies directly to AI extractions. Pass ISO 3166-1 alpha-2 country codes to ensure localized prices and currencies are extracted correctly:

```bash
curl https://webapi.oxylabs.io/v1/scrape \
  -H "Authorization: Bearer $OXYLABS_WEB_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "url": "https://sandbox.oxylabs.io/products/1",
    "output": ["json"],
    "location": "DE",
    "json": {
      "schema": {
        "type": "object",
        "properties": {
          "price": { "type": "number" },
          "currency": { "type": "string" }
        },
        "required": ["price", "currency"]
      }
    }
  }'
```

{% hint style="info" %}
Without `location`, the `currency` you get back is whichever one the page decided to show.
{% endhint %}

## Response schema

Extracted JSON is returned inside the standard envelope under `results[].json`:

{% code expandable="true" %}

```json
{
  "state": "done",
  "results": [
    {
      "json": {
        "currency": "€",
        "price": 91.99
      },
      "metadata": {
        "page": 1,
        "run_js": false,
        "url": "https://sandbox.oxylabs.io/products/1",
        "statuses": {
          "http_request": {
            "code": 200,
            "tag": "HTTP_OK"
          },
          "json_parse": {
            "code": 12000,
            "tag": "PARSE_SUCCESS"
          }
        }
      }
    }
  ],
  "params": {
    "url": "https://sandbox.oxylabs.io/products/1",
    "output": [
      "json"
    ],
    "json": {
      "schema": {
        "properties": {
          "currency": {
            "type": "string"
          },
          "price": {
            "type": "number"
          }
        },
        "required": [
          "price",
          "currency"
        ],
        "type": "object"
      },
      "prompt": null
    },
    "location": "DE",
    "device": "desktop",
    "run_js": null
  },
  "metadata": {
    "timestamp": 1790854381,
    "request_id": "7511387695155825665"
  }
}
```

{% endcode %}

{% hint style="info" %}
Request `output` parameter is a list. You can request `output: ["json", "markdown"]` in a single request to receive both extracted fields and the full text. Unrequested format keys return `null`.
{% endhint %}

### Code examples

{% tabs %}
{% tab title="Python (schema-based)" %}

```python
import os
import json
import requests

API_URL = "https://webapi.oxylabs.io/v1/scrape"
HEADERS = {
    "Authorization": f"Bearer {os.environ['OXYLABS_WEB_API_KEY']}",
    "Content-Type": "application/json"
}

extraction_schema = {
    "type": "object",
    "properties": {
        "headline": {"type": "string"},
        "author": {"type": "string", "nullable": True},
        "publish_date": {
            "type": "string", 
            "description": "ISO 8601 formatted publication date string (YYYY-MM-DD)."
        },
        "word_count": {"type": "integer"}
    },
    "required": ["headline", "publish_date"]
}

payload = {
    "url": "https://example.com/blog/article-1",
    "output": ["json"],
    "json": {
        "schema": extraction_schema
    }
}

response = requests.post(API_URL, headers=HEADERS, json=payload, timeout=120)
response.raise_for_status()

parsed_data = response.json()["results"][0]["json"]
print(json.dumps(parsed_data, indent=2))
```

{% endtab %}

{% tab title="Node.js (prompt-based)" %}

```javascript
const API_URL = "https://webapi.oxylabs.io/v1/scrape";
const HEADERS = {
  "Authorization": `Bearer ${process.env.OXYLABS_WEB_API_KEY}`,
  "Content-Type": "application/json"
};

const payload = {
  url: "https://example.com/product/10928",
  output: ["json"],
  json: {
    prompt: "Extract product title, rating out of 5 stars, review count, and primary features."
  }
};

async function executeAiExtraction() {
  const response = await fetch(API_URL, {
    method: "POST",
    headers: HEADERS,
    body: JSON.stringify(payload)
  });

  if (!response.ok) {
    throw new Error(`HTTP Error ${response.status}: ${await response.text()}`);
  }

  const data = await response.json();
  const extractedJson = data.results[0].json;
  console.log("Extracted Data:", JSON.stringify(extractedJson, null, 2));
}

executeAiExtraction();
```

{% endtab %}
{% endtabs %}


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://developers.oxylabs.io/products/web-api/scrape/ai-parsing.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
