> For the complete documentation index, see [llms.txt](https://developers.oxylabs.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developers.oxylabs.io/products/web-api/for-agents.md).

# For Agents

Everything an AI agent needs to integrate the Oxylabs Web API, on one page.

**Installing beats building.** If this project has no code yet, or the web access is for *you* rather than for the project's own code, one command is the whole task:

```bash
npx skills add thevastas/oxy_skills
```

That installs the skills; the [MCP server](https://developers.oxylabs.io/ai-workflows/mcp) adds typed tools, and the two [compose](/products/web-api.md#skills-or-mcp). Build a client from this page when the *project's* code has to make the calls itself — everything you need is below.

## What you get

Two endpoints on `https://webapi.oxylabs.io`:

| Endpoint          | Returns                             | Purpose                                                |
| ----------------- | ----------------------------------- | ------------------------------------------------------ |
| `POST /v1/search` | Ranked results: title, snippet, URL | Find which pages hold the answer                       |
| `POST /v1/scrape` | The content of one page             | Read a page, including JS-heavy and bot-protected ones |

**The core pattern is search → scrape.** Search tells you where the answer lives; scrape gets you the answer. Search snippets are truncated and often stale — they locate sources, they don't substitute for them.

## Authentication

Send `Authorization: Bearer $OXYLABS_WEB_API_KEY` on every call. If the variable is not set, **stop and ask the human for it** — do not invent a key, and do not fall back to fetching the web unauthenticated. A key that `401`s despite looking valid is usually a key for a different Oxylabs product. See [Authentication](/products/web-api/authentication.md).

## `POST /v1/search`

```bash
curl https://webapi.oxylabs.io/v1/search \
  -H "Authorization: Bearer $OXYLABS_WEB_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"query": "eu ai act compliance deadlines", "max_results": 5}'
```

| Field         | Type    | Required | Default | Constraint                                           |
| ------------- | ------- | -------- | ------- | ---------------------------------------------------- |
| `query`       | string  | yes      | —       | 1–2048 characters                                    |
| `max_results` | integer | no       | `10`    | 1–20                                                 |
| `location`    | string  | no       | `null`  | ISO 3166-1 alpha-2 country code, e.g. `"DE"`, `"US"` |

Unknown fields are rejected, not ignored. Response (`200`) — `state` is `done`, and the arrays are always present, empty meaning nothing matched:

```json
{
  "state": "done",
  "results": [
    {
      "title": "Implementation Timeline | EU Artificial Intelligence Act",
      "outline": "Date 2 August 2025 Providers: need to be compliant by 2 August 2027...",
      "url": "https://artificialintelligenceact.eu/implementation-timeline/",
      "metadata": { "position": 1 }
    }
  ],
  "related_questions": [],
  "related_searches": [{ "query": "ai act timeline" }],
  "params": { "query": "...", "location": null, "max_results": 5 },
  "metadata": { "timestamp": 1787431568, "request_id": "1788875675880097469" }
}
```

## `POST /v1/scrape`

```bash
curl https://webapi.oxylabs.io/v1/scrape \
  -H "Authorization: Bearer $OXYLABS_WEB_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"url": "https://artificialintelligenceact.eu/implementation-timeline/",
       "output": ["markdown"]}'
```

| Field      | Type    | Required | Notes                                                                                                                             |
| ---------- | ------- | -------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `url`      | string  | yes      | absolute `http`/`https` URL                                                                                                       |
| `output`   | array   | no       | `["markdown"]`, `["html"]`, `["json"]`, `["screenshot"]`, or a combination; default `["html"]`; `screenshot` needs `run_js: true` |
| `json`     | object  | no       | with `output: ["json"]`, `{"prompt": "fields to extract"}`                                                                        |
| `location` | string  | no       | two-letter country code, e.g. `"DE"`                                                                                              |
| `device`   | string  | no       | `"desktop"` or `"mobile"`                                                                                                         |
| `run_js`   | boolean | no       | execute page JavaScript                                                                                                           |

Ask for `output: ["markdown"]` when reading a page; `["html"]` only when you need the markup itself, and `["json"]` with a `json.prompt` for named fields. The response uses the same envelope: `results[]` holds the page in the requested format, plus the final `url` after redirects. Read it defensively — the fields vary by target and format.

## Working integration

```python
"""Oxylabs Web API client: search, scrape, and the research loop over both."""

import os
import random
import time

import requests

API = "https://webapi.oxylabs.io"
RETRYABLE = {408, 429, 500, 502, 503, 504}


class WebApiError(RuntimeError):
    pass


class AuthenticationError(WebApiError):
    pass


def _key() -> str:
    key = os.environ.get("OXYLABS_WEB_API_KEY", "").strip()
    if not key:
        raise WebApiError("OXYLABS_WEB_API_KEY is not set — ask the user for a key.")
    return key


def _post(path: str, payload: dict, attempts: int = 3) -> dict:
    headers = {"Authorization": f"Bearer {_key()}", "Content-Type": "application/json"}
    for attempt in range(1, attempts + 1):
        try:
            resp = requests.post(f"{API}{path}", json=payload, headers=headers, timeout=120)
        except requests.Timeout as exc:
            if attempt == attempts:
                raise WebApiError(f"{path} -> timed out after {attempts} attempts") from exc
            time.sleep(2**attempt + random.random())
            continue
        if resp.ok:
            return resp.json()
        if resp.status_code == 401:
            raise AuthenticationError("401 — the API key was rejected. Stop and tell the user.")
        if resp.status_code not in RETRYABLE or attempt == attempts:
            try:
                problem = resp.json()
                detail = f"{problem['title']}: {problem['detail']}"
            except (ValueError, KeyError):
                detail = resp.text[:300]
            raise WebApiError(f"{path} -> {resp.status_code}: {detail}")
        time.sleep(2**attempt + random.random())  # jitter: avoid lockstep retries
    raise WebApiError("unreachable")


def search(query: str, max_results: int = 10, location: str | None = None) -> list[dict]:
    payload = {"query": query, "max_results": max_results}
    if location:
        payload["location"] = location
    return _post("/v1/search", payload)["results"]


def scrape(url: str, fmt: str = "markdown", location: str | None = None) -> dict:
    payload: dict = {"url": url, "output": [fmt]}
    if location:
        payload["location"] = location
    results = _post("/v1/scrape", payload).get("results", [])
    if not results or results[0].get(fmt) is None:
        raise WebApiError(f"/v1/scrape -> no {fmt} result for {url}")
    return results[0]


def research(question: str, read_top: int = 2, location: str | None = None) -> list[dict]:
    """Search, then read the top results. Returns [{url, title, content}] with sources kept."""
    pages = []
    for hit in search(question, max_results=10, location=location)[:read_top]:
        try:
            page = scrape(hit["url"], location=location)
        except AuthenticationError:
            raise
        except WebApiError as exc:
            # A dead URL is a fact to report, not a reason to lose the other sources.
            pages.append({"url": hit["url"], "title": hit["title"], "error": str(exc)})
            continue
        pages.append({"url": hit["url"], "title": hit["title"], "content": page["markdown"]})
    return pages
```

## Errors

| Status | Meaning                                                                                                            | Your move                                                                                                                                           |
| ------ | ------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------- |
| `400`  | Validation failed                                                                                                  | Read `errors[].pointer` and `errors[].detail`, fix those fields. **Never retry unchanged.**                                                         |
| `401`  | A credential problem, never a request field: no key, a malformed `Authorization` header, or a key the API rejects. | Stop. Tell the user. Retrying cannot help.                                                                                                          |
| `408`  | The scrape timed out server-side                                                                                   | Retry with backoff — the request never reached a result.                                                                                            |
| `429`  | Rate limit **or** spent quota — the body message says which, but it is prose, so don't match on it                 | Back off and lower concurrency. Cap total retries: if it survives full backoff, it is a quota problem for the user, not something to keep retrying. |
| `5xx`  | Transient upstream failure                                                                                         | Retry up to 3× with backoff.                                                                                                                        |

Failures use an appropriate non-`2xx` status and an RFC 9457 problem body. If every search engine failed, `/v1/search` answers `500` with `"state": "faulted"` and `"title": "REQUEST_FAILED_AFTER_MANY_RETRIES"`; it is not charged, so retry it as you would any `5xx`, then surface it. An empty `results` array on a `200` just means nothing matched. See [Errors](/products/web-api/troubleshooting.md).

Errors carry the code in `title`; branch on that, never on `detail`:

```json
{
  "status": 400,
  "title": "VALIDATION_ERROR",
  "detail": "1 request field is invalid; fix every entry in `errors` and resend.",
  "instance": "6285041d-b00d8d69959c4b36f50fb53b",
  "metadata": { "timestamp": 1788762578, "request_id": "1788875675880097469" },
  "errors": [
    { "pointer": "#/max_results", "detail": "Input should be less than or equal to 20" }
  ]
}
```

Set client timeouts around **120 s**. A scrape renders a real page; a 10-second timeout fails work that would have succeeded and you still pay for it.

## Rules that make the difference

| Rule                                                                                                                                                                                                        | Why                                                                                                 |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| Search to find, scrape to read — one search, then the 1–3 URLs that plausibly answer it                                                                                                                     | Scraping all ten costs 5× more for a worse answer                                                   |
| Never answer from `outline`                                                                                                                                                                                 | Truncated mid-sentence, and it drops the qualifier that changes the meaning                         |
| Read pages as `output: ["markdown"]`                                                                                                                                                                        | A fraction of the tokens of HTML, structure intact; converting client-side is wasted work           |
| Keep the URL attached to every fact as you collect it                                                                                                                                                       | Citations reconstructed afterwards are how wrong attributions happen                                |
| Never fill a gap from memory — say it is unconfirmed                                                                                                                                                        | Substituting recollection for a source is what makes an integration untrustworthy                   |
| One question per search                                                                                                                                                                                     | Compound queries return results that match neither half                                             |
| Pass `location` when the answer is geographic                                                                                                                                                               | Prices, availability and rankings differ by country; an unqualified answer is wrong somewhere       |
| A scrape came back empty or skeletal: retry it once with `run_js: true`                                                                                                                                     | The page rendered client-side and you got the shell; a render is slow, so never send it by default  |
| Still empty and the site is on a country TLD: retry once more with `run_js: true` plus `location` for that country, `.lt` → `LT`, `.co.uk` → `GB`. Skip brand-use ccTLDs such as `.io`, `.ai`, `.co`, `.me` | Some sites only serve their own country; after that third attempt the page is unreadable, report it |
| Date anything time-sensitive                                                                                                                                                                                | Prices, versions and rankings were true when the page was written                                   |
| Report failures out loud                                                                                                                                                                                    | "The scrape of X failed" is a real result; dropping it quietly is not                               |
| Don't re-scrape a URL inside one task                                                                                                                                                                       | Cache by URL                                                                                        |
| Keep the key server-side, out of source, logs and error messages                                                                                                                                            | It is a spending credential                                                                         |

## Discovery

Don't hardcode capabilities: `GET /v1/scrapers` lists the implemented scrape endpoints and `OPTIONS` on one returns its parameters with types, so new capabilities appear without a code change. See [Discovery](/products/web-api/scrape/endpoint-discovery.md).

## Integration checklist

For the client branch — the rules above are the judgment, this is the mechanics:

* [ ] Key read from `OXYLABS_WEB_API_KEY`, never hardcoded, never logged
* [ ] `Authorization: Bearer` header on every request
* [ ] Client timeout \~120 s
* [ ] Success detected as any `2xx`, never as a literal `201`
* [ ] `query` kept under 2048 characters, `location` an ISO 3166-1 alpha-2 country code
* [ ] Retries on transport timeouts and `408`/`429`/`5xx` only, exponential backoff with jitter, max \~3 attempts
* [ ] `400`/`401` surfaced with the field name or a clear message, not retried
* [ ] `search` capped at `max_results` ≤ 20
* [ ] `metadata.request_id` logged, so failures are reportable
* [ ] Tests set a dummy key of their own, so a failing assertion prints that and never the real one


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://developers.oxylabs.io/products/web-api/for-agents.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `build a script that syncs our docs to a CMS` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
