# Extract

`POST https://api.serpkite.com/v1/extract` · Credits: 1 per URL that comes back (0.5 from cache); failed URLs are free

Up to 20 URLs (HTML or PDF) per call as clean Markdown, text or HTML, with query-ranked highlights, links and images. Failed URLs are listed with a reason and cost nothing.

Extract reads many pages in one call: the same fetch and Markdown conversion as [`/v1/webpage`](https://serpkite.com/docs/endpoints/webpage), PDFs included, for up to 20 URLs at once, fetched in parallel. Duplicate URLs are dropped.

With `query` and `highlights`, each page also gets its passages most relevant to the query, ranked with BM25 (a lexical ranker: deterministic, no model involved). Use `max_tokens` to cap each page instead.

The response is `200` even when some or all URLs fail: each failure is listed in `failed` with a code (`invalid_request`, `upstream_error`, `upstream_timeout`, `empty_page`) and costs nothing.

## Request body

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `urls` **required** | string[] |  | 1–20 public http(s) URLs of HTML pages or PDFs. |
| `format` | string | `markdown` | What each result carries: `markdown`, `text` or the raw `html`. One of: `markdown`, `text`, `html`. |
| `query` | string |  | What `highlights` are ranked against (up to 4,096 characters). |
| `highlights` | integer | `0` | Passages per page ranked against `query`, 0–10 (about 600 characters each). |
| `max_tokens` | integer |  | Trim each page's Markdown or text to about this many tokens (100–100,000). |
| `include_links` | boolean | `false` | Webpage: also return the page's outbound links (url, text). |
| `include_images` | boolean | `false` | Webpage: also return the page's image URLs. |
| `max_age` | integer |  | Accept a cached page up to this many seconds old, for 0.5 credit. |
| `country` | string |  | Fetch through an exit in this country (two-letter ISO code). |
| `timeout` | integer |  | Seconds to wait for the pages, 1–90 (default 50). A page still loading is listed in `failed` (`upstream_timeout`) and not charged. |

## Example request

cURL:

```bash
curl https://api.serpkite.com/v1/extract \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://en.wikipedia.org/wiki/Web_crawler","https://www.rfc-editor.org/rfc/rfc9309.html"],"query":"robots.txt crawl delay","highlights":3}'
```

TypeScript:

```ts
import { SerpKite } from "serpkite";

const sk = new SerpKite(); // reads SERPKITE_API_KEY
const res = await sk.extract({ urls: ["https://en.wikipedia.org/wiki/Web_crawler","https://www.rfc-editor.org/rfc/rfc9309.html"], query: "robots.txt crawl delay", highlights: 3 });
console.log(res.results.map((r) => r.title), res.failed.length);
```

Python:

```python
from serpkite import SerpKite

sk = SerpKite()  # reads SERPKITE_API_KEY
res = sk.extract(["https://en.wikipedia.org/wiki/Web_crawler","https://www.rfc-editor.org/rfc/rfc9309.html"], query="robots.txt crawl delay", highlights=3)
print([r.title for r in res.results], len(res.failed))
```

## Response fields

| Field | Type | Description |
| --- | --- | --- |
| `request` | object | `urls` (deduplicated), `format`, `query`. |
| `results[]` | array | One entry per page read: `url` (final, after redirects), `status_code`, `title`, `published_at`, `metadata`, `cached`, `markdown` / `text` / `html`, and `links`, `images`, `highlights` (`text`, `score`, `heading`) when asked for. |
| `failed[]` | array | URLs that could not be read: `url`, `error` (`code`, `message`). Not charged. |
| `meta` | object | `request_id`, `credits_used`, `latency_ms`, `succeeded`, `failed`. |

## Example response

```json
{
  "request": {
    "endpoint": "extract",
    "urls": [
      "https://en.wikipedia.org/wiki/Web_crawler",
      "https://www.rfc-editor.org/rfc/rfc9309.html"
    ],
    "format": "markdown",
    "query": "robots.txt crawl delay"
  },
  "results": [
    {
      "url": "https://www.rfc-editor.org/rfc/rfc9309.html",
      "status_code": 200,
      "title": "RFC 9309: Robots Exclusion Protocol",
      "metadata": {
        "title": "RFC 9309: Robots Exclusion Protocol",
        "content_type": "text/html"
      },
      "cached": false,
      "markdown": "# Robots Exclusion Protocol\n\n…",
      "highlights": [
        {
          "text": "Crawlers SHOULD NOT use the cached version for more than 24 hours, unless the robots.txt is unreachable…",
          "score": 6.214,
          "heading": "Robots Exclusion Protocol > Caching"
        }
      ]
    }
  ],
  "failed": [
    {
      "url": "https://en.wikipedia.org/wiki/Web_crawler",
      "error": {
        "code": "upstream_timeout",
        "message": "the page did not load in time"
      }
    }
  ],
  "meta": {
    "request_id": "req_01J8ZK4M6Q2V7",
    "credits_used": 1,
    "latency_ms": 2210,
    "succeeded": 1,
    "failed": 1
  }
}
```

## Errors

| Status | Code | Meaning |
| --- | --- | --- |
| 400 | `invalid_request` | A parameter is missing or invalid. |
| 401 | `unauthorized` | The API key is missing, invalid or revoked. |
| 402 | `insufficient_credits` | Your balance is too low. Buy a pack or wait for the monthly free grant. |
| 429 | `rate_limited` | Too many requests per second for your plan. Retry after the Retry-After header. |
| 503 | `upstream_error` | Google could not be fetched or parsed. Not billed; retry after Retry-After. |

All errors: https://serpkite.com/docs/errors

## Notes

- Pages are fetched in parallel, 5 at a time; the call takes as long as the slowest page or `timeout`.
- The remote MCP server has an `extract` tool.

## Related

- [Site ingestion guide](https://serpkite.com/docs/guides/site-ingestion)
- [Webpage](https://serpkite.com/docs/endpoints/webpage)
- [Map](https://serpkite.com/docs/endpoints/map)