# Crawl

`POST https://api.serpkite.com/v1/crawl` · Credits: 1 per page read (0.5 from cache); the limit is reserved up front and the unused part refunded

Async crawl of one site: up to 1,000 pages and 10 link hops, from links and sitemaps, best first with a query, include/exclude path regexes, robots.txt and Crawl-delay respected. Poll GET /v1/crawl/{id} or get a signed crawl.completed webhook.

Crawl reads a whole site, or one section of it, in the background. It starts at `url`, follows links on the same host (optionally subdomains) up to `max_depth` hops, and unless `sitemap` is `skip` also reads the URLs from the site's sitemaps. Each page comes back as Markdown (or text) with its metadata, through the same fetch as [`/v1/webpage`](https://serpkite.com/docs/endpoints/webpage), PDFs included.

It is polite by default: every host's `robots.txt` is honoured (RFC 9309 rules and `Crawl-delay`; a `robots.txt` that answers `5xx` stops the crawl at the start page), URLs are canonicalised (fragments and tracking parameters dropped, `www.`, trailing slashes and redirects folded) so each page is read once, and at most 5 pages are fetched at a time. Breadth first, or best first with `query`: pages whose URL and link text match it are read first.

`POST /v1/crawl` returns `202` with a task. Poll `GET /v1/crawl/{id}` (`progress` while it runs, `result` when it ends; kept 24 hours) or receive the signed [`crawl.completed`](https://serpkite.com/docs/webhooks#other-events) webhook. `DELETE /v1/crawl/{id}` cancels it: a queued crawl is refunded at once, a running one stops at its next step and is charged for the pages read.

## Request body

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `url` **required** | string |  | Start page. Always read; path filters and robots.txt apply to the pages found from it. |
| `limit` | integer | `25` | Maximum pages read, 1–1,000. This many credits are reserved up front. |
| `max_depth` | integer | `2` | Link hops from `url`, 0–10 (sitemap pages count as 1 hop; 0 reads only `url`). |
| `include_paths` | string[] |  | Up to 20 regular expressions matched against the URL path; follow only paths matching one, e.g. `^/docs/`. |
| `exclude_paths` | string[] |  | Up to 20 regular expressions matched against the URL path; never follow paths matching one. |
| `include_subdomains` | boolean | `false` | Also crawl subdomains of the start page's host. |
| `sitemap` | string | `include` | `include` also seeds the crawl with the site's sitemap URLs (after the start page's links), `only` reads the start page and sitemap URLs without following links, `skip` follows links only. One of: `include`, `only`, `skip`. |
| `query` | string |  | Read the pages most relevant to these words first (best-first crawl, up to 512 characters). |
| `ignore_query_parameters` | boolean | `false` | Treat URLs that differ only in their query string as one page. |
| `format` | string | `markdown` | What each page carries. One of: `markdown`, `text`. |
| `include_links` | boolean | `false` | Return each page's outbound links. |
| `max_tokens` | integer |  | Trim each page to about this many tokens (100–100,000). |
| `max_age` | integer |  | Accept cached pages up to this many seconds old, for 0.5 credit each. |
| `webhook_url` | string |  | Receives the signed `crawl.completed` event. Defaults to the account webhook URL. |

## Example request

cURL:

```bash
curl https://api.serpkite.com/v1/crawl \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://docs.python.org/3/library/","limit":25,"max_depth":2,"include_paths":["^/3/library/asyncio"]}'
```

TypeScript:

```ts
import { SerpKite } from "serpkite";

const sk = new SerpKite(); // reads SERPKITE_API_KEY
const res = await sk.crawl({ url: "https://docs.python.org/3/library/", limit: 25, max_depth: 2, include_paths: ["^/3/library/asyncio"] });
const done = await sk.waitForCrawl(res);
console.log(done.status, done.result?.pages.length);
```

Python:

```python
from serpkite import SerpKite

sk = SerpKite()  # reads SERPKITE_API_KEY
res = sk.crawl("https://docs.python.org/3/library/", limit=25, max_depth=2, include_paths=["^/3/library/asyncio"])
done = sk.wait_for_crawl(res.id)
print(done.status, len(done.result.pages) if done.result else 0)
```

## Response fields

| Field | Type | Description |
| --- | --- | --- |
| `id` | string | The crawl's ID. `poll_url` is `https://api.serpkite.com/v1/crawl/{id}`. |
| `status` | string | `queued`, `running`, `completed`, `failed` or `canceled`. |
| `credits_reserved / credits_used` | number | The reservation (`limit` credits) and what the pages read cost; the difference is refunded when the crawl ends. |
| `progress` | object | While running: `pages_done`, `pages_failed`, `pages_queued`, `pages_discovered`, `limit`. |
| `result` | object | When it ends: `url`, `pages[]` (`url`, `depth`, `title`, `markdown` or `text`, `published_at`, `metadata`, `links`), `failed[]` (`url`, `error`) and `stats` (`pages`, `failed`, `seconds`, `discovered`, `queued`, `duplicates`, `sitemap_urls`, `robots`, `robots_disallowed`, `robots_blocked`, `crawl_delay_ms`, `stopped`). |
| `error` | object \| null | Set when the crawl failed (`no_pages`: nothing could be read, refunded) or was canceled. |

## Example response

```json
{
  "id": "01926a3e-7b2c-7c4e-9f1a-2b3c4d5e6f70",
  "kind": "crawl",
  "status": "completed",
  "created_at": "2026-10-03T12:00:00Z",
  "started_at": "2026-10-03T12:00:01Z",
  "completed_at": "2026-10-03T12:00:41Z",
  "credits_reserved": 25,
  "credits_used": 18,
  "webhook_status": null,
  "progress": {
    "pages_done": 18,
    "pages_failed": 0,
    "pages_queued": 0,
    "pages_discovered": 18,
    "limit": 25
  },
  "result": {
    "url": "https://docs.python.org/3/library/",
    "pages": [
      {
        "url": "https://docs.python.org/3/library/asyncio.html",
        "depth": 1,
        "title": "asyncio — Asynchronous I/O",
        "markdown": "# asyncio — Asynchronous I/O\n\nasyncio is a library to write concurrent code…"
      }
    ],
    "failed": [],
    "stats": {
      "pages": 18,
      "failed": 0,
      "seconds": 40,
      "discovered": 18,
      "queued": 0,
      "duplicates": 1,
      "sitemap_urls": 0,
      "robots": "found",
      "robots_disallowed": 0,
      "robots_blocked": 0,
      "crawl_delay_ms": 0,
      "stopped": "done"
    }
  },
  "error": null
}
```

## Errors

| Status | Code | Meaning |
| --- | --- | --- |
| 400 | `invalid_request` | A parameter is missing or invalid. |
| 401 | `unauthorized` | The API key is missing, invalid or revoked. |
| 402 | `insufficient_credits` | Your balance is too low. Buy a pack or wait for the monthly free grant. |
| 429 | `rate_limited` | Too many requests per second for your plan. Retry after the Retry-After header. |
| 503 | `upstream_error` | Google could not be fetched or parsed. Not billed; retry after Retry-After. |

All errors: https://serpkite.com/docs/errors

## Notes

- A crawl runs for at most 30 minutes and returns up to 16 MB of pages; `result.stats.stopped` says why it ended (`done`, `limit`, `time_limit`, `size_limit`, `too_many_failures`, `canceled`).
- At most 20 crawls can be queued or running per account (`429` beyond that). A crawl interrupted by a deploy goes back to the queue and restarts from the beginning.
- The remote MCP server has `crawl` and `crawl_result` tools.

## Related

- [Site ingestion guide](https://serpkite.com/docs/guides/site-ingestion)
- [Map](https://serpkite.com/docs/endpoints/map)
- [Webhooks](https://serpkite.com/docs/webhooks)