Crawl
Async crawl of one site: up to 1,000 pages and 10 link hops, from links and sitemaps, best first with a query, include/exclude path regexes, robots.txt and Crawl-delay respected. Poll GET /v1/crawl/{id} or get a signed crawl.completed webhook.
POST
https://api.serpkite.com/v1/crawl
Credits: 1 per page read (0.5 from cache); the limit is reserved up front and the unused part refunded
Crawl reads a whole site, or one section of it, in the background. It starts at url, follows links on the same host (optionally subdomains) up to max_depth hops, and unless sitemap is skip also reads the URLs from the site's sitemaps. Each page comes back as Markdown (or text) with its metadata, through the same fetch as /v1/webpage, PDFs included.
It is polite by default: every host's robots.txt is honoured (RFC 9309 rules and Crawl-delay; a robots.txt that answers 5xx stops the crawl at the start page), URLs are canonicalised (fragments and tracking parameters dropped, www., trailing slashes and redirects folded) so each page is read once, and at most 5 pages are fetched at a time. Breadth first, or best first with query: pages whose URL and link text match it are read first.
POST /v1/crawl returns 202 with a task. Poll GET /v1/crawl/{id} (progress while it runs, result when it ends; kept 24 hours) or receive the signed crawl.completed webhook. DELETE /v1/crawl/{id} cancels it: a queued crawl is refunded at once, a running one stops at its next step and is charged for the pages read.
Request body
| Parameter | Type | Default | Description |
|---|---|---|---|
url required | string | Start page. Always read; path filters and robots.txt apply to the pages found from it. | |
limit | integer | 25 | Maximum pages read, 1–1,000. This many credits are reserved up front. |
max_depth | integer | 2 | Link hops from url, 0–10 (sitemap pages count as 1 hop; 0 reads only url). |
include_paths | string[] | Up to 20 regular expressions matched against the URL path; follow only paths matching one, e.g. ^/docs/. |
|
exclude_paths | string[] | Up to 20 regular expressions matched against the URL path; never follow paths matching one. | |
include_subdomains | boolean | false | Also crawl subdomains of the start page's host. |
sitemap | string | include | include also seeds the crawl with the site's sitemap URLs (after the start page's links), only reads the start page and sitemap URLs without following links, skip follows links only. One of: include, only, skip. |
query | string | Read the pages most relevant to these words first (best-first crawl, up to 512 characters). | |
ignore_query_parameters | boolean | false | Treat URLs that differ only in their query string as one page. |
format | string | markdown | What each page carries. One of: markdown, text. |
include_links | boolean | false | Return each page's outbound links. |
max_tokens | integer | Trim each page to about this many tokens (100–100,000). | |
max_age | integer | Accept cached pages up to this many seconds old, for 0.5 credit each. | |
webhook_url | string | Receives the signed crawl.completed event. Defaults to the account webhook URL. |
Example request
Authenticate with Authorization: Bearer $SERPKITE_API_KEY (GET requests may pass
?api_key= instead). The official SDKs read
SERPKITE_API_KEY for you. See Authentication.
curl https://api.serpkite.com/v1/crawl \
-H "Authorization: Bearer $SERPKITE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://docs.python.org/3/library/","limit":25,"max_depth":2,"include_paths":["^/3/library/asyncio"]}'Response fields
| Field | Type | Description |
|---|---|---|
id | string | The crawl's ID. poll_url is https://api.serpkite.com/v1/crawl/{id}. |
status | string | queued, running, completed, failed or canceled. |
credits_reserved / credits_used | number | The reservation (limit credits) and what the pages read cost; the difference is refunded when the crawl ends. |
progress | object | While running: pages_done, pages_failed, pages_queued, pages_discovered, limit. |
result | object | When it ends: url, pages[] (url, depth, title, markdown or text, published_at, metadata, links), failed[] (url, error) and stats (pages, failed, seconds, discovered, queued, duplicates, sitemap_urls, robots, robots_disallowed, robots_blocked, crawl_delay_ms, stopped). |
error | object | null | Set when the crawl failed (no_pages: nothing could be read, refunded) or was canceled. |
Every billed response also carries credit and latency headers
(X-Credits-Used, X-Credits-Remaining, X-Request-Id…).
Example response
{
"id": "01926a3e-7b2c-7c4e-9f1a-2b3c4d5e6f70",
"kind": "crawl",
"status": "completed",
"created_at": "2026-10-03T12:00:00Z",
"started_at": "2026-10-03T12:00:01Z",
"completed_at": "2026-10-03T12:00:41Z",
"credits_reserved": 25,
"credits_used": 18,
"webhook_status": null,
"progress": {
"pages_done": 18,
"pages_failed": 0,
"pages_queued": 0,
"pages_discovered": 18,
"limit": 25
},
"result": {
"url": "https://docs.python.org/3/library/",
"pages": [
{
"url": "https://docs.python.org/3/library/asyncio.html",
"depth": 1,
"title": "asyncio — Asynchronous I/O",
"markdown": "# asyncio — Asynchronous I/O\n\nasyncio is a library to write concurrent code…"
}
],
"failed": [],
"stats": {
"pages": 18,
"failed": 0,
"seconds": 40,
"discovered": 18,
"queued": 0,
"duplicates": 1,
"sitemap_urls": 0,
"robots": "found",
"robots_disallowed": 0,
"robots_blocked": 0,
"crawl_delay_ms": 0,
"stopped": "done"
}
},
"error": null
}Errors
Errors use one shape: {"error":{"code","message","request_id"}}. Errors are never billed.
Full list in Errors.
| Status | Code | Meaning |
|---|---|---|
| 400 | invalid_request | A parameter is missing or invalid. |
| 401 | unauthorized | The API key is missing, invalid or revoked. |
| 402 | insufficient_credits | Your balance is too low. Buy a pack or wait for the monthly free grant. |
| 429 | rate_limited | Too many requests per second for your plan. Retry after the Retry-After header. |
| 503 | upstream_error | Google could not be fetched or parsed. Not billed; retry after Retry-After. |
Notes
- A crawl runs for at most 30 minutes and returns up to 16 MB of pages;
result.stats.stoppedsays why it ended (done,limit,time_limit,size_limit,too_many_failures,canceled). - At most 20 crawls can be queued or running per account (
429beyond that). A crawl interrupted by a deploy goes back to the queue and restarts from the beginning. - The remote MCP server has
crawlandcrawl_resulttools.