Skip to content

New Official SDKs for TypeScript, Python and Go

SerpKite
Get API key
Docs menu / Crawl

Crawl

Async crawl of one site: up to 1,000 pages and 10 link hops, from links and sitemaps, best first with a query, include/exclude path regexes, robots.txt and Crawl-delay respected. Poll GET /v1/crawl/{id} or get a signed crawl.completed webhook.

POST https://api.serpkite.com/v1/crawl Credits: 1 per page read (0.5 from cache); the limit is reserved up front and the unused part refunded
View as Markdown

Crawl reads a whole site, or one section of it, in the background. It starts at url, follows links on the same host (optionally subdomains) up to max_depth hops, and unless sitemap is skip also reads the URLs from the site's sitemaps. Each page comes back as Markdown (or text) with its metadata, through the same fetch as /v1/webpage, PDFs included.

It is polite by default: every host's robots.txt is honoured (RFC 9309 rules and Crawl-delay; a robots.txt that answers 5xx stops the crawl at the start page), URLs are canonicalised (fragments and tracking parameters dropped, www., trailing slashes and redirects folded) so each page is read once, and at most 5 pages are fetched at a time. Breadth first, or best first with query: pages whose URL and link text match it are read first.

POST /v1/crawl returns 202 with a task. Poll GET /v1/crawl/{id} (progress while it runs, result when it ends; kept 24 hours) or receive the signed crawl.completed webhook. DELETE /v1/crawl/{id} cancels it: a queued crawl is refunded at once, a running one stops at its next step and is charged for the pages read.

Request body

ParameterTypeDefaultDescription
url requiredstringStart page. Always read; path filters and robots.txt apply to the pages found from it.
limitinteger25Maximum pages read, 1–1,000. This many credits are reserved up front.
max_depthinteger2Link hops from url, 0–10 (sitemap pages count as 1 hop; 0 reads only url).
include_pathsstring[]Up to 20 regular expressions matched against the URL path; follow only paths matching one, e.g. ^/docs/.
exclude_pathsstring[]Up to 20 regular expressions matched against the URL path; never follow paths matching one.
include_subdomainsbooleanfalseAlso crawl subdomains of the start page's host.
sitemapstringincludeinclude also seeds the crawl with the site's sitemap URLs (after the start page's links), only reads the start page and sitemap URLs without following links, skip follows links only. One of: include, only, skip.
querystringRead the pages most relevant to these words first (best-first crawl, up to 512 characters).
ignore_query_parametersbooleanfalseTreat URLs that differ only in their query string as one page.
formatstringmarkdownWhat each page carries. One of: markdown, text.
include_linksbooleanfalseReturn each page's outbound links.
max_tokensintegerTrim each page to about this many tokens (100–100,000).
max_ageintegerAccept cached pages up to this many seconds old, for 0.5 credit each.
webhook_urlstringReceives the signed crawl.completed event. Defaults to the account webhook URL.

Example request

Authenticate with Authorization: Bearer $SERPKITE_API_KEY (GET requests may pass ?api_key= instead). The official SDKs read SERPKITE_API_KEY for you. See Authentication.

curl https://api.serpkite.com/v1/crawl \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://docs.python.org/3/library/","limit":25,"max_depth":2,"include_paths":["^/3/library/asyncio"]}'

Response fields

FieldTypeDescription
idstringThe crawl's ID. poll_url is https://api.serpkite.com/v1/crawl/{id}.
statusstringqueued, running, completed, failed or canceled.
credits_reserved / credits_usednumberThe reservation (limit credits) and what the pages read cost; the difference is refunded when the crawl ends.
progressobjectWhile running: pages_done, pages_failed, pages_queued, pages_discovered, limit.
resultobjectWhen it ends: url, pages[] (url, depth, title, markdown or text, published_at, metadata, links), failed[] (url, error) and stats (pages, failed, seconds, discovered, queued, duplicates, sitemap_urls, robots, robots_disallowed, robots_blocked, crawl_delay_ms, stopped).
errorobject | nullSet when the crawl failed (no_pages: nothing could be read, refunded) or was canceled.

Every billed response also carries credit and latency headers (X-Credits-Used, X-Credits-Remaining, X-Request-Id…).

Example response

200 OK · illustrative
{
  "id": "01926a3e-7b2c-7c4e-9f1a-2b3c4d5e6f70",
  "kind": "crawl",
  "status": "completed",
  "created_at": "2026-10-03T12:00:00Z",
  "started_at": "2026-10-03T12:00:01Z",
  "completed_at": "2026-10-03T12:00:41Z",
  "credits_reserved": 25,
  "credits_used": 18,
  "webhook_status": null,
  "progress": {
    "pages_done": 18,
    "pages_failed": 0,
    "pages_queued": 0,
    "pages_discovered": 18,
    "limit": 25
  },
  "result": {
    "url": "https://docs.python.org/3/library/",
    "pages": [
      {
        "url": "https://docs.python.org/3/library/asyncio.html",
        "depth": 1,
        "title": "asyncio — Asynchronous I/O",
        "markdown": "# asyncio — Asynchronous I/O\n\nasyncio is a library to write concurrent code…"
      }
    ],
    "failed": [],
    "stats": {
      "pages": 18,
      "failed": 0,
      "seconds": 40,
      "discovered": 18,
      "queued": 0,
      "duplicates": 1,
      "sitemap_urls": 0,
      "robots": "found",
      "robots_disallowed": 0,
      "robots_blocked": 0,
      "crawl_delay_ms": 0,
      "stopped": "done"
    }
  },
  "error": null
}

Errors

Errors use one shape: {"error":{"code","message","request_id"}}. Errors are never billed. Full list in Errors.

StatusCodeMeaning
400invalid_requestA parameter is missing or invalid.
401unauthorizedThe API key is missing, invalid or revoked.
402insufficient_creditsYour balance is too low. Buy a pack or wait for the monthly free grant.
429rate_limitedToo many requests per second for your plan. Retry after the Retry-After header.
503upstream_errorGoogle could not be fetched or parsed. Not billed; retry after Retry-After.

Notes

  • A crawl runs for at most 30 minutes and returns up to 16 MB of pages; result.stats.stopped says why it ended (done, limit, time_limit, size_limit, too_many_failures, canceled).
  • At most 20 crawls can be queued or running per account (429 beyond that). A crawl interrupted by a deploy goes back to the queue and restarts from the beginning.
  • The remote MCP server has crawl and crawl_result tools.