Skip to content

New Official SDKs for TypeScript, Python and Go

SerpKite
Get API key
Docs menu / Extract

Extract

Up to 20 URLs (HTML or PDF) per call as clean Markdown, text or HTML, with query-ranked highlights, links and images. Failed URLs are listed with a reason and cost nothing.

POST https://api.serpkite.com/v1/extract Credits: 1 per URL that comes back (0.5 from cache); failed URLs are free
View as Markdown

Extract reads many pages in one call: the same fetch and Markdown conversion as /v1/webpage, PDFs included, for up to 20 URLs at once, fetched in parallel. Duplicate URLs are dropped.

With query and highlights, each page also gets its passages most relevant to the query, ranked with BM25 (a lexical ranker: deterministic, no model involved). Use max_tokens to cap each page instead.

The response is 200 even when some or all URLs fail: each failure is listed in failed with a code (invalid_request, upstream_error, upstream_timeout, empty_page) and costs nothing.

Request body

ParameterTypeDefaultDescription
urls requiredstring[]1–20 public http(s) URLs of HTML pages or PDFs.
formatstringmarkdownWhat each result carries: markdown, text or the raw html. One of: markdown, text, html.
querystringWhat highlights are ranked against (up to 4,096 characters).
highlightsinteger0Passages per page ranked against query, 0–10 (about 600 characters each).
max_tokensintegerTrim each page's Markdown or text to about this many tokens (100–100,000).
include_linksbooleanfalseWebpage: also return the page's outbound links (url, text).
include_imagesbooleanfalseWebpage: also return the page's image URLs.
max_ageintegerAccept a cached page up to this many seconds old, for 0.5 credit.
countrystringFetch through an exit in this country (two-letter ISO code).
timeoutintegerSeconds to wait for the pages, 1–90 (default 50). A page still loading is listed in failed (upstream_timeout) and not charged.

Example request

Authenticate with Authorization: Bearer $SERPKITE_API_KEY (GET requests may pass ?api_key= instead). The official SDKs read SERPKITE_API_KEY for you. See Authentication.

curl https://api.serpkite.com/v1/extract \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://en.wikipedia.org/wiki/Web_crawler","https://www.rfc-editor.org/rfc/rfc9309.html"],"query":"robots.txt crawl delay","highlights":3}'

Response fields

FieldTypeDescription
requestobjecturls (deduplicated), format, query.
results[]arrayOne entry per page read: url (final, after redirects), status_code, title, published_at, metadata, cached, markdown / text / html, and links, images, highlights (text, score, heading) when asked for.
failed[]arrayURLs that could not be read: url, error (code, message). Not charged.
metaobjectrequest_id, credits_used, latency_ms, succeeded, failed.

Every billed response also carries credit and latency headers (X-Credits-Used, X-Credits-Remaining, X-Request-Id…).

Example response

200 OK · illustrative
{
  "request": {
    "endpoint": "extract",
    "urls": [
      "https://en.wikipedia.org/wiki/Web_crawler",
      "https://www.rfc-editor.org/rfc/rfc9309.html"
    ],
    "format": "markdown",
    "query": "robots.txt crawl delay"
  },
  "results": [
    {
      "url": "https://www.rfc-editor.org/rfc/rfc9309.html",
      "status_code": 200,
      "title": "RFC 9309: Robots Exclusion Protocol",
      "metadata": {
        "title": "RFC 9309: Robots Exclusion Protocol",
        "content_type": "text/html"
      },
      "cached": false,
      "markdown": "# Robots Exclusion Protocol\n\n…",
      "highlights": [
        {
          "text": "Crawlers SHOULD NOT use the cached version for more than 24 hours, unless the robots.txt is unreachable…",
          "score": 6.214,
          "heading": "Robots Exclusion Protocol > Caching"
        }
      ]
    }
  ],
  "failed": [
    {
      "url": "https://en.wikipedia.org/wiki/Web_crawler",
      "error": {
        "code": "upstream_timeout",
        "message": "the page did not load in time"
      }
    }
  ],
  "meta": {
    "request_id": "req_01J8ZK4M6Q2V7",
    "credits_used": 1,
    "latency_ms": 2210,
    "succeeded": 1,
    "failed": 1
  }
}

Errors

Errors use one shape: {"error":{"code","message","request_id"}}. Errors are never billed. Full list in Errors.

StatusCodeMeaning
400invalid_requestA parameter is missing or invalid.
401unauthorizedThe API key is missing, invalid or revoked.
402insufficient_creditsYour balance is too low. Buy a pack or wait for the monthly free grant.
429rate_limitedToo many requests per second for your plan. Retry after the Retry-After header.
503upstream_errorGoogle could not be fetched or parsed. Not billed; retry after Retry-After.

Notes

  • Pages are fetched in parallel, 5 at a time; the call takes as long as the slowest page or timeout.
  • The remote MCP server has an extract tool.