Skip to content

New Official SDKs for TypeScript, Python and Go

SerpKite
Get API key
Docs menu / Ingest a website or a set of URLs

Ingest a website or a set of URLs

Discover URLs with map, read selected HTML pages and PDFs with extract, or crawl a site asynchronously. SDK examples, partial failures, polling, cancellation and credit reservations.

View as Markdown

Use SerpKite’s site APIs to turn public web content into documents for a knowledge base, a research pipeline or an agent. Choose the endpoint by how much of the site you need:

Starting point Endpoint Result
One known URL /v1/webpage One page’s Markdown and metadata
Up to 20 known URLs /v1/extract Page content, optional ranked passages and individual failures
A site whose URLs you want to discover /v1/map URLs from sitemaps and the start page’s links
A site section you want to read /v1/crawl An asynchronous task containing the pages it read

All requests use the same Bearer API key. The site endpoints take JSON objects over POST; their options differ from the search endpoints. For example, format: "markdown" on extract or crawl selects the content field inside a JSON response. It does not turn the whole response into a Markdown string.

Discover URLs, then extract selected pages

Map discovers URLs without extracting every page. Its search option ranks URL paths and known titles against your terms; it does not search the pages’ full contents. include_paths and exclude_paths are regular expressions matched against the URL path.

The following Python example discovers Python documentation about asyncio, reads at most 20 matching pages and prints passages relevant to a question. Install the Python SDK and set SERPKITE_API_KEY first.

from serpkite import SerpKite

sk = SerpKite()
discovered = sk.map(
    "https://docs.python.org/3/",
    search="asyncio",
    include_paths=[r"^/3/library/"],
    limit=20,
)
urls = [row.url for row in discovered.results]

if urls:
    extracted = sk.extract(
        urls,
        query="How do asyncio task groups handle cancellation?",
        highlights=3,
        max_tokens=4000,
        include_links=True,
        timeout=50,
    )
    for page in extracted.results:
        print(page.url, page.title)
        for passage in page.highlights or []:
            print(passage.text)
    for failure in extracted.failed:
        print("Could not read:", failure.url, failure.error.code)
else:
    print("No matching URLs were discovered.")

The same workflow in TypeScript:

import { SerpKite } from "serpkite";

const sk = new SerpKite();
const discovered = await sk.map({
  url: "https://docs.python.org/3/",
  search: "asyncio",
  include_paths: ["^/3/library/"],
  limit: 20,
});
const urls = discovered.results.map((row) => row.url);

if (urls.length) {
  const extracted = await sk.extract({
    urls,
    query: "How do asyncio task groups handle cancellation?",
    highlights: 3,
    max_tokens: 4000,
    include_links: true,
    timeout: 50,
  });
  for (const page of extracted.results) {
    console.log(page.url, page.highlights);
  }
  for (const failure of extracted.failed) {
    console.error(failure.url, failure.error.code);
  }
}

An extract response is 200 even when some or all URLs fail. Check both results and failed; only successful pages are charged. URLs can redirect, so store the final results[].url with each document. For more than 20 URLs, divide your list into groups of 20 and respect your rate limit.

highlights are passages ranked with BM25, a lexical ranker, rather than generated summaries. Extract adds them alongside the page content. max_tokens caps Markdown or text per page. PDFs use their text layer; scanned PDFs without readable text produce no content. See Extract for formats and limits.

Crawl a section in the background

Crawl follows links and sitemap URLs and reads pages for you. Start with a small limit, then poll the task. The start URL is always read; path filters scope the pages discovered from it. Set the start URL to the section you intend to ingest.

from serpkite import SerpKite

sk = SerpKite()
task = sk.crawl(
    "https://docs.python.org/3/library/",
    limit=25,
    max_depth=2,
    include_paths=[r"^/3/library/"],
    query="asyncio cancellation",
    max_tokens=4000,
)
finished = sk.wait_for_crawl(task.id)

print(finished.status, finished.credits_used)
if finished.result:
    for page in finished.result.pages:
        print(page.url, page.markdown)
    print("Stopped because:", finished.result.stats.stopped)
if finished.error:
    print(finished.error.code, finished.error.message)

query prioritizes promising URLs and link text; it does not exclude every irrelevant page. max_depth is the number of link hops, up to 10; sitemap URLs count as one hop. sitemap: "only" reads the start page and sitemap URLs without following page links, while sitemap: "skip" follows links without sitemap discovery.

For direct HTTP use:

curl https://api.serpkite.com/v1/crawl \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url":"https://docs.python.org/3/library/",
    "limit":25,
    "max_depth":2,
    "include_paths":["^/3/library/"],
    "max_tokens":4000
  }'

The 202 response contains id and poll_url. Poll GET /v1/crawl/{id} until status is completed, failed or canceled. While running, progress reports counts; after it ends, result.pages, result.failed and result.stats describe what was read. A completed crawl can contain fewer than limit pages or individual failures, so inspect the result before indexing it.

Polling is free. Results are kept for 24 hours: save the documents in your own store before they expire. A webhook_url on the create request delivers a signed crawl.completed event; large results are omitted from the webhook with result_omitted: true, so fetch its poll_url. See Webhooks.

Cancel with DELETE /v1/crawl/{id} or sk.cancel_crawl(task.id). A queued crawl is refunded; a running crawl stops at its next checkpoint and charges for pages already read. Canceling an unknown task returns 404; a finished task returns 409.

Budget and retries

Operation Credits
Map 1 per nonempty call
Extract 1 per successfully read URL, or 0.5 from cache
Crawl 1 per successfully read page, or 0.5 from cache
Crawl polling and cancellation 0

Crawl reserves limit credits when queued and refunds the unused part when it ends. For example, limit: 25 reserves 25 credits; reading 18 live pages settles to 18 and refunds 7. Reservations count against your balance, key limit and spend cap.

max_age lets extract and crawl accept cached pages. Failed and empty pages cost nothing. Map does not accept max_age. The batch lane accepts search verticals and webpage requests; it does not accept extract, map or crawl.

The SDKs retry extract and crawl creation only on 429. After a lost response, automatically repeating either request can charge for another extraction or start another crawl. The batch API’s Idempotency-Key support does not extend to these endpoints.

Crawls honor robots.txt and Crawl-delay, which can reduce coverage or slow a task. SerpKite reads public, logged-out content; private addresses and authenticated pages are unavailable.

Last updated: 2026-10-04