Skip to content

New Official SDKs for TypeScript, Python and Go

SerpKite
Get API key
Docs menu / Webpage to Markdown

Webpage to Markdown

Fetch any public URL and get clean Markdown plus title, description and metadata.

POST https://api.serpkite.com/v1/webpage Credits: 1 per page Try in playground
View as Markdown

Fetch any public URL and get clean Markdown and plain text plus metadata (title, description, language, canonical, author, published time). Built for feeding pages to an LLM. This is the one endpoint without a results list.

To fetch the top search results in the same call as the search, use include_content on /v1/search instead.

Request body

Send a JSON object with Content-Type: application/json. The same parameters also work as a query string on GET /v1/webpage. JSON arrays are rejected: to run many queries at once, use POST /v1/batches at half price. Unknown parameters return 400 invalid_request with a message that names the replacement; see strict validation.

ParameterTypeDefaultDescription
url requiredstringPublic http(s) URL to fetch. Required.
formatstringjsonmarkdown returns the page Markdown as text/markdown. One of: json, markdown.
include_htmlbooleanfalseWebpage: also return the raw HTML.
max_ageintegerAccept a cached result up to this many seconds old. Cache hits cost 50% of the credits.

Example request

Authenticate with Authorization: Bearer $SERPKITE_API_KEY (GET requests may pass ?api_key= instead). The official SDKs read SERPKITE_API_KEY for you. See Authentication.

curl https://api.serpkite.com/v1/webpage \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://en.wikipedia.org/wiki/Search_engine_results_page","format":"markdown"}'

Response fields

FieldTypeDescription
requestobjectThe normalised request, with defaults filled in: endpoint, engine, q, country, language, location, num, page, device, autocorrect…
urlstringFinal URL after redirects.
status_codeintegerHTTP status of the fetched page.
markdownstringMain content as Markdown.
textstringMain content as plain text.
htmlstringRaw HTML, only with include_html: true.
metadataobjecttitle, description, language, canonical, site_name, image, published_time, author.
metaobjectrequest_id, credits_used, cached, cached_at, engine (the provider that answered), route (provider attempts, see Search providers), latency_ms, parse_quality (ok, partial, empty), resolved_urls.

Every billed response also carries credit and latency headers (X-Credits-Used, X-Credits-Remaining, X-Request-Id…).

Example response

200 OK · illustrative
{
  "request": {
    "endpoint": "webpage",
    "engine": "google",
    "url": "https://en.wikipedia.org/wiki/Search_engine_results_page",
    "format": "json"
  },
  "url": "https://en.wikipedia.org/wiki/Search_engine_results_page",
  "status_code": 200,
  "markdown": "# Search engine results page\n\nA **search engine results page** (**SERP**) is a webpage that is displayed by a search engine in response to a query by a user…",
  "text": "Search engine results page. A search engine results page (SERP) is a webpage that is displayed by a search engine in response to a query by a user…",
  "metadata": {
    "title": "Search engine results page - Wikipedia",
    "language": "en",
    "canonical": "https://en.wikipedia.org/wiki/Search_engine_results_page",
    "site_name": "Wikipedia"
  },
  "meta": {
    "request_id": "req_01J8ZK4M6Q2V7",
    "credits_used": 1,
    "cached": false,
    "engine": "google",
    "latency_ms": 1034
  }
}

Errors

Errors use one shape: {"error":{"code","message","request_id"}}. Errors are never billed. Full list in Errors.

StatusCodeMeaning
400invalid_requestA parameter is missing or invalid.
401unauthorizedThe API key is missing, invalid or revoked.
402insufficient_creditsYour balance is too low. Buy a pack or wait for the monthly free grant.
429rate_limitedToo many requests per second for your plan. Retry after the Retry-After header.
503upstream_errorGoogle could not be fetched or parsed. Not billed; retry after Retry-After.

Notes

  • Only public pages are fetched: no logins, no paywalled content.
  • Pages that can't be fetched return 503 upstream_error and are not billed.