Skip to content

New Official SDKs for TypeScript, Python and Go

SerpKite
Get API key

POST /v1/webpage

Webpage to Markdown API

Send a URL, get the page's main content as clean Markdown plus title, description, author, published date and canonical URL. Navigation, cookie banners and ads are stripped, so your LLM reads the article, not the chrome.

Get a free API key Try in playground
1 credit per URL 2,500 free credits, no card Failed calls are free

Overview

What the Webpage to Markdown API does

The Webpage to Markdown API fetches a public URL through our proxy network, extracts the main content and returns it as Markdown. You also get text, metadata (title, description, language, canonical, site_name, image, published_time, author), the final url after redirects and the upstream status_code.

It's the second half of most agent search loops: search Google, then read the best results. If you only need the top results of a search, include_content=1…5 on /v1/search does both in one call.

What you get

markdown string
Main content as Markdown: headings, lists, tables, links and code blocks kept.
text string
Plain-text version, for embeddings or keyword search.
metadata object
title, description, language, canonical, site_name, image, published_time, author.
url string
Final URL after redirects.
status_code integer
HTTP status of the fetched page.
html string
Raw HTML, only with include_html=true.
meta object
request_id, credits_used, cached, engine, latency_ms, resolved_urls and parse_quality. parse_quality=empty means the call was not billed.

Request

curl https://api.serpkite.com/v1/webpage \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://en.wikipedia.org/wiki/Search_engine_results_page","format":"json"}'

Response (illustrative, placeholder domains)

200 OK · application/json
{
  "url": "https://en.wikipedia.org/wiki/Search_engine_results_page",
  "status_code": 200,
  "markdown": "# Search engine results page\n\nA **search engine results page** (SERP) is a webpage that is displayed by a search engine in response to a query by a user…\n\n## Components\n\n- Organic results\n- Sponsored results\n- Knowledge panels\n",
  "metadata": {
    "title": "Search engine results page - Wikipedia",
    "description": "Webpage displayed by a search engine in response to a query",
    "language": "en",
    "canonical": "https://en.wikipedia.org/wiki/Search_engine_results_page",
    "site_name": "Wikipedia"
  },
  "meta": {
    "request_id": "req_01J8ZK4M6Q2V7",
    "credits_used": 1,
    "engine": "google",
    "cached": false,
    "latency_ms": 812,
    "resolved_urls": true,
    "parse_quality": "ok"
  }
}

Every response also carries X-Credits-Used, X-Credits-Remaining, X-Cache and X-Tokens-Estimate headers. Run this query in the playground

Search, then read

A two-step RAG pipeline

Search Google for candidates, then pull each page as Markdown. Markdown is typically a fraction of the tokens of raw HTML, and the metadata gives you a title, canonical URL and published date for citations.

  • format=markdown returns text/markdown directly, no JSON unwrapping.
  • X-Tokens-Estimate tells you the size before you add it to a prompt.
  • include_content on /v1/search fetches the top 1–5 results in the same call.
  • Public pages only: nothing behind a login or paywall.
from serpkite import SerpKite

sk = SerpKite()  # reads SERPKITE_API_KEY

def search_and_read(query: str, k: int = 3) -> list[dict]:
    """Search Google, then fetch the top k results as Markdown (1 + k credits)."""
    serp = sk.search(query, fields="results.title,results.link")
    docs = []
    for r in serp.results[:k]:
        page = sk.webpage(r.link)
        docs.append({"url": page.url, "title": page.metadata.title, "markdown": page.markdown})
    return docs

# Shortcut: one call does the same with include_content (+1 credit per fetched page)
# sk.search(query, include_content=3)  → results[i].content holds the page Markdown

Reference

Parameters

The parameters /v1/webpage accepts, in the JSON body or the query string.

Webpage to Markdown API parameters
Name Type Default Description
url required string – Public http(s) URL to fetch.
format string json json returns markdown inside a JSON object; markdown returns text/markdown directly.
include_html boolean false Also return the raw HTML in html.
max_age integer – Accept a cached result up to this many seconds old. Cache hits cost 50% of the credits.

Full reference, error codes and headers are in the API docs. Send a JSON array of up to 100 request objects to run them in one call.

Use cases

What people build with the Webpage to Markdown API

01

RAG ingestion

Turn URLs into clean chunks for your vector store without writing scrapers.

02

Agent browsing

Give an agent a read_page tool that returns compact Markdown.

03

Article summarization

Summarize news or blog posts found with the Google News API.

04

Docs mirroring

Snapshot documentation pages as Markdown for offline or LLM use.

Pricing

Webpage to Markdown API pricing

Cost per call

1 credit

per URL

  • From $0.60 per 1,000 calls at volume.
  • Credits never expire. No subscription.
  • Failed, empty and blocked calls are refunded.
  • 2,500 free credits = 2,500 calls, then 1,000 credits a month.

For reference: Serper lists Google searches at $1.00 per 1k on its entry tier and $0.30 at its best tier, and its credits expire after 6 months. SerpKite credits cost $1.00 to $0.60 per 1k.

See all pricing
Webpage to Markdown API cost per pack
Pack Price Per 1k credits Per 1k calls Calls per pack
Starter $10 $1.00 $1.00 10,000
Growth $50 $0.80 $0.80 62,500
Pro $300 $0.60 $0.60 500,000

FAQ

Webpage to Markdown API: common questions

Start building

Start with the Webpage to Markdown API

2,500 free credits, then 1,000 every month. No credit card.