Contents Extraction

Pull clean Markdown or HTML from any URL—ideal for LLM ingestion.

View as MarkdownOpen in Claude

Hand the Contents API a list of URLs and get back clean Markdown (or HTML) plus page metadata—no headless browser, no parsing. It’s the fastest way to turn arbitrary web pages into LLM-ready text without standing up your own crawler.


What You’ll Build

A one-call extractor that turns any URL into clean Markdown—title, body, and page metadata in a single response. The formats parameter lets you ask for Markdown, raw HTML, or both, and crawl_timeout caps the wait per page.


Try It Live

Run a real Contents API request right here—no setup. Open the Try It panel below, paste your API key, drop in a URL, and send it against the live endpoint.

POST
/v1/contents
curl -X POST https://ydc-index.io/v1/contents \
-H "X-API-Key: <apiKey>" \
-H "Content-Type: application/json" \
-d '{
"urls": [
"https://en.wikipedia.org/wiki/Main_Page"
],
"formats": [
"html",
"metadata"
]
}'
Response
[
{
"url": "https://en.wikipedia.org/wiki/Main_Page",
"title": "Wikipedia, the free encyclopedia",
"html": "Wikipedia was just a dream.\ndiv class=\"frb-subheader\">\n<span class=\"frb-replaced\">December 2</span>: Readers <span class=\"frb-replaced\">in the United States</span> deserve an explanation.\n</div>\n</div>\n<div class=\"frb-message-content\">\n<p>\nPlease don't skip this 1-minute read. It's <span class=\"frb-replaced\">Tuesday</span>, <span class=\"frb-replaced\">December 2</span>, and if you're like us, you've used Wikipedia countless times. To settle an argument with a friend. To satisfy a curiosity. Whether it's 3 in the morning or afternoon, Wikipedia is useful in your life. Please give <span class=\"frb-replaced\">$2.75</span>.\n</p>\n<p>\nWikipedia's been around since 2001. Back then, it was just a wildly ambitious, probably impossible dream. But it came together piece by piece—created by people, not machines. Wikipedia's not perfect, but it's always been free thanks to everyday readers.\n</p>\n<p>\nOnly 2% ever donate. But that small group makes a big difference. When you support Wikipedia, you're standing up for something simple",
"metadata": {
"site_name": "Wikipedia",
"favicon_url": "https://api.ydc-index.io/favicon?domain=en.wikipedia.org&size=128"
}
}
]

Prerequisites

pip install youdotcom # Python ≥ 3.10
npm install @youdotcom-oss/sdk # Node ≥ 20

Walkthrough

contents.py
"""Contents — fetch clean Markdown from any URL via the You.com Contents API."""
import sys
from youdotcom import You, models
# take URL from command line, or use a default
url = sys.argv[1] if len(sys.argv) > 1 else "https://en.wikipedia.org/wiki/Retrieval-augmented_generation"
# initialize the client with your API key
with You() as you:
pages = you.contents(
urls=[url],
formats=[models.ContentsFormats.MARKDOWN],
crawl_timeout=15,
)
# print the title and the first 500 chars of the markdown body
for page in pages:
print(page.title)
print(page.url)
print()
print((page.markdown or "")[:500] + "...")
export YDC_API_KEY="your-api-key-here"
python contents.py "https://en.wikipedia.org/wiki/Retrieval-augmented_generation"

Example Output

# Retrieval-augmented generation
Retrieval-augmented generation (RAG) is a technique that grants generative
artificial intelligence models information retrieval capabilities. It modifies
interactions with a large language model so that the model responds to user
queries with reference to a specified set of documents…
(Returned alongside title, url, and metadata.site_name = "Wikipedia". Full
Markdown body is ~12,000 chars.)

Next Steps


Resources