
What Is in a Web Search API JSON Response? Fields, Parsing, and Pitfalls
TLDR: A web search API returns structured JSON instead of a rendered page, so you skip the scraping and parsing layer entirely. The You.com Web Search API returns a results object with web and news arrays, where each result carries a url, title, description, snippets, and metadata such as publication dates. This guide walks the response shape, defensive parsing code, how the fields change with extraction modes, and the parsing bugs that hit production.
The point of a search API is that you never parse HTML. You send a query, you get JSON back, and every field you need sits at a predictable path. That contract is what makes search APIs usable from RAG pipelines, agents, and data jobs without a scraping stack. You.com designed its Web Search API around a single endpoint that returns clean, structured JSON, which makes it a good worked example for how to consume any search API response correctly. The JSON specification itself is short, and worth a read if you handle API payloads daily.
What Shape Does the Response Take?
The Web Search API endpoint is a single POST to ydc-index.io with path /v1/search. The response contains a results object with a web array, a news array when the query has news intent, and a knowledge array when you request licensed knowledge results. A minimal response looks like this, trimmed to the core fields.
{
"results": {
"web": [
{
"url": "https://example.com/docs/streaming",
"title": "Streaming responses",
"description": "A short summary of the page",
"snippets": ["one keyword-centered fragment", "another fragment"],
"page_age": "2026-08-02T14:10:00Z",
"thumbnail_url": "https://example.com/og.png",
"favicon_url": "https://you.com/favicon?domain=example.com"
}
],
"news": [
{
"url": "https://news.example.com/release",
"title": "Vendor ships new release",
"description": "Summary of the article",
"snippets": ["lead paragraph fragment"],
"page_age": "2026-09-28T09:00:00Z"
}
]
}
}
Three things to internalize. The url, title, and description are the core fields you can always build on. The snippets array holds short, keyword-centered fragments of the page, and is what you get by default without any extraction options. The news array only appears when the query has news intent, because the API classifies intent server-side, and you do not toggle a news mode.
How Do You Parse It Without Breaking?
Parse defensively, because every optional array in the contract is a future KeyError in your pipeline.
import json, urllib.request, os
def search(query: str, count: int = 5):
req = urllib.request.Request(
"https://ydc-index.io/v1/search",
data=json.dumps({"query": query, "count": count}).encode(),
headers={
"X-API-Key": os.environ["YDC_API_KEY"],
"Content-Type": "application/json",
},
method="POST",
)
with urllib.request.urlopen(req, timeout=15) as resp:
data = json.loads(resp.read().decode())
results = data.get("results") or {}
web = results.get("web") or []
news = results.get("news") or []
return {
"web": [
{
"url": r.get("url") or "",
"title": (r.get("title") or "").strip(),
"text": " ".join(r.get("snippets") or [])
or (r.get("description") or ""),
}
for r in web
],
"news_count": len(news),
}
The pattern to copy: default every field with or-empty values, never assume the news array exists, and derive the text you actually need, here snippets falling back to description, at the boundary. Downstream code then never sees a missing field, it sees an empty string. The Web Search API in Python guide extends this into a full client with retries.
What Fields Change With Extraction Modes?
The default response carries snippets. Two extraction modes change what text comes back per result, and they change the response fields.
- No extraction parameter: each result carries snippets, the short keyword-centered fragments shown above.
- extraction_mode "highlights": each result gains a contents.highlights field holding the passages from that page that address your query, and snippets are omitted. This is the token-efficient option for agents and RAG prompts.
- extraction_mode "full_page": each result gains contents.markdown or contents.html carrying the whole page. This is for archiving and full-document analysis, and it multiplies response size accordingly.
The mistake to avoid is switching extraction modes and letting your parser look for the old field. If your code reads snippets and you turn highlights on, you get empty strings everywhere, silently. The Web Search API guide documents the extraction object and every mode.
What Are the Common Parsing Bugs?
Assuming the news array exists. It is classified by query intent, not requested. A pipeline that indexes results.news without a default crashes exactly once, on the first non-news query. Detection: always default the array, and log how often it is present versus absent.
Confusing description with snippets. They are different fields with different shapes, a string versus an array, and code that treats them interchangeably breaks on one or the other. Detection: normalize once at the boundary, as in the parser above.
Ignoring pagination. The API paginates with an offset parameter between zero and nine, and a count parameter that caps results per section at one hundred. A job that reads one page and reports "all results" is wrong by construction. Detection: page until the API stops returning new results, and record the total you actually collected.
Assuming provider-agnostic fields. Response shapes differ across search APIs. Exa, for example, returns results with id and publishedDate fields as of its current documentation, while You.com uses page_age. Neither is wrong, they are different contracts. Detection: build a thin normalization layer per provider and test it against each provider's documented schema.
How Do You Keep Parsing Stable Across Providers?
Normalize early, in one place. Map every provider's response into your own internal shape, url, title, text, published, before anything else touches the data. When you swap or add a provider, you write one adapter, and the rest of the pipeline cannot tell the difference. This is the same pattern our guide to choosing a search API for AI agents recommends testing against, because the cheapest time to discover a field mismatch is before you commit to a vendor.
Two extra habits keep the adapter layer honest over time. First, pin a small set of canned responses from each provider as test fixtures, so a vendor-side schema change fails your test suite instead of your production pipeline. Second, version the adapter alongside the vendor's documented schema: when a provider adds a field or renames one, the change should show up as a deliberate commit in your adapter, not as a surprise in a dashboard. Teams that skip this treat every vendor changelog as an outage waiting to happen, and teams that do it keep vendors replaceable, which is the position you want to be in when priorities or prices change.
Related Guides
- Web Search API in Python: A Practical Guide With the You.com SDK
- How to Call the You.com Web Search API With cURL
- How to Use the You.com Web Search API in Go
- Web Search API: Programmatic Access to Real-Time Web Data
FAQ
Do all web search APIs return the same JSON fields? No. Field names, nesting, and text options differ per provider, which is why a normalization layer pays for itself the first time you evaluate or swap a vendor. Check each provider's documented response schema before parsing.
Why is my results.news array missing? The You.com Web Search API classifies query intent server-side and only includes news results when the query has news intent. There is no client-side toggle, and a missing news array on a non-news query is normal behavior, not an error.
What is the difference between description and snippets? Description is a short summary string of the page. Snippets is an array of short, keyword-centered fragments from the page, returned by default. Different shapes, so normalize them once at the boundary.
How large can one response get? It depends on your count, capped at one hundred per section, and your extraction mode. Full-page extraction multiplies response size because it ships the entire page content per result, so request it only where you need whole documents.
LI Test
LI Test
Share Article:
