Blog
 / 
September 29, 2026

How to Restrict a Search API to Specific Sites With the You.com Web Search API

How to Restrict a Search API to Specific Sites With the You.com Web Search API

TLDR: The You.com Web Search API accepts three domain steering parameters: include_domains filters results to a list of domains, exclude_domains removes domains, and boost_domains prefers them without excluding anything else. Each accepts up to 500 domains, and include_domains and boost_domains cannot be combined in one request. This guide shows how to use them, when to pick boost over include, and the silent failure modes to monitor.

Sometimes a search query is fine and the index is the problem. A RAG pipeline that only trusts certain sources, a brand-monitoring job that tracks coverage on specific outlets, and a research tool that must never cite a competitor's marketing page all need the same thing: results from some domains and only those domains. On a search results page a person does this with the site: operator, which Google documents in its search refinement guide. Programmatically, you use the domain parameters of the You.com Web Search API. You.com returns the same structured results either way, so the filter happens before the results reach your code, not after.

Which Parameters Restrict Results to Specific Sites?

Three request-body parameters control domains, and they do different jobs.

  • include_domains: a hard filter. Results come only from the domains in the list. Use it when a wrong-domain result is worse than a missing result, such as trusted-source pipelines.
  • exclude_domains: a hard exclusion. Results never come from the listed domains. Use it to remove known noise, such as SEO-heavy aggregator pages, from open queries.
  • boost_domains: a soft preference. Listed domains rank higher when relevant, but other domains can still appear. Use it for research tasks that want to favor primary sources without missing context.

Each list accepts up to 500 domains. One combination is rejected: include_domains and boost_domains cannot appear in the same request, because a hard filter makes a soft preference meaningless. If you need both behaviors, run two requests and merge, or pick the one that matches your failure mode.

How Do You Call It in Python?

The domain parameters ride in the same JSON body as the query. The example below restricts a query to a set of documentation and standards domains, which is the trusted-source pattern for a research agent.

import os, json, urllib.request

def site_restricted_search(query: str, include: list, exclude: list = None,
                           count: int = 10) -> dict:
    url = "https://ydc-index.io/v1/search"
    body = {"query": query, "count": count}

    if include:
        body["include_domains"] = include
    if exclude:
        body["exclude_domains"] = exclude

    req = urllib.request.Request(
        url,
        data=json.dumps(body).encode(),
        headers={
            "X-API-Key": os.environ["YDC_API_KEY"],
            "Content-Type": "application/json",
        },
        method="POST",
    )
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:
            return json.loads(resp.read().decode())
    except urllib.error.HTTPError as e:
        return {"error": f"HTTP {e.code}", "body_hint": e.read()[:200]}

result = site_restricted_search(
    query="python urllib retry timeout best practice",
    include=["docs.python.org", "developer.mozilla.org", "peps.python.org"],
)
web = result.get("results", {}).get("web", [])
print(f"got {len(web)} results, all from allowed domains by contract")

The error branch matters here more than usual, because one of the silent failure modes below starts as a client-side mistake that surfaces as an HTTP error. Reading a fragment of the error body when the status is not 2xx turns a mysterious empty result into a diagnosable one. For the raw request shape with no wrapper, see our cURL guide to the Web Search API and the Python guide for the SDK path.

When Do You Want Boost Domains Instead of Include Domains?

The include-versus-boost choice is a decision about which failure you prefer, and that makes it a decision framework rather than a style preference.

Choose include_domains when a result from outside the list is a correctness problem. A compliance-adjacent research pipeline that may only cite regulatory bodies is the canonical case. The tradeoff is recall: an include filter that is too narrow returns nothing, and the pipeline must handle empty results explicitly.

Choose boost_domains when coverage matters more than purity. A technical research agent that should prefer official documentation but may still cite a high-quality blog post or a GitHub issue is the canonical case. The tradeoff is trust: a boosted domain does not dominate, so results can still include domains you would rather avoid, and downstream filtering is your responsibility.

A practical middle pattern is two requests: one boosted open query for coverage, and one include-filtered query for the authoritative answer, then present the filtered set as the primary sources and the boosted set as further reading. That is the same pattern our news pipeline guide uses for outlet steering.

What Fails Silently?

Domain filtering fails quietly, which is the dangerous kind of failure for an unattended pipeline. Four modes cover most of it.

The typo domain. You pass "developer.mozilla.or" or "docs.python.co" and the filter faithfully matches nothing. The query returns results from other domains anyway if the filter logic is a boost, or nothing if it is an include. Detection: validate the domain list before the first run, and log the exact request body on every search. A one-time assertion that each domain resolves, using the URL parsing rules from the URL API on MDN as the normal model, catches most of these in a preflight check.

Subdomain assumptions. Whether a domain entry matches subdomains is the kind of behavior that differs between providers, and your pipeline should not depend on an assumption it has not tested. Detection: run a small probe set once, query for content you know lives on a subdomain, and record the observed behavior in your integration test so a provider behavior change is caught instead of absorbed.

Over-restriction. A tight include list plus a narrow query returns zero results, and an agent that treats empty as "no information exists" will confidently say the wrong thing. Detection: log the result count per query alongside the filter set, and alert when a filter configuration's median result count drops near zero. Empty results from a filtered query mean the filter is too tight, not that the world ran out of answers.

The rejected combination. include_domains and boost_domains in one request are rejected rather than silently ignored, which is good, but the rejection is an HTTP error your code has to handle rather than a note in the response. Detection: the probe above already reads the error body. Treat a 400 on a domain-filtered request as a configuration bug, and check for the include plus boost combination first.

Related Guides

FAQ

How many domains can I filter by? Each of include_domains, exclude_domains, and boost_domains accepts up to 500 domains per request. The limits are per parameter, so an exclude list and an include list can coexist within their own limits.

Can I use include_domains and boost_domains together? No. The request is rejected when both appear, because a hard include filter already decides which domains can appear, which leaves a preference ranking nothing to boost.

Do the domain filters work with the MCP server too? The you-search tool on the You.com MCP server supports exclude_domains with the same 500-domain limit, plus inline site: filters in the query text. The two cannot be combined in one call. See our MCP server setup guide for the full tool schema.

Why does my filtered query return zero results? Usually the filter is tighter than the corpus. A narrow query combined with a short include list can exhaust the matching pages. Widen the query, expand the domain list, or switch the include filter to a boost, and log the result count per filter configuration so the regression is visible over time.

    Share Article:

  1. LI Test

  2. LI Test

Related resources.

No items found.
No items found.