inside the database

20 page types. Real URLs. Every domain.

To build this database we traversed the live link structure of each domain and analyzed over 10 billion pages through a multi-step AI classification pipeline. The result is not a list of guessed paths — it is the actual URL a site links to for each page type, per domain, across tiers of 10, 15 and 30 million domains.

The 20 page types

Classified by what an agent should do with them

Availability depends on each site’s structure — not every domain has all 20. Typical policy grouping shown; you define your own rules.

loginCredential surfaces. The #1 page class agents should never enter autonomously.
checkout / cartPurchase flows. Deny unless the agent is explicitly authorized to transact.
paymentCard and billing entry points, including third-party payment gateways.
account / signupRegistration and account-management surfaces — identity side effects.
pricingThe page agents hunt for most in research tasks. One lookup, zero browsing.
documentationDocs and developer guides — safe, high-value reading targets.
contactVerified contact page URL — including forms an agent may be told to avoid.
aboutCompany background for entity resolution and vendor research.
leadership / teamExecutive pages for due diligence and sales-intelligence agents.
careersJob listings — hiring-signal extraction without crawling the whole site.
blog / newsContent hubs for monitoring and summarization tasks.
statusService-status pages — the right target for uptime-checking agents.
legal / privacyTerms and privacy policies — compliance agents read these constantly.
partnersPartner and integration directories for ecosystem mapping.
case studiesProof-point pages for competitive and sales research.
support / FAQHelp centers — the safe alternative when an agent would otherwise poke at forms.
downloadsFile-download surfaces — often policy-relevant (executables, PDFs).
developer / APIAPI portals — where an agent should integrate instead of scraping.
press / mediaNewsroom pages for monitoring agents.
product / storeCatalog entry points — browse-safe, transact-gated.
Schema

Every record carries the full context

Page types are the core, but policy needs context: what kind of site is this, how popular, where, in what language, for whom?

page_typesUp to 20 type=URL pairs, verified from the site’s live link structure.
IAB v2 + v3 categoriesTier 1–4 content categories — know it’s a bank before the agent acts like it’s a blog.
web_filtering_category59-category filtering taxonomy — adult, gambling, malware and other block-worthy classes.
popularity ranksGlobal and country-level rank groups from real-world browsing data.
country + languageJurisdiction and language signals for residency-aware agent policies.
personas + OpenPageRank2000+ user personas and link-authority scores for targeting and trust weighting.

Delivery: CSV via the client portal (one-time tiers) or JSON over the lookup API. Custom slices — by industry, geography, or page type — can be built from the full 102M-domain repository. Ask us.

Verify the quality yourself

100 well-known domains, full schema, real URLs — free, no signup.

Download the Sample