AI Agent Allowlist
Home Page-Types Database Agent Guardrails 2026 Incidents API Docs Pricing
Resources
Use Cases Industries & Buyers Learn: Core Concepts Implementation Guides Comparisons Schema & Data Reference FAQ Glossary
Why It Matters
2026 Agent Incidents Category Targeting Database Refreshes Contact Customer Login
Download Free Sample
compare: website categories vs page types

Categories Say What a Site Is. Page Types Say Where the Function Lives.

A content-category taxonomy answers "what kind of site is this" — software, finance, news, retail. A page-type map answers a completely different question: "where on this specific site is the login, the checkout, the pricing page." Agent policy needs both answers, at the same time, on the same request, and treating either one as sufficient on its own leaves a real gap in coverage.

0IAB content categories (v2/v3)
0Web-filtering categories
0Page types per domain
0Domains carrying both schemas

Category alone would not have caught the DseWiki hijack. A wiki domain classified under a benign "reference" or "community" category still exposes a wiki_edit endpoint that page-type policy denies specifically. Several 2026 incidents crossed exactly this gap — our database and egress rules would have prevented almost all of them.

See the incident-by-incident prevention analysis
The category is not the page

A site's category tells you almost nothing about which page an agent just requested

Web categorization is a well-established discipline: classify a domain into a content taxonomy such as IAB's 700+ categories across four tiers, or a coarser web-filtering scheme, so that policy can be written in terms of "software," "finance," "adult content," or "gambling" rather than in terms of individual URLs. It answers a question about the site as a whole, and it answers it well — a single category label, assigned once per domain (with periodic refresh), tells a filtering or research system what kind of content or business the domain represents.

That label says nothing about which specific page on the domain an agent is about to request. A domain categorized as "Software" carries that label whether the URL in question is its public blog post, its documentation, its pricing page, or its login form. Category answers "what is this site," uniformly, across every page the domain serves. It was never built to answer "is this particular URL a place my agent should be submitting credentials," because that question operates at a level of granularity category was not designed to reach.

Page-type classification answers exactly that narrower question, and only that question. It says nothing about whether the domain overall is a bank, a blog, or a SaaS vendor; it says, for one verified URL on that domain, which of 28 functional types the page serves — login, checkout, pricing, documentation, and so on. A page-type classifier has no opinion on whether the site as a whole belongs in a "Business" or "Technology" category, because that is not the question it was built to answer either.

Put the two side by side and neither one is wrong or redundant; they simply answer different halves of the same policy question. A team that has only a category feed can write a rule like "allow research agents onto Software-category domains," but that rule has no way to keep the same agent off that domain's login page. A team that has only a page-type map can deny logins and checkouts everywhere, but has no way to additionally say "and be more cautious across Finance-category domains specifically." Combined, the two produce a policy neither one could express alone.

The mismatch is easiest to see in scale terms. A domain-level category assignment is, by construction, one label (or one small set of labels across tiers) per domain, which means a site with ten thousand pages gets the same category value on every one of them. Page-type coverage runs in the opposite direction: it is deliberately sparse and specific, naming only the handful of pages per domain — typically well under two dozen — that correspond to one of the 28 defined functional types, and saying nothing about the other several thousand pages that do not fit any of those types. Neither schema was built to do the other's job at the other's resolution, and no amount of refining one closes the gap the other is meant to cover.

Two questions, defined

What kind of site, and where on it

what is this site

  Category taxonomy

One label (or set of labels across IAB v2/v3 Tier 1–4, plus a 59-category web-filtering scheme) assigned per domain, describing the nature of its content or business.

example.com → IAB: Business > Software; Filtering: Technology/SaaS
where on this site

  Page-type map

28 verified page types per domain, each a specific URL that answers a specific functional question about that one page, independent of what the domain as a whole is about.

example.com/login → page_type: login (identity group, deny)
The comparison

Category taxonomy and page-type schema, dimension by dimension

DimensionCategory taxonomy (IAB / web-filtering)Page-type schema (28 types)
Question answeredWhat kind of site is this, overallWhat function does this specific URL serve
Unit of classificationThe domain (sometimes subdomain)The individual page
DepthUp to 4 tiers (IAB), or one of 59 filtering categories28 named types across identity, commerce, content-write, and research/read groups
Distinguishes login from blog on the same domain?No — both carry the same domain-level categoryYes — each has its own verified URL and type
Distinguishes a bank's site from a blog?Yes — that is exactly its jobNo — a login page is a login page regardless of the domain's business
Typical policy use"Restrict research agents to Business/Software domains""Deny checkout, cart, and identity pages everywhere, allow pricing and docs"
Refresh cadencePeriodic, per-domain, as content and business focus shiftVerified per URL, refreshed alongside the domain database
What each one misses aloneAny distinction between pages on the same domainAny signal about the domain's overall subject matter or industry

The two schemas are not competing standards for the same job. They are complementary axes on the same domain record: every one of the 40 million+ domains in the database carries both an IAB category, a web-filtering category, and a set of verified page types, so a policy can combine "what kind of site" with "what kind of page" in a single rule.

Why the combination matters

A rule neither schema can write alone

Consider a compliance-monitoring agent tasked with tracking terms-of-service and privacy-policy changes across a company's vendor list, plus a wider set of Financial Services-category domains for regulatory-change awareness. A category-only policy can scope the second half of that task — "watch Financial Services domains" — but has no way to say "and only touch their legal and press pages, never their login or account pages," because category does not resolve to individual page types at all. A page-type-only policy can enforce the safety half — "allow legal and press, deny login and account" — on any domain, but has no way to scope the monitoring specifically to Financial Services as a vertical, because page type carries no industry signal.

Rule expressedCategory alonePage type aloneCombined schema
Scope monitoring to Financial Services domainspossiblenot possiblepossible
Deny login/account pages on those domainsnot possiblepossiblepossible
Allow legal/press pages on those domains onlypartialpartialpossible
Apply the same safety rule everywhere, regardless of industrynot possiblepossiblepossible

Only the combined schema can express all four rules in the same policy document. This is not a marginal improvement; it is the difference between a policy that can be scoped by industry and a policy that cannot, and between a policy that is safe by page type and one that only sounds safe because it never had to name a specific dangerous page.

The gap shows up most clearly during an audit or compliance review, when someone asks a two-part question: "which domains was the agent authorized to visit, and what could it do once it got there." A category-only record answers the first half well and cannot answer the second at all — it has no page-level field to point to. A page-type-only record answers the second half precisely and has nothing to say about whether the domain list itself was appropriately scoped to the task's industry or subject matter. Only a record carrying both fields lets a reviewer trace a single navigation decision back to both the domain-level reason it was in scope and the page-level reason it was allowed or refused, which is the level of detail most governance reviews are actually asking for once they get past the first question.

Worked example

A market-research agent, four domains, category plus page type

A market-research agent is scoped to survey News and Business-category domains for coverage of a product category, pulling article and press content while staying off anything transactional or credentialed. Here is how category and page type resolve together on four representative requests.

RequestCategoryPage typeResult
/2026/09/market-report on Domain ANewsblogallow — in scope, safe type
/press/product-launch on Domain BBusiness > Softwarepressallow — in scope, safe type
/subscribe on Domain A (paywall prompt)Newssubscribeflag — in scope, action type
/account/login on Domain CRetail — out of task scope entirelylogindeny — wrong category AND unsafe type

The fourth row shows both schemas agreeing for different reasons: Domain C is out of scope on category grounds (it is not News or Business), and its login page would have been denied by page-type policy regardless of category. The third row shows why the two schemas cannot be collapsed into one: subscribe is in-scope by category (a News domain, exactly the kind of site the task cares about) but is still an action page type, so it is flagged for human sign-off rather than silently allowed or silently denied. Category told the agent this domain matters; page type told it exactly how carefully to proceed once there.

A common mistake

Treating category as a proxy for page-level safety

category-only mistake

  "It's a Business-category domain, so it's fine"

A policy that allows an agent onto any domain in a "trusted" category, with no further check, treats the entire domain as equally safe — its blog post and its account-settings page get the same verdict, because category cannot distinguish between them. The agent proceeds past a login or checkout with no additional resistance, since nothing in the category schema was ever built to flag that specific page.

combined-schema correction

  Category scopes the task, page type gates the action

Category answers "should this domain be in the agent's task at all," and page type answers "is this specific page one the agent may act on." A domain passing the first test still has its login, checkout, and upload pages denied by the second, independent of how trustworthy the domain's overall category appears.

What the combined record looks like

One domain, both schemas, in the same row

The free sample CSV shows this directly: every one of its 100 rows carries a domain alongside its page types, its IAB v2 and v3 categories across all four tiers, its web-filtering category, personas, OpenPageRank, country, and popularity rank groups — one record, two schemas, no need to join separate datasets.

1domainThe single key both schemas attach to.
2page_typesWhich of the 28 verified page types this domain has, and their URLs.
3IAB v2 / v3, Tier 1–4700+ content categories describing what the site is about.
4Web Filtering CategoryOne of 59 categories for filtering-style policy.
5Personas, OpenPageRankAudience and popularity signal for prioritization.
6num_distinct_page_typesHow thoroughly this domain has been mapped.

See the full column reference, including exact IAB tier structure and the web-filtering category list, on the page-types database page, or pull all twelve columns directly from the free 100-domain sample. Teams building their own policy engine on top of this record typically load it once into whatever table or index already backs their existing category-based rules, then add a second lookup keyed on the exact URL rather than the domain, so the two checks can run in the same request path without a separate service call for each schema.

Building the combined rule

Four steps to a policy that uses both schemas

Write the safety baseline in page types

Deny login, signup, checkout, cart, upload, and the rest of the identity and commerce groups by default, on any domain, before category enters the decision.

Layer category scoping on top

Add "and only proceed on Financial Services / Software / News category domains" as a narrowing filter for the specific task, never as a replacement for the page-type baseline.

Never let category override a page-type deny

A "trusted" category is not a reason to allow a login or checkout page. The safety rule wins regardless of what industry the domain belongs to.

Log both fields per decision

Record the category and the page type on every navigation decision, so a later review can see both why a domain was in scope and why a specific page was allowed or denied.

# combined-schema policy: scope by category, deny by page type
policy: combined_schema_v1
default: deny
rules:
  - match: { page_type: [login, signup, password_reset, checkout, cart, upload] }
    action: deny  # page-type baseline, applies to every category
  - match: { iab_category: "Financial Services", page_type: [legal, press, about] }
    action: allow  # category-scoped task, safety rule already applied above
  - match: { page_type: [pricing, documentation, blog] }
    action: allow
Integration

One lookup returns both signals

A single check against the database or API returns the page type for enforcement and the domain's category fields for scoping, in one response:

GET https://www.aiagentallowlist.com/api/check?url=https://example.com/login
{ "result": "deny", "id": "login" }

Full-field responses, including IAB and filtering categories per domain, are part of the licensed database rather than the lightweight per-URL check endpoint; see the API docs for request shapes and the pricing page for lookup and license tiers.

FAQ

Categories vs page types, answered

No. Category is assigned at the domain level and describes the site's content or business as a whole. It has no mechanism for resolving to a specific URL like a login form, which is exactly the job page-type classification does instead.
No. A credential surface carries the same risk regardless of the domain's category. The correct policy denies identity and commerce page types by default across every category, and uses category only to scope which domains a task should touch at all.
Most teams use whichever taxonomy their existing policy tooling already speaks: IAB's 700+ categories across four tiers for granular content classification, or the 59-category web-filtering scheme for coarser, filtering-style rules. Both ship with every domain record, so you are not forced to choose in advance.
No, and it should not. Page types add page-level granularity to a policy; they do not carry industry or content-subject information the way a category feed does. The two are additive, and the sample CSV ships both together specifically so you do not need to join separate sources.
Both are refreshed as part of the same domain database update cycle. Database license customers can add a monthly refresh at 30% of the license price per year; see the pricing page for exact terms.
Related reading

See the same schema from other angles

For the same category taxonomy applied to blocking human access to risky AI tools rather than agent navigation, see Web Filtering Database, the sibling product covering 100M+ domains by filtering category.

Get both schemas on the same domain record

28 page types, 700+ IAB categories, 59 filtering categories, 40M+ domains. Start with the free sample.

See Pricing