AI Agent Allowlist
Home Page-Types Database Agent Guardrails 2026 Incidents API Docs Pricing
Resources
Use Cases Industries & Buyers Learn: Core Concepts Implementation Guides Comparisons Schema & Data Reference FAQ Glossary
Why It Matters
2026 Agent Incidents Category Targeting Database Refreshes Contact Customer Login
Download Free Sample
sizing the license to the workload

10M, 15M, 30M, 40M+. Which tier resolves your agent's actual traffic?

Domain count alone does not tell you what a tier covers. This page works through the traffic-coverage math for each license tier and gives a concrete way to decide which one fits a given agent workload's domain diversity.

4License tiers, 10M to 40M+
99.99%Of active internet traffic at full coverage
$14,999Entry tier, one-time, 10M domains
30%Optional annual refresh, of license price

Coverage tier is not the whole story. The JFrog Artifactory covert channel was stopped by egress pattern rules on plugin installs and WebDAV writes, not by domain-level coverage — a reminder that tier sizing complements the other three layers rather than replacing them. See the 2026 agent incidents, prevented: how the Hugging Face breach could have been stopped.

The four tiers

One-time license pricing, exactly as sold

All four tiers are perpetual, one-time licenses with no per-lookup fee once deployed on-prem. An optional refresh keeps the data current.

Entry10M domains$14,999

One-time. Covers the majority of an agent's real-world navigation traffic for most vendor-research and procurement workloads.

Common pickMid15M domains$24,999

One-time. The tier most teams land on once actual traffic logs show how diverse their agent's destinations really are.

Broad30M domains$49,999

One-time. For agents with open-ended discovery workloads: market research, lead enrichment, broad competitor monitoring.

Maximum40M+ domainsCustom

Custom and OEM cuts on request. Approaches 99.99% of active internet traffic; for platforms redistributing coverage to their own customers.

Optional monthly refresh for any tier is priced at 30% of the license price per year. Full pricing, including the self-serve lookup API tiers for teams not ready for an on-prem license, is on the pricing page. Every tier carries the same schema and the same product described on the page-types database overview — the difference between tiers is domain count alone, never field completeness or classification quality.

Two purchase paths

On-prem license vs. the self-serve lookup API

The tiers on this page describe the on-prem database license specifically. A second, separate path exists for teams that want to start smaller or do not want to operate the database themselves.

The self-serve lookup API is priced by monthly call volume rather than by domain count: Pro at $99/month for 90,000 lookups, Pro Plus at $249/month for 225,000, then $499/month for 450,000, $999/month for 900,000, and $1,997/month for 2,000,000 lookups at roughly $1.00 per 1,000. Smaller plans run closer to $1.10 per 1,000 lookups. Because the API resolves one URL per call against the full underlying 40M+ repository, it does not require choosing a coverage tier at all — every plan sees the same full coverage, and the plan choice is purely about call volume. Custom volume arrangements exist beyond 10M lookups a month. Fair use on every plan covers live lookups made ahead of real navigation decisions, not bulk enumeration of the underlying dataset through repeated calls.

The tier decision on this page matters specifically for teams choosing the on-prem path: running the database locally, with no per-call fee and no round-trip to an external API, in exchange for choosing how much of the 40M+ repository to license up front. Teams uncertain which path fits often start on the lookup API to validate the integration and their actual call volume, then move to an on-prem license once volume or latency requirements justify it.

The coverage math

Domain count is not the same as traffic coverage

Internet traffic, including agent navigation traffic, is heavily concentrated. A relatively small number of domains accounts for a large share of all requests, and the tail of rarely visited domains is extremely long. This is why coverage does not scale linearly with domain count.

10Mentry
Resolves the concentrated head of agent navigation traffic
15Mmid
Extends into the upper-mid tail: emerging vendors, regional sites, niche SaaS
30Mbroad
Reaches deep into the long tail: small businesses, single-product sites, less-linked domains
40M+maximum
99.99% of active internet traffic — the full repository

The practical consequence is that the jump from 10M to 15M domains closes a meaningfully larger coverage gap, for most workloads, than the jump from 30M to 40M+ does in absolute traffic terms — because the additional 5M domains at the low end are still reasonably well-linked and commonly encountered, while the additional domains beyond 30M are increasingly obscure, low-traffic, or narrowly regional. This is not a reason to under-buy; it is a reason to size the tier to your workload's actual domain diversity rather than assuming bigger is proportionally better for every use case.

It helps to think of the four tiers as four points along a curve of diminishing marginal coverage per additional million domains, rather than as four equal-sized slices of the internet. Early tiers pick up the domains an agent is statistically most likely to visit: major SaaS vendors, well-known e-commerce platforms, established media and documentation sites. Later tiers pick up domains that are individually far less likely to be visited by any given agent, but collectively still matter for workloads with genuinely broad or unpredictable discovery patterns. Neither end of the curve is "wrong" to buy; the right point on it depends entirely on where your own workload sits.

Sizing method

How to decide which tier fits your workload

Three questions, worked through in order, generally settle the tier decision without guesswork.

How closed is the domain set?

An agent restricted to a known, finite list of enterprise vendors and internal tools has a small, predictable domain footprint. This workload is usually well served by the 10M entry tier, since the vendors it touches are almost always well-linked, well-known domains near the head of the traffic distribution.

How much open-ended discovery does it do?

An agent doing market research, lead enrichment, or competitor discovery encounters domains nobody pre-selected, including smaller and more regional sites. This pushes the sizing decision toward the 15M or 30M tier, depending on how far into the long tail the discovery work actually reaches.

Is the deployment redistributing coverage downstream?

A platform, gateway, or OEM shipping page-type coverage as part of its own product to its own customers needs to plan for its customers' combined domain diversity, not just one workload. This is the scenario the 40M+ and custom/OEM tier is built for.

These three questions are ordered deliberately: closedness first, because it is the fastest to answer and often ends the analysis on its own; discovery depth second, because it is where most workloads actually land between the 10M and 30M tiers; and redistribution last, because it changes the buyer's unit of analysis entirely, from one agent's traffic to a whole customer base's combined traffic.

Worked example

Sizing a procurement-research agent's tier

A composite, illustrative scenario, not a specific customer, showing the method in practice.

QuestionAnswer for this workloadImplication
What is the agent's job?Research vendors for a specific software category, gather pricing and case studies, never checkout.Domain-heavy but page-type-light: mostly pricing, documentation, case_studies and about pages.
How many vendors are in the category?Perhaps 200-500 known vendors, plus an unknown number the agent discovers via search and comparison sites it visits.The known vendors alone would fit comfortably in even the smallest tier; the discovery tail is the real sizing question.
How far does discovery actually reach?Traffic logs from a pilot run show most discovered domains are established SaaS companies with reasonable web presence, not obscure long-tail sites.10M or 15M tier likely resolves the large majority of what this agent actually visits.
Is there a compliance reason to want maximal coverage regardless?No specific regulatory requirement in this scenario.15M tier chosen as a margin of safety over the 10M entry tier, without paying for 30M+ breadth this workload will rarely use.

The same method applied to a broader discovery workload, such as the ones described on manual curation vs. the database, would more often land on the 30M tier instead, because open-ended vendor discovery reaches further into the long tail than a scoped procurement category does.

Notice what the worked example does not do: it does not start from "how much can we afford" or "what does the biggest competitor buy." It starts from the agent's actual job, moves to the domain set that job realistically touches, and only then asks which tier resolves that domain set at an acceptable rate. Sizing decisions made in the opposite order — picking a tier first and hoping the workload fits it — are exactly how teams end up either overpaying for breadth they never use or under-provisioning coverage for a workload that quietly outgrew its original scope.

Refresh economics

One-time license vs. keeping it current

The tier decision and the refresh decision are separate, and it is worth treating them that way rather than bundling them into one choice.

License only, no refresh

The domain set and page-type URLs are correct as of the delivery date and do not update afterward. Reasonable for a fixed-duration project, a one-time migration, or a team that plans to re-license periodically rather than subscribe to ongoing refresh. Over time, the drift dynamics described on manual curation vs. the database begin to apply to the licensed copy too, just starting from a much larger and more accurate baseline than a hand-list ever would.

License plus annual refresh (30% of license price/year)

New domains enter the repository, existing page-type URLs are re-verified against current link structure, and popularity ranks and personas update as traffic patterns shift. This is the right default for any production deployment expected to run for more than a few months, since it converts the same drift risk a hand-list faces into a scheduled, budgeted update rather than a silent decay.

The refresh decision interacts with the tier decision in one practical way worth flagging: a smaller tier with an active refresh often stays more accurate over time than a larger tier left static, because the refresh keeps re-verifying what it already covers rather than adding untouched breadth. A team choosing between spending its budget on a bigger tier or on refreshing a smaller one should weigh how much its workload's domain diversity is actually growing against how much its existing domains' page structures are actually changing — these are different kinds of drift, and the tier and the refresh address them separately.

Common mistake

Buying the largest tier "just in case"

The instinct to buy 40M+ coverage regardless of workload is understandable but usually not the efficient choice, for two concrete reasons.

First, cost scales with tier, and the 40M+ tier is priced custom specifically because it is aimed at OEM and redistribution use cases with a fundamentally different economics than a single agent deployment. Paying for that breadth when a 10M or 15M tier would resolve the overwhelming majority of a scoped workload's actual traffic is not caution; it is an unnecessary line item that a traffic-log-based sizing exercise, like the worked example above, would have avoided. The same reasoning applies in reverse: a team that genuinely needs OEM-scale breadth but under-buys to save on the sticker price ends up paying for a partial re-license later once the gap becomes obvious in production, which usually costs more than sizing correctly the first time.

Second, larger is not automatically safer in a way that changes the security posture. The layers described on agent guardrails — the host list, the egress rules, and default-deny — already cover domains outside whatever tier is licensed. A domain the database has not classified falls to default-deny rather than being silently allowed, so under-provisioning the tier does not create a silent security gap the way an incomplete hand-list does; it simply means slightly more domains fall to the pattern-based egress rules or default-deny instead of a direct page-type lookup. Sizing the tier correctly is a cost-efficiency decision layered on top of a security posture that holds regardless of tier.

There is a related mistake worth naming separately: assuming that a larger tier substitutes for the other three enforcement layers rather than sitting alongside them. Even at the 40M+ tier, approaching 99.99% of active internet traffic, the database only ever answers "what page type is this URL," never "should this specific agent, for this specific task, be allowed here right now." That second question is what egress rules, host-list denial, and default-deny exist to answer for cases the page-type data alone cannot resolve, and no coverage tier changes that division of responsibility.

A useful gut check before signing off on a tier: pull a week of actual outbound agent traffic, or a representative sample if a week is not available yet, and check what fraction of distinct domains would have resolved against the 10M tier alone. Teams are often surprised in both directions — a narrowly scoped internal research agent frequently clears 99% on the smallest tier, while a broad web-shopping agent touching long-tail merchant sites can need the 30M tier just to keep default-deny fallbacks to a manageable minority of requests. Run that check before the purchase, not after the first production incident report.

Related reference pages

Where to go next

Schema

See the full field-by-field reference behind every domain row in any tier on the page-type schema reference.

Login detail

Regardless of tier, see how the highest-stakes page type, login, is specifically found and verified on login page detection.

Build vs. buy

The labor-cost case for any tier over a hand-maintained list, worked in detail, on manual curation vs. the database.

For domain-only categorization at a much larger scale (100M+ domains) without agent-specific page-type data, see the sibling product Web Filtering Database — a useful reference point for what coverage numbers look like in a product designed around raw domain reach rather than page-level verification depth.

FAQ

Coverage tiers

Can I upgrade from a smaller tier to a larger one later?

Licensing is one-time per tier; moving to a larger tier later is a new license at the difference in scope rather than an automatic upgrade path, so it is worth sizing carefully up front using the method on this page.

Does a smaller tier mean weaker security, or just less coverage?

Just less direct coverage. Domains outside the licensed tier still fall under the egress rules library and default-deny, so the security posture does not have a gap at tier boundaries the way an incomplete hand-list would.

Is the 40M+ tier meant for a single company's internal agent deployment?

Usually not economically. It is priced custom specifically for OEM and platform-redistribution use cases where the buyer's own customers' combined domain diversity, not one agent's traffic, determines the required breadth.

How do I estimate my own workload's domain diversity before buying?

The most reliable method is reviewing an existing traffic log or a pilot run's request history, if one exists, and checking how many distinct domains appear and how concentrated they are. The free 100-domain sample CSV can also be checked against a handful of your own agent's real destinations as a smaller-scale sanity check.

What happens if my agent visits a domain outside the licensed tier?

It is treated as unclassified for page-type purposes. The egress rules library can still catch a risky URL shape on it, and default-deny applies to whatever is left unresolved.

Is the refresh mandatory?

No, it is optional at 30% of the license price per year. Without it, the licensed data is correct as of delivery and does not update, which is acceptable for fixed-duration projects but not recommended for ongoing production use.

Size the right tier for your workload.

Download the free sample or review exact tier pricing.

Download sample CSV See pricing