The classification database behind the filter

120 million domains, sorted so a school can make a decision

Every block page, every allowed link and every audit report a district produces traces back to one thing: a list that says what a website is. This page is the full account of that list — what a single record holds, where new domains come from, how the 57+ category taxonomy is shaped around school decisions rather than advertising taxonomies, how often it changes, and the four ways districts pull it into whatever they already run.

120M+Domains classified
57+Content categories
16,328+AI tools tracked
DailyUpdate cycle

13 fields per recordFiltering category, IAB v2 and v3 trees, country, popularity signals
Rebuilt every dayNew registrations screened, existing domains re-checked on rotation
Four delivery formatsAPI, CSV, DNS blocklist, or the full database on disk
CIPA pillars kept separateThe statutory floor is labelled apart from district discretion
Anatomy of a row

What the database actually holds for a single domain

People picture a blocklist as two columns: a domain and a verdict. That shape breaks the first time a school needs to allow a site for staff and restrict it for fourth graders. Our unit of storage is a described domain, not a judgement about one.

Description, not verdict

A record answers the question "what is this site?" and stops there. It does not carry a block flag, because a block flag would encode our opinion of your policy into your data. A high school library and a K-2 building can pull the identical row and reach opposite conclusions, which is exactly as it should be.

That separation has a practical payoff at renewal time. When a district changes its stance on, say, AI homework helpers, nothing about the underlying data has to change — the policy that reads the data changes. Districts that have lived through a blocklist migration know how much work that saves.

Alongside the primary web filtering category, each row carries the domain's position in the IAB v2 and IAB v3 content taxonomies, down to four tiers where the tree goes that deep. Those trees matter to districts that already own analytics or ad-management tooling keyed to IAB labels: the same export drops straight into those systems without a translation layer someone has to maintain.

Two popularity signals round the record out. OpenPageRank gives a rough measure of a domain's authority on the open web, and the popularity rank groups indicate whether a site sits in the global top thousand, the long tail, or somewhere between — globally and within its own country. Those numbers are what let a network administrator triage: a miscategorisation on a top-1000 domain deserves attention this afternoon, while an obscure parked domain can wait for the next review cycle.

Fields carried on every classified domain
DomainThe fully qualified domain name — example.com
Web Filtering CategoryThe school-facing label, drawn from the 57+ category taxonomy
IAB v2 Tier 1–4Position in the IAB v2 content tree, four levels deep where defined
IAB v3 Tier 1–4The same placement in the newer IAB v3 taxonomy
PersonasAudience descriptors associated with the site's visitors
OpenPageRankAuthority score on a 0–10 scale — 10.0 at the very top
CountryPrimary country association for the domain
Global Popularity Rank GroupTraffic band worldwide — for example 1-1000
Country Popularity Rank GroupThe same banding computed within the domain's own country

Full types, null behaviour and example values are documented in the data dictionary. The same field set is present whether you consume the data by API, CSV or full download.

Pipeline

How a domain gets found, read and labelled

Roughly 300,000 domains are registered every day. A classification database that only grows when a customer complains will always be one incident behind, so discovery has to be a standing process rather than a reaction.

1

Discovery: registration feeds, crawl frontiers and live traffic

New domains enter the queue from three directions. Newly registered domain feeds catch sites the moment they exist, often before any content is on them. Crawl expansion pulls in domains linked from pages already known to us, which is how regional and non-English sites surface without anyone hand-seeding a country list. And uncategorised lookups from live filtering deployments push a domain to the front of the queue: if a student in a district somewhere tried to reach it this morning, it is more urgent than a domain nobody has visited.

2

Collection: fetch what the site actually serves

The classifier works from what a browser would receive, not from the domain string. Page text, title and meta description, visible navigation, on-page media cues and the site's own structural signals are collected. Guessing from a name is where cheap blocklists fail spectacularly — wholly innocuous domains contain unfortunate substrings, and plenty of genuinely harmful sites are named after fruit.

3

Classification: machine labelling against the taxonomy

Collected signals are scored against the full category set, producing a primary web filtering category plus the IAB placements. Confidence matters as much as the label: a strong, unambiguous signal on an adult site is treated very differently from a marginal call on a small business page, and the low-confidence tail is what gets routed to review rather than published unchecked.

4

Layering: a domain can be several things at once

A video platform hosting educational lectures alongside content nobody wants in a middle school is genuinely both. Rather than forcing a single winner, labels are layered so a filter can act on the intersection: allow the platform, block the categories within it that a given grade band should not reach. Single-label databases push that entire problem onto the district's exception list, where it becomes somebody's weekly chore.

5

Publication: into the daily build, and out through every format

Accepted classifications land in the next daily build. From there the same records populate the API, the CSV exports, the DNS blocklist artefacts and the full database download simultaneously — there is no premium tier of the data and no lagging secondary copy. What the API says at noon is what the CSV you downloaded that morning says.

Why this order matters to a school: the gap between a site appearing and a site being classified is the entire window in which a filter fails silently. Discovery driven by registration feeds and live lookups — rather than by complaints — is what keeps that window measured in hours instead of months. You can test the current state of any single domain yourself with the domain checker.

Taxonomy

57+ categories, organised around the decisions a school actually makes

Most content taxonomies were designed for advertisers, who care about what a page is selling. Schools care about something different: whether a student should be there, and at what age. The category structure is shaped around three tiers of decision.

Tier one: never in doubt

Adult and obscene content, child sexual abuse material, graphic violence and gore, drug marketplaces, self-harm and pro-suicide communities, hate and extremist material. These carry the categories that map onto CIPA's statutory requirements and the harms no district has ever debated in a board meeting.

Granularity still matters here, because "adult" is not one thing. Separating explicit content from mature-but-not-explicit material lets a district block both for elementary students while allowing a high school health curriculum to function.

Blocked in every template

Tier two: it depends on the grade

Social media, streaming and video, games, chat and messaging, forums, shopping, dating, AI tools. Nothing in CIPA touches any of these. They are where district policy lives, and where the argument between the technology office and the teaching staff usually happens.

Fine-grained categories are what let that argument end well. A district that can block competitive gaming while leaving educational simulations reachable has a different conversation than one whose only lever is a single "Games" switch.

Tuned per grade band

Tier three: the everyday web

News, reference, education, health, business, sports, government, libraries and museums, software and technology. The overwhelming majority of the 120 million classified domains sit here, and the job of the taxonomy is to keep them out of the way.

This tier is quietly the most important for classroom credibility. Over-blocking teachers into a corner produces workarounds — personal hotspots, unmanaged devices — that do far more damage to a filtering programme than any single missed site.

Open by default

Why granularity is not a vanity metric

The number of categories in a filtering database is easy to treat as marketing arithmetic. It is not. Each category is a lever, and a district can only make distinctions its data supports. With a coarse taxonomy, the choices available to a middle school principal are "block the whole of social media" or "allow the whole of social media" — and since neither is acceptable, what actually happens is an exception list that grows for three years until nobody remembers why half the entries are there.

A 57-plus category structure changes the shape of that work. Policy stops being a pile of individual domains and becomes a small set of category decisions that can be written down, explained to a school board, defended to a parent, and handed to a successor. The exception list still exists, but it stays short enough to review.

The full category listing, including how each one is defined and typical school treatment, is documented on the filtering categories and taxonomy page and in the category reference in the docs.

Freshness

Daily updates, and the two different ways stale data fails

A classification database is a perishable good. The web it describes changes underneath it every hour, and a copy that was accurate in September is quietly wrong by February in two opposite directions at once.

120M+Domains under classification
Full corpus
DailyRebuild and publish cycle
Every build day
~300KNew registrations screened daily
Continuous intake
57+Categories kept in sync
One taxonomy
Two failure modes, one cause

Under-blocking and over-blocking both come from age

Districts tend to worry about one of these and get surprised by the other. Both are symptoms of the same thing: a database describing a web that has moved on.

What went staleHow it failsWhat the district sees
A brand-new adult site registered last weekUnder-blockContent reaches a student that the certification says is blocked
A proxy or VPN service that launched this monthUnder-blockStudents route around the filter entirely; logs go quiet
An AI essay generator that did not exist last termUnder-blockAcademic-integrity policy is unenforceable because the tool is invisible
A dormant domain that changed hands and contentEither, unpredictablyA site allowed for years starts serving something else entirely
A legitimate site classified from an old parked pageOver-blockA teacher's lesson plan breaks mid-period with no obvious reason
A vendor domain that migrated to a new subdomainOver-blockAn approved edtech platform half-works; helpdesk tickets pile up
A category definition that shifted with the webOver-blockPolicy blocks more than the board ever agreed to, invisibly

Daily rebuilds address both columns at once. New domains enter classification within the same cycle they are discovered, and existing records are re-checked on rotation so a site that changed hands does not keep its old label for a school year. That second half is the one vendors rarely talk about, and it is the half that causes the confusing tickets.

A category that did not exist five years ago

The AI tools blocklist: 16,328+ domains, tracked separately

Generative AI produced an entire class of sites at a speed no general web taxonomy was built to absorb. It is maintained as its own structured dataset inside the database rather than being flattened into a single "AI" label.

Structured, not a single bucket

The 16,328+ tracked AI-tool domains are organised into 18 categories and more than 165 subcategories. That structure exists because "AI" is not a policy position. A district may want a research assistant available in eleventh grade, a paraphrasing tool blocked everywhere, and an image generator allowed only in the art lab — three different answers that a single label cannot express.

Where it overlaps student safety

Some corners of this dataset are not academic-integrity questions at all. Deepfake and face-swap services, voice-cloning tools and AI companion chatbots raise harassment, image-abuse and privacy problems that sit uncomfortably close to what CIPA exists to address. Districts consistently find these subcategories more urgent than the essay writers once they can see them enumerated.

Deepfake & face-swap Voice cloning Companion chat

Why it has to be tracked daily

AI tools appear, rebrand, get acquired and relaunch on new domains faster than any other segment of the web. A quarterly list of AI sites is a historical document. Screening newly registered domains every day is the only way this dataset stays useful, and it is why the AI blocklist rides the same daily build as everything else.

Same daily cycle

Worth saying plainly: nothing in CIPA requires blocking AI tools. This dataset exists because districts asked for visibility into a category they could not see, not because a statute demands it. The detailed breakdown lives on the AI tools blocklist page, with the academic-integrity angle covered separately under blocking AI cheating tools and the safety angle under deepfake and NSFW AI filtering.

Consumption

Four ways to get the data, and the honest trade-off in each

The same classifications are available as an API, a CSV download, a DNS blocklist, or the full database. None of them is the right answer for everyone, and a vendor that pretends otherwise is selling you their architecture rather than solving your problem.

REST API

Query a domain, get its record back. Nothing to store, nothing to refresh, and every answer reflects the current build the moment you ask.

The trade-off: you take on a network dependency in the request path. If your filter needs a verdict in single-digit milliseconds under classroom load, live lookups need caching in front of them, and you should plan what happens when the call times out.

Best for real-time lookups

CSV download

Flat files you load into whatever you already run — a filtering appliance, a database, a spreadsheet a director actually opens.

The trade-off: the data is only as current as your last import, so the value depends entirely on whether the refresh is automated. A CSV pipeline nobody scheduled becomes the stale-data problem described above.

Best for existing stacks

DNS blocklist

Category-derived zones you point a resolver at. This is the lightest possible deployment: no agents, no inline appliance, network-wide coverage in an afternoon.

The trade-off: DNS operates on domains, so it cannot make per-path distinctions. It blocks a site, not a section of one, and it does not follow a device off the district network by itself.

Best for fast coverage

Full database download

The entire classified corpus, on your infrastructure. No per-query cost, no external dependency, and complete freedom to index and join it however you like.

The trade-off: it is a real dataset with real storage and refresh obligations, and it goes stale the day you stop pulling updates. Best suited to teams who already run infrastructure and want no third-party call in the hot path.

Best for full control
ConsiderationAPICSVDNS blocklistFull download
Always reflects the latest build YesOn importOn zone refreshOn sync
Works with no internet dependency at query time No YesResolver-local Yes
Deployable without touching endpointsDependsDepends YesDepends
Supports per-URL or per-path decisions Yes Yes Domain only Yes
Storage and maintenance burden on youNoneLowLowSubstantial
Typical time from contract to first blockHoursDaysHoursDays to weeks

Most districts end up using two: a bulk format for the everyday corpus and the API for lookups the bulk copy cannot answer. Mechanics for each are in the integration guide, and the DNS-first path is covered in depth on the DNS filtering for education page.

Accuracy

Quality assurance, and what happens when we get one wrong

Any database of this size contains mistakes. The question worth asking a vendor is not whether errors exist but what the path is from noticing one to having it fixed — and how long that path takes.

How quality is defended before publication

  • Low-confidence classifications are routed to review instead of being published as if they were certain
  • High-traffic domains get proportionally more scrutiny — an error on a top-ranked site affects far more students than one on an obscure domain
  • Sensitive categories carry a higher bar in both directions, because a false positive on adult content is a phone call and a false negative is an incident
  • Existing records are re-checked on rotation, so a domain that changed hands does not keep a label earned years ago
  • Category definitions are reviewed as the web shifts, so a label still means in June what it meant in September

Popularity ranking does real work in this process. It is the signal that decides what gets human attention when there is more to check than there is time to check it.

Recategorisation requests from districts

When a district finds a domain in the wrong place, that report is a signal we cannot generate any other way. A network administrator who knows a particular vendor's course platform got flagged as streaming is telling us something no crawler noticed.

Requests are reviewed against the same criteria as any other classification, and an accepted change flows into the next daily build for every customer — not as a private override attached to one district's account. That distinction matters: private overrides let the same error persist for everyone else, and they quietly fork your copy of the data away from ours.

While a review is in progress, districts are not stuck. Local allow and block entries in your own filtering layer take precedence immediately, so a teacher blocked out of a legitimate resource can be unblocked in minutes and the database catches up behind them.

Before filing anything, the domain checker will show you exactly what the current build says about a domain — which frequently reveals that the classification is correct and the surprise is coming from a policy rule further down the stack.

Coverage

Where the 120 million domains actually are

Coverage is usually quoted as one number, which hides the thing districts need to know: whether the parts of the web their students actually reach are well described.

Global, not US-only

Every record carries a country association, and popularity is banded both globally and within-country. That second banding is what makes the database usable outside the largest markets: a site that ranks nowhere worldwide can be the most-visited news outlet in its own country, and a filter that only understands global rank will treat it as an obscure unknown.

For districts with multilingual student populations, this is not a theoretical concern. Students reach content in the languages they speak at home, and those domains rarely appear in an English-language top-sites list.

Across legacy and new TLDs

The classic namespace — .com, .org, .net, .edu — is the easy part and everyone covers it. The interesting coverage question is the long tail: country-code domains and the hundreds of newer generic TLDs, which are cheap, fast to register, and consequently over-represented among the sites schools most want to catch.

Because discovery runs off registration feeds rather than a curated seed list, new TLDs enter the pipeline on the same footing as .com. There is no namespace waiting to be added later.

Depth where it counts

Coverage of the popular web is close to complete, and coverage of the deep tail is where every database in this market thins out. We would rather say that plainly than claim a number nobody can verify.

The practical consequence is that unknown-domain handling is a policy decision your filter has to make, and it should be a conscious one. Districts that block unknowns outright get tighter compliance and more helpdesk traffic; districts that allow them get the reverse.

Statute vs. policy

How the data maps onto CIPA — and where it stops

CIPA covers three things and only three things: obscene material, child sexual abuse material, and material harmful to minors. Everything else a district blocks is a local decision. The database keeps that boundary visible instead of blurring it.

What is being blockedBasisWhat that means in practice
Obscene materialCIPA statutory requirementCovered by the adult and obscenity categories, blocked for all users under any policy template
Child sexual abuse materialCIPA statutory requirementIllegal without exception; no configuration in which these categories can be permitted
Material harmful to minorsCIPA statutory requirementBroader than obscenity; multi-label classification lets a mixed site be restricted for students while remaining reachable for staff
Social media platformsDistrict discretionNot mentioned in the statute; commonly restricted by grade band for instructional reasons
Streaming, video and gamesDistrict discretionOften bandwidth or attention policy rather than safety policy — and worth labelling as such
AI tools and chatbotsDistrict discretionAcademic integrity and student-safety judgements the district makes for itself
Proxies, VPNs and anonymisersDistrict discretion, compliance-adjacentNot required by name, but permitting them undermines every other control you rely on
Malware and phishing domainsDistrict discretion, security-drivenOutside CIPA entirely; blocked because the alternative is a ransomware incident

The reason we keep insisting on this: districts that blur the statutory floor into their policy preferences end up unable to answer a parent who asks why a site is blocked. "The law requires it" is a strong answer when true and a corrosive one when it is not. Keeping the two layers separate in the data makes the honest answer easy to give.

How the categories translate into a defensible technology protection measure and E-Rate evidence is covered on the CIPA-compliant web filter page. If you are earlier in the process, the plain-English guide to what CIPA is and the compliance checklist are the better starting points.

Integration

Dropping the data into what you already run

This is a database, not an appliance. It is designed to make the filtering stack you already own better informed rather than to replace it — which is usually the cheaper and less disruptive path for a district mid-contract.

Firewalls and secure web gateways

Most enterprise firewalls accept external dynamic lists or equivalent URL-object feeds. Category-derived lists are published in the formats those systems expect, so a district can enrich an existing gateway's native categories with ours instead of migrating platforms. Practical note: check your platform's per-list entry ceiling before you point it at a large category, and split by subcategory if you hit it.

DNS resolvers and RPZ

For resolver-based filtering, category zones load as response policy zones on standard DNS infrastructure. This is the fastest route to network-wide coverage and needs nothing installed on a single endpoint. Set the refresh interval to match the daily build cycle, because a zone pulled once and forgotten is the single most common source of stale-data complaints we see.

Custom and in-house systems

Districts and edtech vendors building their own filtering, dashboards or reporting call the API directly or load the full download into their own store. Because every record carries IAB v2 and v3 placements alongside the school-facing category, the same dataset serves both a filtering decision and an analytics rollup without a second vocabulary to reconcile.

Authentication, rate limits, file layouts, refresh scheduling and worked examples for each path are in the integration guide. Licensing terms and redistribution boundaries are set out under licensing, and availability commitments under the service level agreement.

Boundaries

What this database deliberately does not do

Scope is a design decision, and the omissions here are all intentional. Knowing them up front saves a district from discovering mid-deployment that it bought half of what it assumed.

Not in scope, by design

  • It does not enforce anything. The database describes domains; your filter, firewall or resolver does the blocking. Nothing here inspects traffic or terminates a connection.
  • It holds no student data. No names, no devices, no browsing histories, no identifiers of any kind. It is a description of the public web, and it stays that way.
  • It does not monitor individuals. Keystroke logging, screen capture and student-level surveillance are not features we omitted for a later release — they are not products we build.
  • It does not classify per page. The unit is the domain. Distinguishing one article from another on the same site is a job for your filter's inline inspection, not for this dataset.
  • It does not judge quality. Categories say what a site is about, not whether it is accurate, well made, or appropriate for a given lesson. That call belongs to educators.
  • It does not write your policy. Templates and defaults help, but no vendor can decide for a community what its middle schoolers should reach.

Why drawing the line here is the point

A dataset that stays a dataset is portable. A district can change filtering vendors, move from an on-premise appliance to a cloud gateway, or bring reporting in-house, and the classification layer comes along unchanged. That portability is only possible because the data does not assume anything about the system reading it.

The privacy boundary is equally deliberate. Because no student information is ever involved in producing or consuming a classification, an entire category of procurement questions simply does not apply — there is no data-sharing agreement to negotiate over browsing records, because there are no browsing records. Districts under state student-privacy statutes tend to find this the shortest review they run all year.

And the surveillance boundary reflects a view about what filtering is for. A filter that blocks a harmful site protects a student. A system that reports on what each child looked at all day changes the relationship between a school and its students, and it is not obvious that the trade is worth making. We build the first thing.

Questions

Database questions we get asked most

It means 120 million classified domain records. Some are large sites with thousands of pages, some are single-page operations, and some are domains that exist mainly to be registered. The count reflects how many domains have a category attached, which is the number that determines whether your filter recognises a domain when a student requests it. It is not a claim about how many distinct organisations are on the web.

Your filter decides — and it should be a decision you made deliberately rather than a default you inherited. Blocking unknowns is the conservative posture and generates helpdesk traffic; allowing them is friendlier and leaves a gap. Either way, an uncategorised lookup from a live deployment pushes that domain up our classification queue, so the same domain is typically resolved within the next build cycle rather than staying unknown indefinitely.

Yes, and it is the feature that makes the database usable in a school. Sites are frequently more than one thing, and forcing a single label means either over-blocking a useful resource or letting through the part of it you wanted stopped. Layered labels let a policy act on the combination — permit the platform, restrict the categories inside it — without the district maintaining a hand-built exception list to paper over the gap.

This page is the overview: what a record contains, how the pipeline works, how the taxonomy is shaped, how you consume the data, and what it deliberately excludes. The coverage page goes narrower and deeper on one question — how much of the web is actually described and how that breaks down. Start here for the shape of the thing, go there when you need the numbers behind a specific claim.

Usually not. Most appliances ship with their own categorisation, and its depth varies enormously — particularly outside the English-language web and in fast-moving segments like AI tools. Districts commonly layer our categories on top of a gateway's native ones through an external list feed, which closes the gaps without a platform migration. If your current categories already handle everything you throw at them, you do not need us, and we would rather say so.

Accepted changes flow into the next daily build and reach every customer at once. In the meantime your local allow and block rules override the database immediately, so nobody has to wait on us to unblock a teacher who needs a resource this period. The database catching up afterwards is what stops that local override from becoming permanent technical debt.

If you have limited infrastructure staff, the DNS blocklist is almost always the right starting point: it covers the whole network, installs nothing on endpoints, and can be running the day you get access. Add API lookups later if you need per-request decisions your resolver cannot make. The full download makes sense when you have a team that wants the data in-house and no external dependency in the request path — it is the most capable option and the most work.

Test the database before you trust it

Look up the domains you already argue about internally — the edtech platform that keeps getting blocked, the AI site that keeps getting through — and see what the current build says about each one.

Check a Domain See Pricing Talk to Us