Every block page, every allowed link and every audit report a district produces traces back to one thing: a list that says what a website is. This page is the full account of that list — what a single record holds, where new domains come from, how the 57+ category taxonomy is shaped around school decisions rather than advertising taxonomies, how often it changes, and the four ways districts pull it into whatever they already run.
People picture a blocklist as two columns: a domain and a verdict. That shape breaks the first time a school needs to allow a site for staff and restrict it for fourth graders. Our unit of storage is a described domain, not a judgement about one.
A record answers the question "what is this site?" and stops there. It does not carry a block flag, because a block flag would encode our opinion of your policy into your data. A high school library and a K-2 building can pull the identical row and reach opposite conclusions, which is exactly as it should be.
That separation has a practical payoff at renewal time. When a district changes its stance on, say, AI homework helpers, nothing about the underlying data has to change — the policy that reads the data changes. Districts that have lived through a blocklist migration know how much work that saves.
Alongside the primary web filtering category, each row carries the domain's position in the IAB v2 and IAB v3 content taxonomies, down to four tiers where the tree goes that deep. Those trees matter to districts that already own analytics or ad-management tooling keyed to IAB labels: the same export drops straight into those systems without a translation layer someone has to maintain.
Two popularity signals round the record out. OpenPageRank gives a rough measure of a domain's authority on the open web, and the popularity rank groups indicate whether a site sits in the global top thousand, the long tail, or somewhere between — globally and within its own country. Those numbers are what let a network administrator triage: a miscategorisation on a top-1000 domain deserves attention this afternoon, while an obscure parked domain can wait for the next review cycle.
example.com10.0 at the very top1-1000Full types, null behaviour and example values are documented in the data dictionary. The same field set is present whether you consume the data by API, CSV or full download.
Roughly 300,000 domains are registered every day. A classification database that only grows when a customer complains will always be one incident behind, so discovery has to be a standing process rather than a reaction.
New domains enter the queue from three directions. Newly registered domain feeds catch sites the moment they exist, often before any content is on them. Crawl expansion pulls in domains linked from pages already known to us, which is how regional and non-English sites surface without anyone hand-seeding a country list. And uncategorised lookups from live filtering deployments push a domain to the front of the queue: if a student in a district somewhere tried to reach it this morning, it is more urgent than a domain nobody has visited.
The classifier works from what a browser would receive, not from the domain string. Page text, title and meta description, visible navigation, on-page media cues and the site's own structural signals are collected. Guessing from a name is where cheap blocklists fail spectacularly — wholly innocuous domains contain unfortunate substrings, and plenty of genuinely harmful sites are named after fruit.
Collected signals are scored against the full category set, producing a primary web filtering category plus the IAB placements. Confidence matters as much as the label: a strong, unambiguous signal on an adult site is treated very differently from a marginal call on a small business page, and the low-confidence tail is what gets routed to review rather than published unchecked.
A video platform hosting educational lectures alongside content nobody wants in a middle school is genuinely both. Rather than forcing a single winner, labels are layered so a filter can act on the intersection: allow the platform, block the categories within it that a given grade band should not reach. Single-label databases push that entire problem onto the district's exception list, where it becomes somebody's weekly chore.
Accepted classifications land in the next daily build. From there the same records populate the API, the CSV exports, the DNS blocklist artefacts and the full database download simultaneously — there is no premium tier of the data and no lagging secondary copy. What the API says at noon is what the CSV you downloaded that morning says.
Why this order matters to a school: the gap between a site appearing and a site being classified is the entire window in which a filter fails silently. Discovery driven by registration feeds and live lookups — rather than by complaints — is what keeps that window measured in hours instead of months. You can test the current state of any single domain yourself with the domain checker.
Most content taxonomies were designed for advertisers, who care about what a page is selling. Schools care about something different: whether a student should be there, and at what age. The category structure is shaped around three tiers of decision.
Adult and obscene content, child sexual abuse material, graphic violence and gore, drug marketplaces, self-harm and pro-suicide communities, hate and extremist material. These carry the categories that map onto CIPA's statutory requirements and the harms no district has ever debated in a board meeting.
Granularity still matters here, because "adult" is not one thing. Separating explicit content from mature-but-not-explicit material lets a district block both for elementary students while allowing a high school health curriculum to function.
Social media, streaming and video, games, chat and messaging, forums, shopping, dating, AI tools. Nothing in CIPA touches any of these. They are where district policy lives, and where the argument between the technology office and the teaching staff usually happens.
Fine-grained categories are what let that argument end well. A district that can block competitive gaming while leaving educational simulations reachable has a different conversation than one whose only lever is a single "Games" switch.
News, reference, education, health, business, sports, government, libraries and museums, software and technology. The overwhelming majority of the 120 million classified domains sit here, and the job of the taxonomy is to keep them out of the way.
This tier is quietly the most important for classroom credibility. Over-blocking teachers into a corner produces workarounds — personal hotspots, unmanaged devices — that do far more damage to a filtering programme than any single missed site.
The number of categories in a filtering database is easy to treat as marketing arithmetic. It is not. Each category is a lever, and a district can only make distinctions its data supports. With a coarse taxonomy, the choices available to a middle school principal are "block the whole of social media" or "allow the whole of social media" — and since neither is acceptable, what actually happens is an exception list that grows for three years until nobody remembers why half the entries are there.
A 57-plus category structure changes the shape of that work. Policy stops being a pile of individual domains and becomes a small set of category decisions that can be written down, explained to a school board, defended to a parent, and handed to a successor. The exception list still exists, but it stays short enough to review.
The full category listing, including how each one is defined and typical school treatment, is documented on the filtering categories and taxonomy page and in the category reference in the docs.
Generative AI produced an entire class of sites at a speed no general web taxonomy was built to absorb. It is maintained as its own structured dataset inside the database rather than being flattened into a single "AI" label.
The 16,328+ tracked AI-tool domains are organised into 18 categories and more than 165 subcategories. That structure exists because "AI" is not a policy position. A district may want a research assistant available in eleventh grade, a paraphrasing tool blocked everywhere, and an image generator allowed only in the art lab — three different answers that a single label cannot express.
Some corners of this dataset are not academic-integrity questions at all. Deepfake and face-swap services, voice-cloning tools and AI companion chatbots raise harassment, image-abuse and privacy problems that sit uncomfortably close to what CIPA exists to address. Districts consistently find these subcategories more urgent than the essay writers once they can see them enumerated.
AI tools appear, rebrand, get acquired and relaunch on new domains faster than any other segment of the web. A quarterly list of AI sites is a historical document. Screening newly registered domains every day is the only way this dataset stays useful, and it is why the AI blocklist rides the same daily build as everything else.
Worth saying plainly: nothing in CIPA requires blocking AI tools. This dataset exists because districts asked for visibility into a category they could not see, not because a statute demands it. The detailed breakdown lives on the AI tools blocklist page, with the academic-integrity angle covered separately under blocking AI cheating tools and the safety angle under deepfake and NSFW AI filtering.
The same classifications are available as an API, a CSV download, a DNS blocklist, or the full database. None of them is the right answer for everyone, and a vendor that pretends otherwise is selling you their architecture rather than solving your problem.
Query a domain, get its record back. Nothing to store, nothing to refresh, and every answer reflects the current build the moment you ask.
The trade-off: you take on a network dependency in the request path. If your filter needs a verdict in single-digit milliseconds under classroom load, live lookups need caching in front of them, and you should plan what happens when the call times out.
Flat files you load into whatever you already run — a filtering appliance, a database, a spreadsheet a director actually opens.
The trade-off: the data is only as current as your last import, so the value depends entirely on whether the refresh is automated. A CSV pipeline nobody scheduled becomes the stale-data problem described above.
Category-derived zones you point a resolver at. This is the lightest possible deployment: no agents, no inline appliance, network-wide coverage in an afternoon.
The trade-off: DNS operates on domains, so it cannot make per-path distinctions. It blocks a site, not a section of one, and it does not follow a device off the district network by itself.
The entire classified corpus, on your infrastructure. No per-query cost, no external dependency, and complete freedom to index and join it however you like.
The trade-off: it is a real dataset with real storage and refresh obligations, and it goes stale the day you stop pulling updates. Best suited to teams who already run infrastructure and want no third-party call in the hot path.
| Consideration | API | CSV | DNS blocklist | Full download |
|---|---|---|---|---|
| Always reflects the latest build | Yes | On import | On zone refresh | On sync |
| Works with no internet dependency at query time | No | Yes | Resolver-local | Yes |
| Deployable without touching endpoints | Depends | Depends | Yes | Depends |
| Supports per-URL or per-path decisions | Yes | Yes | Domain only | Yes |
| Storage and maintenance burden on you | None | Low | Low | Substantial |
| Typical time from contract to first block | Hours | Days | Hours | Days to weeks |
Most districts end up using two: a bulk format for the everyday corpus and the API for lookups the bulk copy cannot answer. Mechanics for each are in the integration guide, and the DNS-first path is covered in depth on the DNS filtering for education page.
Any database of this size contains mistakes. The question worth asking a vendor is not whether errors exist but what the path is from noticing one to having it fixed — and how long that path takes.
Popularity ranking does real work in this process. It is the signal that decides what gets human attention when there is more to check than there is time to check it.
When a district finds a domain in the wrong place, that report is a signal we cannot generate any other way. A network administrator who knows a particular vendor's course platform got flagged as streaming is telling us something no crawler noticed.
Requests are reviewed against the same criteria as any other classification, and an accepted change flows into the next daily build for every customer — not as a private override attached to one district's account. That distinction matters: private overrides let the same error persist for everyone else, and they quietly fork your copy of the data away from ours.
While a review is in progress, districts are not stuck. Local allow and block entries in your own filtering layer take precedence immediately, so a teacher blocked out of a legitimate resource can be unblocked in minutes and the database catches up behind them.
Before filing anything, the domain checker will show you exactly what the current build says about a domain — which frequently reveals that the classification is correct and the surprise is coming from a policy rule further down the stack.
Coverage is usually quoted as one number, which hides the thing districts need to know: whether the parts of the web their students actually reach are well described.
Every record carries a country association, and popularity is banded both globally and within-country. That second banding is what makes the database usable outside the largest markets: a site that ranks nowhere worldwide can be the most-visited news outlet in its own country, and a filter that only understands global rank will treat it as an obscure unknown.
For districts with multilingual student populations, this is not a theoretical concern. Students reach content in the languages they speak at home, and those domains rarely appear in an English-language top-sites list.
The classic namespace — .com, .org, .net, .edu — is the easy part and everyone covers it. The interesting coverage question is the long tail: country-code domains and the hundreds of newer generic TLDs, which are cheap, fast to register, and consequently over-represented among the sites schools most want to catch.
Because discovery runs off registration feeds rather than a curated seed list, new TLDs enter the pipeline on the same footing as .com. There is no namespace waiting to be added later.
Coverage of the popular web is close to complete, and coverage of the deep tail is where every database in this market thins out. We would rather say that plainly than claim a number nobody can verify.
The practical consequence is that unknown-domain handling is a policy decision your filter has to make, and it should be a conscious one. Districts that block unknowns outright get tighter compliance and more helpdesk traffic; districts that allow them get the reverse.
CIPA covers three things and only three things: obscene material, child sexual abuse material, and material harmful to minors. Everything else a district blocks is a local decision. The database keeps that boundary visible instead of blurring it.
| What is being blocked | Basis | What that means in practice |
|---|---|---|
| Obscene material | CIPA statutory requirement | Covered by the adult and obscenity categories, blocked for all users under any policy template |
| Child sexual abuse material | CIPA statutory requirement | Illegal without exception; no configuration in which these categories can be permitted |
| Material harmful to minors | CIPA statutory requirement | Broader than obscenity; multi-label classification lets a mixed site be restricted for students while remaining reachable for staff |
| Social media platforms | District discretion | Not mentioned in the statute; commonly restricted by grade band for instructional reasons |
| Streaming, video and games | District discretion | Often bandwidth or attention policy rather than safety policy — and worth labelling as such |
| AI tools and chatbots | District discretion | Academic integrity and student-safety judgements the district makes for itself |
| Proxies, VPNs and anonymisers | District discretion, compliance-adjacent | Not required by name, but permitting them undermines every other control you rely on |
| Malware and phishing domains | District discretion, security-driven | Outside CIPA entirely; blocked because the alternative is a ransomware incident |
The reason we keep insisting on this: districts that blur the statutory floor into their policy preferences end up unable to answer a parent who asks why a site is blocked. "The law requires it" is a strong answer when true and a corrosive one when it is not. Keeping the two layers separate in the data makes the honest answer easy to give.
How the categories translate into a defensible technology protection measure and E-Rate evidence is covered on the CIPA-compliant web filter page. If you are earlier in the process, the plain-English guide to what CIPA is and the compliance checklist are the better starting points.
This is a database, not an appliance. It is designed to make the filtering stack you already own better informed rather than to replace it — which is usually the cheaper and less disruptive path for a district mid-contract.
Most enterprise firewalls accept external dynamic lists or equivalent URL-object feeds. Category-derived lists are published in the formats those systems expect, so a district can enrich an existing gateway's native categories with ours instead of migrating platforms. Practical note: check your platform's per-list entry ceiling before you point it at a large category, and split by subcategory if you hit it.
For resolver-based filtering, category zones load as response policy zones on standard DNS infrastructure. This is the fastest route to network-wide coverage and needs nothing installed on a single endpoint. Set the refresh interval to match the daily build cycle, because a zone pulled once and forgotten is the single most common source of stale-data complaints we see.
Districts and edtech vendors building their own filtering, dashboards or reporting call the API directly or load the full download into their own store. Because every record carries IAB v2 and v3 placements alongside the school-facing category, the same dataset serves both a filtering decision and an analytics rollup without a second vocabulary to reconcile.
Authentication, rate limits, file layouts, refresh scheduling and worked examples for each path are in the integration guide. Licensing terms and redistribution boundaries are set out under licensing, and availability commitments under the service level agreement.
Scope is a design decision, and the omissions here are all intentional. Knowing them up front saves a district from discovering mid-deployment that it bought half of what it assumed.
A dataset that stays a dataset is portable. A district can change filtering vendors, move from an on-premise appliance to a cloud gateway, or bring reporting in-house, and the classification layer comes along unchanged. That portability is only possible because the data does not assume anything about the system reading it.
The privacy boundary is equally deliberate. Because no student information is ever involved in producing or consuming a classification, an entire category of procurement questions simply does not apply — there is no data-sharing agreement to negotiate over browsing records, because there are no browsing records. Districts under state student-privacy statutes tend to find this the shortest review they run all year.
And the surveillance boundary reflects a view about what filtering is for. A filter that blocks a harmful site protects a student. A system that reports on what each child looked at all day changes the relationship between a school and its students, and it is not obvious that the trade is worth making. We build the first thing.
It means 120 million classified domain records. Some are large sites with thousands of pages, some are single-page operations, and some are domains that exist mainly to be registered. The count reflects how many domains have a category attached, which is the number that determines whether your filter recognises a domain when a student requests it. It is not a claim about how many distinct organisations are on the web.
Your filter decides — and it should be a decision you made deliberately rather than a default you inherited. Blocking unknowns is the conservative posture and generates helpdesk traffic; allowing them is friendlier and leaves a gap. Either way, an uncategorised lookup from a live deployment pushes that domain up our classification queue, so the same domain is typically resolved within the next build cycle rather than staying unknown indefinitely.
Yes, and it is the feature that makes the database usable in a school. Sites are frequently more than one thing, and forcing a single label means either over-blocking a useful resource or letting through the part of it you wanted stopped. Layered labels let a policy act on the combination — permit the platform, restrict the categories inside it — without the district maintaining a hand-built exception list to paper over the gap.
This page is the overview: what a record contains, how the pipeline works, how the taxonomy is shaped, how you consume the data, and what it deliberately excludes. The coverage page goes narrower and deeper on one question — how much of the web is actually described and how that breaks down. Start here for the shape of the thing, go there when you need the numbers behind a specific claim.
Usually not. Most appliances ship with their own categorisation, and its depth varies enormously — particularly outside the English-language web and in fast-moving segments like AI tools. Districts commonly layer our categories on top of a gateway's native ones through an external list feed, which closes the gaps without a platform migration. If your current categories already handle everything you throw at them, you do not need us, and we would rather say so.
Accepted changes flow into the next daily build and reach every customer at once. In the meantime your local allow and block rules override the database immediately, so nobody has to wait on us to unblock a teacher who needs a resource this period. The database catching up afterwards is what stops that local override from becoming permanent technical debt.
If you have limited infrastructure staff, the DNS blocklist is almost always the right starting point: it covers the whole network, installs nothing on endpoints, and can be running the day you get access. Add API lookups later if you need per-request decisions your resolver cannot make. The full download makes sense when you have a team that wants the data in-house and no external dependency in the request path — it is the most capable option and the most work.
Look up the domains you already argue about internally — the edtech platform that keeps getting blocked, the AI site that keeps getting through — and see what the current build says about each one.