Domain & Category Coverage

The Data Behind the Filter

Every block page and every allow decision your students see traces back to one question: does the filter actually know the site? Ours does — more than 120 million domains classified into 57+ content categories, refreshed daily, with new domains labeled as they appear on the web.

120M+Domains classified
57+Content categories
DailyUpdates
16,328+AI-tool domains
Why Coverage Matters

A filter is only as good as the data behind it

Every request a student's device makes gets compared to a database of known domains and their categories — and when a domain isn't in that database, the filter has to guess. In a school district, guessing is where CIPA compliance quietly breaks down.

The web moves faster than any list

Hundreds of thousands of new domains are registered every single day, existing sites change what they host without changing their name, and entire categories of tool — AI chatbots and generators chief among them — have gone from nonexistent to ubiquitous in the space of two school years. A coverage database that was comprehensive eighteen months ago can be meaningfully behind today, and the gap shows up first as unrated traffic in your logs.

One domain, multiple categories

A single domain can legitimately belong to more than one category at once — a retailer that sells both school supplies and alcohol, a video platform that hosts both curriculum content and unmoderated live chat, a "productivity" AI tool that also generates images on request. A filter that can only assign one label per domain is forced to pick a side, and whichever side it picks, it gets something wrong for somebody's policy.

The standard we hold our own database to

Broad enough to recognize the sites your students actually visit, granular enough to label what those sites actually do, and fresh enough that "new this week" doesn't mean "unrated this month." The rest of this page walks through exactly how that database is built, kept current, and turned into the block and allow decisions your policy engine makes every day.

Coverage Dashboard

Numbers you can hold us to

Not marketing round numbers — this is the working shape of the database that powers every filtering decision.

120M+

Domains classified

The active, categorized domain database

57+

Content categories

Purpose-built for K-12 policy decisions

24/7

Automated crawling

Discovery never stops running

16,328+

AI-tool domains

Chatbots, generators, and AI extensions

Multi-label domains

Carry two or more categories at once

Database freshness

Pushed to production daily

Human review layer

Analyst-audited edge cases and appeals

New-domain triage

Classified on first meaningful appearance

1

Discovery

Crawlers, registration feeds, and DNS activity surface new domains before they show up in meaningful traffic.

2

Content analysis

Automated systems inspect page content, structure, and hosting signals to determine what the domain actually is.

3

Multi-label assignment

A domain that fits more than one category gets more than one label, not a forced single answer.

4

Review & correction

Analysts audit low-confidence classifications, edge cases, and customer-submitted appeals.

5

Publish to the policy engine

The finished classification ships into the live database that every filtering decision reads from.

CIPA-critical content Required

Obscenity, child sexual abuse material, and material harmful to minors — the categories CIPA itself requires be blocked.

Safety & risk Sensitive

Self-harm, suicide, violence, weapons, and extremist content that warrants a policy conversation, not a default.

Security threats

Malware, phishing, botnets, and newly registered domains showing suspicious hosting patterns.

Media & entertainment

Streaming, gaming, social media, and video sharing, broken out by platform behavior, not just domain name.

AI tools Fast-moving

Chatbots, image and text generators, homework helpers, and AI browser extensions — 16,328+ domains and counting.

Education & reference Instructional

LMS platforms, research databases, reference sites, and vetted classroom tools kept reliably reachable.

Multi-category labeling

One domain, several labels

Real sites don't sort neatly into one bucket. A general retailer that carries firearms accessories, a news outlet with an unmoderated comment section, a "creative writing" AI tool that will also draft an essay word-for-word — a single-category filter has to pick one truth and quietly ignore the rest.

Our database assigns every domain as many categories as actually apply, so your policy engine can make a real decision instead of an approximation: allow the retailer, restrict the weapons-adjacent pages; allow the news article, flag the comment stream; allow research use of an AI tool while logging or restricting generation features.

That is the difference between a filter that reacts to a domain name and one that understands what a domain does.

Live database lookup
quicksketch-ai.example
AI Tools Image Generation Creative & Design Unrated Content Risk
First seen14 days ago
Category count4 active labels
Review statusAnalyst-confirmed
Policy exampleBlock generation, allow gallery
Step 1

Continuous crawling

The web doesn't stand still, so discovery doesn't either — crawling and registration monitoring run around the clock, not on a batch schedule.

Step 2

New domain intake

Freshly registered domains and sites newly seen in real traffic enter a triage queue rather than waiting for a scheduled sweep.

Step 3

Same-day classification

Most new domains receive an initial category the same day they're identified — ahead of any meaningful student traffic reaching them.

Step 4

Daily push to production

The live filtering database updates every day, not every quarter, so classification work turns into protection on the same timeline.

Compliance

Why coverage is a CIPA question

The Children's Internet Protection Act requires filtering visual depictions that are obscene, child pornography, or harmful to minors. That requirement is written in terms of content — but every filter enforces it by matching against a category database.

The gap nobody notices

If a site carrying harmful content isn't in the database, or isn't labeled correctly, the CIPA requirement isn't actually being met — regardless of how the policy is configured. Nobody gets an alert when a domain slips through uncategorized. The request simply resolves, the student loads the page, and the gap only becomes visible later, if at all.

Compliance is only as strong as the database

A district can have a technically correct filtering policy and still fail the spirit of CIPA because the underlying database hasn't caught up with what's actually on the web. Deep, current coverage is what turns a written policy into an enforced one — and makes a vendor's compliance claims verifiable rather than aspirational.

Scenario120M+ domain databaseThin / legacy database
A new AI chatbot launchesClassified and filterable within days of appearing in trafficUnrated and silently allowed for weeks or months
A student finds a lesser-known siteAlready carries an assigned category from prior crawlingFalls into "uncategorized," often allowed by default
A site sells alcohol and also blogsBoth categories applied; policy restricts one, permits the otherSingle label misses one side of the site entirely
A phishing domain registers this morningEnters the classification queue immediately, often blocked same dayWaits for the next periodic list refresh, which can be weeks out
IT audits filtering logs for the boardDetailed, current category data supports a clear reportLogs show generic "uncategorized" hits that are hard to explain
District policy example
campussports-hub.example
Sports & News Live Chat / Social Video Streaming
AllowArticles, scores, video recaps
RestrictLive chat & comment threads
Applied byElementary + secondary profiles
ResultOne domain, two enforced outcomes
Built for school policy

Multi-category coverage, built for how schools actually decide

Schools don't need one blanket rule for the entire internet — they need different rules for elementary, middle, and high school, and for instructional time versus open periods. That only works if the underlying data supports it. A sports news domain with an unmoderated comment section shouldn't have to be entirely blocked or entirely allowed; multi-category labeling lets a policy permit the articles and restrict the chat, for the exact age group where that distinction matters.

The same logic applies across the database: a shopping site that also carries age-restricted products, a video platform mixing curriculum and unmoderated uploads, a search engine with an image mode that behaves differently than its text mode. Coverage depth is what makes those nuanced policies enforceable instead of theoretical.

In Practice

What daily updates actually mean on a Tuesday

Daily updates aren't a feature you interact with — they're the reason the gap between "something new appeared on the internet" and "our filter knows what it is" stays measured in hours instead of months.

Overnight: automatic discovery

The crawling pipeline picks up several thousand domains that either registered for the first time or crossed the threshold into meaningful traffic. None of that required anyone to notice a problem first — it happened because discovery runs continuously, not because a support ticket triggered it.

Mid-morning: student traffic appears

A handful of those domains are already showing up in student traffic: a new AI writing tool that spread through a group chat over the weekend, a gaming site that changed hosting providers, a shopping page that added a firearms accessories section. Each one moves through content analysis and multi-label assignment the same day.

Afternoon: live in production

Those classifications are live in the production database. Your policy engine doesn't need a manual override request, a support ticket, or a wait for "the next update cycle" — the AI tool is already labeled and filterable, the re-hosted gaming site still resolves to its correct category, and the shopping site now carries both labels.

Invisible unless something goes wrong

None of that is visible from the IT office unless something goes wrong, which is exactly the point. This is how the gap between "something new appeared on the internet" and "our filter knows what it is" stays measured in hours instead of months.

The database currently classifies more than 120 million domains across 57+ content categories. That number changes daily as new domains are discovered, reviewed, and published to the live filtering database — it isn't a static list that ages between refreshes.

An unrated domain gets handled according to your district's policy for uncategorized content, which most schools configure conservatively. At the same time, that domain enters the discovery and classification pipeline, so a truly new site typically doesn't stay unrated for long.

Daily. Discovery and crawling run continuously, and classified domains are pushed to the production database every day rather than on a weekly, monthly, or quarterly cycle.

Yes, and most policy-relevant domains do. A retailer that also sells age-restricted products, a news site with an unmoderated comment section, or an AI tool that both writes essays and generates images will carry every category that applies, not just one.

AI tools are tracked as their own dedicated category — currently 16,328+ domains — because they don't behave like traditional websites. A single AI tool can be used for legitimate research, homework help, or plagiarism, so it's classified separately and often paired with additional labels describing what the tool actually does.

Automated content analysis makes the first pass on every domain, but low-confidence results, edge cases, and customer-submitted appeals are routed to human analysts for review and correction before anything is published to the live database.

CIPA requires filtering of obscene content, child sexual abuse material, and material harmful to minors. That requirement is enforced technically by matching domains against a category database — so if a site carrying that content isn't in the database or isn't labeled correctly, the requirement isn't actually being met no matter how the policy is written.

Yes. Filtering logs and policy reporting show the category or categories a domain carried at the time of the request, so a block or allow decision can be traced back to the underlying classification rather than treated as a black box.

It's the opposite in practice. Broader, multi-category coverage gives the policy engine more precise information to work with, so decisions rely less on broad catch-all rules for uncategorized traffic — which is usually where over-blocking and under-blocking both come from.

Put the coverage claims to the test

Run a sample of your own traffic logs against the database, or talk through how multi-category coverage would apply to your district's policies.