Every block page and every allow decision your students see traces back to one question: does the filter actually know the site? Ours does — more than 120 million domains classified into 57+ content categories, refreshed daily, with new domains labeled as they appear on the web.
Every request a student's device makes gets compared to a database of known domains and their categories — and when a domain isn't in that database, the filter has to guess. In a school district, guessing is where CIPA compliance quietly breaks down.
Hundreds of thousands of new domains are registered every single day, existing sites change what they host without changing their name, and entire categories of tool — AI chatbots and generators chief among them — have gone from nonexistent to ubiquitous in the space of two school years. A coverage database that was comprehensive eighteen months ago can be meaningfully behind today, and the gap shows up first as unrated traffic in your logs.
A single domain can legitimately belong to more than one category at once — a retailer that sells both school supplies and alcohol, a video platform that hosts both curriculum content and unmoderated live chat, a "productivity" AI tool that also generates images on request. A filter that can only assign one label per domain is forced to pick a side, and whichever side it picks, it gets something wrong for somebody's policy.
Broad enough to recognize the sites your students actually visit, granular enough to label what those sites actually do, and fresh enough that "new this week" doesn't mean "unrated this month." The rest of this page walks through exactly how that database is built, kept current, and turned into the block and allow decisions your policy engine makes every day.
Not marketing round numbers — this is the working shape of the database that powers every filtering decision.
The active, categorized domain database
Purpose-built for K-12 policy decisions
Discovery never stops running
Chatbots, generators, and AI extensions
Carry two or more categories at once
Pushed to production daily
Analyst-audited edge cases and appeals
Classified on first meaningful appearance
Crawlers, registration feeds, and DNS activity surface new domains before they show up in meaningful traffic.
Automated systems inspect page content, structure, and hosting signals to determine what the domain actually is.
A domain that fits more than one category gets more than one label, not a forced single answer.
Analysts audit low-confidence classifications, edge cases, and customer-submitted appeals.
The finished classification ships into the live database that every filtering decision reads from.
Obscenity, child sexual abuse material, and material harmful to minors — the categories CIPA itself requires be blocked.
Self-harm, suicide, violence, weapons, and extremist content that warrants a policy conversation, not a default.
Malware, phishing, botnets, and newly registered domains showing suspicious hosting patterns.
Streaming, gaming, social media, and video sharing, broken out by platform behavior, not just domain name.
Chatbots, image and text generators, homework helpers, and AI browser extensions — 16,328+ domains and counting.
LMS platforms, research databases, reference sites, and vetted classroom tools kept reliably reachable.
Real sites don't sort neatly into one bucket. A general retailer that carries firearms accessories, a news outlet with an unmoderated comment section, a "creative writing" AI tool that will also draft an essay word-for-word — a single-category filter has to pick one truth and quietly ignore the rest.
Our database assigns every domain as many categories as actually apply, so your policy engine can make a real decision instead of an approximation: allow the retailer, restrict the weapons-adjacent pages; allow the news article, flag the comment stream; allow research use of an AI tool while logging or restricting generation features.
That is the difference between a filter that reacts to a domain name and one that understands what a domain does.
The web doesn't stand still, so discovery doesn't either — crawling and registration monitoring run around the clock, not on a batch schedule.
Freshly registered domains and sites newly seen in real traffic enter a triage queue rather than waiting for a scheduled sweep.
Most new domains receive an initial category the same day they're identified — ahead of any meaningful student traffic reaching them.
The live filtering database updates every day, not every quarter, so classification work turns into protection on the same timeline.
The Children's Internet Protection Act requires filtering visual depictions that are obscene, child pornography, or harmful to minors. That requirement is written in terms of content — but every filter enforces it by matching against a category database.
If a site carrying harmful content isn't in the database, or isn't labeled correctly, the CIPA requirement isn't actually being met — regardless of how the policy is configured. Nobody gets an alert when a domain slips through uncategorized. The request simply resolves, the student loads the page, and the gap only becomes visible later, if at all.
A district can have a technically correct filtering policy and still fail the spirit of CIPA because the underlying database hasn't caught up with what's actually on the web. Deep, current coverage is what turns a written policy into an enforced one — and makes a vendor's compliance claims verifiable rather than aspirational.
| Scenario | 120M+ domain database | Thin / legacy database |
|---|---|---|
| A new AI chatbot launches | Classified and filterable within days of appearing in traffic | Unrated and silently allowed for weeks or months |
| A student finds a lesser-known site | Already carries an assigned category from prior crawling | Falls into "uncategorized," often allowed by default |
| A site sells alcohol and also blogs | Both categories applied; policy restricts one, permits the other | Single label misses one side of the site entirely |
| A phishing domain registers this morning | Enters the classification queue immediately, often blocked same day | Waits for the next periodic list refresh, which can be weeks out |
| IT audits filtering logs for the board | Detailed, current category data supports a clear report | Logs show generic "uncategorized" hits that are hard to explain |
Schools don't need one blanket rule for the entire internet — they need different rules for elementary, middle, and high school, and for instructional time versus open periods. That only works if the underlying data supports it. A sports news domain with an unmoderated comment section shouldn't have to be entirely blocked or entirely allowed; multi-category labeling lets a policy permit the articles and restrict the chat, for the exact age group where that distinction matters.
The same logic applies across the database: a shopping site that also carries age-restricted products, a video platform mixing curriculum and unmoderated uploads, a search engine with an image mode that behaves differently than its text mode. Coverage depth is what makes those nuanced policies enforceable instead of theoretical.
Daily updates aren't a feature you interact with — they're the reason the gap between "something new appeared on the internet" and "our filter knows what it is" stays measured in hours instead of months.
The crawling pipeline picks up several thousand domains that either registered for the first time or crossed the threshold into meaningful traffic. None of that required anyone to notice a problem first — it happened because discovery runs continuously, not because a support ticket triggered it.
A handful of those domains are already showing up in student traffic: a new AI writing tool that spread through a group chat over the weekend, a gaming site that changed hosting providers, a shopping page that added a firearms accessories section. Each one moves through content analysis and multi-label assignment the same day.
Those classifications are live in the production database. Your policy engine doesn't need a manual override request, a support ticket, or a wait for "the next update cycle" — the AI tool is already labeled and filterable, the re-hosted gaming site still resolves to its correct category, and the shopping site now carries both labels.
None of that is visible from the IT office unless something goes wrong, which is exactly the point. This is how the gap between "something new appeared on the internet" and "our filter knows what it is" stays measured in hours instead of months.
Run a sample of your own traffic logs against the database, or talk through how multi-category coverage would apply to your district's policies.