Public check

Privacy and limitations

Describes the code as deployed on 2026-08-06.

What we store

The domain you checked, the time you checked it, the contract version the checks ran under, and the result each check landed on — for every check in the list further down this page, which state it ended in (allowed, blocked, present, absent, unknown, and so on) and which rule decided that state. Not the full URL you typed, not the page content, not your IP address, not anything that identifies you. Every one of these is an observation about a public website, not about the person looking at it.

We also store counts describing the shape of what the check fetched: the scanned site's own HTTP status, the size of the document and how much of it was visible text, how long the scan took, how many JSON-LD blocks it declared and how many failed to parse, which schema.org vocabulary terms those blocks used, how many hreflang alternates and sitemap URLs it declared, and how many internal links, external links, and images it contains (and how many of those images have no alt text or an empty one). These are counts and vocabulary terms, never the values themselves — we record that a page declared, say, three schema.org types and twelve images, not what those images show or what the schema markup says. We also store a guess at the publishing platform and what that guess was based on.

We use this to know which domains get checked and what the check is finding across everyone who runs it — the only way to build a distribution honest thresholds can be calibrated against. It is not a per-visitor history: nothing ties a row back to you or to any browser session. What is stored is the outcome of each check, not its workings — the evidence quotes, the raw file contents, and the reasoning behind each row are built for the page you see and the PDF you can save, and none of that is written down here.

This table lives in a database pinned to the EU jurisdiction — not just hosted in an EU region as a default, but restricted at the platform level to run and store data within the EU.

Logging is best-effort: a failure to write a row never surfaces to you and never affects your result.

How long, and in what form

Decided, being built: what leaves this database for any external use or publication will be aggregates only, never a row naming a specific domain, and the domain on each row will be replaced with a salted hash 30 days after it's written — the measurements stay usable for the distribution above, the link back to which domain they came from does not. That mechanism is not live yet. Until it ships, domain rows are retained as written, without an automatic deletion or hashing window. This page will say so plainly once the hashing is in place rather than continuing to describe it as pending.

Analytics on this site

This site (not the check specifically — every page on cror.link) uses Google Analytics 4 (Google LLC) to measure visits: page views, referrer, approximate location, and device and browser information. It runs through Cloudflare Zaraz, which loads and proxies the tag through cror.link's own domain rather than loading it directly from Google's servers. Google Analytics sets first-party cookies (_ga and_ga_*) to distinguish visits. Data collected this way is retained under Google's own retention setting for this property, not a window we control from our side.

There is currently no consent prompt gating this — it runs on every page load. We also use Google Search Console to see how Google's own crawler and search results treat this site; that involves no script on the page and no visitor data, only what Google already computes from crawling and indexing a public site.

What the check fetches

The URL you submit, and up to three files it points to: robots.txt, a declared sitemap, and /llms.txt. All four are public resources — the same ones an ordinary crawler would request. Nothing behind a login, nothing you haven't already published.

Each fetch identifies itself honestly with the User-AgentCrorLinkSourceCheck/1.0. It never poses as a browser or as a named crawler belonging to another company.

What the check itself talks to

Before fetching, the check resolves your submitted hostname through Cloudflare's public DNS-over-HTTPS resolver (1.1.1.1), to confirm it doesn't point at a private or internal address. That request carries the hostname only — never the full URL, never anything you typed beyond the domain name, never the page content.

This is the only call the check itself makes to any service outside Cloudflare's own infrastructure — no AI provider, no advertising network. The site as a whole also runs Google Analytics, described above, which this check does not add to or depend on.

Your IP address

Processed, not stored. Your IP is used transiently throughcf-connecting-ip to limit how many checks one address can run per minute (five), so the check stays available to everyone. This runs through Cloudflare's Workers rate-limiting counters, which Cloudflare's own documentation describes as approximate and short-lived rather than an accounting system — we cannot query, export, or look up what any address has done. Your IP is never written to the database, and nothing there connects it to a domain you checked.

What we log

Standard Cloudflare Workers request logs — timestamp, status code, response time — kept for reliability and abuse response, the same as any web service. These logs are not reviewed for URLs you checked, and we have not built any tooling to search them by URL, domain, or visitor. If that changes, this notice will say so before it happens.

The export

The Save as PDF option is built by your browser's print function from the result already on your screen. It is not generated by our server, and it does not pass back through it. We do not receive a copy.

Where the security boundary is honest about its limits

The check refuses to fetch private, local, or reserved addresses — including on every hop of a redirect chain, not only the URL you submitted. Before fetching, it also resolves the hostname and checks that resolved address against the same list.

That second check is a best-effort signal, not a guarantee: the address a DNS lookup returns is not provably the address the fetch itself connects to a moment later. This is a known class of gap (sometimes called DNS rebinding), and we are not claiming to have closed it — only to have narrowed it. If this matters for your use of the check, treat it as a research tool for public resources, not as a boundary you can rely on for anything sensitive.

What the check cannot tell you

It reads declarations — what a page and its related files say to automated systems. It does not observe whether any crawler visited, whether the page is indexed, or whether any AI system has read, cited, or recommended it. The result page states this in full next to every result; it is not buried here.

If this changes

This notice is written against the code, not against a plan for it. If we ever add storage, telemetry, or a new third-party call, this page will describe it — with what is collected, for how long, and why — before it ships, not after.

Questions: hello@cror.link.

Back to the check