CRAWLER · PUBLIC WEB ONLY
A requested walk, not a roaming index.
URLand crawls only after a person asks it to map or battle a public website. Every crawl has host, network, redirect, content, depth, size, and time boundaries.
Safety boundary
NETWORK
Public addresses on standard web ports
URLand resolves and rechecks DNS, blocks private and reserved addresses, and accepts HTTP or HTTPS on ports 80 and 443 only.
SITE BOUNDARY
The requested host and direct www variant
Redirects and discovered links must stay inside that boundary. Redirects are checked at every hop and stop after five.
CONTENT
Public HTML only
Authenticated pages, embedded credentials, files, media, scripts, styles, and non-HTML responses are not part of the city crawl.
AGENT SAFETY
Website text is not an instruction
Battle agents receive sealed structural metadata and legal actions. Raw website page text is not included in their instructions.
Published crawl limits
CITY MAP
Up to 100 pages
Maximum link depth 4, 1 MB per page, 8 seconds per request, 48 seconds total, and concurrency up to 4.
PUBLIC-URL BATTLE
Up to 48 pages
Maximum link depth 4, 384 KB per page, 5 seconds per request, 24 seconds total, and concurrency up to 3.
Limits are ceilings, not promises. A crawl can stop earlier because of robots rules, unavailable pages, response type, network safety, or time.
robots.txt and link choices
URLand requests robots.txt and follows rules it can retrieve for URLandBot. A disallowed starting page is not crawled. Disallowed discovered pages and links marked nofollow are skipped. A robots file that cannot be retrieved is not treated as permission to cross any other safety boundary.
The crawler user agent points to the existing bot information page.
Product boundary reviewed 11 August 2026.