Dark web · guide

Deep web vs dark web, and which one you are in now

By Alberto Gulotta · Updated · 17 min read

Deep web vs dark web is one question with two answers that are not close to each other. The deep web is every page a search engine does not hold — your webmail, your bank, your employer’s intranet. The dark web is a network you have to install software to reach, and you are not in it by accident.

Drawn table comparing surface web, deep web and dark web: what each names and what it takes to reach it
The three words, side by side, from Google Search Central and the Tor Project, both read on 23 September 2026. Figure drawn by AI Tools Primer.
“A full ninety-five per cent of the deep Web is publicly accessible information — not subject to fees or subscriptions.”

That is from the paper that named the deep web, published in the Journal of Electronic Publishing in August 2001 on data collected in March 2000. It is the opposite of the definition in circulation today, which is that the deep web is the paywalled, password-protected part of the internet. Both cannot be right, and only one of them can be read back to a source.

Which one are you in right now?

Almost certainly the deep web, and it is not a dramatic place to be. What is the deep web, in one sentence: every page a search engine does not hold in its index. Every page that asked you to sign in today was in it. The word describes a property of an index, not of a page — nothing has been hidden and nothing is suspicious. It is simply not in the list.

Five ordinary things, and which of the three words each one belongs to
The thingWhich wordWhy, and who says so
The page you are readingSurface web A search engine has crawled it and holds it in an index. Nothing is asked of you to open it.
Your webmail, open in another tabDeep web It is behind a sign-in, and Google names password protection as one of the two ways “to keep a web page out of Google”.
A page its owner blocked in robots.txtNeither, reliably Google: that file “is not a mechanism for keeping a web page out of Google”, and the address “can still appear in search results, but the search result won’t have a description”.
A newspaper’s anonymous tip boxDark web It is the Tor Project’s own worked example: “your local newspaper decides to set up an Onion Service (using SecureDrop) to receive anonymous tips”.
Facebook, reached over TorDark web The Tor Project lists “more secure ways to reach popular websites like Facebook” among what onion services are used for.

Three tests, and you can run all of them yourself

Not definitions to memorise: things you can check on the thing in front of you.

What makes a page deep: three ways out of an index

Drawn table of robots.txt, noindex and password protection, and which of them keeps a page out of Google
What actually makes a page part of the deep web. Every quoted phrase is Google’s own, from its documentation for site owners. Figure drawn by AI Tools Primer.

This is where most short explanations go wrong, and it is checkable against the documentation of the company that runs the index. There are three common ways a page stays out of Google, and they do not have the same effect. A rule in a file called robots.txt asks crawlers not to fetch the page: Google’s own guide says that file “is used mainly to avoid overloading your site with requests; it is not a mechanism for keeping a web page out of Google”, and that a page blocked that way can still have its address listed, only “the search result won’t have a description”.

The two that do work are a noindex rule and a password. Of the first, Google writes that when its crawler sees the rule it “will drop that page entirely from Google Search results, regardless of whether other sites link to it” — with a catch worth knowing, because it is the commonest mistake in this area: “for the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file”. Block the crawler and it never reads the instruction telling it to forget you. The second, a password, is the one almost everybody reading this is behind right now.

So the deep web is three arrangements with different consequences, and only one of them — the sign-in — is what people picture. For the neighbouring confusion, what the padlock in the address bar does and does not promise.

What makes a network dark: onion services

What is the dark web, then? Not a deeper layer of the same web: a different network laid over the ordinary one. The Tor Project — which builds it — defines the pieces without any drama: “Onion services are services that can only be accessed over Tor.” They use “the special-use top level domain (TLD) .onion (instead of .com, .net, .org, etc.)”, and your browser cannot resolve one of those addresses, because there is nothing in the ordinary internet’s address system to resolve it against.

Three properties follow, and they explain most of what looks strange about the place. First, there is no server address to find: “Onion services are an overlay network on top of TCP/IP, so in some sense IP addresses are not even meaningful to Onion Services: they are not even used in the protocol.” Second, the address is not a name somebody chose — it “looks weird and random because it’s the identity public key of the Onion Service”. Third, the connection is encrypted the whole way: “Onion service traffic is encrypted from the client to the onion host. This is like getting strong SSL/HTTPS for free.”

It is worth saying what these services are, because the word suggests a high street of shops and the documentation describes something else. The Tor Project’s own list: “metadata-free chat and file sharing, safer interaction between journalists and their sources like with SecureDrop or OnionShare, safer software updates, and more secure ways to reach popular websites like Facebook”. A software update service has an onion address for the same reason a newspaper does: so that using it reveals nothing. For the software itself, whether Tor Browser is safe, and safe from whom, and bridges for what happens when a network blocks it.

How big is each one, and who counted

Drawn table of the three numbers used about the deep web and dark web, with who measured each one
The three numbers, and who counted them: the 2001 paper and the Tor Project’s published daily data. Figure drawn by AI Tools Primer.

Here is the part that is usually asserted and almost never sourced. The famous figure — that the deep web is around 90% of the internet — has an ancestor worth meeting. In August 2001 the Journal of Electronic Publishing published a paper by Michael K. Bergman, “The Deep Web: Surfacing Hidden Value”, on data “collected between March 13 and 30, 2000”. Its headline finding was that “public information on the deep Web is currently 400 to 550 times larger than the commonly defined World Wide Web”, with “7,500 terabytes of information compared to nineteen terabytes” on the surface, across “more than 200,000 deep Web sites”.

Three things about that paper change how the number should be read. It was measuring searchable databases, not pages behind passwords: “Deep Web sources store their content in searchable databases that only produce results dynamically in response to a direct request.” It described a deep web that was mostly open — “a full ninety-five per cent” of it “publicly accessible information — not subject to fees or subscriptions”, which is the reverse of today’s definition. And it was a marketing document, something the journal said in its own introduction: it is “designed as a marketing tool” for a product, and the editors ran it anyway because of what it showed about the structure of the web. None of that makes it worthless: it makes it a twenty-five-year-old measurement of a different thing, which is not what a round number in a bullet list suggests.

The other side has a number that is neither old nor anonymous, and it is published every day by the people who run the network. Tor Metrics puts out a file counting “the number of unique .onion addresses for version 3 onion services in the network per day”, extrapolated “from aggregated statistics on unique version 3 .onion addresses reported by single relays acting as onion-service directories”. Read from that file on 23 September 2026, the twenty-eight days ending 20 September carry between 620,821 and 711,536 a day, averaging 672,540; the last day in the file, 20 September 2026, is 620,821. Those are addresses rather than sites, one service can hold several, and the publisher labels the figure an estimate — but it has a source, a method and a date, which is three more than the percentages have.

One consequence worth carrying away, because it is the reason people arrive here. When a company writes to say your details have been “found on the dark web”, that sentence means something narrow and checkable: the data turned up somewhere reachable only over Tor. When the same sentence says deep web, it means almost nothing — your details are on the deep web by definition, and so are everybody’s, because that is where an account lives. The two words are not interchangeable in that message, and the difference decides how worried to be.

Is any of this illegal?

No legal advice here, and none of the sources read for this page gives any. What they do show is that the technology is documented in public by the organisation that builds it, that its worked example is a newspaper receiving anonymous tips, and that its listed uses include software updates and reaching Facebook. A network is not a crime and neither is a word; what happens on it is a separate question, and it is not this page’s.

The practical version of the question is usually different anyway, and it has a page already: if something has written to tell you your information was found somewhere, start at what a data breach means and what to do in the first hour, and if somebody is already using the details, identity theft has deadlines attached to it.

Where to start

Three ways on, depending on what brought you to the two words.

“Something says my data is on the dark web.”
That is a breach, and it is survivable — Data breach
“Am I on the deep web right now?”
Almost certainly, and the three tests are in which one are you in
“Where does the 90% figure come from?”
A paper from 2001, measuring something else — how big is each one

If something has already happened

The reader this page is mostly written for did not arrive out of curiosity: a bank, an employer or a subscription service sent them a message. These four are the practical sequence, and none of them costs anything.

The technology, one page each

What the software does, written from the documentation of the people who build it rather than from a summary of a summary. No addresses on any of these, by a rule that is not up for discussion.

What is visible without any of this

The deep web is where your accounts live, which makes it the part of this subject that touches you daily. These four are about what is already visible, and what a browser gives away before any network sees it.

Questions people also ask

Is the dark web illegal?

This page gives no legal opinion, and none of its sources does either. What they show is that the network is documented in public by the nonprofit that builds it, that its own worked example is a newspaper collecting anonymous tips through SecureDrop, and that its listed uses include software updates and reaching Facebook more securely.

Is DuckDuckGo dark web?

No. The Tor Project’s definition of an onion service is one “that can only be accessed over Tor”, and duckduckgo.com opens in any ordinary browser, with no Tor involved. It is a search engine on the ordinary web.

What is considered a deep web?

Any page an index does not hold. In practice that means three arrangements: a sign-in, a subscription, or a noindex rule in the page itself. A block in robots.txt does not reliably count — Google says that file “is not a mechanism for keeping a web page out of Google”.

Is the deep web bigger than the dark web?

Yes, by any measure anybody publishes, and the two figures are not comparable. The deep web figure in circulation traces to a 2001 paper measuring searchable databases in March 2000. The dark web has a live count instead: 620,821 to 711,536 unique .onion addresses a day in the four weeks to 20 September 2026.

Not covered here. It publishes no addresses. No .onion links, no directories, no mirrors, not as illustration and not as curiosity — the Tor Project’s own protocol page prints an example address and it has deliberately not been copied across.

It does not explain how to get there, which is what most of the better-known writing on these two words spends its second half doing. Explaining how a mechanism works is one thing; a set of steps is another, and this is not the page for it.

It does not say how dangerous the place is, either. That is a different question with a different set of answers, it will get its own page, and answering it in passing here would mean repeating what the monitoring industry says about the product it sells.

And it does not sell anything. Twelve pages about these two words were read while writing this one, leaving aside the documentation and the 2001 paper, and ten of the twelve carry an offer — monitoring, a free trial or a demo. One of the remaining two is a university selling degrees in the subject. There is nothing here to buy, no subscription recommended, and no commission earned by anybody from a word of it. What holds instead is simple: every number on this page names who counted it, and when.

Sources

  1. Michael K. Bergman, “The Deep Web: Surfacing Hidden Value” — Journal of Electronic Publishing, volume 7 issue 1, August 2001, on data collected between 13 and 30 March 2000. This is the paper that named the deep web and produced the size estimate the round percentages descend from; it is read here for the estimate itself (400 to 550 times larger than the surface web, 7,500 terabytes against nineteen, more than 200,000 deep web sites), for what it was measuring (searchable databases returning results dynamically), and for the finding that ninety-five per cent of the deep web was publicly accessible and not subject to fees or subscriptions. Read through the Internet Archive because the journal’s own addresses do not serve it: quod.lib.umich.edu returns about 47 words and a security check, journals.publishing.umich.edu about 204 words of a bot challenge. The journal’s introduction notes that the paper is designed as a marketing tool for its author’s employer, and that is stated on this page too — web.archive.org, read 23 September 2026.
  2. Tor Project, “How do Onion Services work?” — the protocol described by the organisation that maintains it (onion services as services reachable only over Tor; the overlay network in which IP addresses are not used by the protocol; the address as the identity public key of the service; encryption from client to onion host; and the worked example of a newspaper running SecureDrop to receive anonymous tips). The page contains an example .onion address, which is not reproduced anywhere here — community.torproject.org, read 23 September 2026.
  3. Tor Project, “What are .onion sites and onion services?” — the same organisation’s short definition, read for the .onion top level domain and for what these services are used for beyond websites: metadata-free chat and file sharing, contact between journalists and sources through SecureDrop or OnionShare, safer software updates, and reaching popular sites such as Facebook — support.torproject.org, read 23 September 2026.
  4. Tor Metrics, “Unique .onion addresses (version 3 only)” — the daily count published by the Tor Project, with the page’s own description of what the number is and how it is produced. The figures on this page are read from the CSV behind that graph, not from the picture: 620,821 to 711,536 a day across the twenty-eight days ending 20 September 2026, average 672,540, last day in the file 20 September 2026. The same organisation as the two sources above, and counted as one source rather than three — metrics.torproject.org, read 23 September 2026.
  5. Google Search Central, “Robots.txt Introduction and Guide” — the documentation of the company that runs the index, read for what robots.txt does and does not do: it is used mainly to avoid overloading a site with requests, it is not a mechanism for keeping a web page out of Google, and a page blocked with it can still have its address listed without a description — developers.google.com, read 23 September 2026.
  6. Google Search Central, “Block Search Indexing with noindex” — the other half of the same rule: a noindex rule makes Google drop the page entirely from search results regardless of inbound links, and it only works if robots.txt is not blocking the crawler from reading the rule in the first place — developers.google.com, read 23 September 2026.

Written by Alberto Gulotta

Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of taking computers apart, starting with a Commodore 64 — the long version is on the about page.

Something wrong on this page? Write to aitoolsprimer@gmail.com and it gets fixed.

Written on 23 September 2026.

Independence and limits

No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here, and no company has seen this page before you did.

This is general information, not professional advice. Where a page touches money, health, safety or the law, it names its source and the date it was read — and your situation may still differ. See the privacy page and the cookie policy.