Skip to content

Google Dorking

Google (and other search engines) have already indexed much of what an organization exposes without realizing it: login panels, config files, internal documents, backups, credentials in plain text. “Dorking” is using advanced search operators to find exactly that. It’s 100% passive recon —just queries to the search engine— and among the most rewarding: critical findings without touching the target.

site:target.com limit to a domain (and subdomains)
inurl:admin term in the URL
intitle:"index of" term in the title (directory listings)
filetype:pdf / ext:sql by file extension
intext:"password" term in the body
cache:target.com cached version
-site:www.target.com exclude (see other subdomains)
"exact phrase" literal match
* wildcard

They combine: site:target.com filetype:pdf confidential.

# open directory listings
site:target.com intitle:"index of"
# indexed sensitive files
site:target.com ext:sql | ext:bak | ext:log | ext:env | ext:config
site:target.com filetype:xls | filetype:csv intext:password
# login / admin panels
site:target.com inurl:login | inurl:admin | inurl:dashboard
# documents with metadata (see recon-metadatos)
site:target.com filetype:pdf | filetype:docx | filetype:xlsx
# errors leaking info / paths
site:target.com intext:"sql syntax near" | "fatal error"
# subdomains and environments
site:*.target.com -www
# keys / secrets in indexed code
site:target.com intext:"api_key" | "BEGIN RSA PRIVATE KEY"
Bing (own operators: ip:, feed:) Yandex (very strong for images/dorks)
DuckDuckGo, Brave GitHub code search (see recon-code)
# dork aggregators
Exploit-DB's Google Hacking Database (GHDB): thousands of cataloged dorks
# leaked documents with metadata
site:target.com filetype:pdf -> download and extract metadata (recon-metadatos)
# exposed cloud buckets (see cloud)
site:s3.amazonaws.com target "target" in Google / grayhatwarfare

Dorking queries the search engine, not the target → very stealthy. Still, don’t automate hundreds of queries from your IP (Google will block you with CAPTCHAs); use tools that respect the rate or do it manually for what matters.

Prevent the sensitive from being indexed: robots.txt does not hide (it only asks not to index, and reveals paths); the right way is not to expose sensitive files, use authentication, X-Robots-Tag: noindex headers, and monitor with the same dorks what’s indexed about your org. Removing already-indexed content from Google requires deleting the resource + requesting removal.

  • site: to map what’s indexed of the domain and subdomains
  • Directory listings (intitle:"index of")
  • Sensitive files by extension (sql, env, bak, log, config)
  • Documents (pdf/docx/xlsx) for metadata
  • Login/admin panels (inurl:)
  • Errors leaking paths or stack
  • Secrets/keys in indexed content
  • Check GHDB for stack-specific dorks