Skip to content

Historical Archives (Wayback)

The Internet doesn’t forget: the Wayback Machine and other archives keep old versions of websites. For an attacker, that’s a goldmine: endpoints no longer linked but still working, old vulnerable parameters, files that were “removed” but not deleted, secrets in previous code versions, and subdomains/technologies the company no longer uses but left exposed. A site’s past reveals surface its present hides.

  • Historical URLs and endpoints: paths, parameters, and APIs no longer linked.
  • Old parameters: candidates for fuzzing (many are still active and vulnerable).
  • Removed files: backups, docs, configs pulled from the index but not the server.
  • Technology changes: what they used before (old versions still reachable).
  • Deleted content: text, emails, names the company removed.
Wayback Machine (archive.org) the main historical archive
archive.today / archive.ph point-in-time snapshots
Common Crawl massive web corpus
Google/Bing cache recent cached versions
# gau: URLs from Wayback + Common Crawl + others
gau target.com | sort -u > urls.txt
# waymore: more thorough (Wayback + CC + URLScan + VirusTotal)
waymore -i target.com -mode U
# waybackurls (classic)
waybackurls target.com
# Wayback's direct API
curl -s "http://web.archive.org/cdx/search/cdx?url=target.com*&output=text&fl=original&collapse=urlkey"
# filter those with parameters (candidates for fuzzing/IDOR/SQLi/XSS)
cat urls.txt | grep "=" | sort -u
# extract only parameter names (to fuzz values)
cat urls.txt | grep -oE '[?&][a-zA-Z0-9_]+=' | sort -u
# check which are still alive today
cat urls.txt | httpx -mc 200,301,302,403 -silent
# hunt for juicy files in the history
cat urls.txt | grep -iE '\.(bak|sql|zip|env|config|log|json)$'

Live URLs with parameters are direct targets for the Web section tests (SQLi, XSS, IDOR, open redirect).

# a specific snapshot of a page (see the code/secrets of that time)
https://web.archive.org/web/2021*/target.com/config.js

Useful to recover old JavaScript with endpoints/keys no longer in the current version.

You can’t erase the archive’s past, but you can: rotate any secret that was ever in a public version, truly remove (not just unlink) sensitive resources from the server, disable old endpoints/parameters, and assume everything you ever published is still accessible to anyone who looks.

  • Extract historical URLs (gau/waymore/waybackurls)
  • Filter URLs with parameters (fuzzing candidates)
  • Check which are still alive (httpx)
  • Hunt for removed files (bak/sql/env/config)
  • Recover old JS with endpoints/secrets
  • Detect abandoned technologies/subdomains
  • Feed the Web tests with live endpoints
  • Review archive.today for point-in-time snapshots