Skip to content

Document Metadata

The documents an organization publishes (PDFs, Word, Excel, PowerPoint, images) carry hidden metadata: who created them, with what software, on which machine, sometimes internal paths, usernames, printers, and even GPS coordinates in photos. Collecting a company’s public documents and extracting their metadata reveals internal users, naming conventions, software in use, and network structure —all useful for the attack phase.

Author / Last editor -> internal usernames (jdoe, a.perez)
Software and version -> which office suite/OS they use (vectors, exploits)
Computer name -> internal hostname convention
File paths -> folder structure, network drives
Printers / servers -> internal infrastructure
Dates -> activity, reused templates
GPS (photo EXIF) -> physical locations

The extracted usernames feed password spraying (see Password Spraying) and phishing directly.

# dorking by file type (see recon-dorking)
site:target.com filetype:pdf | filetype:docx | filetype:xlsx | filetype:pptx
# bulk-download the found documents
# FOCA (Windows): the classic tool -> downloads a domain's docs and extracts metadata
# metagoofil / metafinder: automate download + extraction
metagoofil -d target.com -t pdf,docx,xlsx -o ./docs
# exiftool: extracts metadata from any file (images, PDF, Office)
exiftool document.pdf
exiftool -r ./docs/ | grep -iE 'author|creator|producer|company'
# images (EXIF, GPS)
exiftool photo.jpg | grep -i gps
1. Author names -> internal user list -> deduce emails (recon-email)
2. Software/versions -> look for CVEs of that office suite/OS
3. Paths and hostnames -> understand the internal network (useful in post-exploitation)
4. GPS in photos -> locations (physical red team)

Just as you extract others’, don’t leak yours in reports/documents:

exiftool -all= file.pdf # remove all metadata

Prevention: clean metadata before publishing any document (sanitization tools, “document inspector” policies), configure the office suite not to embed author/paths, train on what files reveal, and periodically audit which public org documents leak metadata (run FOCA/metagoofil against your own domain).

  • Public documents collected (dorking by filetype)
  • Metadata extracted (exiftool/metagoofil/FOCA)
  • Author names → internal user list
  • Software and versions → look for CVEs
  • Internal paths and hostnames noted
  • GPS in images (if applicable)
  • Users cross-referenced with recon-email/ad-spray
  • My own documents cleaned of metadata