Document Metadata
The documents an organization publishes (PDFs, Word, Excel, PowerPoint, images) carry hidden metadata: who created them, with what software, on which machine, sometimes internal paths, usernames, printers, and even GPS coordinates in photos. Collecting a company’s public documents and extracting their metadata reveals internal users, naming conventions, software in use, and network structure —all useful for the attack phase.
What metadata reveals
Section titled “What metadata reveals”Author / Last editor -> internal usernames (jdoe, a.perez)Software and version -> which office suite/OS they use (vectors, exploits)Computer name -> internal hostname conventionFile paths -> folder structure, network drivesPrinters / servers -> internal infrastructureDates -> activity, reused templatesGPS (photo EXIF) -> physical locationsThe extracted usernames feed password spraying (see Password Spraying) and phishing directly.
Collect the documents
Section titled “Collect the documents”# dorking by file type (see recon-dorking)site:target.com filetype:pdf | filetype:docx | filetype:xlsx | filetype:pptx# bulk-download the found documentsExtract metadata
Section titled “Extract metadata”# FOCA (Windows): the classic tool -> downloads a domain's docs and extracts metadata# metagoofil / metafinder: automate download + extractionmetagoofil -d target.com -t pdf,docx,xlsx -o ./docs# exiftool: extracts metadata from any file (images, PDF, Office)exiftool document.pdfexiftool -r ./docs/ | grep -iE 'author|creator|producer|company'# images (EXIF, GPS)exiftool photo.jpg | grep -i gpsFrom metadata to attack
Section titled “From metadata to attack”1. Author names -> internal user list -> deduce emails (recon-email)2. Software/versions -> look for CVEs of that office suite/OS3. Paths and hostnames -> understand the internal network (useful in post-exploitation)4. GPS in photos -> locations (physical red team)Clean your own metadata (attacker OPSEC)
Section titled “Clean your own metadata (attacker OPSEC)”Just as you extract others’, don’t leak yours in reports/documents:
exiftool -all= file.pdf # remove all metadataFor the defense
Section titled “For the defense”Prevention: clean metadata before publishing any document (sanitization tools, “document inspector” policies), configure the office suite not to embed author/paths, train on what files reveal, and periodically audit which public org documents leak metadata (run FOCA/metagoofil against your own domain).
Testing checklist
Section titled “Testing checklist”- Public documents collected (dorking by filetype)
- Metadata extracted (exiftool/metagoofil/FOCA)
- Author names → internal user list
- Software and versions → look for CVEs
- Internal paths and hostnames noted
- GPS in images (if applicable)
- Users cross-referenced with recon-email/ad-spray
- My own documents cleaned of metadata