Skip to content

Introduction to Reversing

Reverse engineering (reversing) is the art of understanding how a program works without its source code, starting only from the compiled binary. It’s used to analyze malware, find vulnerabilities, bypass protections (licenses, DRM), understand proprietary protocols, and solve CTF challenges. It consists of translating machine code back into something understandable —assembly and, with decompilers, pseudo-C— and reasoning about its logic.

- malware analysis: understand what a suspicious executable does (see malware)
- vulnerability hunting: find bugs in software without source (pwn)
- protection bypass: licenses, DRM, anti-cheat, checks
- interoperability: understand proprietary formats/protocols
- CTF: solve "crackme" challenges (find the password/flag)

The two main, complementary approaches:

Static analysis (see rev-static) examine the binary WITHOUT running it
-> disassemble/decompile, read the logic, strings, imports
-> safe (you don't run malware), but more laborious
Dynamic analysis (see rev-dynamic) run the binary and observe it
-> debugger, traces, monitor syscalls/APIs, see memory live
-> fast for understanding behavior, but you run the code (risk with malware)

In practice they combine: static to map the logic, dynamic to confirm and for what’s hard to follow statically (obfuscated code, runtime decryption).

file binary # type, architecture, stripped? static?
strings binary # readable strings: messages, paths, URLs, hints
strings -e l binary # unicode strings (Windows)
checksec --file=binary # mitigations (if ELF)
nm binary / readelf -s # symbols (if not stripped)
ltrace / strace ./binary # library/system calls (quick dynamic)

strings is always the first step: it often reveals passwords, error messages, function names, or the flag itself.

ELF Linux (executables, .so)
PE Windows (.exe, .dll)
Mach-O macOS/iOS
# also: .NET (IL, decompilable almost to source with dnSpy/ILSpy),
# Java (bytecode -> jd-gui), Python (pyc -> decompyle3), etc.

Bytecode languages (.NET, Java, Python) are much easier to reverse (nearly recovering source); C/C++ compiled to native is the hardest.

# decompilers/disassemblers (static)
Ghidra free, powerful, pseudo-C decompiler (NSA) -> the free standard
IDA Pro the commercial standard; IDA Free limited
Binary Ninja, radare2/cutter, Hopper
# debuggers (dynamic)
gdb + pwndbg/GEF (Linux), x64dbg (Windows), lldb
# utilities
strings, file, nm, objdump, readelf, ltrace/strace
# bytecode: dnSpy/ILSpy (.NET), jadx/jd-gui (Java/Android)
1. file + strings + checksec -> first idea
2. open in Ghidra -> locate main / the check function
3. read the logic (pseudo-C): how does it validate the password/flag?
4. if clear -> deduce the correct input
5. if obfuscation/encryption -> dynamic (gdb): set breakpoints, see memory at runtime
6. (if applicable) patch the binary to bypass the check

Reversing is made harder —not prevented— with protections covered in rev-antidebug:

- code and string obfuscation
- packers / executable encryption (decrypted at runtime)
- anti-debugging / anti-VM / analysis detection
- integrity checks (anti-tampering)
# to protect intellectual property or malware; for the analyst, obstacles to bypass
  • file/strings/checksec as first contact
  • Distinguish static vs dynamic and when to use each
  • Recognize the format (ELF/PE/Mach-O/bytecode)
  • Open a binary in Ghidra and locate main/logic
  • Read pseudo-C and follow the validation logic
  • ltrace/strace to see calls quickly
  • Combine static + dynamic on obfuscated code
  • Solve a basic crackme (deduce/patch)