Reconnaissance: Information Gathering
CAP, ACID vs BASE, latency numbers, back-of-envelope estimation, single points of failure — the vocabulary every system designer thinks in.
Core Philosophy: Attacks are won or lost before a single exploit is fired. Reconnaissance — methodically learning everything about a target — is what separates someone who pokes at the obvious from someone who finds the forgotten door. You attack what you can find, so the practitioner who maps the target most thoroughly has the biggest advantage. Recon is patient, unglamorous, and the highest-leverage skill in the offensive toolkit.
Part 1: The Problem
Beginners want to jump straight to exploitation — find the SQL injection, fire the payload. But you cannot exploit what you haven’t found. A real application is far bigger than its homepage: forgotten subdomains, old admin panels, staging servers, exposed files, APIs nobody documented. Real attackers and skilled bug bounty hunters spend the majority of their time here, in reconnaissance, because the bug that pays is usually on the asset nobody remembered existed.
This section teaches you to see the whole target, not just the front door.
Part 2: The Concept — Passive vs Active Recon
Reconnaissance splits into two modes, and the distinction matters both tactically and legally.
| Mode | What it is | Does it touch the target? | Detectable? |
|---|---|---|---|
| Passive recon | Gathering information from third-party sources about the target | No — you never interact with the target’s systems | Essentially invisible to the target |
| Active recon | Directly interacting with the target’s systems to learn about them | Yes — you send requests to the target | Yes — appears in the target’s logs |
Passive recon is reading public information: search engines, DNS records, public code repositories, certificate transparency logs, archived versions of pages. You’re touching other people’s systems (search engines, public databases), not the target.
Active recon is sending requests to the target itself — probing its servers, enumerating its content, scanning its ports. This interacts with systems you must be authorized to test.
PASSIVE RECON ACTIVE RECON
You ──► search engines You ──────────► TARGET
You ──► public DNS data (direct requests, probing,
You ──► cert transparency enumeration — shows up in
You ──► archived pages the target's logs)
(target never sees you)
⚖️ Scope still governs active recon. Within an authorized engagement or an in-scope bug bounty target, active recon is part of the job. Outside that, sending probing requests to a target is unauthorized access. Passive recon on public sources is generally safe; active recon needs authorization.
Part 3: What You’re Looking For
Recon hunts for a handful of high-value categories of information:
1. The attack surface — every asset the target owns.
- Subdomains.
app.example.com,api.example.com,dev.example.com,vpn.example.com… Each subdomain is potentially a separate application with its own bugs. Forgotten ones (old.,staging.,test.) are gold because they’re often un-patched. - IP ranges and hosts. The actual servers behind the names.
- Ports and services. What’s listening (this overlaps with Phase 3’s scanning).
2. The technology stack — what it’s built with. What web server, framework, language, CMS, and libraries the target uses, and their versions. A specific version often maps directly to specific known vulnerabilities. This is called fingerprinting.
3. Content and endpoints — the hidden parts of the app.
- Directories and files not linked from anywhere (
/admin,/backup,/.git,/api/v1). - API endpoints.
- Parameters that pages accept.
This is content discovery, often done by directory brute-forcing — testing thousands of common paths to see which return a response.
4. Human and organizational information (OSINT).OSINT (Open-Source Intelligence) is intelligence gathered from public sources: employee names and emails, technologies mentioned in job postings, leaked credentials in past breaches, information in public code repositories. A job ad saying “must know AWS and Django” just told you the tech stack.
Part 4: The Recon Toolkit
A first set of tools for the job. Add these to your Tools & Reference Cheatsheet page.
| Tool / technique | Purpose |
|---|---|
| Search engines | Passive discovery. Search operators (e.g. site:, filetype:) narrow results to a target — sometimes called “Google dorking.” |
whois | Domain registration details. |
dig / nslookup | DNS records — IPs, mail servers, name servers, subdomains. |
| Certificate transparency logs | Public logs of issued TLS certificates — a rich source of subdomains. |
| Subdomain enumeration tools | Automate finding subdomains (e.g. Amass, Subfinder). |
| Web archive | Archived past versions of a site — reveals old endpoints, removed pages. |
| Directory brute-forcers | Content discovery — testing common paths (e.g. ffuf, dirb, Gobuster). |
| Tech fingerprinting | Identify the stack (e.g. the Wappalyzer browser extension, WhatWeb). |
| Nmap | Port and service discovery (covered fully in Phase 3.1). |
You don’t need all of these on day one. Learn a couple per category and expand.
Part 5: A Recon Methodology
Recon without a method becomes aimless clicking. Use a repeatable flow:
1. SCOPE Confirm exactly what is authorized. Write it down.
│
2. PASSIVE Gather all you can WITHOUT touching the target:
│ whois, DNS, cert logs, search engines,
│ archives, public repos, OSINT.
3. ASSETS Build the full list of in-scope assets:
│ subdomains, hosts, IP ranges.
4. ACTIVE For in-scope assets, probe directly:
│ port/service scan, tech fingerprinting,
│ content discovery / directory brute-force.
5. MAP Assemble everything into a target map:
│ every app, endpoint, technology, version.
6. PRIORITIZE Pick the most promising attack surface to test
first (old/forgotten assets, admin areas, APIs).
Then — and only then — does targeted exploitation (sections 2.4 onward) begin. The map you build here is what every later technique consults.
Part 6: Why Recon Wins
A closing point worth fully absorbing: in bug bounty especially, recon is the differentiator. Thousands of hunters look at the same well-known main application and test the same obvious inputs. The hunter who finds the bug is often the one who discovered legacy-api.example.com — a forgotten asset that nobody patched in years — through better recon.
The same is true in professional pentesting: thorough recon means you test the whole attack surface, not just the polished front-facing app, and that’s where serious findings hide.
Recon is also continuous, not a one-time phase. Organizations constantly add new subdomains, deploy new services, ship new features. Skilled hunters monitor their targets for new attack surface appearing over time (a technique developed fully in Track A, section 5A.3).
The lesson: be patient and thorough in recon. The exciting exploitation later is built entirely on the unglamorous mapping you do here.
📓 Key Terms
| Term | Plain meaning |
|---|---|
| Reconnaissance (recon) | Systematically gathering information about a target. |
| Passive recon | Gathering info without ever touching the target’s systems. |
| Active recon | Gathering info by directly interacting with the target. |
| OSINT | Open-Source Intelligence — information from public sources. |
| Subdomain | A sub-section of a domain (api.example.com), often a separate app. |
| Fingerprinting | Identifying the technologies and versions a target uses. |
| Content discovery | Finding unlinked directories, files, and endpoints. |
| Directory brute-forcing | Testing many common paths to discover hidden content. |
| Attack surface map | The assembled picture of all of a target’s testable points. |
🧪 Hands-On Lab
⚖️ Passive recon on public sources is fine to practice broadly. Active recon (Tasks 4–5) is done only against your own lab VM or an explicitly authorized target.
Task 1 — Passive recon on a domain. Pick a large, well-known company. Using only passive sources, find: its registration info (whois), its DNS records (dig), and as many subdomains as you can via certificate transparency logs. You never touched their servers — this is all public data.
Task 2 — Fingerprint a tech stack. Install the Wappalyzer browser extension. Visit several websites and see what it reports — web server, framework, analytics, CMS. You’re reading the technology stack passively.
Task 3 — Practice search operators. Learn 4–5 search engine operators (site:, filetype:, intitle:, inurl:). Use them to find specific things on a large public site — e.g. all PDF files on a domain. This is “dorking,” and it’s pure passive recon.
Task 4 — Active recon on your lab. Against your Phase 0.5 victim VM, run a directory brute-forcer (e.g. ffuf or Gobuster) with a common wordlist. See which paths it discovers. This is content discovery — done legally, on a machine you own.
Task 5 — Build a target map. For your lab victim VM, assemble a single Notion page: its IP, open ports/services, identified technologies and versions, and discovered directories/endpoints. This is your first attack surface map — the deliverable recon always produces.
⚠️ Common Mistakes
- Skipping recon to rush to exploitation. The bug is usually on the asset you didn’t bother to find. Recon is the work.
- Running active recon on out-of-scope or unauthorized targets. Directory brute-forcing and scanning are active — they hit the target and appear in logs. Authorization is mandatory.
- Ignoring forgotten assets. The polished main app is the most-tested thing on the internet. Old subdomains, staging servers, and legacy APIs are where real findings live.
- Not recording findings. Recon produces a lot of information. Without an organized map, you lose track and re-do work. Document as you go.
- Treating recon as one-and-done. Attack surface grows over time. Good practitioners re-run recon and monitor for new assets.
✅ Recap & What’s Next
- Reconnaissance is systematically mapping a target; you can only exploit what you first find.
- It splits into passive (no contact with the target — generally safe) and active (direct probing — needs authorization), and hunts for the attack surface, the tech stack, hidden content, and OSINT.
- Recon is the highest-leverage offensive skill — thorough mapping, especially of forgotten assets, is what separates a finder from a poker.
Next (2.2): With a target mapped, you need one environment to intercept, inspect, and manipulate every request you send. That tool is Burp Suite — the web hacker’s daily workbench.
⁂ Back to all modules