Home
Cybersecurity & AI Security / Part 13 — Reconnaissance: Information Gathering

Reconnaissance: Information Gathering

CAP, ACID vs BASE, latency numbers, back-of-envelope estimation, single points of failure — the vocabulary every system designer thinks in.


Core Philosophy: Attacks are won or lost before a single exploit is fired. Reconnaissance — methodically learning everything about a target — is what separates someone who pokes at the obvious from someone who finds the forgotten door. You attack what you can find, so the practitioner who maps the target most thoroughly has the biggest advantage. Recon is patient, unglamorous, and the highest-leverage skill in the offensive toolkit.

Part 1: The Problem

Beginners want to jump straight to exploitation — find the SQL injection, fire the payload. But you cannot exploit what you haven’t found. A real application is far bigger than its homepage: forgotten subdomains, old admin panels, staging servers, exposed files, APIs nobody documented. Real attackers and skilled bug bounty hunters spend the majority of their time here, in reconnaissance, because the bug that pays is usually on the asset nobody remembered existed.

This section teaches you to see the whole target, not just the front door.

Part 2: The Concept — Passive vs Active Recon

Reconnaissance splits into two modes, and the distinction matters both tactically and legally.

Mode What it is Does it touch the target? Detectable?
Passive reconGathering information from third-party sources about the targetNo — you never interact with the target’s systemsEssentially invisible to the target
Active reconDirectly interacting with the target’s systems to learn about themYes — you send requests to the targetYes — appears in the target’s logs

Passive recon is reading public information: search engines, DNS records, public code repositories, certificate transparency logs, archived versions of pages. You’re touching other people’s systems (search engines, public databases), not the target.

Active recon is sending requests to the target itself — probing its servers, enumerating its content, scanning its ports. This interacts with systems you must be authorized to test.

text
   PASSIVE RECON                    ACTIVE RECON
   You ──► search engines           You ──────────► TARGET
   You ──► public DNS data          (direct requests, probing,
   You ──► cert transparency         enumeration — shows up in
   You ──► archived pages            the target's logs)
   (target never sees you)
⚖️ Scope still governs active recon. Within an authorized engagement or an in-scope bug bounty target, active recon is part of the job. Outside that, sending probing requests to a target is unauthorized access. Passive recon on public sources is generally safe; active recon needs authorization.

Part 3: What You’re Looking For

Recon hunts for a handful of high-value categories of information:

1. The attack surface — every asset the target owns.

2. The technology stack — what it’s built with. What web server, framework, language, CMS, and libraries the target uses, and their versions. A specific version often maps directly to specific known vulnerabilities. This is called fingerprinting.

3. Content and endpoints — the hidden parts of the app.

This is content discovery, often done by directory brute-forcing — testing thousands of common paths to see which return a response.

4. Human and organizational information (OSINT).OSINT (Open-Source Intelligence) is intelligence gathered from public sources: employee names and emails, technologies mentioned in job postings, leaked credentials in past breaches, information in public code repositories. A job ad saying “must know AWS and Django” just told you the tech stack.

Part 4: The Recon Toolkit

A first set of tools for the job. Add these to your Tools & Reference Cheatsheet page.

Tool / technique Purpose
Search enginesPassive discovery. Search operators (e.g. site:, filetype:) narrow results to a target — sometimes called “Google dorking.”
whoisDomain registration details.
dig / nslookupDNS records — IPs, mail servers, name servers, subdomains.
Certificate transparency logsPublic logs of issued TLS certificates — a rich source of subdomains.
Subdomain enumeration toolsAutomate finding subdomains (e.g. Amass, Subfinder).
Web archiveArchived past versions of a site — reveals old endpoints, removed pages.
Directory brute-forcersContent discovery — testing common paths (e.g. ffuf, dirb, Gobuster).
Tech fingerprintingIdentify the stack (e.g. the Wappalyzer browser extension, WhatWeb).
NmapPort and service discovery (covered fully in Phase 3.1).

You don’t need all of these on day one. Learn a couple per category and expand.

Part 5: A Recon Methodology

Recon without a method becomes aimless clicking. Use a repeatable flow:

text
1. SCOPE     Confirm exactly what is authorized. Write it down.
                 │
2. PASSIVE   Gather all you can WITHOUT touching the target:
                 │  whois, DNS, cert logs, search engines,
                 │  archives, public repos, OSINT.
3. ASSETS    Build the full list of in-scope assets:
                 │  subdomains, hosts, IP ranges.
4. ACTIVE    For in-scope assets, probe directly:
                 │  port/service scan, tech fingerprinting,
                 │  content discovery / directory brute-force.
5. MAP       Assemble everything into a target map:
                 │  every app, endpoint, technology, version.
6. PRIORITIZE  Pick the most promising attack surface to test
                 first (old/forgotten assets, admin areas, APIs).

Then — and only then — does targeted exploitation (sections 2.4 onward) begin. The map you build here is what every later technique consults.

Part 6: Why Recon Wins

A closing point worth fully absorbing: in bug bounty especially, recon is the differentiator. Thousands of hunters look at the same well-known main application and test the same obvious inputs. The hunter who finds the bug is often the one who discovered legacy-api.example.com — a forgotten asset that nobody patched in years — through better recon.

The same is true in professional pentesting: thorough recon means you test the whole attack surface, not just the polished front-facing app, and that’s where serious findings hide.

Recon is also continuous, not a one-time phase. Organizations constantly add new subdomains, deploy new services, ship new features. Skilled hunters monitor their targets for new attack surface appearing over time (a technique developed fully in Track A, section 5A.3).

The lesson: be patient and thorough in recon. The exciting exploitation later is built entirely on the unglamorous mapping you do here.

📓 Key Terms

Term Plain meaning
Reconnaissance (recon)Systematically gathering information about a target.
Passive reconGathering info without ever touching the target’s systems.
Active reconGathering info by directly interacting with the target.
OSINTOpen-Source Intelligence — information from public sources.
SubdomainA sub-section of a domain (api.example.com), often a separate app.
FingerprintingIdentifying the technologies and versions a target uses.
Content discoveryFinding unlinked directories, files, and endpoints.
Directory brute-forcingTesting many common paths to discover hidden content.
Attack surface mapThe assembled picture of all of a target’s testable points.

🧪 Hands-On Lab

⚖️ Passive recon on public sources is fine to practice broadly. Active recon (Tasks 4–5) is done only against your own lab VM or an explicitly authorized target.

Task 1 — Passive recon on a domain. Pick a large, well-known company. Using only passive sources, find: its registration info (whois), its DNS records (dig), and as many subdomains as you can via certificate transparency logs. You never touched their servers — this is all public data.

Task 2 — Fingerprint a tech stack. Install the Wappalyzer browser extension. Visit several websites and see what it reports — web server, framework, analytics, CMS. You’re reading the technology stack passively.

Task 3 — Practice search operators. Learn 4–5 search engine operators (site:, filetype:, intitle:, inurl:). Use them to find specific things on a large public site — e.g. all PDF files on a domain. This is “dorking,” and it’s pure passive recon.

Task 4 — Active recon on your lab. Against your Phase 0.5 victim VM, run a directory brute-forcer (e.g. ffuf or Gobuster) with a common wordlist. See which paths it discovers. This is content discovery — done legally, on a machine you own.

Task 5 — Build a target map. For your lab victim VM, assemble a single Notion page: its IP, open ports/services, identified technologies and versions, and discovered directories/endpoints. This is your first attack surface map — the deliverable recon always produces.

⚠️ Common Mistakes

✅ Recap & What’s Next

Next (2.2): With a target mapped, you need one environment to intercept, inspect, and manipulate every request you send. That tool is Burp Suite — the web hacker’s daily workbench.

⁂ Back to all modules