Open Source Intelligence is the practice of producing intelligence from publicly available information. Two words in that sentence do most of the work. Public means legally accessible without circumventing a control — not merely "obtainable". Intelligence means processed and analysed, not collected. A folder of scraped data is not intelligence, however large it is.
In offensive security, OSINT is the phase that precedes any interaction with a target. It exists because an organisation's externally visible footprint is almost always larger than the organisation believes, and because that gap between what they think they expose and what they actually expose is where engagements succeed.
Information is not intelligence
The distinction matters more than it sounds:
| Term | Meaning |
|---|---|
| Data | Raw, unprocessed values — a list of resolved IP addresses |
| Information | Data with context — those IPs, mapped to hostnames and open ports |
| Intelligence | Information that has been analysed to answer a question — "this staging host is the weakest externally reachable entry point, and here is why" |
Most OSINT work that disappoints stops at the second row. The tools produce information efficiently; only the analyst produces intelligence.
The intelligence cycle
Borrowed from formal intelligence practice, and worth following literally because it enforces the two steps that get skipped:
| Stage | What happens | Failure mode if skipped |
|---|---|---|
| Direction | Define the question the collection must answer | Aimless collection; volume mistaken for value |
| Collection | Gather from selected sources | — |
| Processing | Normalise, deduplicate, resolve into a queryable form | Findings buried in inconsistent raw output |
| Analysis | Correlate across sources, discard noise, draw conclusions | A list of facts with no meaning attached |
| Dissemination | Report in a form the reader can act on | Correct work that changes nothing |
| Feedback | New questions from analysis drive the next round | Single-pass collection that misses the obvious follow-up |
The passive/active boundary
This is the single most important operational concept in OSINT, and the one most often blurred by tooling.
| Category | Definition | Examples |
|---|---|---|
| Passive | No packet reaches the target's infrastructure; you query third-party datasets | Certificate transparency logs, Shodan, Censys, web archives, search engines, WHOIS, breach indexes |
| Semi-passive | Traffic that looks like normal use and reaches the target indirectly | Recursive DNS resolution, loading the public website in a browser |
| Active | Direct, non-normal interaction with the target | Port scanning, DNS brute forcing, directory fuzzing, SMTP address verification |
The reason to care is that these carry different authorisation requirements and
different detection profiles. A single flag moves a tool across the line —
amass enum -passive and amass enum -active differ by
one word and by whether the client's DNS logs record you. Knowing which mode you
are in at every moment is a basic professional competence, not a detail.
A useful test: if the target switched off every one of their systems right now, would this command still return results? If yes, it is passive.
Source categories
| Category | Typical sources |
|---|---|
| Infrastructure | DNS, WHOIS, certificate transparency, BGP/ASN registries, internet-wide scan indexes |
| Content | Search engines, web archives, public code repositories, document metadata |
| Human | Professional networks, company sites, conference material, job listings, press |
| Corporate | Company registries, filings, trademark databases, acquisition announcements |
| Exposure | Breach indexes, paste sites, leaked credential notifications |
| Geospatial | Satellite and street imagery, mapping data, geotagged media |
Source reliability
Not all public information is equally trustworthy, and treating it as though it were is how a report gets challenged. The NATO admiralty grading system is a reasonable mental model: rate the source and the content separately.
| Question | Why it matters |
|---|---|
| Is this primary or derived? | A certificate transparency log is primary. An aggregator repeating it is derived and may be stale |
| When was it collected? | All OSINT is a snapshot. A Shodan record may be weeks old; an archived page may be years old |
| Is it self-reported? | A location on a social profile is a claim, not a fact. So is document metadata, which is trivially forged |
| Is it corroborated? | One source is a lead. Two independent sources is a finding |
| Could it be deliberate? | Mature organisations seed canary records and honeypot hosts specifically to detect reconnaissance |
Legal and ethical constraints
"Publicly accessible" and "lawful to collect" are different tests, and the second is the one that matters.
- Authorisation is not optional. Even passive research against a named organisation should sit inside a signed engagement. Passive activity that later informs an intrusion is part of that intrusion.
- Personal data has a legal regime. GDPR, India's DPDP Act, and equivalents require a lawful basis for processing. Public availability is not a lawful basis. Necessity for a legitimate, documented purpose can be.
- Minimise deliberately. Collect only what answers a stated requirement. Special categories — health, religion, politics, sexuality, trade union membership — should not be collected at all in a security assessment.
- Accessing is not the same as finding. Discovering an exposed database is a finding. Reading its contents is unauthorised access, whatever the file permissions say.
- Stolen data stays stolen. Breach corpora containing plaintext credentials are the proceeds of a crime. Counting exposure is legitimate; sourcing the passwords is not.
- Retention has an end date. Your working data describes exactly how to attack your client. Encrypt it, and destroy it on the schedule written into the engagement.
A practical heuristic: if you would be uncomfortable justifying a specific lookup to the person it concerns, with the sentence "this was necessary to assess my client's risk", then it was not necessary.
Why organisations lose this fight
OSINT works because of structural asymmetries that are hard for a defender to close:
- Certificate transparency is mandatory. Every publicly trusted certificate is published by design. Internal-sounding hostnames leak as a consequence of doing TLS correctly.
- Archives are permanent. Removing a page removes it from the live site, not from the record.
- Asset inventories decay. Cloud resources are created faster than they are catalogued, and nobody decommissions a staging environment on schedule.
- Metadata is invisible to its authors. Nobody sees the author field before publishing a PDF.
- Employees have a professional incentive to be visible. A detailed LinkedIn profile is good for a career and good for an attacker.
- Reconnaissance is largely undetectable. Reading a third-party dataset generates no log entry anywhere in the target's estate.
What good looks like
A competent OSINT phase is characterised less by the tools used than by four properties: a stated question driving every collection step, multiple independent sources behind every material finding, a source and date attached to every claim, and an analysis section that says what the findings mean rather than restating them.
For the operational version of all of this — the phase-by-phase workflow, the actual commands, and how the outputs chain together — see OSINT From a Single Domain: A Red Team Methodology. Tool-level reference for every tool involved is in theToolkit Vulns & Misconfigs.