Guide

How to assess a website's external attack surface

Written so you can do the first half by hand, with a terminal and a browser, and understand what an automated run is telling you before you decide whether to pay for one. No product required to follow it.

Somebody who wants into a website does not need to break anything. They need to know what is there. Every name that resolves, every host that answers, every file the server is willing to hand over — that inventory is the reconnaissance phase, and everything after it is a series of attempts against what the inventory said.

This is a method for building that inventory yourself. It works entirely from outside, needs no credentials, and can be done by hand with a terminal and a browser. That matters: understanding what a scanner does is much easier when you have done the first half by hand, and the manual version is what tells you whether an automated one is telling you the truth.

Work in this order. Each layer is cheaper to check than the one below it and more likely to find something.

Start with the cheapest layer: DNS

Before any of this touches your website, ask DNS. DNS answers from public nameservers, is cached by everybody, and tells an attacker more about your infrastructure than anything you publish deliberately.

Four record types to look at first:

A and AAAA
Where the site actually is. Behind a shared host, a CDN or a cloud load balancer, this is a provider rather than a server — which is useful to know, because it changes what is reachable on the other side.
MX
Where your mail goes. A hostname here is a hostname that can be a takeover target, and mail providers are frequently forgotten until something bounces.
TXT
The messiest and most informative record type. SPF, DKIM, DMARC, MTA-STS, TLS-RPT and Google verification all live here, and so does the accumulated evidence of which third-party services have at some point needed to verify your domain.
CNAME
A subdomain pointing somewhere else — a SaaS provider, a CDN, a customer-specific hostname. This is the shape of a subdomain takeover, and it is worth walking deliberately: any CNAME pointing at a hosting provider is only safe for as long as somebody is actively claiming it.

SPF, DKIM and DMARC have their own guide and are worth reading separately because they have their own failure modes:

SPF, DKIM and DMARC: a practical checklist

Beyond the apex domain, the question worth asking is which subdomains exist. Certificate transparency logs are public and free and searchable, and they record every name a certificate has ever covered — which is frequently a longer list than the inventory anyone keeps. A subdomain found this way that nobody remembers is the single most common genuine finding in this whole process.

Transport security, read from outside

Complete a TLS handshake against the hostname and read what you get. No request is sent to the application at this stage, so this step cannot affect your site in any way.

Look at, in order:

  • Validity. Expired, or close to it. This is the finding that generates the most support tickets and the least security.
  • Hostname match. Test every hostname the site answers on, including the apex, www, and anything else you believe is canonical. A certificate valid for one and served on the other fails on whichever the visitor used.
  • Chain. Whether the intermediates are actually served. A machine that has cached them will succeed where a new visitor fails.
  • Protocols and ciphers. Whether deprecated protocol versions and weak cipher suites are still accepted. Read the TLS checker page for what each finding means.
  • Plain HTTP. Does http:// redirect to https://? If not, the first request of every visit is unprotected.

Two things to write down while you are here, because they are invisible later: whether a CDN or proxy terminates TLS in front of you, and which certificate names are published for your domain in transparency logs. The first means your answer may not describe every visitor. The second is your attack surface, published, permanently.

HTTP response headers and cookies

Fetch a page and read the response headers. This is the cheapest high-value check on the whole list, because the fixes are configuration lines and the evidence is a single request.

The headers that carry weight: Strict-Transport-Security, Content-Security-Policy, X-Content-Type-Options, framing control, Referrer-Policy, Permissions-Policy, the cross-origin policies, and cache control on anything per-user.

Then the cookies, which are headers too and the ones most often wrong. For every Set-Cookie, three attributes:

  • Secure — the cookie will not be sent over plain HTTP. A session cookie without it is a session cookie an on-path attacker can read.
  • HttpOnly — the cookie is invisible to JavaScript. Without it, any injected script can read the session and send it anywhere, which turns a single cross-site scripting flaw into account takeover.
  • SameSite — the cookie is not sent on cross-site requests. The default has varied across browsers and versions, so set it explicitly rather than relying on it.

Most of these have a guide of their own covering what each prevents and the configuration mistakes that make a strong-looking header do nothing.

Subdomains, files and reachable services

This is where an external view starts to cost something, because it involves making more requests. It is also where the genuinely serious findings usually are.

Files. Ask for the things that should never be served and see what comes back. The list is short and worth walking: .git/ and .svn/, .env, wp-config.php.bak, backup.sql, phpinfo.php, composer.lock, package-lock.json, source maps, and the various .bak and .old and ~ variants. A reachable configuration backup containing database credentials and an authentication salt is a total compromise, and it is found by asking for a handful of obvious filenames.

Endpoints. What the site exposes to a caller: API documentation, a GraphQL endpoint, an admin panel, a sitemap listing internal URLs, a robots.txt that names directories you did not know existed. Discovery here means crawling within a depth and page budget, which is what a verified-ownership assessment does — and why the deeper checks are gated on proving you own the domain.

WordPress and other CMSs have their own external posture and their own page: what a WordPress scan can and cannot see. Read that one before buying anything for a CMS site, because the boundary there is sharper than it looks.

What an external assessment cannot see

This is the part that decides whether any of the above is worth doing, so it is worth being blunt about. From outside, unauthenticated, you cannot see:

  • Anything behind a login. Not the admin, not member areas, not authenticated API endpoints. A clean external view says nothing about any of it.
  • Application logic. Whether a request is authorised for the identity making it, whether input is handled correctly, whether a workflow can be skipped. These are the flaws that matter most and the ones no external scan will find.
  • Your dependencies. Which library versions you ship, whether one has a known vulnerability, whether it is reachable. A package-lock.json being readable is a finding; interpreting it against a vulnerability database is a different product.
  • The file system. Whether a file on disk was modified. Only files the server publishes are visible.
  • What you have not discovered. An external assessment assesses the domains you give it. It cannot enumerate what else you own.

So the honest summary of this whole process: an external assessment is a real and worthwhile layer, it catches a category of problem that is genuinely common — misconfiguration and accidental exposure — and it is not a substitute for testing the application or for a human penetration test. If somebody tells you a scan proves a site is secure, that is not true of any scan, and the methodology says so in the terms the engine enforces.