Guide

Website security reporting for agencies and managed service providers

Producing the findings is the straightforward part. Getting a document that changes somebody's behaviour is the hard part — including the authorisation question for a site you maintain but do not own.

A security report that nobody acts on has failed, whatever it found. That is a harder problem than finding the issues, and it is the one this guide is about — because the reader is usually a consultant or an engineer whose deliverable is a document that goes to somebody else's inbox.

It also has a constraint most security writing ignores: the person reading the report is not you, does not share your vocabulary, and has a competing project. Every design decision below follows from that.

What the report has to contain

Six things, in this order. The order is the argument: a reader who stops after the first section should still have got the important part.

  1. What was assessed and when. The exact hostnames, the date, and — this is the one that is routinely missing — the level of access the assessment ran at. A report that does not say whether it was authenticated, and what that meant, cannot be read in context. If a fundamental part of the surface was not assessed, that belongs on page one, not in an appendix.
  2. A short list of what to do first. Three to five items, ordered by what an attacker would actually do with them. Not the finding count, and not a table of everything. This section is what a reader forwards to a developer.
  3. Evidence, per finding. The observation that produced the finding, so the reader can verify it independently. This is the single most important property of the whole document — see the next section.
  4. Each finding, with its specific fix. The change to make, not the name of the control that is missing. "Add HttpOnly to the session cookie, here is the framework setting that does it" is actionable. "Implement secure cookie attributes" is a line in a report nobody keeps.
  5. What was not covered. Everything outside the method, stated plainly. Not as a disclaimer but as a scope section, because the reader's next question will be "does this mean we are fine?" and the honest answer is a list of what is not known.
  6. The method. A page or an appendix. What was sent, what was not sent, the bounds. This is what lets somebody re-run the assessment and compare.

What does not need to be in it: a score in isolation, a severity-coloured chart of findings by category, a letter grade. A number with no context is a number the reader will argue with rather than act on, and an argument about the score is time taken away from fixing the thing it describes.

Evidence, not assertions

A finding is a claim. Evidence is what turns it into something a reader can check, and checking is what makes them trust the rest of it — including the findings that are more convenient.

Concrete examples of what "evidence" means at each level:

Findings stated as assertions versus the same findings with evidence
Without evidenceWith evidence
Weak TLS configuration. TLS 1.0 and 1.1 accepted on port 443; the strongest suite offered is TLS_RSA_WITH_AES_128_CBC_SHA.
Session cookies are not secure. Set-Cookie: sid=…; Path=/; HttpOnly returned by /login — no Secure, no SameSite.
DMARC is not enforced. _dmarc.example.com TXT "v=DMARC1; p=none; rua=mailto:[email protected]".
A backup file is exposed. GET /wp-config.php.bak returns 200, application/octet-stream, 8 KB, beginning <?php … DB_NAME.

The difference is not verbosity. It is that the second column lets the reader confirm the finding in thirty seconds without asking you, and lets them confirm the fix the same way afterwards.

Three rules for calibrating severity honestly, which is what keeps a report credible across a portfolio of client sites:

  • Separate what was confirmed from what was inferred. A response header you read is confirmed. A header pattern consistent with a misconfiguration is inferred. Grouping them into one ranked list is how a report ends up with eleven items at the top, none of which is the one that mattered.
  • Grade against consequence, not against effort to fix. An expired TLS certificate is an outage waiting for a renewal process to fail; it is urgent for a different reason than a medium-severity header gap, and the ordering should say so.
  • Record the negative findings. "TLS is correctly configured" is a result. A report that only lists problems gives the reader no way to know what was checked, and makes every re-run look like a fresh exercise rather than a comparison.

Assessing a site you do not own

This is the part with legal exposure, and it is settled before the tool is opened rather than during the engagement.

You need written authorisation from the domain's owner, and the authorisation should name the domains, the date range, and permission for the requests the method makes. Keeping it means keeping it: in the engagement record, not in an inbox.

The practical reason an agency should be comfortable doing this is that the method is narrow. Requests are read-only, no credentials are used, no state-changing request is sent, and the scanner refuses private and reserved addresses so it cannot reach anything not publicly exposed. Being able to send a client the methodology — rather than a promise that nothing bad will happen — is what gets the signature. The methodology page is written to be sent to a client.

Three practical points:

  • Warn the client's provider. Assessment traffic looks like ordinary visitor traffic and can trip rate limiting, a WAF rule or intrusion detection. A scan that takes a client's site down is a relationship problem, and it is entirely avoidable by sending one email.
  • Use the domain verification step. It unlocks the full check surface, and it leaves a durable artefact in the client's own DNS that evidences the client authorised the work. Put it in the handover documentation and it becomes part of the estate record.
  • Contract on the cadence, not the scan. A one-off report is a document. A recurring assessment with a re-run after each change is a service, and it is the version that is defensible when somebody asks what the estate's posture actually is.

One thing to be straight about with a client: this is not a penetration test and the report should say so on page one. An agency that presents an external assessment as a penetration test has set up a misapprehension that surfaces at the worst moment, and a client who later discovers their penetration test did not cover the application has a legitimate complaint.

Proving the fix worked

The step that makes an assessment worth repeating, and the one most often skipped.

Re-running after a fix produces a comparison, and the comparison is a different kind of document from the original report: it is a statement about what changed rather than a list of things that are wrong. It is also the only evidence an agency has that its remediation advice worked, which matters both commercially and when a client's auditor asks.

What makes the comparison trustworthy:

  • The same method and the same level of access. Comparing a verified-ownership report against an earlier basic one compares two different assessments. Both state their level; check that they match.
  • The same hostname set. Add a subdomain and the surface changed. Note what was assessed each time.
  • Both reports retained. The pair is the evidence. A report from eighteen months ago with no baseline is much less useful than two recent ones with a difference.

The cadence that tends to survive client turnover: after every change to anything a domain touches — DNS, the application, the host, the CDN — and on a schedule regardless. Quarterly is a defensible default for sites that do not change often. Write it into the contract as a schedule rather than leaving it as "when somebody remembers", because the value is entirely in the cadence.

A cadence that survives client turnover

The practical problem with a recurring assessment programme is not the scanning. It is that the client forgets why it is happening, the person who set it up leaves, and the second assessment is run by somebody who has never seen the first.

Four things that prevent it:

  1. One document that accumulates. A short standing document — a re-assessment summary with the history, the outstanding items and the current posture — rather than a new full report each cycle that supersedes the last. It answers "are we improving?" without anybody having to diff anything.
  2. Outstanding findings carried forward. An item that was not fixed does not disappear because the next cycle is clean; it appears again with its age. That single property prevents the most common failure, which is a known issue quietly ageing out of every report.
  3. Names and dates on the document. Who ran it, when, and against which assets. Boring, and it is what makes the report usable in an audit two years later.
  4. A scope section the client actually read. What was assessed, and what was not. Reuse the same scope on every cycle so the client recognises it. A scope that changes without explanation reads as either a downgrade or a sales attempt.

And the thing to keep saying: a clean external assessment means nothing was found from outside, on that day, by that method. It is a real result and it is worth having. It is not a certification, and a client who understands that distinction will trust the next report too.