What is continuous penetration testing?
Continuous penetration testing is manual testing delivered on a repeating rhythm instead of as a single dated engagement. That much is uncontroversial. Everything else about the term is set by whoever is selling it, so this guide separates the parts with a published definition from the parts that are pure packaging.
The activity is defined. The delivery model is not.
There are two different things hiding behind the phrase “continuous penetration testing”, and they have very different provenance.
The first is penetration testing itself, which is defined in several places by people with no product to sell. The UK National Cyber Security Centre defines it as “a method for gaining assurance in the security of an IT system by attempting to breach some or all of that system’s security, using the same tools and techniques as an adversary might”. NIST Special Publication 800-115 describes it as “security testing in which assessors mimic real-world attacks to identify methods for circumventing the security features of an application, system, or network”, and adds the characteristic that distinguishes it from scanning: “Most penetration tests involve looking for combinations of vulnerabilities on one or more systems that can be used to gain more access than could be achieved through a single vulnerability.” The NIST glossary adds a phrase worth keeping: assessors work “under specific constraints”. A penetration test is bounded, authorized and agreed in advance.
The second thing is the delivery model, usually sold as PTaaS, penetration testing as a service. No standards body, regulator or national authority defines it. CREST publishes a specification for what makes a penetration test commercially defensible, and NIST defines the activity, but neither defines the subscription. That absence is not an oversight, and CREST is candid about why the specification was needed at all: “Across the globe it is widely acknowledged that the definitions, practices and expectations associated with a penetration test are inconsistent and fluid.”
The practical consequence
Because no external definition constrains the term, two vendors can both truthfully call their offering PTaaS while one runs quarterly manual testing with retest included and the other runs a scanner behind a portal. The specification has to come from the buyer. The rest of this guide is about writing that specification.
The four phases, run on a loop
NIST SP 800-115 breaks a penetration test into four phases: planning, discovery, attack and reporting, with an additional discovery loop running back from attack into discovery whenever a successful step opens new ground. Nothing about running the test more often changes those phases. What changes is which of them get repeated, and how much of the previous cycle carries forward.
In a single annual engagement, all four phases happen inside a fixed window and then stop. Planning is a scoping call, discovery is a fresh enumeration of an estate nobody has looked at in a year, attack is compressed into the days that remain after discovery over-ran, and reporting closes the engagement. The next time anyone looks, the discovery work starts from zero again.
In a program, discovery becomes the continuous layer and mostly stops being a manual activity: asset and surface discovery runs on a schedule, so each manual cycle begins with a current picture rather than a stale one. Planning shrinks to deciding what this cycle covers, which is usually driven by what changed. Attack is where the manual effort concentrates. Reporting splits in two: a running finding register that reflects the current state, and a periodic report that can be handed to somebody who needs a document.
A fifth activity appears that has no place in the single-engagement model at all, because there is no next cycle to put it in: retest. That is the subject of the evidence guide, and it is the single most reliable way to tell a real program from a subscription.
What actually changes when testing repeats
Four things change, and it is worth being precise about them because the marketing tends to claim a fifth and a sixth that are not true.
Testing can follow change instead of the calendar
OWASP’s testing framework puts a verification step immediately after deployment: “After every change has been approved and tested in the QA environment and deployed into the production environment, it is vital that the change is checked to ensure that the level of security has not been affected by the change. This should be integrated into the change management process.” An annual test cannot do this. It happens on a date chosen for budget reasons, not for engineering reasons. A program can carry a trigger list, which is covered in the guide on testing after significant change.
Findings acquire a lifecycle instead of a status
In a one-off engagement, a finding is open or it is closed in your own tracker, and nobody external ever revisits it. In a program, a finding moves through found, triaged, fixed and retested, and only the last of those is evidence that the risk went away. DORA requires financial entities to establish “internal validation methodologies to ascertain that all identified weaknesses, deficiencies or gaps are fully addressed”. PCI DSS v4.0.1 requirement 11.4.4 puts it more bluntly: findings are corrected and “penetration testing is repeated to verify the corrections”.
Evidence accumulates instead of expiring
Commission Implementing Regulation (EU) 2024/2690 requires relevant entities to “document the type, scope, time and results of the tests, including assessment of criticality and mitigating actions for each finding”. A program produces that record on a rhythm. A single test produces it once, and eleven months later the honest answer to “when did you last verify this control” is a date from last year.
Testers accumulate context
This one is real but is rarely written into a contract, so treat it as a benefit to ask for rather than one to assume. A team that tested the same platform four times knows where the odd authorization logic lives and which service has the awkward multi-tenant model. That knowledge is only retained if the provider commits to continuity of personnel. Ask.
What the label does not guarantee
Buying more frequent testing does not buy any of the following, and a proposal that implies otherwise is worth pushing back on.
- Complete coverage. OWASP, quoting Gary McGraw, puts it plainly: “In practice, a penetration test can only identify a small representative sample of all possible security risks in a system.” Twelve samples a year is still sampling. Frequency reduces the age of your evidence; it does not close the gap between sample and population.
- Depth. Cycles are usually shorter than annual engagements, and a short cycle is the wrong shape for work that needs sustained effort, such as chaining a low-severity information leak into an account takeover across three services. Most credible programs therefore keep one longer engagement a year alongside the rolling cycles.
- Automation equivalence. NIST SP 800-137 is careful about the word: “continuous” and “ongoing” mean that risks “are assessed and analyzed at a frequency sufficient to support risk-based security decisions”, and “Data collection, no matter how frequent, is performed at discrete intervals.” Continuous is a statement about frequency, not about constancy, and it certainly does not mean “a machine does it”.
- A transfer of the risk decision. NCSC: vulnerability risk assessment and mitigation “is a business process and should not be wholly outsourced to the test team”. The tester establishes what is broken and how badly. Deciding what to do, in what order, against what appetite, remains an internal job whatever you pay.
- A replacement for the rest of the program. OWASP’s balanced approach combines manual inspection and design review, threat modeling, source code review, penetration testing and automated scanning, and notes that “the relative effort devoted to each technique should shift based on the SDLC phase”. Buying one of those techniques more often does not produce the other four.
How to read a PTaaS proposal
CREST’s Defensible Penetration Test specification names three conditions that have to be satisfied together: the provider holds “appropriate policies, procedures, practices and methodologies”; every individual involved has “appropriate levels of skills, experience and competency”; and the provider and the individuals work “towards a defined and agreed test specification”. Those three make a useful spine for reading any subscription proposal, because a subscription that fails them is not made defensible by running monthly.
| Proposal claim | What to ask | A usable answer |
|---|---|---|
| “Continuous testing” | How many manual tester-days per cycle, and how many cycles per year? | A number for both, in the contract, separate from platform access. |
| “Automated and manual” | Which findings in the sample report came from a tool and which from a person? | A report that labels the source of each finding, and a sample that contains genuine logic findings. |
| “Real-time results” | Is a finding published before or after a human has validated it? | After. Unvalidated tool output arriving in your tracker is a workload, not a service. |
| “Retesting included” | How many retests, within how many days of you marking a finding fixed, and at what point does it become chargeable? | A stated turnaround and an explicit position on volume. |
| “Continuity of testers” | Will the same people run consecutive cycles, and what happens when they do not? | A named lead, and a handover commitment rather than a pool. |
| “Compliance ready” | Which requirement, in which version of which standard, does each deliverable satisfy? | A mapping to clause numbers, not a logo grid. |
The cost that is not on the invoice
A program moves effort inside your organization as well as onto a supplier. This is the most common reason a subscription is quietly abandoned after a year, and it is worth planning for before signing.
- Someone has to triage. A rolling stream of findings needs an owner who can validate, rate and route them within days, not at quarter end. If that person does not exist, the register becomes a backlog and the program produces anxiety rather than assurance.
- Someone has to feed the trigger list. Testing after significant change only works if the testers hear about the change. That means a route from your change process to the provider, not an email when someone remembers.
- Someone has to close findings. NCSC frames penetration testing as a way of “gaining assurance in your organisation’s vulnerability assessment and management processes”. If those processes do not exist, the test has nothing to give assurance about, and a subscription simply measures the same gap more often.
- Someone has to keep the evidence. Records only help at audit time if they are complete, dated and retrievable. That is an administrative job, and it is usually yours rather than the provider’s.
Is it right for you?
The honest answer for a meaningful number of organizations is no, and it depends on two variables more than anything else: how often production changes, and whether a rule obliges you to test.
If you deploy weekly and run internet-facing services, an annual test describes a system that no longer exists by the time the report is read, and a program is the right shape. If your estate is stable, nothing is published to the internet and no regulation obliges you to test, scheduled patching plus vulnerability scanning plus a scoped test when the architecture actually changes will serve you better, and a subscription is an expense without a matching risk. If you currently have nothing scheduled at all, start with a baseline engagement rather than a subscription: you cannot design a cadence for an estate you have not enumerated.
The cadence designer on the home page walks those variables and will return a “you do not need this” verdict where that is the right answer. The next guide draws the line between this and the two things it is most often confused with: vulnerability scanning and bug bounty programs.
Sources
- Penetration testing Definition of penetration testing, the assurance framing, and ownership of risk decisions.
- SP 800-115: Technical Guide to Information Security Testing and Assessment Section 5.2: definition and the four-phase methodology.
- penetration testing (glossary) The “under specific constraints” formulation, sourced to SP 800-53 Rev. 5.
- SP 800-137: Information Security Continuous Monitoring What “continuous” and “ongoing” mean, and discrete-interval data collection.
- CREST Defensible Penetration Test The three elements required for a commercially defensible engagement.
- Web Security Testing Guide: Introduction The balanced approach, and the limits of penetration testing as a sole technique.
- Web Security Testing Guide: The Web Security Testing Framework Phase 5.3, change verification after deployment.
- Regulation (EU) 2022/2554 (DORA) Article 24(5), validation that identified gaps are fully addressed.
- Commission Implementing Regulation (EU) 2024/2690 Annex point 6.5.2(c), the record each test must leave behind.
Professional help
Have the cadence conversation before the pricing conversation
OffSeq scopes penetration testing around change rate and obligations rather than around a package. Bring your deployment frequency, your asset list and your drivers, and the shape of the program falls out of those.
Commercial link. This site is published by OffSeq.
Questions