Mean Time to Remediate Improved Because the Counting Changed
A falling remediation time reads as a program getting faster, and some of it is exactly that. The rest belongs to a population redefined over the year by people fixing something genuinely wrong, each change small enough to land in a configuration screen. The slope on the chart is part delivery and part definition, and the split between them has to be reconstructed.
The number improved and the definition moved
You report mean time to remediate every quarter and the line has been falling. Some of that is real. Gates went in, a backlog got worked, two teams got faster at the finding they see most. The question that arrives eventually is how much of the slope belongs to the work.
The measure averages over a set of findings, and two things fix its value: which findings are in the set, and where the clock starts and stops on each. Both are adjustable, and both get adjusted by people correcting something genuinely wrong at the time.
No one edited the number. The definition moved underneath it. Each of those adjustments would pass a review on its own, and their combined effect stays uncomputed, because computing it means knowing what changed and when. The changes are dated: scanning platforms and repository platforms record configuration changes with a timestamp and an actor. What has to be built is the join between those logs and the metric, so a deduplication rule changed in March and a step in the trend line in April can be read as one event. The five below are ordinary, and each arrived with a good reason attached.
- A deduplication rule collapsed one library flaw into a single entry, because an entry per affected repository made the queue unreadable
- The reported population narrowed to the tiers the program had committed to, because reporting against the rest misled in the other direction
- An accepted-risk status was created, because findings the organization had decided to live with sat in the open queue distorting everything computed from it
- A detection was tuned to the estate, because the team reviewing its output judged that review time better spent elsewhere
- The clock moved from detection to triage, because days waiting on a ticket nobody had created yet were charged to the engineers who fix things
Where the clock starts and where it stops
The start event is a choice among defensible ones: when a scan first reported the finding, when a ticket was created, when somebody triaged it, or the commit that introduced the defect. The intervals between them are real elapsed time, and moving the start forward removes that time from every finding at once.
What happens to history then decides whether anyone sees it. Recompute the back series under the new start event and the whole line shifts down together, which looks like nothing happened. Leave history alone and the chart gets a step at the change date, which reads as a quarter of exceptional delivery.
The stop event carries more variety. A finding leaves the open population for several reasons and the record shows most of them as closed: fixed in code, accepted by somebody with the authority to accept it, duplicate, out of scope, and no longer detected. That last one covers a code change, a rule change, and an archived repository, three different facts about your estate.
- Name the start event, and state whether the back series was recomputed the last time it changed
- Break closures out by reason, so fixed in code carries its own count separate from accepted and from no longer detected
- Say what no longer detected means here, including archived repositories and retired rules
- State whether acceptance stops the clock or removes the finding from the population, since the two move the quarter in opposite directions
- State how a reopened finding is counted, once, where the other definitions live
The population under the average
The largest effect comes from the definition itself. Mean time to remediate is computed over the findings that closed in the period, so closure is the condition for entering the population at all. The aged, difficult, cross-team items are the ones still open, and they sit outside the measure of how long remediation takes.
Check your own data for what follows. A team closing new findings quickly while the aged population sits reports an improving average through a period when that aged population grows. Both movements are real and they describe opposite things, so the average of what closed and the age of what is open have to be read together.
Composition moves the number as well. One dependency upgrade can close several hundred findings in a day, and a quarter containing a batch closure is largely a description of that batch. Deduplication changes the count and the start date together, since the surviving record inherits one instance's detection date. Narrowing to the committed tiers removes the part of the estate that closes slowest. And on a distribution this skewed, a mean, a median and a percentile give materially different answers.
Settle the counting rules before publishing an average. That is a CWS position, and the arithmetic behind it is plain: one set of closure records supports several defensible averages, depending on which repositories are in the set and what closing a finding was taken to mean. Scale widens the spread between them, and the estate behind CWS's single application security discovery ran to roughly 6,300 repositories. Six things travel with the number.
- Findings still open on the last day of the period, and their age distribution
- Findings closed in the period, split by closure reason
- The largest single batch closure and how many findings it carried
- The repository population the average covered, with its size and the criterion that defined it
- Any deduplication rule that changed during the period, and what it did to detection dates
- The statistic itself, named, and unchanged from last quarter unless the change is stated
An average over the findings that closed cannot see the ones that did not.
Why the drift runs one way
This is a considered view, and CWS is the one holding it. Every change described so far could move the number in either direction. The reason the drift runs one way sits in how changes get reviewed.
A change that makes the number worse produces a question at the next review. Somebody explains the movement, the explanation surfaces the change, and the change gets revisited. A change that makes the number better passes without a question, because it reads as the program working, which is what everyone in the chain expected. Flattering adjustments survive at a higher rate than unflattering ones, and every individual decision along the way was taken on its merits.
The same filter picks which measures get reported at all. Where a program tracks four remediation measures, the one that reaches the board pack is the one whose line has been going the right way. The argument applies to any figure published without its definition attached.
A change that makes the number worse gets a second look. A change that makes it better gets filed.
What to publish beside the number
The NIST Secure Software Development Framework asks organizations to plan and implement responses to the vulnerabilities they identify. It leaves the timing and the measurement open, correctly, since the measure has to fit the estate. The definition is yours to write, and that makes it yours to publish.
The most useful addition is a cohort view, and it is a CWS recommendation. Take the findings that arrived in one quarter and report what share closed within thirty days, sixty, and ninety, tracking the cohort as it ages. The denominator is fixed on the day the cohort closes to new arrivals, which is what makes the curve readable a year later.
One rule has to be settled in advance, because it is where the method leaks. A finding inside a cohort can later be ruled a false positive, or merged into another record by deduplication, and each of those either shrinks the cohort or leaves noise inside it. State which happens. CWS argues for removing a confirmed false positive from the cohort and restating the published figures for the quarter it belonged to, and for keeping a deduplicated finding in the cohort under the arrival date of the record that survived. Either rule holds, so long as it is written down before it is needed.
The rest is bookkeeping, and it settles most of the questions the number attracts.
- The population as a count: how many findings closed, how many remain open, across how many repositories, over what dates
- The clock definition in one short paragraph, worded the same way every quarter so a change to it is visible
- A dated log of definition changes, marked on the trend line, with a note saying whether history was recomputed
- Closure reasons broken out, each carrying its own count
- The age distribution of the open population, which moves the other way when the closed-set average is flattered by what it excludes
- The rule for late invalidations, stating what happens to a published cohort when a finding inside it is ruled a false positive or merged into another record
- The cohort curve for the last four quarters, kept alongside the average
Introducing the correction
The new reporting lands in a room where the old line has been shown for a year, so the framing carries more weight than the arithmetic. The old number was computed correctly under the definition in force at the time. What changes is that it now travels with its population and its definition.
Bring both series on one chart with the definition changes marked. A board that can see where the definition moved stops asking whether it moved, and the conversation goes to the estate. Expect the cohort curve to look worse than the old line the first time it appears, because it counts what the old line excluded. Say so in that meeting.
One position worth holding: put changes to the measure's definition through the same approval as a change to the risk register, with a named approver and a date. A definition can be changed in a configuration screen during maintenance work, and that is where the drift enters. Making the change a governed event costs one signature.
- The old series and the new series on one chart, with definition changes marked at the quarter they landed
- A one page statement of the current definition, dated, with the approver named
- The cohort curve, introduced as a methodology addition and explained before anyone asks
- The age of the open population, since that is the number the room will use to check the other one
Publish the definition change in the quarter it happens and it is a methodology note. Publish it a year later and it is a finding somebody else made.