You Will Not Fix All of It. Choose What You Fix First.
Findings arrive each week at a rate set by the size of the estate and the pace of change inside it, and they close at a rate set by the engineers who can be spared. What you are choosing when you triage is an order and a residual: the work that gets done this quarter, and the work that stays open while you do it. Severity ranks the finding, and something else has to rank the estate.
The backlog is a standing condition
Look at two lines across the past four quarters: how many findings arrive each week, and how many close. Where that gap is wide, adding engineers raises the closure line and leaves the arrival line where it was, because arrival is set by the estate and the pace of change inside it.
That changes what triage is for. A queue you expect to drain gets worked from the top, and the ordering is a scheduling convenience. A queue that will still be there in three years is a standing allocation of engineering attention, and everything below the line is a position the organization holds.
CWS has run one application security estate discovery, and it covered roughly 6,300 repositories. The arithmetic that settles this question is cheap at any size: put four quarters of arrivals beside four quarters of closures and see which line is longer. Where the gap is wide, four symptoms follow from it.
- The count at the top of the dashboard has moved one direction for four quarters and the commentary explaining it has stayed the same
- Remediation timelines written into policy are missed routinely, and the misses are recorded in your own tooling
- Engineers close what is quick and the remainder ages, which is a triage decision taken by whoever holds the ticket
- The aged report describes precisely which classes of finding the organization has decided to live with, and it is the only place that decision is written down
What the order is built from
A severity score is computed from the properties of the vulnerability: the class of defect, the access required, what an attacker gains. The base score returns the same answer for every organization carrying that defect, which is what makes it portable. Separating two instances of one defect is what the environmental metrics are for, and supplying them is work that sits with whoever knows the estate.
The public exploitation signals sit in the same position. The Exploit Prediction Scoring System estimates the probability of exploitation activity in the near term. The Known Exploited Vulnerabilities catalog published by CISA records what is already under attack. Both describe the vulnerability. Neither knows whether the vulnerable code path executes in anything you deploy.
The NIST Secure Software Development Framework carries a practice for assessing, prioritizing and remediating vulnerabilities, and leaves the basis to you. Repository tiering settles which systems matter; two findings inside one top-tier repository still need an order between them. Four inputs produce it, and three come from you.
- The rating and the public exploitation signals, which describe the vulnerability and arrive alongside it
- Reachability in the built artifact: whether the vulnerable function sits on a call path that executes in what you ship
- Exposure in your estate: whether untrusted input reaches that path, and what the repository tier already tells you about the system around it
- Cost of closure, the size and the shape of the change that removes the finding
Two instances of one defect score identically and mean different things. The difference is held by your organization.
Cost of closure decides how much you reach
Ordering by cost of closure, separating introduced findings from inherited ones, and keeping a residual register with named re-triage triggers are CWS positions. EPSS and the KEV catalog supply inputs to that order. The order itself is an argument, and it is ours.
Cost of closure is the fourth input, and it governs how much of the backlog you reach. Order the queue on risk alone and the top item is a framework migration while the next forty are a version change somebody could make this afternoon.
Group the queue by the change that closes findings. One dependency upgrade closes every finding that traces to that library, in every repository holding it. A misconfiguration repeated through every service built from a shared template closes once, in the template. Worked as tickets, one engineering change becomes hundreds of decisions.
The far end of the distribution has to leave the queue. A finding whose fix is a framework migration or a change of design is a project, and it needs a design, an estimate and a delivery slot. Left in triage it ages and lands in the aged report as an unremediated critical, which describes the finding correctly and describes your organization badly. Three classes, each owned by a different part of the organization.
- Findings closed by a version or configuration change, worked in batches defined by the fix
- Findings closed by a code change inside one team's boundary, which go to that team with a date
- Findings whose fix crosses a boundary or changes a design, scoped as projects and removed from the queue with a record of where they went
A single version change can close more of the backlog than a month of individual triage.
Stop the arrival before working the stock
A backlog that grows while you work it defeats every measure of progress you report. Split the population on the day you start. Findings introduced after that date are one program. The ones that existed before are another, with their own owner, funding argument and end state.
Remediation policy deserves the same restraint. A clause committing the organization to close every critical finding within fifteen days, written against an arrival rate above what the team closes, produces a permanent record of missed commitments in your own tooling, and an auditor reads that record before your strategy. Narrow the commitment to the scope you can meet and say what falls outside it. Three moves hold the split in place.
- Gate newly introduced findings at the point of change, in the top tier first, narrow enough that engineering accepts it and the security team enforces it every time
- Report introduced and inherited findings as two numbers, so the effect of the gate becomes visible early
- Give the inherited backlog an owner, a budget and an end state, because those three are what put work on a plan
The residual, written down
Everything below the line is a decision the organization has taken. Recorded as an absence, it stays outside the approval that every other decision of that size goes through. Write it as a position: these classes of finding, in these tiers, sit outside this year's remediation scope for this reason, accepted by this person on this date.
The register earns its keep by being specific. Compare two lines. Low severity findings are deferred: a reader can agree with that and come away exactly as informed as before. Dependencies with no reachable call path, in repositories below the top tier, reviewed quarterly and left unscheduled: a reader can agree with that, refuse it, or ask for the count.
Then name the events that pull something back out, because the reason a finding was set aside expires quietly. A dependency with no reachable path acquires one in a release. A vulnerability with no observed exploitation turns up in the CISA catalog. A repository moves up a tier. Four lines make the register usable.
- Name the deferred classes precisely enough that an engineer can tell which of their own findings falls into one
- Record who accepted the residual, by name, with the date
- Put a size on it, because a count of deferred findings by class turns an absence into something a risk committee can vote on
- List the events that trigger a re-triage, and name where each one is observed
The room after the incident
One day something you deprioritized is what gets used. Prepare for that room now, because the conversation you get is fixed by what exists in writing before you walk into it.
Two questions arrive before anyone asks about the finding itself. Was there a basis for the order, written before the decisions it explains. Was it applied consistently, or does the record show comparable findings treated differently, with the reason for the difference reconstructed on the day. An organization that answers both made an allocation under a constraint it had declared.
The third question is the one worth building for: did anything change about that finding between the day it was deferred and the day it was used, and would your process have seen it. A finding that entered the CISA catalog four months before the incident is hard to sit with, and the difference between a hard conversation and a bad one is whether a trigger existed to catch it.
Resist the temptation to argue that a better order would have caught it. A triage order allocates finite attention across a population that exceeds it, and a sound allocation still leaves things open. The defensible claim holds: the order was written before it was applied, the residual was named and sized, somebody with the authority accepted it on a date, and the events that should have moved this finding were being watched. Every part of that answer sits in a record somebody had to have kept.
- Keep the version of the triage rule that produced each queue, stamped with the dates it was in force
- Keep earlier versions of the residual register, because the room asks what was decided at the time
- Log every re-triage event that fired. A gap in that log is the honest answer to what the process missed
A deferred finding with a name and a date against it is a decision. The same finding with neither is an oversight waiting to be found by somebody else.