Skip to content
CWS
CorovaAboutContact
Book a Call
All articles
Service Delivery

The Labels Were Designed in a Workshop. The Data Was Not Consulted.

The taxonomy was signed off before anything scanned. It maps onto the obligations and it reads well in review. Then the first simulation run puts most of the volume in the tier the working group spent the least time on.

CWSAugust 14, 20266 min read

The taxonomy was a hypothesis about the estate

A sensitivity taxonomy gets built the way a policy gets built. A working group convenes with legal, records management, a privacy lead and a representative from each large business unit. It works from the obligations, an existing handling standard, and the tier structure two people in the room remember from a previous employer. What comes out has named tiers, sub-labels and a mapping to each obligation, and it reviews well.

Every input to that room came from the company's account of itself: the org chart, the systems people remember buying, and the classes of record the obligations name. That is a fair account of the company's intentions. Whether it describes the files in the tenant is an open question until something scans.

So the scheme that arrives at implementation is a hypothesis, and a well-made one. It arrives with sign-off attached, which makes it feel settled, and the next line in the plan is configuration where a test would do more.

Three inputs shaped the scheme, and the estate was not one of them.

  • The obligations, which describe classes of record and what has to be true about them, worded to hold across every company they apply to
  • An information handling standard written for a smaller estate, before the current cloud tenants existed
  • The org chart, which names the owner of a business process and says nothing about where that process leaves its documents

What the first run returns

Run the draft as auto-labeling in simulation, with a default label configured for the residual tier, and a distribution comes back. Two mechanisms produce it. Documents whose content matches a tier's conditions get that tier. Everything the conditions leave untouched arrives at the default, which is how a residual category with no conditions of its own ends up holding volume at all. The result is the estate's own account of itself.

Microsoft Purview and Securiti each classify according to the taxonomy they are given, and apply it consistently across the sources they are pointed at. The distribution is a faithful report of what the taxonomy says about the connected data. When it looks wrong, what is being reported is the distance between the design and the estate, and that distance is the useful part.

The disagreement shows up in three shapes, and finding one is a reason to go looking for the other two.

  • The residual tier holds most of the volume, because the default sends it everything the conditions left untouched, and the group spent the least time deciding what those conditions should have caught.
  • One or two tiers come back close to empty, including tiers the room spent its longest arguments wording.
  • A category the estate supplies on its own holds a serious share of the risk. Meeting recordings and their transcripts, or a quarterly export somebody built once and left in a shared folder half the department can reach.
The distribution is the first description of the estate that was written by something other than a person.

Read the disagreement before you redesign anything

The reflex when a tier comes back empty is to delete it, and when a tier comes back full, to split it. Both moves are premature, because each result has more than one cause and the causes lead to different work.

A sparse tier has three explanations worth separating. The data may sit outside the estate entirely, which is a finding in its own right and worth recording with a date. It may sit somewhere the run was pointed away from: a source held back from the agreed scope, an on-premises share, a departmental application that missed the kickoff inventory. Or it is in the connected sources and misses the detection the tier was configured with, because this company writes that identifier in a house format, or keeps it in free text where the pattern expects a structured field.

Sampling separates them and costs an afternoon. Ask the business unit that owns the process for a handful of documents they know contain that class of data, then find those documents in the results. If they were scanned and landed elsewhere, the detection needs work. If they were outside the scanned scope, the scope does. If the team cannot produce them at all, that is informative too.

Three questions get answered before anyone edits the scheme.

  • Was the source in scope, and does that scope match where this business keeps those documents
  • Does the detection match how this company writes the identifier, or how a general pattern assumes it is written
  • Can the owning team produce examples on request, and where did those examples land

Which side gives way

Someone then decides whether the taxonomy changes or the estate does. The answer differs by which part of the scheme is under pressure, and the rule is easier to settle before the meeting than during it.

A tier that exists because an obligation requires the company to know where a class of record lives, and to be able to act on it, stays. A sparse first run changes the detection, the scope, or the sampling. The obligation is unaffected by all three, and deleting the tier removes the only evidence that anyone went looking.

A tier holding most of the estate has stopped doing work. Every control attached to it now applies to almost everything, which makes those controls either the universal baseline or unenforceable. Two honest ways out: split the tier on an attribute a machine can see, or keep it and accept that its controls are the ones you would apply everywhere. Splitting on an attribute only a document author can judge puts you back at the start.

Part of the argument sits under the tiers, in the mechanics, and it is cheaper to have before publication than after a partner calls. Encryption is the control that gives a Confidential tier its meaning, and it is the same control that stops an external reviewer opening the file and stops two people co-authoring it in the browser. Label priority decides which tier wins when two sets of conditions match the same document. A label applied to a site or a team sets a default for whatever lands in it, which is a separate decision from the label on the document itself. Each of those changes what happens to a real file on a real day, so each of them gets a named decision in the same session, written down.

The category the estate revealed is where the scheme gives way outright. Real risk is sitting in a shape the room had no reason to consider, and the labels in front of it describe something else. Adding it is smaller work than it looks: it attaches to a tier that already exists, and what it needs is a detection and a named owner.

Renaming a tier is a documentation change. Redrawing a boundary changes which files get encrypted and which get held back from external sharing, and it lands on a business process. Price the two differently in the room.

Four rules decide which side gives way.

  • A tier that exists to meet an obligation stays, whatever the first run returned
  • A tier holding most of the estate splits on something visible in the data, or its controls become the baseline for everything
  • A tier that turned out to be a wording exercise folds into its neighbor, and the record says why
  • A category the estate revealed gets added, with its owner named in the same meeting
A name can be changed in an afternoon. A boundary carries a control, and the control lands on somebody's Tuesday.

Book the reconciliation before the workshop starts

The estate is going to disagree with the design. The version of this process that works puts that disagreement on the calendar in advance.

This reconciliation ran on all three of the Microsoft Purview and Securiti data security posture engagements CWS has delivered. The taxonomies were competently built and both platforms applied them as written. Whether a taxonomy survives contact with its own distribution is decided by whether anybody in your organization planned for the day that distribution came back.

The change is procedural and small. The workshop output carries a version number and one line in its terms of reference: this draft is provisional until the discovery results have been reviewed against it. A group that agrees to that sentence in week one can reopen its own work without anyone losing standing. A group that ratifies a final scheme on the last day spends the reconciliation defending it, and defending a scheme is slower than amending one.

Taxonomy sequencing, sampling before ratification, and booking the reconciliation in advance are CWS recommendations.

The sequence that gets there runs in a fixed order.

  • Draft with the working group, and mark the document provisional on its cover
  • Book the review session in the same week, before anyone knows what it will be about
  • Run auto-labeling in simulation against the draft, with the default label configured, across a scope wide enough to be representative
  • Sample the sparse tiers with the business units that own the underlying process
  • Reconvene with the distribution on screen, settle each disagreement in that session, and record which way it went
  • Ratify the scheme, give it a version number, and attach enforcement after that
Put the reconciliation in the plan before the workshop starts, and the workshop stops being a place where a hypothesis gets ratified as a fact.

Sources

  • CWS delivery corpus, three data security posture engagements delivered on Microsoft Purview and Securiti
  • Taxonomy sequencing, sampling, and booking the reconciliation before the workshop: CWS recommendations