Direct Answer

OT security metrics are trustworthy only when they define the measured population, evidence source, collection date, exclusions, owner, confidence level, and management decision. Tool counts and unbounded percentages can describe activity, but they cannot prove that a control works across a specific operational scope.

Key Takeaways

  • Measurement maturity and control maturity are different.

  • A discovered device is not automatically a governed asset.

  • Access evidence must show attribution, route, destination, expiry, and revocation.

  • Vulnerability metrics need applicability, reachability, consequence, and treatment evidence.

  • Comparisons should preserve exclusions instead of forcing one universal score.

The Benchmarking Problem: More Data Does Not Always Mean More Confidence

OT security programs are frequently compared through coverage percentages, tool counts, maturity scores, and backlog totals. Those measures can be useful, but they often combine evidence of different quality. A discovered device is not necessarily a governed asset. An enabled logging feature is not proof that data is healthy or used. A closed vulnerability ticket is not proof that the relevant pathway or consequence changed.

This research report proposes an evidence benchmark rather than a universal maturity ranking. It asks what a specific data set can prove, what it cannot prove, how current it is, which population it represents, and which management decision it supports. The approach separates measurement maturity from control maturity so organizations are not rewarded for poor visibility or penalized for better disclosure.

The benchmark is a structured secondary-research synthesis based on public guidance. It is not a statistically representative industry survey and does not claim market-wide performance levels. Its value is a decision method that organizations can apply to local evidence without converting uncertainty into false precision.

Asset evidence

Asset counts are frequently treated as inventory maturity. A stronger measure distinguishes discovered endpoints from governed assets with validated process role, owner, lifecycle state, connectivity, configuration source, and recovery dependency. A site with more unknown devices after a better discovery exercise may appear worse even though its measurement maturity has improved. Reporting should make that distinction explicit.

Asset evidence becomes decision-ready when discovery is reconciled with ownership, process role, lifecycle status, communication paths, authoritative configuration sources, and recovery dependencies. The benchmark should distinguish discovered, validated, governed, and tested populations rather than treating inventory coverage as one number.

Access evidence

Account inventories alone do not establish access control. Decision-grade evidence shows attributable identity, approved device, brokered route, permitted destination and protocol, session record, expiry condition, exception, and tested revocation. Third-party and emergency pathways should be included because they often sit outside standard identity reporting.

Benchmark Interpretation Guide: From Evidence Grade to Management Action

An OT security benchmark is useful only when readers understand what is being compared and what the result can legitimately support. Evidence quality should not be interpreted as a universal maturity score, a prediction of incident likelihood, or proof that one site is safer than another. It is a measure of how well a stated conclusion can be traced to current, scoped, attributable, and testable evidence. The management value comes from identifying which decisions are supported, which remain conditional, and which cannot yet be made with confidence.

Separate measurement maturity from control maturity

A site with stronger discovery, better exception reporting, and more rigorous testing may initially appear weaker because it exposes conditions that another site has not measured. This is not a contradiction. Measurement maturity describes the organization’s ability to see, explain, and validate its condition. Control maturity describes the performance of the controls themselves. The benchmark should report both dimensions and avoid rewarding low disclosure. A site with incomplete visibility should be marked as uncertain, not assumed to be performing well because few defects are recorded.

For each evidence domain, the report should show the population represented, the evidence period, the method used, major exclusions, and the decision attached to the result. An asset inventory built from passive discovery has different decision value from an owner-validated inventory connected to process role, lifecycle, and recovery data. A privileged-access count has different value from a record that proves named identity, approved route, destination restriction, session evidence, expiry, and tested revocation.

Compare like-for-like cohorts

Enterprise averages can conceal structural differences between sites. Benchmarking should therefore group comparable environments by process consequence, architecture pattern, supplier model, lifecycle condition, operating window, and regulatory context. A continuously operated utility process should not be ranked against a laboratory system without explaining the different maintenance and recovery constraints. Cohort comparison allows leaders to identify repeatable practices while preserving the local conditions that determine whether a control can be implemented and verified.

The comparison should also distinguish evidence age. A tested isolation path from eighteen months ago does not have the same decision value as one tested after the latest network, supplier, or identity change. The benchmark should therefore attach a freshness threshold to each evidence type and identify the events that require revalidation, such as commissioning, major firmware change, supplier transition, identity migration, network redesign, or recovery exercise.

Convert low evidence grades into specific improvement work

A low grade should produce an evidence-improvement action, not a generic recommendation to “increase maturity.” If asset evidence is weak, the next action may be owner validation for one critical process. If access evidence is weak, the action may be to reconcile routes, identities, destinations, and revocation tests for the highest-consequence supplier service. If vulnerability evidence is weak, the action may be to validate product applicability and reachability rather than purchase another scanning capability. Each action should name an owner, a bounded scope, a completion date, and the management decision that better evidence will enable.

Unknown should be treated as a legitimate benchmark state. It signals that the evidence does not support a reliable conclusion and prevents an unsupported favorable score. Leaders should prioritize unknowns based on consequence and decision urgency rather than attempting to close every data gap at once.

Use confidence bands instead of false precision

Where a score is used, the report should display a confidence band and the reason for it. High confidence requires current evidence for a defined population, reproducible collection, clear ownership, and independent challenge where appropriate. Moderate confidence may reflect partial coverage, sampling, or a recent operating change. Low confidence reflects stale, unowned, contradictory, or unverified evidence. This approach is more useful than a single percentage because it tells leadership whether the number is fit for a decision.

Govern the benchmark as a repeatable research process

The benchmark owner should maintain a source hierarchy, evidence definitions, cohort rules, review cadence, and change log. When definitions change, prior results should be restated or clearly marked so that trend lines do not compare different measures. Sites should be able to challenge their result with contrary evidence, and reviewers should record why the grade changed. This creates an auditable research process rather than a one-time dashboard exercise.

The executive outcome is not a ranking. It is a portfolio view showing where evidence is strong enough to support action, where management is operating on assumptions, and which evidence improvements will unlock the next control, funding, or risk-acceptance decision.

Interpreting the Six Evidence Domains as One Decision System

Asset evidence establishes the population on which every other conclusion depends. When process role, owner, lifecycle, connectivity, and recovery source are missing, access and vulnerability measures cannot be interpreted reliably. The benchmark should therefore state whether the population is discovered, validated, governed, or tested. Moving from one state to the next should require explicit evidence rather than an assumed percentage.

Access evidence should connect identity to a route and a destination. Lists of privileged accounts do not show persistent tunnels, shared support appliances, certificates, local vendor paths, or emergency workarounds. A strong benchmark reconciles account records with observed communication and supplier obligations, then tests whether access can be revoked without hidden operating dependency.

Vulnerability evidence should measure decision quality rather than backlog size. The useful record includes advisory applicability, deployed configuration, credible pathways, process consequence, threat context, treatment feasibility, exception expiry, and proof. A site that closes many tickets without these elements may be processing work without demonstrating risk reduction.

Detection evidence should be evaluated through exercises tied to consequential behaviors. The benchmark should show whether required data was present and healthy, whether the behavior was recognized, whether triage had process context, whether an authorized person could intervene, and whether defects were corrected. Rule count and alert volume cannot answer those questions.

Recovery evidence is strongest when a known-good source is restored through the real dependency chain and the process is accepted by the authorized owner. Backup completion, file presence, or a tabletop discussion can support preparation, but they should not be graded as tested recovery. The benchmark should identify the exact restore scope, dependencies, observed result, and limitations.

Assurance evidence connects the domains. Independent challenge, consequence-based sampling, negative findings, definition control, and a change log make the benchmark reproducible. Without that layer, a dashboard can change because the method changed rather than because the control condition improved.

Domain interaction

Why it matters

Example management action

Asset + vulnerability

Applicability and consequence depend on a validated asset and process role

Validate the top finding population before changing the treatment queue.

Access + detection

A route matters differently when activity cannot be attributed or observed

Standardize supplier access and exercise session visibility.

Configuration + recovery

A backup cannot be trusted when the authoritative source is unclear

Reconcile known-good projects and run a restore acceptance test.

Exception + assurance

A temporary control can remain invisible without expiry and independent review

Escalate expired exceptions and verify the compensating control.

 

Decision Thresholds and Revalidation Triggers

A benchmark becomes operational when grades are connected to thresholds. Leadership should define which evidence condition permits routine management, which requires corrective action, and which blocks a risk-acceptance or investment decision. Thresholds should also state the events that invalidate the conclusion, such as supplier change, network redesign, identity migration, major firmware update, commissioning, or failed recovery exercise. This prevents a credible result from being reused after its evidence boundary has changed.

Use the Benchmark to Improve the Next Decision

The final test is whether a grade changes an owner, threshold, funding choice, evidence request, or validation plan. If the benchmark does not alter a decision, the organization should simplify the measure or clarify the management question it is intended to answer.

CyberTech Intelligence Perspective

CyberTech Intelligence’s perspective is that evidence quality is itself an executive risk variable. A control may be strong, weak, or unknown; the confidence placed in that conclusion should be reported separately. This prevents a site with limited measurement from appearing stronger than a site that finds and discloses more defects.

Benchmarking should therefore reward reproducibility, transparency, freshness, and decision value. A metric is useful when another reviewer can understand the represented population, repeat the collection method, see exclusions, and identify the management action that follows. A number without those properties is a reporting artifact, not decision evidence.

The conversion opportunity is an evidence-improvement plan. Rather than selling another dashboard, the benchmark should identify the smallest evidence changes that unlock better vulnerability, access, recovery, investment, or risk-acceptance decisions.

CyberTech Intelligence OT Evidence Quality Index

The index grades six evidence domains and one executive interpretation layer. It reports the decision supported and the confidence limitation for each domain.

Evidence domain

What credible evidence must show

Decision enabled

Asset evidence

Owner-validated population, process role, lifecycle, connectivity, configuration source, recovery dependency

Which assets require governance, treatment, lifecycle action, or further validation.

Access evidence

Named identity, approved route, destination, privilege, session record, expiry, tested revocation

Which pathways should remain, be restricted, be redesigned, or be removed.

Vulnerability evidence

Applicability, configuration, reachability, threat context, process consequence, feasible treatment, post-treatment proof

Which finding should move first and which treatment can be executed safely.

Detection evidence

Data health, observable behavior, analytic performance, triage, process context, response authority

Whether the organization can recognize and act on a consequential behavior.

Recovery evidence

Known-good source, dependencies, restore sequence, process validation, authority to return to service

Whether recovery claims can support continuity and risk-acceptance decisions.

Assurance evidence

Sampling logic, independent challenge, negative findings, limitations, corrective action

How much reliance leadership can place on reported control performance.

Executive interpretation

Scope, trend, threshold, confidence, limitation, owner, management action

Which decision should be funded, escalated, accepted, or retested.

 

The index should not be collapsed into a single score without displaying the domain grades and confidence limits. A high asset score cannot compensate for untested recovery, and a strong enterprise average cannot erase a material supplier-path exception.

Unknown is a valid state. It should trigger a consequence-based evidence action, not an assumption of low risk.

Benchmark Interpretation Example: When a Lower-Scoring Site Is Actually Measuring Better

A ten-site infrastructure group compares asset ownership, supplier access, vulnerability decisions, detection exercises, and recovery validation. Site A reports few vulnerabilities and high policy coverage. Site B reports more unsupported assets, several stale supplier routes, and multiple recovery defects. A conventional maturity dashboard ranks Site A higher.

The evidence review shows that Site A’s inventory excludes temporary engineering devices and local vendor paths, its access data is based on account lists rather than observed routes, and its recovery score is derived from backup-job completion. Site B reconciles passive discovery with process owners, tests supplier access, validates vulnerability applicability, and runs restore exercises. Site B looks worse because its measurement maturity is higher.

The benchmark separates control condition from evidence confidence. Site A receives several unknown or low-confidence grades and an evidence-improvement plan. Site B receives credible negative findings and prioritized corrective actions. Leadership can now avoid rewarding low disclosure and can fund evidence improvements where they will change the next decision.

Reported result

Evidence limitation

Correct interpretation

Next action

95% asset coverage

Population excludes temporary devices and owner validation

Coverage cannot support lifecycle or consequence decisions

Validate one critical-process population and disclose exclusions.

100% privileged-account review

Routes, destinations, and session activity not tested

Account review does not prove pathway control

Reconcile identities to observed supplier and maintenance paths.

All backups successful

No restore sequence or process acceptance test

Backup completion does not establish recoverability

Exercise restoration for one critical control function.

More vulnerabilities at Site B

Better discovery and applicability validation

Higher finding count may indicate stronger measurement

Track evidence and control maturity separately.

 

The correct executive outcome is not a site ranking. It is a portfolio view showing which decisions are supported, which remain assumption-driven, and which evidence improvements should be funded first.

Evidence Benchmark Quality Checklist

  • Is the represented population explicit?

  • Is the collection method reproducible?

  • Is the evidence date within a defined freshness threshold?

  • Are exclusions and known blind spots disclosed?

  • Is ownership assigned for the evidence and the resulting action?

  • Does the measure distinguish designed, operating, tested, and independently reviewed evidence?

  • Can a reviewer see the negative findings and contradictions?

  • Is the benchmark cohort genuinely comparable?

  • Will a low or unknown grade produce a bounded evidence action?

  • Is the decision enabled by the metric stated beside it?

A metric that fails these tests should remain diagnostic rather than appearing in an executive benchmark.

90-Day Evidence Benchmark Deployment

Days 1–30: Define cohorts and evidence rules

Choose one evidence domain and one comparable cohort. Define the population, source hierarchy, grade definitions, freshness rules, exclusions, and decision that the benchmark must support.

  • Separate measurement and control maturity.

  • Name the evidence owner and reviewer.

  • Document events that trigger revalidation.

Days 31–60: Collect, challenge, and grade

Collect the evidence for the bounded population, test reproducibility, preserve contradictory findings, and assign confidence. Do not fill unknowns with assumptions or enterprise averages.

  • Compare sites only within defensible cohorts.

  • Record why each grade was assigned.

  • Translate low grades into evidence-improvement actions.

Days 61–90: Link evidence to decisions

Present the benchmark with confidence, limitations, owner, threshold, and management action. Retest one domain after the evidence improvement to confirm that the benchmark leads to a better decision.

  • Publish a definition and change log.

  • Schedule freshness reviews.

  • Scale to another domain only after the first decision is improved.

Common Benchmarking Errors

Ranking sites with incomparable operating conditions

Process consequence, lifecycle, supplier model, architecture, and maintenance windows can make the same control evidence mean different things.

Rewarding low disclosure

Few reported defects may reflect limited discovery or weak ownership. Unknown should not be treated as good performance.

Using percentages without populations

A coverage number is not interpretable until the represented population, exclusions, collection method, and evidence date are visible.

Treating benchmark completion as assurance

A completed scorecard does not prove local control effectiveness. Tests and operating evidence remain necessary.

Conclusion

OT security measurement becomes trustworthy when every metric discloses what it represents, how it was collected, when it was collected, what it excludes, and which decision follows. The OT Evidence Quality Index makes those properties visible without pretending to be a universal industry survey.

The strategic objective is not a higher benchmark score. It is stronger confidence where decisions demand it, explicit uncertainty where evidence is weak, and a prioritized plan for closing the evidence gaps that block action.

OT Evidence and Metrics Review

CyberTech Intelligence can benchmark one OT evidence domain or a defined site cohort and convert the findings into an evidence-improvement plan tied to management decisions.

  • Evidence definitions, source hierarchy, and cohort rules

  • Confidence grading and exclusion analysis

  • Decision-linked benchmark scorecard

  • Prioritized evidence-improvement roadmap

Benchmark Your OT Evidence Quality

Start the conversation

Related CyberTech Intelligence Resources

References and Source Links

1. NIST SP 800-82 Rev. 3: Guide to Operational Technology Security

2. NIST Cybersecurity Framework 2.0

3. Principles of Operational Technology Cybersecurity

4. MITRE ATT&CK for ICS

5. CISA Known Exploited Vulnerabilities Catalog