VCyber Twin · Breach & attack simulation on a digital twin

Your detection rate is a number.

Most teams have never seen it.

Coverage percentages come from counting rules. We get the number the other way round: replay the technique, watch the SIEM, and write down what fired, what stayed silent, and how long each one took.

Run Wazuh? Rule Doctor Lite is free: one read-only Python file that lists the custom rules that never fired, and why — including the ones Wazuh drops at load.

T1110.001 · same attack, three SIEMs
Measured on ATK’s bench alert time from first attempt
Same build · run over run
Measured coverage re-run on the same bench
3 / 24
Techniques caught by an out-of-the-box deployment on our bench. Defaults are a low bar, and nobody publishes how low.
50 → 83 %
Coverage on the same build, run over run, after adding two detection rules.
84
ATT&CK techniques in the library, across 75 scenarios and 84 Sigma rules. Having a rule in the library is not the same as that rule being proven to fire — we track those separately, and say which is which.
3
SIEMs measured natively — Wazuh, Splunk and IBM QRadar — with the same attack and the same clock.

“Nobody publishes how low” is the reason we did. Same bench, same terms, method and limits in the body: Detection Reality Index Vol.1 (the 3 / 24 run) · Vol.2 (IBM QRadar) · Vol.3 (what moves between two audits).

How it works

Nothing touches
your environment.

01 · Model

We build a twin of the stack you want measured — hosts, roles, network zones and the SIEM configuration — as a graph. No agent is installed on your side and no production system is involved.

02 · Replay

Real techniques run against the twin: credential brute force, lateral movement, ransomware behaviour, command and control. Not synthetic log injection — the attack actually executes.

03 · Time it

Every technique is timed against the SIEM's own alert stream. The same SSH brute force (T1110.001) alerted in roughly 0–2 s on Wazuh, ~4 s on Splunk and ~4–8 s on QRadar. Those are clock readings, not estimates.

04 · Close the loop

Where we have detection content for a silent technique, it comes back with the report, and re-running shows the delta. Two honest limits: not every gap we find has a rule behind it, and each rule is labelled either verified — loaded into a live SIEM and observed to fire, or not yet verified. We will not hand you a rule of the second kind and call it a fix.

01 · ModelTwin of your stackhosts · zones · SIEM configBuilt as a graph from what you declare. Nothing installed on your side.
02 · ReplayReal techniques runlibrary: 84 techniques · 75 scenariosThe attack executes against the twin. Not synthetic log injection.
03 · Time itClock on every alert0–2 s · ~4 s · 4–8 sT1110.001 on Wazuh, Splunk and QRadar: clock readings, not estimates.
04 · Close the loopFix, then re-run50 → 83 % after two rulesEach rule labelled verified-to-fire or not yet verified.

What we refuse to do

A zero is not
a measurement.

These are enforced in the software, not promised in a contract. A promise can be forgotten. A gate runs on every job.

“No alert” has three possible causes. Only one of them is your problem.

The monitoring missed it · the monitoring received nothing to miss · the attack never reached the sensor. Before we call anything a gap, the software has to prove the log pipeline was alive inside the measurement window. If it cannot, the cell is labelled UNMEASURED, with the reason printed verbatim — and it is excluded from your gap list, from the coverage denominator, and from the invoice.

We say whose monitoring we measured.

By default the measurement runs on a replica on our bench, built from what you declared. The attack is real, the alert is real, the rule ID and detection time are real — but the thing being measured is the replica, not the system you are running. Every coverage figure carries that label next to it, not in a footnote. If you give us access to your own SIEM, the label changes accordingly.

We do not bill a measurement that could not measure.

Every run records how many techniques were actually measurable. If none were — a dead log pipeline, an expired SIEM licence, an attack that never landed — the run is marked non-billable with the reason, and you can read that ledger yourself.

We publish our own error bars, including the corrections.

We ran both of our modes against the same twin, compared them cell by cell, and published the confusion matrix — along with a correction to our own sales document when a later review showed one of our figures was wrong. A buyer who watches a vendor retract a claim should trust the remaining claims more, not less.


What you get

Two pages that
settle an argument.

What fired. What stayed silent. How long each one took.

A report on one stack profile, in language a technical buyer can check and a non-technical one can act on. It runs on our bench, so nothing goes through your change board and nothing goes through your procurement.

For a provider

A number you can put in a renewal conversation or a competitive pitch — evidence that the tuning you did is worth what you charge for it, rather than an assertion the client is asked to trust.

For an in-house team

A third-party measurement. A number your own team produced is the one a board quietly discounts; the one that travels is the one you did not produce yourself.


Scope

What this is not.

The honest sentence

We configure a build to match one of your stack profiles and measure that. The report characterises that build, not your live estate — and it says so, in the report, in those words. Anyone selling you a bench result as a measurement of production is selling you something they cannot deliver.

Not a penetration test

A pentest answers can this be exploited? This answers would anyone have seen it? The two are complementary, and a pentest report carries no detection telemetry.

Not a scorecard on your vendors

This measures configuration and deployment. It is not a verdict on any product you have partnered with or resold.

Current status

Pilot. Live demo environment, single-tenant per twin, no production customer deployment yet. We would rather tell you that here than have you find out in week three.


Fixed price

Wazuh Fix Pack.
$490, paid after it works.

One scoped problem on your own Wazuh, fixed and tested against your exact version. You apply it on your cluster. You pay only once it runs there.

A rule that loads and never fires

We work out why — shadowed by a sibling rule that matches first, an event that never reaches the manager, no event that matches it, or a rule Wazuh dropped while loading — by replaying the events you send and reading the manager’s own log, not by reading rule IDs. Where the rule can be fixed, you get the fixed rule and the replay that shows it firing.

Alerts that never reach the indexer

mapper_parsing_exception in Filebeat, a field that is an object in one event and a string in the next, documents dropped without a trace. You get the pipeline processor and index template, tested on an indexer of your version against your event samples, with apply and rollback steps.

How paying works

We agree the scope in one email: the rule or the error, your Wazuh version, the samples. We deliver files and steps. When they run on your cluster — the rule fires on the replayed event, or the samples index with no mapping error on the next daily index — we send the invoice (bank transfer or PayPal). If it does not run, you owe nothing. We never need write access to your systems.


Pick one representative stack.

Tell us the shape you see most often across your customers — SIEM, endpoint, and roughly how it is tuned — and we will scope the measurement against that. Pricing follows the number of stack profiles, so it is a short conversation.