# Changelog
- unreleased
* The demo executables (censor-demo, censor-viz, censor-demo-ffi,
and the per-library suites censor-poly, censor-chacha,
censor-sha, censor-sha512, censor-aead, censor-secp), their
cabal flags, the rs-demo Rust crate, and the flake inputs they
pulled in are removed. Target-specific harnesses belong
outside censor, built on the library's suite scaffolds in
Censor.Runner.Report, which are unchanged.
* **Breaking:** Censor.Runner drops Case, runBattery, and
partitionVerdicts, and Censor.Runner.Report drops Suite and
runSuite, all unused. resultToJSON moves from Censor.Runner to
Censor.Runner.Report. Censor.Runner.Manifest newly exports
readRange and readSeed.
* The dlopen CLI now calls its target through an unsafe FFI
import. The import had defaulted to a safe call, which put the
RTS's thread suspend/resume inside every timed region.
* In the CLI's single-target mode the A/A and B/B controls get a
fresh hypothesis instance, as in manifest mode, rather than
continuing the main run's sampler streams.
* The haddocks are trimmed and deduplicated, with stale text
fixed (including a Censor.FFI example that did not compile).
- 0.5.1 (2026-08-23)
* **Copy-source symmetry in fix-vs-random**
(issues/handled/ISSUE4.md).
`fixVsRandom` and `fixVsRandomCtx` take a re-materialisation
function (`s -> IO s`), run on the pinned secret by *both*
classes every sample: class A completes around the fresh copy,
class B discards its copy and completes around its fresh draw.
With a function that performs a real copy (for ByteString
secrets, `\s -> evaluate (BS.copy s)`) the long-lived pinned
value is read exactly once per sample in each class, and the
completion's input is a fresh, same-aged allocation in both --
removing a copy-source locality asymmetry that surfaced as
intermittent CDF/shape rejections (effect CI straddling zero)
under the cycles meter, invisible to the A/A and B/B controls.
`pure` retains the old behaviour for secret types with no
cheap re-materialiser.
* **Secret regeneration in the dlopen CLI.** The CLI's secret
ranges are no longer overlaid from a long-lived buffer that
only class A reads -- the FFI form of the copy-source
asymmetry above. Both classes now refill their secret ranges
in place each sample through one shared generator, reseeded
before every use (`Censor.Rng.reseed`, new): class A to a
seed pinned at setup, class B to a word drawn from its own
stream, so the classes run identical code and differ only in
the seed value written into the generator state. Sampler
streams for a given `--seed` shift accordingly.
* censor-poly grows a pinned-pair (k1 vs k2, 320 B)
discriminator case: neither class draws or re-materialises
per sample, so a rejection there is key-content-dependent
timing alone -- the instrument that separates a copy-source
artefact from a real leak.
- 0.5.0 (2026-08-21)
* **Margin calibration cells and a host power-gain probe** in
censor-validate. The margin calibration cells shift class B by
less than the resolved margin, so their interval null is true
while the sharp null is false -- empirically validating margin
mode's type-I claim under exactly the sub-tolerance systematic
the margin absorbs; margin power cells confirm shifts beyond
the margin are still caught, and a sharp-contrast cell shows
the sub-margin shift rejecting at ~100% without one.
`censor-validate --power-gain [METER]` measures the host
instead of a synthetic meter: a cycle-identical, maximally
power-asymmetric control (zeros vs alternating 01/10 rewrites
of a 64 KiB buffer) bounds the host's power-to-time coupling
gain from above, with an A/A control and, when the bound
resolves usefully, a suggested `--margin`.
* **Interval nulls (margin mode).** `cfgMargin` (CLI `--margin`,
manifest `margin`) replaces the sharp null with the interval
null |E d_clipped| <= delta, where delta resolves at warmup as
the given fraction of the median pooled per-batch reading. The
three magnitude components become interval-null tests anchored
at +-delta (via ppad-eproc 0.5's `configInterval`) and alone
gate the verdict; the sign and CDF components are demoted to
non-gating diagnostics, since they carry no meter-unit
tolerance. A `Reject` then means the mean effect exceeds the
stated tolerance -- a claim the wall meter can actually support
on frequency-scaled hosts, where the sharp null is genuinely
false for host-mediated reasons and a consistent test
eventually rejects it (see the README's noisy-hosts section).
The resolved absolute margin is reported as `resMargin`
(`margin` in the result JSON) and must land in (0, c) for the
warmup clip bound c. `leakShape`/`diagnose` rank only the
gating components on margin-mode runs. Adding `cfgMargin` /
`resMargin` breaks positional construction of `Config` /
`Result`. Requires ppad-eproc >= 0.5.
* **Channel-attribution counters.** Four new `Counter` meters
complete the wall-decomposition ladder: `RefCycles`
(constant-rate reference cycles -- a class difference here with
clean `Cycles` is a frequency/DVFS effect by construction),
`TaskClock` (on-CPU nanoseconds including kernel time on the
task's behalf; a software counter, so it needs no PMU access
and is available in VMs -- but counting kernel time is itself
privileged, so it needs `perf_event_paranoid` at most 1 and
fails with `OpenFailed 13` at the 2 the other counters
tolerate), `BranchMisses` (secret-dependent branch
direction under balanced retired counts), and `CacheMisses`
(memory access patterns). With the existing meters, each
adjacent rung of instructions -> cycles -> ref-cycles ->
task-clock -> wall isolates exactly one cause of a class
difference: code path, microarchitectural latency, frequency,
kernel/interrupt time, scheduling. The dlopen CLI and suites
accept them as `ref-cycles`, `task-clock`, `branch-misses`,
and `cache-misses`. Adding constructors to `Counter` is a
breaking change for exhaustive matches on that type.
- 0.4.4 (2026-08-05)
* **Per-case baseline cost.** `baselineReading` measures the
target on fresh class-A draws at the case's batch and reports
the median and IQR in per-batch meter units -- the scale of
`resEffect` -- so an effect interval can be read as a relative
bound against the cost of a call. Fresh draws make every
measurement genuine work under any batch policy, unlike the
calibration noise probe, whose shared applied thunk makes it
meaningless for pure Haskell targets. `runTraceSuite` and both
dlopen CLI modes probe it per case; it surfaces as a new
optional per-case `baseline` key in the `censor/report-v3`
object and a `baseline med=N iqr=N` segment on the case detail
line.
* Fixes the benchmark suites, which had failed to compile since
`prepare` was added to `Hypothesis` in 0.4.1.
- 0.4.3 (2026-08-01)
* **The driver's pair loop is now flip-oblivious.** `runPair`'s
order branch compiled the two A/B orders to two distinct code
paths, so the per-pair order bit selected which physical code
executed the pair -- coupling the bit into the measured
readings through microarchitectural state and breaking the
conditional exchangeability the null rests on. The A/A
negative controls surfaced this as pervasive timing-meter
rejections on long-duration pure targets, with the
instruction lane clean (issues/ISSUE2.md). The pair now runs
one code path with a single call site per sampler and
measurement: sampler order is selected by indexing a
two-element array with the order word, class labels are
recovered by mask arithmetic after both measurements, and the
order bit is never scrutinised as a `Bool` inside the loop.
Verified branchless by disassembly, in both the main and
warmup loops.
* **The darwin wall meter is tick-resolved.** `wallClock` now
reads `clock_gettime_nsec_np(CLOCK_UPTIME_RAW)` (~42 ns
ticks on Apple Silicon) rather than GHC's
`getMonotonicTimeNSec`, whose `CLOCK_MONOTONIC` macOS
quantises to microseconds. Microsecond atoms concentrated
wall readings so that warmup placed CDF cut points exactly on
quantisation boundaries, where the indicator channel
amplifies nanosecond-scale harness perturbations into
percent-scale class asymmetries.
* Adds a `primitive` dependency (`SmallArray`, backing the
branchless sampler selection).
- 0.4.2 (2026-08-01)
* Makes minor adjustments to the flake setup and some demo binaries.
- 0.4.1 (2026-07-30)
* **`prepare` now runs during batch calibration.** `runPair` has
always run the prologue before either sampler, but
`calibrateBatchReport` drew its probe sample without one, so a
hypothesis whose sampler depends on its prologue calibrated
against uninitialised state. It now runs exactly one prologue
and one class-A draw.
* **Calibration and the main run no longer share a hypothesis
instance.** `runTraceSuite` and the manifest runner built one
instance, calibrated it, then ran it — advancing its sampler
state, and contradicting `TraceCase`'s documented contract that
the builder is re-executed per phase. Each phase now gets its
own instance, as documented.
* **Paired public context reaches the manifest.** A new
`context=off:len` case range (and matching `--context` flag) is
redrawn once per *pair* and written into both classes, so a
manifest can express a shared fresh public input. Previously
`fixVsRandomCtx` existed in the library with no route to it from
the foreign runner, whose public ranges were pinned for the
whole cell.
* **One layout validator for both entry points.** `checkLayout`
requires every range to fit the workspace and the three roles —
secret, public, context — to be pairwise disjoint. The manifest
and the single-case command line now share it; previously the
command line bounds-checked only, so `--secret 0:8 --context
0:8` ran a layout the manifest rejected, and the manifest in
turn accepted a public range that masked the context it
overlapped. A byte belongs to exactly one role, so an overlap
has no coherent reading and is an error rather than a
precedence question.
* **`label=` is presentation only.** It set `mcName`, from which
the per-case seed was derived, so renaming a case for display
silently re-rolled the secret it tested against. `ManifestCase`
now carries `mcId` — the case token as written — and seeds
derive from that.
* **The foreign runner no longer strands its buffers.** The three
per-hypothesis workspaces were raw `mallocBytes` with no
release; phase isolation multiplied that to as many as nine
workspaces per cell. They are now `ForeignPtr`s held by the
sampler closures, reclaimed with the hypothesis.
* **Replicate identity is stable.** `replicates=1` and
`replicates=2` now agree on cell #1: replicate #1 keeps the
unreplicated case's seed, so adding coverage extends a battery
instead of redefining the cell already in it. Each case also
mixes its own name into the run-wide seed, so cases no longer
all begin from identical fixed bytes.
* **`family-alpha on`** runs each of the M expanded cells at
`alpha / M`, making the battery a test at `alpha` family-wise
rather than only reporting the union bound. Off by default.
* **Report schema is now `censor/report-v3`.** The two seeds are
named apart — `samplerSeed` in the header, `orderSeed` in each
case result — where v2 called both `seed`. The resolved order
seed also appears in pretty and plain output, not only in JSON;
it is what replays a case, and a reader should not have to
reach for the JSON to find it.
* **The CDF channel reports per-cut-point detail.** `resCutPoints`
and the `resPeakCdf` field are replaced by `resCdf ::
[(Double, Double)]`, pairing each cut point with its
component's peak log e-value, so a CDF-driven rejection names
the threshold that found the difference. `resPeakCdf` remains,
as a function taking the maximum. In JSON, `cutPoints` becomes
`cdf` as `[cut point, peak]` pairs.
* Documentation corrections. The null is *stronger* than "the two
classes induce the same timing distribution", not weaker:
exchangeability implies equal marginals, and equal marginals
alone do not imply an exchangeable joint law. What delivers it
is the randomised order, given class-blind within-pair nuisance
dynamics — an assumption now stated rather than elided. The
order-bit seed's OS-entropy source is documented as the normal
path, not an unconditional one, since `/dev/urandom` failure
falls back to the monotonic clock. Stale `K=4` claims are
corrected to `K = 4 + n`.
* `censor-viz` uses `runCTWith` instead of a hand-rolled copy of
the driver, which had drifted: it omitted the CDF components and
the prologue while claiming to mirror `runCT`. The AEAD suite
draws its public input (nonce, AAD, plaintext) once per pair and
shares it across both classes, rather than independently per
class.
- 0.4.0 (2026-07-30)
* The null is now stated, and tested, as conditional
*exchangeability* of the measured pair `(ta, tb)` rather than
symmetry of `d = ta - tb`. Exchangeability is what the
randomised A/B order actually delivers; every hedge component
is an antisymmetric functional of the pair, and each such
functional is conditionally symmetric about zero under the
null.
* New **CDF-indicator channel**: one bounded-mean e-process on
`1[ta <= q] - 1[tb <= q]` per warmup-fixed cut point `q` (the
deciles and quartiles of the pooled warmup readings,
deduplicated). This detects dispersion leaks — a target whose
timing variance depends on the secret while its mean does not —
which leave `d` symmetric and were invisible to every previous
component. `K` grows from 4 to `4 + n`; `Result` gains
`resPeakCdf` and `resCutPoints`, and `LeakShape` gains
`CdfShift`.
Previously such a target passed unconditionally. The
variance-only cell in the validation suite moves from H_0 to
the power table accordingly.
The functional family is finite, so this is still not an
omnibus test of distributional equality. That limit is now
documented rather than left implicit.
* `Attribution` is no longer a four-way label. It is a record
carrying the `ControlVerdict` plus *both* control runs in full,
and `TargetLeak` is renamed `ControlsPass`. The A/A and B/B
runs are negative controls: they can convict the harness but
cannot acquit it, so the old name overstated what they
establish. The JSON `attribution` string field becomes a
`controls` object with `verdict`, `aa`, and `bb`.
* Clipping is counted at each of the hedge's three magnitude
bounds, not only the tight one. `Result` gains `resClipped4`
(`|d| > 4c`) and `resClipped16` (`|d| > 16c`) alongside
`resClipped`, and `resultToJSON` emits them as `clipped4` and
`clipped16`. The counts nest, and they qualify different
quantities: `resClipped` bounds what the tight magnitude
component could see, while `resClipped16` qualifies
`resEffect`, `16c` being the bound the effect interval's own
input is clipped to. A consumer quoting that interval
previously had no way to tell a starved tight component from a
truncated interval.
`HighClipRate` accordingly carries a `ClipRates` record with
the fraction at all three bounds rather than a bare `Double`
(`clipFraction`, `clipFraction4`, `clipFraction16` in JSON).
Its trigger is unchanged — the counts nest, so the tight rate
is the largest of the three and one trigger covers the
profile. `Censor.Runner` exports `clipText`, which renders the
profile for the one-line and terminal reports and stays as
terse as it was when nothing clipped past `c`.
* Reports carry suite-wide alpha accounting: `totals` gains
`alpha` and `familyAlpha` (the union bound `M * alpha` over `M`
cells), and multi-case footers render it. Censor's guarantee is
per run; a battery that fails when any cell rejects has the
family bound, not alpha.
* The default order-bit seed is drawn from OS entropy
(`/dev/urandom`) rather than `getMonotonicTimeNSec`; a clock
reading is not independent entropy, and the nuisance process
the order bit must be independent of may itself be periodic in
that same clock. The resolved seed is reported as `resSeed`, so
a randomised run stays exactly replayable.
* The `censor` CLI splits `--seed` (input samplers) from
`--order-seed` (the A/B order bit). They previously shared one
value, weakening the independence the conditional null needs.
* `Censor.Runner.Env` records the conditions a run executed
under: a best-effort host snapshot (OS, arch, kernel release,
CPU model, core count, Linux cpufreq/SMT/perf_event knobs, load
average, UTC timestamp) taken once at run start, backed by a
small C shim. Rides on the report header and into the JSON.
Alongside it, `calibrateBatchReport` records the meter's noise
floor (`Noise`: median and IQR of repeated readings) at the
selected batch, so a `Pass` obtained where dispersion rivals
the effect sizes of interest can be recognised as such.
* `TraceCase` carries an explicit `BatchPolicy` — `SingleShot`,
`RepeatableAuto`, or `FixedBatch n`. `runTraceSuite` no longer
auto-calibrates every case: repeating a pure Haskell target
forces one thunk and then measures no-ops, so a calibrated
batch counted no-ops as work. All bundled demo targets are
pure and now declare `SingleShot`. `--batch N` overrides every
case's policy; plain suites reject the flag. `rhAutoBatch`
becomes `rhPerCaseBatch`, rendering `batch=per-case`.
* `Hypothesis` gains a `prepare` field: a per-pair prologue run
before either sampler and outside the timed region. New
`fixVsRandomCtx` uses it to draw one fresh public context per
pair and build both classes around that same context, so the
tested axis is secret dependence conditional on identical
fresh public input. Previously the completion ran independently
per class, so public inputs were either independently random
across the pair (adding nuisance variance to `d`) or fixed for
the whole run (testing one public input). `FFIHypothesis`
gains the matching `ffiPrepare`.
* New `Censor.Runner.Manifest` and a `--manifest PATH` mode on
the `censor` CLI: a declarative battery of cases, driven by one
dlopen and emitted as a single report carrying every case, its
controls, the host environment, and the family-wise alpha.
Replaces the per-symbol process loop with hand-stitched output,
which duplicated the tool version and config in every script.
Cases support `replicates=N` to run several independent fixed
secrets, since a single pinned secret answers only whether
*that* secret times differently from random. See
`etc/example.manifest`.
* `TraceSuite` takes a heterogeneous case list (an existential
`TraceCase`), so cases with different hypothesis payload types
share one list, and auto-attributes on `Reject`. A run collapses
to a single canonical JSON object — the same one the `json`
renderer emits to stdout, with optional per-case batch, noise,
and trace fields. `renderTrace` is gone.
* The `censor` executable catches up with that surface: `--trace`
writes the same report object the `json` renderer emits, with
the frame trajectory embedded, rather than a separate format.
* `Censor.Runner` exports `shapeName`; `Censor.Runner.Report` now
imports it rather than keeping a second copy that could (and
did) fall out of step with `LeakShape`.
* `censor-secp` presents an 8-case surface: both constant-time
multipliers (`mul`, `mul_wnaf`), the ladder-based `sign_ecdsa`
/ `sign_schnorr`, their precomputed-context wNAF variants, and
`ecdh` / `derive_pub`. The four signing cases hold the message
(and Schnorr aux) fixed and public across both classes, so the
derived RFC6979 / BIP340 nonce is deterministic in the fixed
class and fix-vs-random controls key + nonce rather than the
key alone.
* Report schema bumped to `censor/report-v2`, reflecting the
`controls` object, the new `Result` fields (`seed`,
`cutPoints`, `peaks.cdf`, `clipped4`, `clipped16`), the
`totals` alpha fields, and `"batch":"per-case"`.
- 0.3.1 (2026-07-09)
* The `censor` executable renders through `Censor.Runner.Report`
like the other suites. `--json` and `--summary` become
`--format pretty|json|plain`; verdict and attribution ride on
a single `CaseReport`. Exit codes and the `--trace` /
`--attribute` behavior are unchanged.
- 0.3.0 (2026-07-09)
* `Censor.Runner.Report`: uniform report rendering with a
header/case/footer data model and three renderers:
`pretty` (the ppad terminal aesthetic), `json`
(line-oriented `censor/report-v1` object), and `plain`
(ASCII/CI-safe).
* `Suite` and `TraceSuite` scaffolds package the full
test-suite lifecycle (arg parse, meter selection, header
assembly, `calibrateBatch`, per-case loop, optional trace
JSON dump) so a tool's `Main.hs` reduces to the
hypotheses and the case list.
* `Censor.Runner.DL` and the `censor` executable: a
dlopen-based runner that drives a constant-time
hypothesis against a foreign target resolved by name from
a shared library.
* `diagnose` interprets a `Result` into actionable
`Advisory`s (clip-rate warning, low-power `Pass`,
rejection-driver hint); `summary` and `resultToJSON` give
one-line and JSON views of the same evidence;
`calibrateBatchReport` returns the auto-batch decision
alongside its calibration probes.
- 0.2.0 (2026-07-07)
* Hardware performance-counter meters on Linux
(instructions retired, branches retired, cycles) via
perf_event_open, alongside the existing wall-clock meter.
* Foreign-function target support ('Censor.FFI'): a pinned
workspace shared across pairs lets C, Rust, or Go targets
run under the same sequential CT driver.
* Fix-vs-random hypothesis and 'Attribution': A/A and B/B
diagnostics fold into a four-way attribution of any
rejection (target, sampler, or harness).
* 'Result' now reports an anytime-valid p-value and a
time-uniform confidence interval on the leak's effect
size, with named fields throughout.
* Mixture internals delegated to ppad-eproc; verdicts
unchanged.
* 'Censor.Rng' extracts the splitmix64 sampler PRNG
previously duplicated across demos.
- 0.1.0 (2026-05-31)
* Initial release. Sequential constant-time testing via
anytime-valid e-processes, with wall-clock measurement on macOS
and Linux.