Skip to content
ULAP ONE
Research preview
Research · benchmarks · evidence

Methodology

Build snapshot · Not live

One section per versioned suite contract: what identifies the workload and the data, what each metric means and which direction is better, every pass/fail threshold with its source and its version, the accepted statistical treatment, the known limitations verbatim, and the exact reproduction command where it is safe and available.

Build snapshot Not live Captured 2026-08-10T07:58:14.448Z Source benchmarks/reports Ingested 2026-08-10T07:58:14.432Z 44 JSON files seen
The safety policy for reproduction commands, stated once
A command is printed on this page only when running it satisfies every clause below. Trip one and the command is withheld: the producer is still named, the preconditions are still listed, and the clause that was tripped is quoted, so nothing is silently omitted.
  1. A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
  2. A reproduction command is offered only when running it does NOT drive an endpoint that is serving the live demo.
  3. A reproduction command is offered only when running it does NOT require a new RF capture session.
  4. A reproduction command is offered only when its producer exists in this tree.

Contract census

Suite contracts
11 / 14

Versioned suite contracts, and how many of them at least one indexed artefact is governed by. A contract with no artefact is listed here anyway: its method is declared even where nobody has run it.

Declared thresholds
22

Gates declared across all contracts. Every one carries its own rule, its severity, the claim a failure removes, and the source and version of the threshold. A threshold with no source is not expressible in this model.

Reproduction commands offered
7 / 17

A command is printed only when the safety policy below permits it. Where it is withheld, the producer is still named and the tripped clause is stated — nothing is silently omitted.

Contracts

Jump to a suite

14 of 14

UHF occupancy on a 28-channel over-the-air survey

UHF occupancy on a 28-channel over-the-air survey

served-occupancy-survey@1 contract v1 no artefacts in this snapshot

On real captured RF with both occupied and vacant channels present, does passing a calibrated control floor change what the served model decides — and what does it cost?

Contract version 1 Valid tier: Over the air served-occupancy@1 No artefact in this snapshot

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it. It is also served through an endpoint, so the serving seam is part of the measurement.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
no artefacts none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
no artefacts none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
no artefacts none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
no artefacts none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
no artefacts none statedWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
no artefacts none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
no artefacts none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
no artefacts none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
no artefacts none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
no artefacts none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
no artefacts none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
no artefacts none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
no artefacts none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
no artefacts none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
multiplexes detected on the held-out half
report_multiplexes_detected_after
channels
count
Higher is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one derived-occupied multiplex in UHF 35-48, the half no threshold was tuned on3 no artefacts
Secondary
multiplexes detected, whole survey
multiplexes_detected_after
channels
count
Higher is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one derived-occupied multiplex across all 28 swept channels10 no artefacts
Secondary
occupied records detected, whole survey
occupied_records_detected_after
records
count
Higher is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one capture record on a derived-occupied channel20 no artefacts
Safety
occupied channels the change made worse
regressed_channels
channels
count
Lower is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one derived-occupied channel detected in fewer records after the change1 no artefacts
Safety
vacant channels still called occupied in at least one record
residual_false_positive_channels
channels
count
Lower is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one derived-vacant channel with a non-zero after-arm occupied-record count1 no artefacts
Safety
vacant records called occupied
vacant_false_positive_records_after
records
count
Lower is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one capture record on a derived-vacant channel20 no artefacts

Pass/fail thresholds — and where they came from

Threshold source

Operator rules: zero regressed occupied channels, and a vacant class must exist before any false-positive rate is quoted. Neither is an FCC threshold nor derived from one.

Threshold version operator-2026-08

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking no occupied channel detected less often after the change metric regressed_channels · at most 0

    Did any channel the survey derived as OCCUPIED lose detections when the control floor was applied?

    Why it matters
    A net win that hides a per-channel loss is how a regression ships. Every occupied channel must be detected at least as often after the change as before it, or the headline is an average standing over a casualty.
    Blocks the claim
    that the control floor is a strict improvement on every occupied channel
    On fail
    At least one occupied channel is detected in fewer records after the change. State it wherever the headline is stated; a net improvement does not settle it.
    On absent
    The per-channel derived_truth and before/after record counts needed to detect a regression are not in this artefact.
  • Blocking no vacant channel is still called occupied metric residual_false_positive_channels · at most 0

    Does any channel the survey derived as VACANT still get called OCCUPIED in at least one record?

    Why it matters
    This is the not-clean half of the result and it must be carried at equal prominence with the win. READ THE DIRECTION on UHF 28: 20 of 20 records to 8 of 20 is a FALSE POSITIVE that improved and was not eliminated, not a detection that was lost — the channel is derived VACANT, 0.39 dB over the session floor. The cyclostationary plane fires there at comb z = 4.90 against a Z_PRESENT of 3.5. The threshold was NOT moved and the case is pinned xfail(strict=True). This text used to add that Z_PRESENT was calibrated against synthetic AWGN maxima, the same defect one layer down; measured 2026-08-05 that is false — real receiver noise reaches comb z = 2.289 against the synthetic 2.19 and 0 of 300 null records reach 3.5. The firing is antenna-borne and its cause is open between a genuine emission below the energy floor and front-end intermodulation, separable only by an RX-gain sweep that has not been run.
    Blocks the claim
    that the vacant direction is clean on this survey
    On fail
    Vacant channels are still called occupied on some records. The channel-level false-positive count reaching zero does not settle it; state this beside the headline every time.
    On absent
    The per-channel derived_truth and after-arm record counts needed to find a residual false positive are not in this artefact.
  • Blocking the survey contains vacant channels fact dataset.has_vacant_captures · exactly true

    Are there genuinely vacant captures to measure a false-positive rate on?

    Why it matters
    The 2026-07-29 set had none, so an unconditional OCCUPIED scored full marks on it and the headline drawn from it was withdrawn. This survey has 15.
    Blocks the claim
    any false-positive rate on real RF
    On fail
    The survey contains no vacant channels and no false-positive rate is measurable.
    On absent
    The artefact does not state how many vacant channels the survey contains.
  • Blocking the served default is unchanged fact model.endpoint_weight_hash_proven · exactly true

    Does the running service behave as measured in the AFTER arm?

    Why it matters
    ModelSettings.control_floor_dbfs defaults to None. With no calibrated floor the engine falls back to the legacy CNN verdict and says so via occupancy_source. The AFTER arm is what the service CAN do, not what it does.
    Blocks the claim
    that a deployed SIDSense service currently decides occupancy this way
    On fail
    The endpoint is not proven to serve this checkpoint or this configuration.
    On absent
    No endpoint weight hash exists, so this can be neither confirmed nor denied. The AFTER arm describes an opt-in path, not the shipped default.
  • Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true

    Was the receive antenna specified for the band these captures were taken in?

    Why it matters
    Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
    Blocks the claim
    any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
    On fail
    The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
    On absent
    This artefact does not state whether the antenna matched the band. Absent, not confirmed.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 0 / 0

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
wilson

Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.

What counts as a repetition
A re-capture at another site or session, scored by the same harness against the same checkpoint. Re-scoring the same stored IQ is not a repetition.
Minimum sample count
3

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • any percentile — this suite produces counts and proportions, not distributions
  • a single accuracy folding the occupied and vacant directions together
  • quoting the whole-survey number as a held-out result — the calibration half is in it

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • One site, one session, one antenna. Nothing here generalises to another site or another receiver.
  • The control floor is opt-in. ModelSettings.control_floor_dbfs defaults to None and there is deliberately no default floor, so the shipped service still returns the legacy CNN verdict unless a floor is supplied.
  • The CNN modulation label is reported and never scored here: the captures are DVB-T2 and the model has no DVB-T2 class. This says nothing about the 0.8864 synthetic modulation-class figure, which is a different measurement.
  • The 15 at-floor channels sit at the receiver’s own floor, so a genuinely empty channel and one the antenna cannot hear are indistinguishable. The receiver-noise control was taken on 2026-08-05, on the B210’s unconnected RX chain, and it does not separate those two — it characterises noise structure, not what the antenna can hear. A matched UHF antenna is what would, and it has not been used.

Caveats carried verbatim from the artefacts (0)

No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.

Reproduction

  • Command offered Served occupancy: CNN vs physics detector 0 artefact(s)
    Producer scripts/score-served-occupancy.py verified

    Scores both arms in one process against one unchanged checkpoint.

    Named in the artefact’s own harness field.

    python scripts/score-served-occupancy.py --survey <uhf-survey-dir> --checkpoint external/tvws-sensing/models/checkpoints/best.pt --out benchmarks/reports/served-occupancy-2026-08-05.json --records 20 --split all

    Assembled from the producer’s own argparse (scripts/score-served-occupancy.py lines 119-126). This is NOT A TRANSCRIPT — the artefact does not record its own invocation, so the survey path here is a placeholder and the record counts are the parser defaults. The checkpoint path and its sha256 ARE recorded, under model_identity.

    Preconditions

    • The nairobi-region-2026-08-04-uhf-survey captures. They are on the box and outside this repository.
    • The checkpoint at the recorded sha256 8eca022f…. A sha pins a FILE; it does not prove any endpoint serves it.
    • Ground truth is derived from the same session’s measured floor, so the survey directory must be the one that produced the floor, not a re-capture.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs the survey captures and the checkpoint, neither of which is in the repository.

TVWS occupancy on real over-the-air captures

TVWS occupancy on real over-the-air captures

tvws-ota-occupancy@1 contract v1 3 artefacts

Does the sensing model call a licensed, occupied multiplex OCCUPIED, on real captured RF?

Contract version 1 Valid tier: Over the air ota-eval@1

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it. It is also served through an endpoint, so the serving seam is part of the measurement.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
ABSENT from all none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
ABSENT from all none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
3 / 3scripts/eval-ota-checkpoint.pyWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
3 / 3/otaWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
ABSENT from all none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
3 / 38eca022f09b897a579ef74225881343949a1a4a53b880edabf82225aacbb34ee · 038d32176653d11ae0a998b4a1fa3df5de68a67ca8c741024d375c6fd7737973 · 7aca8bad07930190a63346468170b9e4f3a84ebb9e93c00b5460a503d39ce69aA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
record-level occupancy rate
records_occupied_rate
fraction
fraction
Higher is better.proportion

A proportion of a countable population. A confidence interval is expressible only when the numerator and denominator are whole records.

one real OTA capture record from a channel labelled occupied20 3 / 3
Secondary
channels with correct occupancy
occupancy_correct_channels
channels
count
Higher is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one captured multiplex channel1 3 / 3
Secondary
records with the nearest correct class
records_nearest_correct_class
records
count
Higher is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one real OTA capture record20 3 / 3
Safety
channels wrongly called VACANT
false_vacant_channels
channels
count
Lower is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one licensed multiplex channel whose records are all labelled occupied1 3 / 3
Safety
records wrongly called VACANT
false_vacant_records
records
count
Lower is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one real OTA capture record labelled occupied1 3 / 3

Pass/fail thresholds — and where they came from

Threshold source

Operator safety rule: zero false-vacant channels. Not an FCC threshold and not derived from one.

Threshold version operator-2026-08

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Safety critical no licensed channel called vacant metric false_vacant_channels · at most 0

    Did the model call any occupied, licensed multiplex VACANT?

    Why it matters
    A false-vacant decision authorises a transmission on a channel a licensed incumbent is using. It is the only error in this suite that causes harm outside the system.
    Blocks the claim
    that the model is safe to gate transmission on a vacant-channel decision
    On fail
    FAILED SAFETY GATE: occupied licensed channels were called vacant. This result stands on its own and is not offset by any accuracy figure, here or elsewhere.
    On absent
    The per-channel labels needed to count false-vacant decisions are not in this artefact.
  • Blocking the capture set contains vacant channels fact dataset.has_vacant_captures · exactly true

    Are there any genuinely vacant captures to measure a false-positive rate on?

    Why it matters
    Every capture in this set is labelled occupied, so a model that answers "occupied" unconditionally scores full marks here.
    Blocks the claim
    any false-positive rate on real RF
    On fail
    The capture set contains no vacant channels.
    On absent
    No vacant captures exist in this set, so no false-positive rate on real RF is measurable at all — absent, not zero.
  • Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true

    Was the receive antenna specified for the band these captures were taken in?

    Why it matters
    Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
    Blocks the claim
    any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
    On fail
    The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
    On absent
    This artefact does not state whether the antenna matched the band. Absent, not confirmed.
  • Blocking the endpoint is proven to serve this checkpoint fact model.endpoint_weight_hash_proven · exactly true

    Is there evidence that the service under test served the weights named here?

    Why it matters
    A checkpoint sha256 pins a FILE on disk. The service exposes no weight hash, so nothing in this artefact can bind the file to the endpoint that produced these numbers.
    Blocks the claim
    that the deployed service serves the checkpoint named here
    On fail
    The endpoint is not proven to serve this checkpoint.
    On absent
    No endpoint weight hash exists, so this can be neither confirmed nor denied. The checkpoint hash is provenance, not deployment verification.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 0 / 3

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
wilson

Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.

What counts as a repetition
A re-capture at the same site with the same receive chain, evaluated by the same harness revision against the same checkpoint. Re-running the evaluator over the same stored IQ is not a repetition.
Minimum sample count
20

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • any percentile — this suite produces counts and proportions, not distributions
  • a single accuracy that folds false-vacant and false-occupied decisions together

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • Nothing here establishes anything about vacant channels: no vacant captures exist in this set.
  • One site, one afternoon, one receive chain. Nothing here generalises to another site or another receiver.
  • The captures are DVB-T2 and the model has no DVB-T2 class.

Caveats carried verbatim from the artefacts (6)

  • caveats[0]
    Captures are DVB-T2; the model has no DVB-T2 class. DVB-T is nearest.
    3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json
  • caveats[1]
    Received through a mismatched 860-930 MHz antenna at 470-700 MHz.
    3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json
  • caveats[2]
    Proves nothing about vacant channels — no vacant captures exist here.
    3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json
  • harness_note
    Runs the service's own load_model + InferenceEngine.infer, not a reimplementation. The (H,W,C)->(C,H,W) transpose lives inside InferenceEngine.infer.
    3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json
  • model_identity.note
    checkpoint_sha256 pins a FILE. It does NOT prove that any endpoint serves these weights.
    3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json
  • summary.record_level_note
    Channel counting is all-or-nothing and hides movement: a channel going 0/20 -> 18/20 occupied still reads as one failed channel. Compare checkpoints on the record-level rate as well.
    3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json

Reproduction

  • Command offered TVWS over-the-air eval 3 artefact(s)
    Producer scripts/eval-ota-checkpoint.py verified

    Scores a checkpoint against the real over-the-air captures using the service's own inference path.

    The artefacts name it in their own harness field.

    python scripts/eval-ota-checkpoint.py --checkpoint <path/to/best.pt> --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --confidence-threshold 0.7 --json-out benchmarks/reports/<name>.json

    Assembled from the producer's own argparse (lines 71-83). The capture directory is the one this family's artefacts name in dataset.dir. It is the producer's documented invocation, not a transcript of the run.

    Preconditions

    • The capture set at benchmarks/datasets/ota/nairobi-region-2026-07-29. It exists in this tree.
    • A checkpoint. The artefacts record checkpoint_sha256 and say in their own words that it "pins a FILE. It does NOT prove that any endpoint serves these weights."
    • EVERY capture in this set is labelled occupied. Re-running proves nothing about vacant channels, whatever the numbers say.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs the OTA capture set and a checkpoint, neither of which is on a GitHub runner.

TVWS per-epoch OTA diagnostic

TVWS per-epoch OTA diagnostic

tvws-ota-per-epoch@1 contract v1 4 artefacts

Does synthetic validation accuracy track real-OTA behaviour across training epochs?

Contract version 1 Valid tier: Over the air ota-per-epoch@1

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
ABSENT from all none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
ABSENT from all none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
4 / 4scripts/ota-per-epoch-sweep.pyWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
4 / 4/otaWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
2 / 47A seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
ABSENT from all none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
Spearman: synthetic val accuracy vs real OTA
spearman_val_acc_vs_real_ota
none
scalar
Higher is better.scalar

A single stated value, not a summary of a distribution.

one training epoch5 ABSENT from all

Pass/fail thresholds — and where they came from

Threshold source

None. This suite is diagnostic and has no pass threshold.

Threshold version none

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true

    Was the receive antenna specified for the band these captures were taken in?

    Why it matters
    Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
    Blocks the claim
    any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
    On fail
    The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
    On absent
    This artefact does not state whether the antenna matched the band. Absent, not confirmed.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 0 / 4

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
none

No proportion interval is defined for this suite.

What counts as a repetition
An independent training seed swept over the same capture set.
Independent seeds are required, recorded IN the artefact and not inferred from a file name.
Minimum sample count
5

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • using this sweep as a selection criterion — it is n=5 channels, all occupied

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • Diagnostic, not a selection criterion — n=5 channels, all occupied.
  • Establishes nothing about vacant channels.

Caveats carried verbatim from the artefacts (5)

  • caveats[0]
    Diagnostic, not a selection criterion — n=5 channels, all occupied.
    4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json
  • caveats[1]
    Captures received through a mismatched antenna.
    4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json
  • caveats[2]
    Establishes nothing about vacant channels.
    4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json
  • dataset.note
    5 channels, one site, one mismatched 860-930 MHz antenna at 470-700 MHz; diagnostic only
    4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json
  • harness_note
    drives the service's own load_model + InferenceEngine.infer, same seam as scripts/eval-ota-checkpoint.py
    4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json

Reproduction

  • Command offered TVWS OTA per-epoch sweep 4 artefact(s)
    Producer scripts/ota-per-epoch-sweep.py verified

    Runs the OTA eval across every checkpoint in a directory, one row per epoch.

    The artefacts name it in their own harness field.

    python scripts/ota-per-epoch-sweep.py --checkpoint-dir <dir/of/epoch/checkpoints> --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --records 20 --json-out benchmarks/reports/<name>.json

    Assembled from the producer's own argparse (lines 124-133). It is the producer's documented invocation, not a transcript of the run.

    Preconditions

    • A directory of per-epoch checkpoints. Those are not in this repository.
    • The antenna used for these captures was NOT matched to the band. Every artefact in this family states that in prose in dataset.note, and no boolean in the family carries it — which is why the mismatch is detected from prose and must stay visible on every number.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs the OTA capture set and a directory of checkpoints.

TVWS checkpoint selection, calibration and OOD

TVWS checkpoint selection, calibration and OOD

tvws-selection-calibration@1 contract v1 4 artefacts

Which epoch shipped, how confident is it, and how does it behave off-distribution?

Contract version 1 Valid tier: Over the air tvws-retrain@1tvws-retrain-metrics@1

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
ABSENT from all none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
4 / 4d:8dd29d376164e84a · d:a8f6e162d67b4ec7 · d:808c8c9bc84991ab +1 moreWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
ABSENT from all none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
4 / 4scripts/sensing-select-calibrate-ood.py · external/tvws-sensing/scripts/train_model.pyWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
ABSENT from all none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
4 / 442 · 7A seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
2 / 48eca022f09b897a579ef74225881343949a1a4a53b880edabf82225aacbb34eeA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
Spearman: synthetic val accuracy vs real records
spearman_synthetic_val_acc_vs_real_records
none
scalar
Higher is better.scalar

A single stated value, not a summary of a distribution.

one candidate epoch4 2 / 4

Pass/fail thresholds — and where they came from

Threshold source

None. Selection and calibration are reported, not gated.

Threshold version none

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking the reported score is not a selection score fact selection.held_out_from_selection · exactly true

    Was the set that produced this number also used to choose the checkpoint?

    Why it matters
    The four verified multiplexes were used to SELECT the shipped epoch, so their score for that checkpoint is optimistically biased.
    Blocks the claim
    that this number is an unbiased test score for the shipped checkpoint
    On fail
    This is a selection score, not a test score.
    On absent
    The artefact does not state that the scoring set was held out of selection. Treat the number as a selection score until it does.
  • Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true

    Was the receive antenna specified for the band these captures were taken in?

    Why it matters
    Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
    Blocks the claim
    any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
    On fail
    The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
    On absent
    This artefact does not state whether the antenna matched the band. Absent, not confirmed.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 2 / 4

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
wilson

Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.

What counts as a repetition
An independent training seed run through the same selection procedure.
Independent seeds are required, recorded IN the artefact and not inferred from a file name.
Minimum sample count
4

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • quoting a selection score as a test score

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • There are ZERO real vacant captures, so no false-positive rate on real RF can be measured at all.
  • Calibration temperature is fitted on held-out SYNTHETIC data; applying it to real captures is an extrapolation.

Caveats carried verbatim from the artefacts (6)

  • honesty
    The ota_trajectory is a MEASUREMENT taken per epoch. If a checkpoint is then selected using it, the real set has become a selection set and is no longer an unbiased test of that checkpoint. Report both, and never quote the selection score as a test score.
    2 file(s): tvws-retrain-seed42-2026-08-03.json, tvws-retrain-seed7-2026-08-03.json
  • honesty.calibration_temperature_is_synthetic
    Temperature is fitted on held-out SYNTHETIC data because that is the only labelled data with enough records to fit anything. Applying it to real captures is an extrapolation. The real-RF ECE is reported before and after so the extrapolation is visible.
    2 file(s): tvws-retrain-metrics-seed42-2026-08-03.json, tvws-retrain-metrics-seed7-2026-08-03.json
  • honesty.selection_set_is_not_a_test_set
    The four verified multiplexes were used to SELECT the shipped epoch. Their score for that checkpoint is a selection score and is optimistically biased. The leave-one-channel-out total is the closest to unbiased that four channels allow: each held-out channel is scored by a checkpoint selected without it.
    2 file(s): tvws-retrain-metrics-seed42-2026-08-03.json, tvws-retrain-metrics-seed7-2026-08-03.json
  • honesty.what_no_real_number_here_can_prove
    One site, one afternoon, one mismatched 860-930 MHz antenna at 470-700 MHz, four channels, every capture occupied. There are ZERO real vacant captures, so no false-positive rate on real RF can be measured at all, and a model answering 'occupied' unconditionally scores 80/80 here. Nothing here generalises to another site or another receiver.
    2 file(s): tvws-retrain-metrics-seed42-2026-08-03.json, tvws-retrain-metrics-seed7-2026-08-03.json
  • selection.spearman_note
    Negative or near-zero means selecting by synthetic validation accuracy is uninformative about real RF for this run. It is one run; treat it as a consistency check on the 2026-07-30 result (-0.1087 and -0.316), not as an independent one.
    2 file(s): tvws-retrain-metrics-seed42-2026-08-03.json, tvws-retrain-metrics-seed7-2026-08-03.json
  • synthetic_test.note
    Held-out split of the SAME generator that produced the training set. It cannot detect a leak it shares.
    2 file(s): tvws-retrain-seed42-2026-08-03.json, tvws-retrain-seed7-2026-08-03.json

Reproduction

  • Command withheld TVWS retrain run 2 artefact(s)
    Producer external/tvws-sensing/scripts/train_model.py verified

    Trains the sensing CNN for a retrain round and writes the run record.

    The artefacts name it in their own harness field.

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • A full training run on a GPU, per seed.
    • The two artefacts differ by seed (42 and 7). Two seeds of one training recipe are a variance probe, not two measurements of one thing.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.

  • Command offered TVWS selection / calibration / OOD 2 artefact(s)
    Producer scripts/sensing-select-calibrate-ood.py verified

    Selects a checkpoint, calibrates it, scores OOD behaviour and re-scores against the OTA captures.

    The artefacts name it in harness AND record their own invocation under args.

    python scripts/sensing-select-calibrate-ood.py --run-json benchmarks/reports/tvws-retrain-seed42-2026-08-03.json --checkpoint-dir <dir> --baseline-checkpoint external/tvws-sensing/models/checkpoints/best.pt --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --ota-records 20 --samples-per-class 2000 --seed 42 --confidence-threshold 0.7

    Assembled from the producer's own argparse (lines 679-701). THIS FAMILY IS THE ONLY ONE IN THE CORPUS THAT RECORDS ITS OWN INVOCATION: open the raw artefact and read the "args" object for the exact values that produced it, including the checkpoint directory this command leaves as a placeholder.

    Preconditions

    • The checkpoint directory named in the artefact's own args.checkpoint_dir. It is outside this repository.
    • The artefact records device "cuda": this ran on a GPU, and the selection set it reports is a SELECTION set, not a test set.

    This family records its own invocation in the artefact under args. Open the raw file for the exact values that produced it, rather than relying on the documented invocation above.

    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a checkpoint directory and the OTA capture set.

RF channel quality of the capture set

RF channel quality of the capture set

rf-channel-quality@1 contract v1 2 artefacts

How much signal is actually in each captured channel, relative to a control floor?

Contract version 1 Valid tier: Lab RF Valid tier: Over the air channel-quality@1detector-scores@unversioned

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
ABSENT from all none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
ABSENT from all none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
1 / 2scripts/analyze-618mhz.pyWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
1 / 2benchmarks/datasets/ota/nairobi-region-2026-07-29Without a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
ABSENT from all none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
ABSENT from all none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
estimated in-window SNR
*.est_in_window_snr_db
dB
scalar
Higher is better.scalar

A single stated value, not a summary of a distribution.

one channel, estimated against a scalar control floor20 1 / 2

Pass/fail thresholds — and where they came from

Threshold source

None. This suite characterises the capture set; it does not grade it.

Threshold version none

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true

    Was the receive antenna specified for the band these captures were taken in?

    Why it matters
    Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
    Blocks the claim
    any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
    On fail
    The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
    On absent
    This artefact does not state whether the antenna matched the band. Absent, not confirmed.

Statistical method

Central tendency
median_iqr

Median with the interquartile range (Tukey hinges). The default here because latency and accuracy distributions in this corpus are small-n and not symmetric, and a mean hides that.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 1 / 2

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
none

No proportion interval is defined for this suite.

What counts as a repetition
A recapture at the same site and gain with control IQ retained.
Minimum sample count
20

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • treating the estimated SNR as a measured SNR

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • In-window SNR figures are ESTIMATES: no control IQ was retained, only a scalar floor.
  • Receive-only captures. Nothing was transmitted.

Caveats carried verbatim from the artefacts (6)

  • caveats[0]
    in-window SNR figures are ESTIMATES: no control IQ was retained, only the scalar -59.9 dBFS floor. They assume in-capture noise == that floor at the same port and gain.
    1 file(s): ch39-618mhz-quality-2026-07-31.json
  • caveats[1]
    captures came through a SenseCAP LoRa 860-930 MHz antenna used at 470-700 MHz. Every spectral claim carries that.
    1 file(s): ch39-618mhz-quality-2026-07-31.json
  • caveats[2]
    610 and 626 MHz were never captured, so the doubly-adjacent hypothesis cannot be settled from this dataset at all.
    1 file(s): ch39-618mhz-quality-2026-07-31.json
  • caveats[3]
    n=20 records, one site, one antenna, five channels, all occupied.
    1 file(s): ch39-618mhz-quality-2026-07-31.json
  • caveats[4]
    receive-only RX2 captures. Nothing was transmitted.
    1 file(s): ch39-618mhz-quality-2026-07-31.json
  • target
    uhf39_618MHz — never classified DVB-T at any converged epoch
    1 file(s): ch39-618mhz-quality-2026-07-31.json

Reproduction

  • Command offered Channel quality capture 1 artefact(s)
    Producer scripts/analyze-618mhz.py verified

    Analyses the 618 MHz capture: quality stages, a synthetic SNR ladder and a bandlimit sweep.

    The artefact names it in its own harness field.

    python scripts/analyze-618mhz.py --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --records 20 --confidence-threshold 0.7

    Assembled from the producer's own argparse (lines 511-527). It is the producer's documented invocation, not a transcript of the run.

    Preconditions

    • The capture directory. The default in the script is /ota, the in-container mount.
    • This is ONE channel at ONE site on ONE day. It supports no statement about any other channel, site or day.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It reads an IQ capture that is not in the repository.

  • Command offered Detector scores 1 artefact(s)
    Producer scripts/score-detectors.py verified

    Runs the detector scorer and writes the artefact with a full provenance envelope.

    The 2026-08-04 artefact names it in its own harness field.

    python scripts/score-detectors.py --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --synth-records 100 --seed0 20260803

    Assembled from the producer's own argparse defaults, so it is the documented invocation rather than a transcript of any particular run. The 2026-08-04 artefact separately records its own invocation under args, which is where the exact values that produced it are readable.

    Preconditions

    • The capture set at benchmarks/datasets/ota/nairobi-region-2026-07-29. It exists in this tree.
    • CPU only, about 13 seconds, no GPU and no running service. It is safe to run while the demo is up.
    • The captures came through an antenna that is not matched to the band, and every record in the set is labelled occupied. Re-running measures the detector, never those two facts.

    This family records its own invocation in the artefact under args. Open the raw file for the exact values that produced it, rather than relying on the documented invocation above.

    Continuous integration No CI job

    No workflow in .github/workflows runs this producer, and nothing about the producer stops one: it is a CPU-only numpy job that finishes in about 13 seconds. What it needs is the 5 IQ captures under benchmarks/datasets/ota/, roughly 400 MB that this repository does not carry, so a GitHub runner has nothing to score.

cuPHY LDPC decode latency

cuPHY LDPC decode latency

cuphy-ldpc-decode@1 contract v1 2 artefacts

How long does ONE cuPHY LDPC decode take, timed on the decoder’s own CUDA stream?

Contract version 1 Valid tier: Hardware in loop cuphy-per-slot@2cuphy-per-slot@3

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
2 / 2BG1/Z384/81cb/rate0.5Without a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
2 / 21000 timed decodesA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
2 / 220 · 50Discarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
2 / 2pyaerial LdpcDecoder + cudaEvent per callWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
ABSENT from all none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
ABSENT from all none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
ABSENT from all none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
cuPHY decode p50
cuphy_decode.p50
us
p50
Lower is better.raw_sample_percentile

A percentile computed over the RETAINED RAW SAMPLES. This is the only aggregation from which a p99 or p99.9 is legitimate.

one cuPHY LDPC decode over 81 code blocks100 2 / 2
Secondary
cuPHY decode p99
cuphy_decode.p99
us
p99
Lower is better.raw_sample_percentile

A percentile computed over the RETAINED RAW SAMPLES. This is the only aggregation from which a p99 or p99.9 is legitimate.

one cuPHY LDPC decode over 81 code blocks100 2 / 2
Secondary
decodes over the slot budget
cuphy_decode.deadline_misses
decodes
count
Lower is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one cuPHY LDPC decode over 81 code blocks100 2 / 2

Pass/fail thresholds — and where they came from

Threshold source

slot_budget_us in the artefact (500 us), which is the producer’s stated budget, not a 3GPP requirement.

Threshold version artefact-slot-budget-500us

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking the measurement covers a full slot fact scope.is_a_full_slot · exactly true

    Does this number cover a whole 5G slot, or only the LDPC decode?

    Why it matters
    A slot also carries control channels, DMRS and the rest of the PUSCH/PDSCH chain. A deadline conclusion drawn from the decoder alone is a lower bound, not a slot result.
    Blocks the claim
    that the platform holds a full-slot L1 deadline
    On fail
    LDPC decode only. This is a microbenchmark and a LOWER BOUND on real slot overrun — it cannot support a full-slot claim.
    On absent
    The artefact does not state whether this covers a full slot.
  • Blocking the device is recorded in the artefact fact device.stated · exactly true

    Does the file say which GPU produced the number?

    Why it matters
    Schema v2 carries no device block, so its number cannot be attributed to a machine and cannot be ranked against one that can.
    Blocks the claim
    attributing this number to a particular GPU
    On fail
    No device block: the GPU behind this number is unknowable from the file.
    On absent
    The artefact does not record whether a device block is present.

Statistical method

Central tendency
median_iqr

Median with the interquartile range (Tukey hinges). The default here because latency and accuracy distributions in this corpus are small-n and not symmetric, and a mean hides that.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
raw_samples

The contract permits a percentile, and only from RETAINED RAW SAMPLES. The sample count still gates which percentile is expressible, and a run that retained no samples cannot supply one however the contract is written — see the reality check beside this.

Raw samples retained by 2 / 2

All 2 artefact(s) retain a raw sample distribution, so the contract's permission is backed by the evidence.

Proportion interval
none

No proportion interval is defined for this suite.

What counts as a repetition
A fresh process on the same device, same image, same code configuration and same warmup, at the same repetition count.
Minimum sample count
100

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • reading a throughput claim off a random-LLR run: early termination never fires, so it is the WORST case by construction

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • LDPC decode only — no MAC, no scheduler, no fronthaul, no cell, no OTA.
  • Random Gaussian LLRs: correct for deadline analysis, wrong for throughput claims.

Caveats carried verbatim from the artefacts (5)

  • llr_mode_note
    random: Gaussian LLRs, no valid codeword, early termination never fires -> full iteration count -> WORST case. Correct for deadline analysis, wrong for throughput claims.
    2 file(s): cuphy-per-slot-20260729-005456.json, cuphy-per-slot-20260730-080736.json
  • measures
    GPU elapsed time of ONE cuPHY LDPC decode over 81 code blocks, timed on the decoder's own CUDA stream.
    2 file(s): cuphy-per-slot-20260729-005456.json, cuphy-per-slot-20260730-080736.json
  • scope_warning
    LDPC decode only. A 5G slot also carries control channels, DMRS and the rest of the PUSCH/PDSCH chain, so a miss counted here is a miss BY THE DECODER ALONE and is a LOWER BOUND on real slot overrun.
    2 file(s): cuphy-per-slot-20260729-005456.json, cuphy-per-slot-20260730-080736.json
  • series_note
    cuphy_decode is cuPHY's LDPC decode. pyaerial_decode is the same work wrapped in pyAerial's public decode(), which re-copies the LLR block and the output on every call. Quote cuphy_decode for PHY latency; the difference is Python-binding overhead the real L1 does not pay.
    2 file(s): cuphy-per-slot-20260729-005456.json, cuphy-per-slot-20260730-080736.json
  • device.co_tenancy_visibility
    namespace-limited: measured from inside the Aerial container, so host processes are INVISIBLE here. An empty list does NOT mean the GPU was idle — verify from the host with nvidia-smi before claiming a clean run.
    1 file(s): cuphy-per-slot-20260730-080736.json

Reproduction

  • Command withheld cuPHY per-slot latency 2 artefact(s)
    Producer scripts/cuphy-per-slot-latency.py verified

    Times one pyaerial LdpcDecoder.decode() call per sample with cudaEvent, inside the Aerial container, and writes the per-slot artefact to benchmarks/reports/.

    The script writes to "benchmarks/reports/" (line 293) and its own docstring documents the percentile gate this family carries.

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • A GB10 or H200 with the Aerial cuBB container image available locally.
    • Exclusive use of the GPU for the duration: the harness re-execs itself into the Aerial container and times decodes.
    • The documented invocation is ./scripts/cuphy-per-slot-latency.py --decodes 1000 (from the script header).
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.

cuPHY LDPC under a capped MPS AI tenant

cuPHY LDPC under a capped MPS AI tenant

mps-cotenancy@1 contract v1 1 artefacts

What happens to cuPHY LDPC invocation-average latency while a capped MPS-managed AI client loads the same GPU?

Contract version 1 Valid tier: Hardware in loop mps-sweep@3

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
1 / 1caps=[100,50,25] sizes=[8192]Without a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
1 / 112 invocations/level x 100 inner decodes, 240s AI loadA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
ABSENT from all none statedWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
ABSENT from all none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
ABSENT from all none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
ABSENT from all none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
p50 of the invocation averages
*.p50
us
p50
Lower is better.invocation_average_percentile

A percentile computed over INVOCATION AVERAGES. Each sample is already a mean, so the tail of the underlying distribution has been averaged away before the percentile was taken. No tail claim survives this.

one cuphy_ex_ldpc invocation, itself already an average across 100 inner decodes12 1 / 1
Secondary
block errors
*.block_errors
blocks
count
Lower is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one decoded code block1 1 / 1
Safety
uncontrolled tenants on the GPU
unmanaged_tenants_at_start
processes
count
Lower is better.count

A count of events. Whole numbers; no interpolation is meaningful.

one process holding a CUDA context outside the MPS server1 1 / 1

Pass/fail thresholds — and where they came from

Threshold source

slot_budget_us in the artefact, applied to invocation averages. Not a deadline guarantee.

Threshold version artefact-slot-budget-500us

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking every GPU client was MPS-managed fact mps.all_clients_managed · exactly true

    Was the GPU free of processes outside the MPS server for the whole run?

    Why it matters
    Uncontrolled default-mode tenants share the same SMs. With any present, the measurement is of the whole machine, not of the capped client.
    Blocks the claim
    that this run demonstrates isolation between the RAN and AI clients
    On fail
    Uncontrolled tenants were on the GPU for this run. No isolation claim is available from it, and the latency it reports is not attributable to the capped client.
    On absent
    The artefact does not census the processes on the GPU.
  • Blocking the GPU was verified clean at the end of the run fact mps.gpu_clean_at_end · exactly true

    Was the GPU free of uncontrolled tenants when the run finished?

    Why it matters
    When end_of_run_gpu_check is null, an empty at-end tenant list is not-collected. Empty is not clean.
    Blocks the claim
    that the GPU was clean for the duration of this run
    On fail
    The end-of-run check ran and did not come back clean.
    On absent
    The end-of-run co-tenancy check did not run, so the empty at-end tenant list proves nothing.
  • Blocking every planned cell was measured fact mps.partial_run · exactly false

    Did the sweep complete the cells it planned?

    Why it matters
    Caps that were never measured are ABSENT, not passing. No envelope may be read across a partial sweep.
    Blocks the claim
    a largest-passing-cap envelope
    On fail
    PARTIAL SWEEP. The caps that were never measured are absent, not passing, so no envelope may be read from this document.
    On absent
    The artefact does not state whether the sweep completed.

Statistical method

Central tendency
median_iqr

Median with the interquartile range (Tukey hinges). The default here because latency and accuracy distributions in this corpus are small-n and not symmetric, and a mean hides that.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 0 / 1

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
none

No proportion interval is defined for this suite.

What counts as a repetition
A fresh run on the same host with the same cap set, load sizes, invocation count and AI load duration — and a clean GPU census at both ends.
Minimum sample count
12

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • any tail percentile: each sample is already an average across 100 inner decodes
  • reading an isolation guarantee from a provisioning knob

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • CUDA_MPS_ACTIVE_THREAD_PERCENTAGE PROVISIONS but does not RESERVE: kernels from different clients may still execute on the same SM.
  • CUDA_MPS_CLIENT_PRIORITY is a documented HINT, not a guarantee.
  • cuPHY LDPC only: no MAC, scheduler, fronthaul, cell or OTA.

Caveats carried verbatim from the artefacts (6)

  • cap_semantics
    CUDA_MPS_ACTIVE_THREAD_PERCENTAGE PROVISIONS but does not RESERVE: the documentation is explicit that kernels from different clients may still execute on the same SM. cuPHY is 100%; the AI client is capped.
    1 file(s): mps-sweep-2026-08-03-livebox.json
  • partial_run_note
    PARTIAL SWEEP: 2 of 4 planned cells completed. The caps that were never measured are absent, not passing, so no envelope may be read from this document and largest_observed_passing_ai_cap_percent is withheld.
    1 file(s): mps-sweep-2026-08-03-livebox.json
  • priority_semantics
    CUDA_MPS_CLIENT_PRIORITY is a documented HINT, not a guarantee. A null result for this arm is expected, not anomalous.
    1 file(s): mps-sweep-2026-08-03-livebox.json
  • run_error
    AI load client exited before MPS join: container=ulap-cuphy-ai-2936710-50-8192; state=exited; exit_code=1; logs_tail='Traceback (most recent call last):\n File "<string>", line 4, in <module>\ntorch.AcceleratorError: CUDA error: CUDA-capable device(s) is/are busy or unavailable\nSearch for `cudaErrorDevicesUnavailable\' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.\nFor more detailed error information, run with CUDA_LOG_FILE=stderr'
    1 file(s): mps-sweep-2026-08-03-livebox.json
  • scope
    cuPHY LDPC invocation-average latency under a capped MPS-managed PyTorch matmul client; no MAC, scheduler, fronthaul, cell, or OTA
    1 file(s): mps-sweep-2026-08-03-livebox.json
  • timing_warning
    Each sample is cuphy_ex_ldpc's average across 100 inner decodes. This is not a per-slot latency distribution and cannot support p99.9.
    1 file(s): mps-sweep-2026-08-03-livebox.json

Reproduction

  • Command withheld MPS co-tenancy sweep 1 artefact(s)
    Producer scripts/cuphy-mps-sweep.py verified

    Sweeps an MPS active-thread cap on an AI client while cuPHY LDPC runs uncapped.

    Writes to "benchmarks/reports/" and emits partial_run / reportable_clean_mps_run.

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • A scheduled clean-GPU window. The producer's own docstring says: "Use a scheduled clean-GPU window, or move every competing client into the same MPS server, before removing the gate."
    • Both managed clients must join the same private MPS server, which means starting an MPS control daemon.
    • The Aerial cuBB image and the AI load image (amini/wg3-sr-worker:latest) present locally.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.

GPU partitioning mechanism probe

GPU partitioning mechanism probe

gpu-partition-mechanism@1 contract v1 2 artefacts

Does the partitioning mechanism itself do what its documentation says?

Contract version 1 Valid tier: Hardware in loop mps-cap@unversionedmig-partition@unversioned

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
ABSENT from all none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
ABSENT from all none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
1 / 2scripts/mps-cap-test.shWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
ABSENT from all none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
ABSENT from all none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
ABSENT from all none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
fp16 matmul throughput
*.tflops
TFLOPS
scalar
Higher is better.scalar

A single stated value, not a summary of a distribution.

one 4096-square fp16 matmul run1 1 / 2

Pass/fail thresholds — and where they came from

Threshold source

None. The probe characterises a mechanism; it grades nothing.

Threshold version none

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

This contract declares no gate. It reports; it does not grade. That is not a pass.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 0 / 2

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
none

No proportion interval is defined for this suite.

What counts as a repetition
A re-run of a harness that does not currently exist.
Minimum sample count
1

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • reading a concurrency result from sequential runs

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • Concurrent non-interference was NOT demonstrated: two instances produced no output when all three ran at once.
  • The host is not this box.

Caveats carried verbatim from the artefacts (4)

  • NOT_measured.cuphy_ldpc_inside_a_mig_instance
    not attempted in this window
    1 file(s): mig-h200-2026-08-03.json
  • NOT_measured.three_way_concurrent_isolation
    Instances 2 and 3 produced NO output when all three ran at once, though each works sequentially. The concurrent non-interference claim is therefore UNPROVEN and is not made. Cause not diagnosed; the window had to close to restore a user-facing service.
    1 file(s): mig-h200-2026-08-03.json
  • note
    Unmanaged default-mode tenants were present and untouched; they are why the cuPHY co-tenancy sweep could not isolate. This test scopes the question to the mechanism by putting both clients inside one private MPS server.
    1 file(s): mps-cap-result.json
  • teardown
    compute instances and GPU instances destroyed, MIG mode returned to Disabled, akili trio restarted and verified serving in 56 s with 95772 MiB restored
    1 file(s): mig-h200-2026-08-03.json

Reproduction

  • Command withheld GPU partitioning probe 1 artefact(s)
    Producer scripts/cotenancy/mig-h200-probe.py verified

    Regenerates the observable half of this family and refuses the destructive half unless an operator explicitly asks and the GPU is free.

    Emits benchmarks/reports/mig-h200-observed-<date>.json.

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • The default `--mode observe` is read-only and safe, but it needs ssh to the H200, which this UI cannot offer as a one-click command.
    • `--mode measure` requires a compute-free H200 and an explicit operator flag. It is a privileged one-way device reconfiguration, not a run.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.

  • Command withheld MPS thread-cap enforcement 1 artefact(s)
    Producer scripts/mps-cap-test.sh verified

    Runs both clients inside a private MPS server to test whether an active-thread cap partitions the GB10.

    Line 16 writes benchmarks/reports/mps-cap-result.json by name.

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • It starts its own MPS control daemon (CUDA_MPS_PIPE_DIRECTORY) and quits it on cleanup.
    • A python at $HOME/.venvs/wg3-test/bin/python, or PY= pointing at one.
    • The script states it "adds only its own processes and never touches an existing service" — the exclusive-control clause is still tripped, because an MPS control daemon is a device-wide object.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.

LDPC latency under a stepped AI load

LDPC latency under a stepped AI load

ldpc-step-load@1 contract v1 2 artefacts

How does LDPC decode latency move as an AI tenant’s load steps up on the same GPU?

Contract version 1 Valid tier: Hardware in loop window-clean@bare-array

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
ABSENT from all none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
ABSENT from all none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
1 / 2scripts/cuphy-cotenancy-sweep.pyWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
ABSENT from all none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
ABSENT from all none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
ABSENT from all none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
p90 decode latency at load level
*.p90
us
p90
Lower is better.raw_sample_percentile

A percentile computed over the RETAINED RAW SAMPLES. This is the only aggregation from which a p99 or p99.9 is legitimate.

one LDPC decode at one AI load level15 2 / 2

Pass/fail thresholds — and where they came from

Threshold source

None recorded in the artefact.

Threshold version none

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

This contract declares no gate. It reports; it does not grade. That is not a pass.

Statistical method

Central tendency
median_iqr

Median with the interquartile range (Tukey hinges). The default here because latency and accuracy distributions in this corpus are small-n and not symmetric, and a mean hides that.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
raw_samples

The contract permits a percentile, and only from RETAINED RAW SAMPLES. The sample count still gates which percentile is expressible, and a run that retained no samples cannot supply one however the contract is written — see the reality check beside this.

Raw samples retained by 0 / 2

The contract permits raw-sample percentiles and 2 of 2 artefact(s) retain NO raw sample distribution. For those runs the percentile in the file was computed by the producer and cannot be recomputed, re-checked or extended here; the evaluation raises a blocking statistical finding for each one.

Proportion interval
none

No proportion interval is defined for this suite.

What counts as a repetition
A second sweep on a recorded host with a recorded revision. runA and runB differ only by filename and mtime, which is not identity.
Minimum sample count
15

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • a tail percentile beyond p90 at n=15

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • LDPC decode only: no fronthaul, no MAC, no scheduler, no cell.
  • runA and runB are distinguishable only by filename and mtime.

Caveats carried verbatim from the artefacts (0)

No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.

Reproduction

  • Command withheld LDPC step-load window 2 artefact(s)
    Producer scripts/cuphy-cotenancy-sweep.py verified

    Steps an AI load against cuPHY LDPC and records one object per load level.

    Names window-clean-gpu-runA.json in its own comments (lines 277-279).

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • Exclusive GPU access while a stepped AI load runs against cuPHY.
    • These two artefacts are a BARE JSON ARRAY with no envelope: the producer must be changed to emit a schema_version, a measurement instant and a device block before a re-run is comparable with anything.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.

Synthetic held-out classification accuracy

Synthetic held-out classification accuracy

heldout-accuracy@1 contract v1 6 artefacts

How accurately does the checkpoint classify a held-out split of the SAME generator?

Contract version 1 Valid tier: Synthetic heldout@unversioned

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
6 / 6/repo/models/checkpoints/best.pt · /ck/best.pt · /ma/tvws-sensing/best.ptWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
6 / 6single pass over the held-out splitA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
ABSENT from all none statedWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
ABSENT from all none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
6 / 642A seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
ABSENT from all none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
overall accuracy
overall_accuracy
fraction
fraction
Higher is better.proportion

A proportion of a countable population. A confidence interval is expressible only when the numerator and denominator are whole records.

one synthetic held-out sample200 6 / 6
Secondary
random chance floor
random_chance_floor
fraction
scalar
No direction. This number is not rankable: "better" is undefined for it.scalar

A single stated value, not a summary of a distribution.

the class prior1 6 / 6

Pass/fail thresholds — and where they came from

Threshold source

None. This suite reports; it does not gate.

Threshold version none

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking the endpoint is proven to serve this checkpoint fact model.endpoint_weight_hash_proven · exactly true

    Is there evidence that the service under test served the weights named here?

    Why it matters
    model_label is a human ASSERTION about what the URL serves; checkpoint_sha256 pins a FILE ON DISK. Both can be wrong together if the operator points --checkpoint at one model and --url at another.
    Blocks the claim
    that the deployed service serves the checkpoint named here
    On fail
    The endpoint is not proven to serve this checkpoint.
    On absent
    The service exposes no weight hash, so nothing here can bind the file to the endpoint. Provenance, not proof.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 0 / 6

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
wilson

Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.

What counts as a repetition
A run at an independent seed over a split with the SAME dataset, split and preprocessing hashes. Two files whose names differ is not a repetition.
Independent seeds are required, recorded IN the artefact and not inferred from a file name.
Minimum sample count
200

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • reading this as an OTA result
  • ranking two arms whose dataset, split and preprocessing hashes are not proven identical

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • synthetic-only; NOT OTA; NOT FCC.
  • A held-out split of the SAME generator that produced the training set cannot detect a leak it shares.

Caveats carried verbatim from the artefacts (1)

  • scope
    synthetic-only; NOT OTA; NOT FCC; NOT paper 94.2%
    6 file(s): heldout-committed-8eca022f-2026-07-31.json, heldout-committed-8eca022f-IMAGE-APP-SRC-trap-2026-07-31.json, heldout-committed-8eca022f-prefixgen-17517e57-2026-07-31.json, heldout-v1-828dc46e-gen-9dc4f892-2026-07-31.json, heldout-v1-828dc46e-headgen-c23711ab-2026-07-31.json, heldout-v1-828dc46e-prefixgen-17517e57-2026-07-31.json

Reproduction

  • Command offered Held-out split eval 6 artefact(s)
    Producer external/tvws-sensing/scripts/evaluate_heldout.py verified

    Scores a checkpoint against a freshly generated held-out split.

    Lines 165-166 emit random_chance_floor and held_out_test_samples.

    python external/tvws-sensing/scripts/evaluate_heldout.py --checkpoint models/checkpoints/best.pt --samples-per-class 1800 --seed 42 --snr-min -10 --snr-max 30 --batch-size 32

    Assembled from the producer's own argparse defaults (lines 49-55). It is the producer's documented invocation, NOT a transcript of the invocation that produced these artefacts — none of them record one.

    Preconditions

    • A checkpoint at the --checkpoint path. The corpus records checkpoint paths but no artefact binds a checkpoint sha256 to a served endpoint.
    • The split is REGENERATED from the seed, not read from disk. Re-running with the same seed does not prove the same data: only a dataset hash would, and this producer does not emit one.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a model checkpoint, which is not in the repository.

Sensing generalization across synthetic conditions

Sensing generalization across synthetic conditions

sensing-generalization@1 contract v1 15 artefacts

How does the served sensing model behave as synthetic conditions move away from the training distribution?

Contract version 1 Valid tier: Deployed service Valid tier: Synthetic sensing-generalization@2sensing-generalization@3

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it. It is also served through an endpoint, so the serving seam is part of the measurement.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
ABSENT from all none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
ABSENT from all none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
ABSENT from all none statedWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
1 / 1563c971b4d0fc967ec500675afe13d6765dd3178bWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
ABSENT from all none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
1 / 159dc4f892f5ff34beccb01ad2d6641abcf59765bc4df0e645fc11597ecda19edeFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
13 / 154242 · 9137A seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
1 / 158eca022f09b897a579ef74225881343949a1a4a53b880edabf82225aacbb34eeA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
15 / 15http://localhost:8103/api/v1/sense · http://localhost:8002/api/v1/senseFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
per-condition accuracy
*.accuracy
fraction
fraction
Higher is better.proportion

A proportion of a countable population. A confidence interval is expressible only when the numerator and denominator are whole records.

one synthetic sample at one condition, classified through the service200 15 / 15

Pass/fail thresholds — and where they came from

Threshold source

random_floor in the artefact — the chance level, not a pass mark.

Threshold version artefact-random-floor

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking the endpoint is proven to serve this checkpoint fact model.endpoint_weight_hash_proven · exactly true

    Is there evidence that the service under test served the weights named here?

    Why it matters
    model_label is a human ASSERTION about what the URL serves; checkpoint_sha256 pins a FILE ON DISK. Both can be wrong together if the operator points --checkpoint at one model and --url at another.
    Blocks the claim
    that the deployed service serves the checkpoint named here
    On fail
    The endpoint is not proven to serve this checkpoint.
    On absent
    The service exposes no weight hash, so nothing here can bind the file to the endpoint. Provenance, not proof.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 0 / 15

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
wilson

Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.

What counts as a repetition
The same arm at an independent seed against the same endpoint and generator hash. 0.125 is seed 4242 only; seed 9137 is 0.050.
Independent seeds are required, recorded IN the artefact and not inferred from a file name.
Minimum sample count
200

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • a single number averaged across the conditions
  • quoting one seed as though it were the measurement

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • This is NOT over-the-air validation and does not substitute for it. A model can pass every condition here and still collapse on air.

Caveats carried verbatim from the artefacts (4)

  • class_excluded
    Unknown (5) — not a generatable ground truth
    15 file(s): sensing-generalization-new-control-seed4242.json, sensing-generalization-new-fixA-seed4242.json, sensing-generalization-new-fixB-seed4242.json, sensing-generalization-new-legacy-seed4242.json, sensing-generalization-old-control-seed4242.json, sensing-generalization-old-control-seed9137.json, sensing-generalization-old-fixA-seed4242.json, sensing-generalization-old-fixA-seed9137.json, sensing-generalization-old-fixB-DEPLOYED-seed4242.json, sensing-generalization-old-fixB-seed4242.json, sensing-generalization-old-fixB-seed9137.json, sensing-generalization-old-legacy-seed4242.json, sensing-generalization-old-legacy-seed9137.json, sensing-generalization-runA.json, sensing-generalization-runB.json
  • measured_through
    live service POST /api/v1/sense (deployed weights)
    15 file(s): sensing-generalization-new-control-seed4242.json, sensing-generalization-new-fixA-seed4242.json, sensing-generalization-new-fixB-seed4242.json, sensing-generalization-new-legacy-seed4242.json, sensing-generalization-old-control-seed4242.json, sensing-generalization-old-control-seed9137.json, sensing-generalization-old-fixA-seed4242.json, sensing-generalization-old-fixA-seed9137.json, sensing-generalization-old-fixB-DEPLOYED-seed4242.json, sensing-generalization-old-fixB-seed4242.json, sensing-generalization-old-fixB-seed9137.json, sensing-generalization-old-legacy-seed4242.json, sensing-generalization-old-legacy-seed9137.json, sensing-generalization-runA.json, sensing-generalization-runB.json
  • ota_caveat
    This is NOT over-the-air validation and does not substitute for it. Real OTA carries front-end nonlinearity, phase noise, ADC quantisation and real interference that no generator reproduces. A model can pass every condition here and still collapse on air. The OTA holdout in this repo is a 4 KB README stub with no IQ data.
    15 file(s): sensing-generalization-new-control-seed4242.json, sensing-generalization-new-fixA-seed4242.json, sensing-generalization-new-fixB-seed4242.json, sensing-generalization-new-legacy-seed4242.json, sensing-generalization-old-control-seed4242.json, sensing-generalization-old-control-seed9137.json, sensing-generalization-old-fixA-seed4242.json, sensing-generalization-old-fixA-seed9137.json, sensing-generalization-old-fixB-DEPLOYED-seed4242.json, sensing-generalization-old-fixB-seed4242.json, sensing-generalization-old-fixB-seed9137.json, sensing-generalization-old-legacy-seed4242.json, sensing-generalization-old-legacy-seed9137.json, sensing-generalization-runA.json, sensing-generalization-runB.json
  • model_identity.note
    model_label is a human ASSERTION about what --url serves. checkpoint_sha256 pins a FILE ON DISK — it does NOT prove the endpoint is serving that file; nothing here can, because the service exposes no weight hash. Both can therefore be wrong together if the operator points --checkpoint at one model and --url at another. Treat them as provenance, not proof, and keep candidate models on their own port so the endpoint corroborates the label. generator_sha256 IS exact: this project generates its training data in-process, so that hash pins the training distribution itself. If self_describing is false the arm is knowable only from the filename or an external manifest — see benchmarks/reports/SENSING-ARMS-MANIFEST.md.
    1 file(s): sensing-generalization-old-fixB-DEPLOYED-seed4242.json

Reproduction

  • Command withheld Sensing generalization arm 15 artefact(s)
    Producer scripts/sensing-generalization.py verified

    Generates per-class samples and scores them through the deployed sensing endpoint.

    Argument parser carries --per-class, --seed, --output, --url.

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT drive an endpoint that is serving the live demo.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • The sensing service reachable at --url. On this host that endpoint serves the live demo, so a sweep against it is load on a running service.
    • The arm identity is only in the FILE NAME for this family (control / fixA / fixB / legacy / old / new). The producer must be changed to write the arm and the checkpoint sha256 INTO the artefact before two arms can be compared at all.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs the sensing service running and a model checkpoint that is not in the repository.

Training history

Training history

training-history@1 contract v1 3 artefacts

What did the training curve do, and which epoch was best on the validation split?

Contract version 1 Valid tier: Synthetic training-history@unversioned

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
ABSENT from all none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
ABSENT from all none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
ABSENT from all none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
ABSENT from all none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
ABSENT from all none statedWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
ABSENT from all none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
ABSENT from all none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
ABSENT from all none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
ABSENT from all none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
ABSENT from all none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
ABSENT from all none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
2 / 342 · 7A seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
ABSENT from all none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
ABSENT from all none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
best validation accuracy
best_val_accuracy
fraction
scalar
Higher is better.proportion

A proportion of a countable population. A confidence interval is expressible only when the numerator and denominator are whole records.

one synthetic validation sample at the best epoch1 3 / 3

Pass/fail thresholds — and where they came from

Threshold source

None.

Threshold version none

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking the reported number is not a selection score fact selection.held_out_from_selection · exactly true

    Was the split that produced this number also used to choose the epoch?

    Why it matters
    best_val_accuracy is the maximum over the epochs it selected from. It is optimistically biased by construction.
    Blocks the claim
    that this is an unbiased estimate of the checkpoint’s accuracy
    On fail
    This is a selection score, not a test score.
    On absent
    The artefact does not state that the scoring split was held out of selection. Treat it as a selection score.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 3 / 3

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
wilson

Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.

What counts as a repetition
An independent training seed with the same generator hash.
Independent seeds are required, recorded IN the artefact and not inferred from a file name.
Minimum sample count
1

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • quoting best_val_accuracy as a test result

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • A validation maximum is a selection score, not a test score.
  • The arm identity exists only in the filename.

Caveats carried verbatim from the artefacts (0)

No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.

Reproduction

  • Command withheld Training history 3 artefact(s)
    Producer external/tvws-sensing/scripts/train_model.py verified

    Trains the sensing CNN and dumps the per-epoch history.

    Line 380 writes "synthetic_history": history.to_dict().

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • A full training run on a GPU. These are training curves, not a benchmark: a validation accuracy from a training loop is a SELECTION signal, not a test result.
    • The producer must be changed to write a measurement instant and a dataset hash into the history file; it currently writes neither.
    Continuous integration No CI job

    No workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.

Bench harness run

Bench harness run

bench-harness-run@1 contract v1 no artefacts in this snapshot

What was the environment, revision and dataset set that a harness run executed against?

Contract version 1 Valid tier: Deployed service bench-run-metadata@unversioned No artefact in this snapshot

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
no artefacts none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
no artefacts none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
no artefacts none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
no artefacts none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
no artefacts none statedWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
no artefacts none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
no artefacts none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
no artefacts none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
no artefacts none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
no artefacts none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
no artefacts none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
no artefacts none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
no artefacts none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
no artefacts none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
tests recorded in this run
tests_recorded
tests
count
No direction. This number is not rankable: "better" is undefined for it.count

A count of events. Whole numbers; no interpolation is meaningful.

one graded test1 no artefacts

Pass/fail thresholds — and where they came from

Threshold source

Per-test thresholds live in the individual tests, not in this metadata record.

Threshold version per-test

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

This contract declares no gate. It reports; it does not grade. That is not a pass.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
not_expressible

PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.

Raw samples retained by 0 / 0

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
none

No proportion interval is defined for this suite.

What counts as a repetition
A run at the same git sha, image digests and dataset hashes, executing the SAME test set.
Minimum sample count
1

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • an aggregate pass/fail across a heterogeneous test set

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • The two harness runs on disk contain DIFFERENT test sets; no aggregate pass/fail is expressible across them.
  • started_at is a run start and is never a measurement instant.

Caveats carried verbatim from the artefacts (0)

No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.

Reproduction

  • Command withheld Bench harness run 0 artefact(s)
    Producer scripts/run-bench.sh verified

    Creates the run directory, runs the harness and post-processes the report.

    Line 64 creates reports/run-$RUN_ID; line 225 runs benchmarks/reports/make_report.py.

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT drive an endpoint that is serving the live demo.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • The demo stack up. The harness drives the running services, which on this host are serving the live demo.
    • CI runs a reduced form of this on GitHub runners with TVWS_FORCE_CPU=1 and SR_USE_CPU=true — see .github/workflows/wg3-ci.yml, job e2e-bench-quick.
    Continuous integration .github/workflows/wg3-ci.yml · e2e-bench-quick

    Step ./scripts/e2e/test-bench-quick.sh — benchmarks/reports/run-*/ uploaded as the bench-reports-${{ github.run_id }} artifact. It runs the harness on GitHub runners with TVWS_FORCE_CPU=1 and SR_USE_CPU=true, so it produces service-tier evidence and cannot produce any GPU, lab-RF or over-the-air evidence.

Bench harness per-test result

Bench harness per-test result

bench-per-test@1 contract v1 no artefacts in this snapshot

What did one graded test measure, and did it record a status at all?

Contract version 1 Valid tier: Deployed service per-test-result@unversioned No artefact in this snapshot

Workload and dataset identity — as the artefacts state it

Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.

Each identity field, how many of this suite's artefacts state it, the distinct values observed, and what its absence costs.
FieldStated inDistinct valuesIf absent
workload id
facts["workload.id"]
no artefacts none statedWithout a workload id, two runs are not known to have answered the same question.
workload hash
identity.workloadHash
no artefacts none statedWithout a workload hash, a shared workload is asserted rather than proven.
measurement protocol
facts["protocol.id"]
no artefacts none statedA different repetition or timing protocol produces a different quantity, however similar the field names are.
warmup discarded
facts["protocol.warmup_discarded"]
no artefacts none statedDiscarding a different number of warmup iterations changes the distribution that was measured.
harness
harness
no artefacts none statedWithout the producing harness named in the artefact, the measurement path is unknown.
harness revision
harnessRevision
no artefacts none statedWithout a revision, a difference between runs may be a change in the measurement rather than in the thing measured.
dataset id
facts["dataset.id"]
no artefacts none statedWithout a dataset id the inputs are unnamed.
dataset hash
identity.datasetHash
no artefacts none statedThe same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs.
split id
facts["dataset.split_id"]
no artefacts none statedA different split is a different test. Accuracy across two splits is not a comparison.
preprocessing id
facts["preprocessing.id"]
no artefacts none statedPreprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement.
generator sha256
identity.generatorSha256
no artefacts none statedFor generated data the generator IS the dataset. Without its hash, a regenerated split is a different split.
seed
identity.seed
no artefacts none statedA seed without a generator hash reproduces a procedure, not a dataset.
checkpoint sha256
identity.checkpointSha256
no artefacts none statedA checkpoint hash pins a FILE. It does not prove that any endpoint served those weights.
serving endpoint
identity.endpoint
no artefacts none statedFor a served model the endpoint is part of the measurement: a different endpoint may be a different model.

Metric definitions and direction

Each metric this contract declares: its role, key, unit, which direction is better, how it was aggregated, what one sample is, the minimum sample count, and how many of this suite's artefacts carry it.
MetricUnitDirectionAggregationOne sample isMin nCarried by
Primary
the test’s own primary metric
primary_metric
none
scalar
No direction. This number is not rankable: "better" is undefined for it.scalar

A single stated value, not a summary of a distribution.

defined by the individual test1 no artefacts

Pass/fail thresholds — and where they came from

Threshold source

The individual test’s own threshold, recorded beside its status.

Threshold version per-test

Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.

  • Blocking opt-in components are not counted as failures fact test.all_components_engaged · exactly true

    Were the components in this test actually asked to run?

    Why it matters
    In sr_model_matrix two of the four entries are opt-in and were never driven this run — ESPCN runs only in mode=fast, and the quality worker needs a POST to reload. Counting them as failures is the same defect as counting a failure as a pass, inverted.
    Blocks the claim
    reading this test’s entry count as a failure count
    On fail
    Components in this test were never engaged; their entries are absent, not failed.
    On absent
    The artefact does not record whether every component was engaged, so entries that did not run cannot be told apart from entries that failed.

Statistical method

Central tendency
none

No central-tendency summary is defined for this suite.

Mean ± SD
Not permitted

Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.

Percentile source
approved_histogram

Percentiles may be read only off the named, approved histogram.

Raw samples retained by 0 / 0

The contract does not source percentiles from raw samples, so retention is not required of these artefacts.

Proportion interval
none

No proportion interval is defined for this suite.

What counts as a repetition
The same test id in a run at the same revision and dataset hashes.
Minimum sample count
1

Below this the number does not mean what its name says.

Treatments this suite's numbers can never support

  • reading a null metric as zero
  • reading a percentile with no source field as though it came from raw samples

Known limitations

Scope limits — claims this suite can never support, whatever its numbers say

  • A result with no status is ABSENT: neither a pass nor a failure.
  • A pass recorded against a NaN metric is a pass-by-absence, not a real pass.

Caveats carried verbatim from the artefacts (0)

No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.

Reproduction

  • Command withheld Bench harness per-test result 0 artefact(s)
    Producer benchmarks/bench_harness.py verified

    Writes one result.json per test under reports/run-<id>/per-test/<test>/.

    The per-test layout is created by the harness scripts/run-bench.sh invokes.

    No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT drive an endpoint that is serving the live demo.

    The producer is named above. The command is not printed here because the safety policy clause below was tripped.

    Preconditions

    • The demo stack up, as for the parent run.
    • A per-test result with no status field reads as ABSENT here, and the harness must be changed to always write a status before absence and failure can be told apart.
    Continuous integration .github/workflows/wg3-ci.yml · e2e-bench-quick

    Step ./scripts/e2e/test-bench-quick.sh — benchmarks/reports/run-*/ uploaded as the bench-reports-${{ github.run_id }} artifact. It runs the harness on GitHub runners with TVWS_FORCE_CPU=1 and SR_USE_CPU=true, so it produces service-tier evidence and cannot produce any GPU, lab-RF or over-the-air evidence.

Amini Amini Infratech for the Global South

Every link and every data fetch on this console is origin-relative. One origin fans out by path: / gateway, /video/, /grafana/, /prom/, /api/. An absolute http://localhost:NNNN URL works on exactly one machine, the one it was written on, and is blocked as mixed content the moment the page is served over https, which is how every operator actually reaches this stack.