Versioned suite contracts, and how many of them at least one indexed artefact is governed by. A contract with no artefact is listed here anyway: its method is declared even where nobody has run it.
Methodology
Build snapshot · Not liveOne section per versioned suite contract: what identifies the workload and the data, what each metric means and which direction is better, every pass/fail threshold with its source and its version, the accepted statistical treatment, the known limitations verbatim, and the exact reproduction command where it is safe and available.
- A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
- A reproduction command is offered only when running it does NOT drive an endpoint that is serving the live demo.
- A reproduction command is offered only when running it does NOT require a new RF capture session.
- A reproduction command is offered only when its producer exists in this tree.
Contract census
Gates declared across all contracts. Every one carries its own rule, its severity, the claim a failure removes, and the source and version of the threshold. A threshold with no source is not expressible in this model.
A command is printed only when the safety policy below permits it. Where it is withheld, the producer is still named and the tripped clause is stated — nothing is silently omitted.
Contracts
Jump to a suite
14 of 14- UHF occupancy on a 28-channel over-the-air survey no artefacts
- TVWS occupancy on real over-the-air captures 3 runs
- TVWS per-epoch OTA diagnostic 4 runs
- TVWS checkpoint selection, calibration and OOD 4 runs
- RF channel quality of the capture set 2 runs
- cuPHY LDPC decode latency 2 runs
- cuPHY LDPC under a capped MPS AI tenant 1 runs
- GPU partitioning mechanism probe 2 runs
- LDPC latency under a stepped AI load 2 runs
- Synthetic held-out classification accuracy 6 runs
- Sensing generalization across synthetic conditions 15 runs
- Training history 3 runs
- Bench harness run no artefacts
- Bench harness per-test result no artefacts
UHF occupancy on a 28-channel over-the-air survey
UHF occupancy on a 28-channel over-the-air survey
served-occupancy-survey@1 contract v1 no artefacts in this snapshotOn real captured RF with both occupied and vacant channels present, does passing a calibrated control floor change what the served model decides — and what does it cost?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it. It is also served through an endpoint, so the serving seam is part of the measurement.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | no artefacts | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | no artefacts | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | no artefacts | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | no artefacts | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | no artefacts | none stated | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | no artefacts | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | no artefacts | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | no artefacts | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | no artefacts | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | no artefacts | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | no artefacts | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | no artefacts | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | no artefacts | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | no artefacts | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary multiplexes detected on the held-out half report_multiplexes_detected_after | channels count | Higher is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one derived-occupied multiplex in UHF 35-48, the half no threshold was tuned on | 3 | no artefacts |
| Secondary multiplexes detected, whole survey multiplexes_detected_after | channels count | Higher is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one derived-occupied multiplex across all 28 swept channels | 10 | no artefacts |
| Secondary occupied records detected, whole survey occupied_records_detected_after | records count | Higher is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one capture record on a derived-occupied channel | 20 | no artefacts |
| Safety occupied channels the change made worse regressed_channels | channels count | Lower is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one derived-occupied channel detected in fewer records after the change | 1 | no artefacts |
| Safety vacant channels still called occupied in at least one record residual_false_positive_channels | channels count | Lower is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one derived-vacant channel with a non-zero after-arm occupied-record count | 1 | no artefacts |
| Safety vacant records called occupied vacant_false_positive_records_after | records count | Lower is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one capture record on a derived-vacant channel | 20 | no artefacts |
Pass/fail thresholds — and where they came from
Operator rules: zero regressed occupied channels, and a vacant class must exist before any false-positive rate is quoted. Neither is an FCC threshold nor derived from one.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking no occupied channel detected less often after the change metric regressed_channels · at most 0
Did any channel the survey derived as OCCUPIED lose detections when the control floor was applied?
- Why it matters
- A net win that hides a per-channel loss is how a regression ships. Every occupied channel must be detected at least as often after the change as before it, or the headline is an average standing over a casualty.
- Blocks the claim
- that the control floor is a strict improvement on every occupied channel
- On fail
- At least one occupied channel is detected in fewer records after the change. State it wherever the headline is stated; a net improvement does not settle it.
- On absent
- The per-channel derived_truth and before/after record counts needed to detect a regression are not in this artefact.
- Blocking no vacant channel is still called occupied metric residual_false_positive_channels · at most 0
Does any channel the survey derived as VACANT still get called OCCUPIED in at least one record?
- Why it matters
- This is the not-clean half of the result and it must be carried at equal prominence with the win. READ THE DIRECTION on UHF 28: 20 of 20 records to 8 of 20 is a FALSE POSITIVE that improved and was not eliminated, not a detection that was lost — the channel is derived VACANT, 0.39 dB over the session floor. The cyclostationary plane fires there at comb z = 4.90 against a Z_PRESENT of 3.5. The threshold was NOT moved and the case is pinned xfail(strict=True). This text used to add that Z_PRESENT was calibrated against synthetic AWGN maxima, the same defect one layer down; measured 2026-08-05 that is false — real receiver noise reaches comb z = 2.289 against the synthetic 2.19 and 0 of 300 null records reach 3.5. The firing is antenna-borne and its cause is open between a genuine emission below the energy floor and front-end intermodulation, separable only by an RX-gain sweep that has not been run.
- Blocks the claim
- that the vacant direction is clean on this survey
- On fail
- Vacant channels are still called occupied on some records. The channel-level false-positive count reaching zero does not settle it; state this beside the headline every time.
- On absent
- The per-channel derived_truth and after-arm record counts needed to find a residual false positive are not in this artefact.
- Blocking the survey contains vacant channels fact dataset.has_vacant_captures · exactly true
Are there genuinely vacant captures to measure a false-positive rate on?
- Why it matters
- The 2026-07-29 set had none, so an unconditional OCCUPIED scored full marks on it and the headline drawn from it was withdrawn. This survey has 15.
- Blocks the claim
- any false-positive rate on real RF
- On fail
- The survey contains no vacant channels and no false-positive rate is measurable.
- On absent
- The artefact does not state how many vacant channels the survey contains.
- Blocking the served default is unchanged fact model.endpoint_weight_hash_proven · exactly true
Does the running service behave as measured in the AFTER arm?
- Why it matters
- ModelSettings.control_floor_dbfs defaults to None. With no calibrated floor the engine falls back to the legacy CNN verdict and says so via occupancy_source. The AFTER arm is what the service CAN do, not what it does.
- Blocks the claim
- that a deployed SIDSense service currently decides occupancy this way
- On fail
- The endpoint is not proven to serve this checkpoint or this configuration.
- On absent
- No endpoint weight hash exists, so this can be neither confirmed nor denied. The AFTER arm describes an opt-in path, not the shipped default.
- Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true
Was the receive antenna specified for the band these captures were taken in?
- Why it matters
- Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
- Blocks the claim
- any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
- On fail
- The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
- On absent
- This artefact does not state whether the antenna matched the band. Absent, not confirmed.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 0 / 0The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- wilson
Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.
- What counts as a repetition
- A re-capture at another site or session, scored by the same harness against the same checkpoint. Re-scoring the same stored IQ is not a repetition.
- Minimum sample count
- 3
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- any percentile — this suite produces counts and proportions, not distributions
- a single accuracy folding the occupied and vacant directions together
- quoting the whole-survey number as a held-out result — the calibration half is in it
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- One site, one session, one antenna. Nothing here generalises to another site or another receiver.
- The control floor is opt-in. ModelSettings.control_floor_dbfs defaults to None and there is deliberately no default floor, so the shipped service still returns the legacy CNN verdict unless a floor is supplied.
- The CNN modulation label is reported and never scored here: the captures are DVB-T2 and the model has no DVB-T2 class. This says nothing about the 0.8864 synthetic modulation-class figure, which is a different measurement.
- The 15 at-floor channels sit at the receiver’s own floor, so a genuinely empty channel and one the antenna cannot hear are indistinguishable. The receiver-noise control was taken on 2026-08-05, on the B210’s unconnected RX chain, and it does not separate those two — it characterises noise structure, not what the antenna can hear. A matched UHF antenna is what would, and it has not been used.
Caveats carried verbatim from the artefacts (0)
No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.
Reproduction
- Command offered Served occupancy: CNN vs physics detector 0 artefact(s)Producer scripts/score-served-occupancy.py verified
Scores both arms in one process against one unchanged checkpoint.
Named in the artefact’s own harness field.
python scripts/score-served-occupancy.py --survey <uhf-survey-dir> --checkpoint external/tvws-sensing/models/checkpoints/best.pt --out benchmarks/reports/served-occupancy-2026-08-05.json --records 20 --split allAssembled from the producer’s own argparse (scripts/score-served-occupancy.py lines 119-126). This is NOT A TRANSCRIPT — the artefact does not record its own invocation, so the survey path here is a placeholder and the record counts are the parser defaults. The checkpoint path and its sha256 ARE recorded, under model_identity.
Preconditions
- The nairobi-region-2026-08-04-uhf-survey captures. They are on the box and outside this repository.
- The checkpoint at the recorded sha256 8eca022f…. A sha pins a FILE; it does not prove any endpoint serves it.
- Ground truth is derived from the same session’s measured floor, so the survey directory must be the one that produced the floor, not a re-capture.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs the survey captures and the checkpoint, neither of which is in the repository.
TVWS occupancy on real over-the-air captures
TVWS occupancy on real over-the-air captures
tvws-ota-occupancy@1 contract v1 3 artefactsDoes the sensing model call a licensed, occupied multiplex OCCUPIED, on real captured RF?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it. It is also served through an endpoint, so the serving seam is part of the measurement.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | ABSENT from all | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | ABSENT from all | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | 3 / 3 | scripts/eval-ota-checkpoint.py | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | 3 / 3 | /ota | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | ABSENT from all | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | 3 / 3 | 8eca022f09b897a579ef74225881343949a1a4a53b880edabf82225aacbb34ee · 038d32176653d11ae0a998b4a1fa3df5de68a67ca8c741024d375c6fd7737973 · 7aca8bad07930190a63346468170b9e4f3a84ebb9e93c00b5460a503d39ce69a | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary record-level occupancy rate records_occupied_rate | fraction fraction | Higher is better. | proportion A proportion of a countable population. A confidence interval is expressible only when the numerator and denominator are whole records. | one real OTA capture record from a channel labelled occupied | 20 | 3 / 3 |
| Secondary channels with correct occupancy occupancy_correct_channels | channels count | Higher is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one captured multiplex channel | 1 | 3 / 3 |
| Secondary records with the nearest correct class records_nearest_correct_class | records count | Higher is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one real OTA capture record | 20 | 3 / 3 |
| Safety channels wrongly called VACANT false_vacant_channels | channels count | Lower is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one licensed multiplex channel whose records are all labelled occupied | 1 | 3 / 3 |
| Safety records wrongly called VACANT false_vacant_records | records count | Lower is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one real OTA capture record labelled occupied | 1 | 3 / 3 |
Pass/fail thresholds — and where they came from
Operator safety rule: zero false-vacant channels. Not an FCC threshold and not derived from one.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Safety critical no licensed channel called vacant metric false_vacant_channels · at most 0
Did the model call any occupied, licensed multiplex VACANT?
- Why it matters
- A false-vacant decision authorises a transmission on a channel a licensed incumbent is using. It is the only error in this suite that causes harm outside the system.
- Blocks the claim
- that the model is safe to gate transmission on a vacant-channel decision
- On fail
- FAILED SAFETY GATE: occupied licensed channels were called vacant. This result stands on its own and is not offset by any accuracy figure, here or elsewhere.
- On absent
- The per-channel labels needed to count false-vacant decisions are not in this artefact.
- Blocking the capture set contains vacant channels fact dataset.has_vacant_captures · exactly true
Are there any genuinely vacant captures to measure a false-positive rate on?
- Why it matters
- Every capture in this set is labelled occupied, so a model that answers "occupied" unconditionally scores full marks here.
- Blocks the claim
- any false-positive rate on real RF
- On fail
- The capture set contains no vacant channels.
- On absent
- No vacant captures exist in this set, so no false-positive rate on real RF is measurable at all — absent, not zero.
- Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true
Was the receive antenna specified for the band these captures were taken in?
- Why it matters
- Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
- Blocks the claim
- any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
- On fail
- The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
- On absent
- This artefact does not state whether the antenna matched the band. Absent, not confirmed.
- Blocking the endpoint is proven to serve this checkpoint fact model.endpoint_weight_hash_proven · exactly true
Is there evidence that the service under test served the weights named here?
- Why it matters
- A checkpoint sha256 pins a FILE on disk. The service exposes no weight hash, so nothing in this artefact can bind the file to the endpoint that produced these numbers.
- Blocks the claim
- that the deployed service serves the checkpoint named here
- On fail
- The endpoint is not proven to serve this checkpoint.
- On absent
- No endpoint weight hash exists, so this can be neither confirmed nor denied. The checkpoint hash is provenance, not deployment verification.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 0 / 3The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- wilson
Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.
- What counts as a repetition
- A re-capture at the same site with the same receive chain, evaluated by the same harness revision against the same checkpoint. Re-running the evaluator over the same stored IQ is not a repetition.
- Minimum sample count
- 20
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- any percentile — this suite produces counts and proportions, not distributions
- a single accuracy that folds false-vacant and false-occupied decisions together
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- Nothing here establishes anything about vacant channels: no vacant captures exist in this set.
- One site, one afternoon, one receive chain. Nothing here generalises to another site or another receiver.
- The captures are DVB-T2 and the model has no DVB-T2 class.
Caveats carried verbatim from the artefacts (6)
- caveats[0]
Captures are DVB-T2; the model has no DVB-T2 class. DVB-T is nearest.
3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json - caveats[1]
Received through a mismatched 860-930 MHz antenna at 470-700 MHz.
3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json - caveats[2]
Proves nothing about vacant channels — no vacant captures exist here.
3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json - harness_note
Runs the service's own load_model + InferenceEngine.infer, not a reimplementation. The (H,W,C)->(C,H,W) transpose lives inside InferenceEngine.infer.
3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json - model_identity.note
checkpoint_sha256 pins a FILE. It does NOT prove that any endpoint serves these weights.
3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json - summary.record_level_note
Channel counting is all-or-nothing and hides movement: a channel going 0/20 -> 18/20 occupied still reads as one failed channel. Compare checkpoints on the record-level rate as well.
3 file(s): ota-eval-fixB.json, ota-eval-fixC-final.json, ota-eval-fixC.json
Reproduction
- Command offered TVWS over-the-air eval 3 artefact(s)Producer scripts/eval-ota-checkpoint.py verified
Scores a checkpoint against the real over-the-air captures using the service's own inference path.
The artefacts name it in their own harness field.
python scripts/eval-ota-checkpoint.py --checkpoint <path/to/best.pt> --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --confidence-threshold 0.7 --json-out benchmarks/reports/<name>.jsonAssembled from the producer's own argparse (lines 71-83). The capture directory is the one this family's artefacts name in dataset.dir. It is the producer's documented invocation, not a transcript of the run.
Preconditions
- The capture set at benchmarks/datasets/ota/nairobi-region-2026-07-29. It exists in this tree.
- A checkpoint. The artefacts record checkpoint_sha256 and say in their own words that it "pins a FILE. It does NOT prove that any endpoint serves these weights."
- EVERY capture in this set is labelled occupied. Re-running proves nothing about vacant channels, whatever the numbers say.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs the OTA capture set and a checkpoint, neither of which is on a GitHub runner.
TVWS per-epoch OTA diagnostic
TVWS per-epoch OTA diagnostic
tvws-ota-per-epoch@1 contract v1 4 artefactsDoes synthetic validation accuracy track real-OTA behaviour across training epochs?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | ABSENT from all | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | ABSENT from all | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | 4 / 4 | scripts/ota-per-epoch-sweep.py | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | 4 / 4 | /ota | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | 2 / 4 | 7 | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | ABSENT from all | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary Spearman: synthetic val accuracy vs real OTA spearman_val_acc_vs_real_ota | none scalar | Higher is better. | scalar A single stated value, not a summary of a distribution. | one training epoch | 5 | ABSENT from all |
Pass/fail thresholds — and where they came from
None. This suite is diagnostic and has no pass threshold.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true
Was the receive antenna specified for the band these captures were taken in?
- Why it matters
- Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
- Blocks the claim
- any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
- On fail
- The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
- On absent
- This artefact does not state whether the antenna matched the band. Absent, not confirmed.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 0 / 4The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- none
No proportion interval is defined for this suite.
- What counts as a repetition
- An independent training seed swept over the same capture set.
Independent seeds are required, recorded IN the artefact and not inferred from a file name. - Minimum sample count
- 5
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- using this sweep as a selection criterion — it is n=5 channels, all occupied
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- Diagnostic, not a selection criterion — n=5 channels, all occupied.
- Establishes nothing about vacant channels.
Caveats carried verbatim from the artefacts (5)
- caveats[0]
Diagnostic, not a selection criterion — n=5 channels, all occupied.
4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json - caveats[1]
Captures received through a mismatched antenna.
4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json - caveats[2]
Establishes nothing about vacant channels.
4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json - dataset.note
5 channels, one site, one mismatched 860-930 MHz antenna at 470-700 MHz; diagnostic only
4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json - harness_note
drives the service's own load_model + InferenceEngine.infer, same seam as scripts/eval-ota-checkpoint.py
4 file(s): ota-per-epoch-2026-07-30.json, ota-per-epoch-packed-2026-07-30.json, ota-per-epoch-packed-seed7-2026-07-30.json, ota-per-epoch-seed7-2026-07-30.json
Reproduction
- Command offered TVWS OTA per-epoch sweep 4 artefact(s)Producer scripts/ota-per-epoch-sweep.py verified
Runs the OTA eval across every checkpoint in a directory, one row per epoch.
The artefacts name it in their own harness field.
python scripts/ota-per-epoch-sweep.py --checkpoint-dir <dir/of/epoch/checkpoints> --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --records 20 --json-out benchmarks/reports/<name>.jsonAssembled from the producer's own argparse (lines 124-133). It is the producer's documented invocation, not a transcript of the run.
Preconditions
- A directory of per-epoch checkpoints. Those are not in this repository.
- The antenna used for these captures was NOT matched to the band. Every artefact in this family states that in prose in dataset.note, and no boolean in the family carries it — which is why the mismatch is detected from prose and must stay visible on every number.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs the OTA capture set and a directory of checkpoints.
TVWS checkpoint selection, calibration and OOD
TVWS checkpoint selection, calibration and OOD
tvws-selection-calibration@1 contract v1 4 artefactsWhich epoch shipped, how confident is it, and how does it behave off-distribution?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | ABSENT from all | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | 4 / 4 | d:8dd29d376164e84a · d:a8f6e162d67b4ec7 · d:808c8c9bc84991ab +1 more | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | ABSENT from all | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | 4 / 4 | scripts/sensing-select-calibrate-ood.py · external/tvws-sensing/scripts/train_model.py | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | ABSENT from all | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | 4 / 4 | 42 · 7 | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | 2 / 4 | 8eca022f09b897a579ef74225881343949a1a4a53b880edabf82225aacbb34ee | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary Spearman: synthetic val accuracy vs real records spearman_synthetic_val_acc_vs_real_records | none scalar | Higher is better. | scalar A single stated value, not a summary of a distribution. | one candidate epoch | 4 | 2 / 4 |
Pass/fail thresholds — and where they came from
None. Selection and calibration are reported, not gated.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking the reported score is not a selection score fact selection.held_out_from_selection · exactly true
Was the set that produced this number also used to choose the checkpoint?
- Why it matters
- The four verified multiplexes were used to SELECT the shipped epoch, so their score for that checkpoint is optimistically biased.
- Blocks the claim
- that this number is an unbiased test score for the shipped checkpoint
- On fail
- This is a selection score, not a test score.
- On absent
- The artefact does not state that the scoring set was held out of selection. Treat the number as a selection score until it does.
- Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true
Was the receive antenna specified for the band these captures were taken in?
- Why it matters
- Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
- Blocks the claim
- any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
- On fail
- The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
- On absent
- This artefact does not state whether the antenna matched the band. Absent, not confirmed.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 2 / 4The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- wilson
Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.
- What counts as a repetition
- An independent training seed run through the same selection procedure.
Independent seeds are required, recorded IN the artefact and not inferred from a file name. - Minimum sample count
- 4
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- quoting a selection score as a test score
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- There are ZERO real vacant captures, so no false-positive rate on real RF can be measured at all.
- Calibration temperature is fitted on held-out SYNTHETIC data; applying it to real captures is an extrapolation.
Caveats carried verbatim from the artefacts (6)
- honesty
The ota_trajectory is a MEASUREMENT taken per epoch. If a checkpoint is then selected using it, the real set has become a selection set and is no longer an unbiased test of that checkpoint. Report both, and never quote the selection score as a test score.
2 file(s): tvws-retrain-seed42-2026-08-03.json, tvws-retrain-seed7-2026-08-03.json - honesty.calibration_temperature_is_synthetic
Temperature is fitted on held-out SYNTHETIC data because that is the only labelled data with enough records to fit anything. Applying it to real captures is an extrapolation. The real-RF ECE is reported before and after so the extrapolation is visible.
2 file(s): tvws-retrain-metrics-seed42-2026-08-03.json, tvws-retrain-metrics-seed7-2026-08-03.json - honesty.selection_set_is_not_a_test_set
The four verified multiplexes were used to SELECT the shipped epoch. Their score for that checkpoint is a selection score and is optimistically biased. The leave-one-channel-out total is the closest to unbiased that four channels allow: each held-out channel is scored by a checkpoint selected without it.
2 file(s): tvws-retrain-metrics-seed42-2026-08-03.json, tvws-retrain-metrics-seed7-2026-08-03.json - honesty.what_no_real_number_here_can_prove
One site, one afternoon, one mismatched 860-930 MHz antenna at 470-700 MHz, four channels, every capture occupied. There are ZERO real vacant captures, so no false-positive rate on real RF can be measured at all, and a model answering 'occupied' unconditionally scores 80/80 here. Nothing here generalises to another site or another receiver.
2 file(s): tvws-retrain-metrics-seed42-2026-08-03.json, tvws-retrain-metrics-seed7-2026-08-03.json - selection.spearman_note
Negative or near-zero means selecting by synthetic validation accuracy is uninformative about real RF for this run. It is one run; treat it as a consistency check on the 2026-07-30 result (-0.1087 and -0.316), not as an independent one.
2 file(s): tvws-retrain-metrics-seed42-2026-08-03.json, tvws-retrain-metrics-seed7-2026-08-03.json - synthetic_test.note
Held-out split of the SAME generator that produced the training set. It cannot detect a leak it shares.
2 file(s): tvws-retrain-seed42-2026-08-03.json, tvws-retrain-seed7-2026-08-03.json
Reproduction
- Command withheld TVWS retrain run 2 artefact(s)Producer external/tvws-sensing/scripts/train_model.py verified
Trains the sensing CNN for a retrain round and writes the run record.
The artefacts name it in their own harness field.
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- A full training run on a GPU, per seed.
- The two artefacts differ by seed (42 and 7). Two seeds of one training recipe are a variance probe, not two measurements of one thing.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.
- Command offered TVWS selection / calibration / OOD 2 artefact(s)Producer scripts/sensing-select-calibrate-ood.py verified
Selects a checkpoint, calibrates it, scores OOD behaviour and re-scores against the OTA captures.
The artefacts name it in harness AND record their own invocation under args.
python scripts/sensing-select-calibrate-ood.py --run-json benchmarks/reports/tvws-retrain-seed42-2026-08-03.json --checkpoint-dir <dir> --baseline-checkpoint external/tvws-sensing/models/checkpoints/best.pt --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --ota-records 20 --samples-per-class 2000 --seed 42 --confidence-threshold 0.7Assembled from the producer's own argparse (lines 679-701). THIS FAMILY IS THE ONLY ONE IN THE CORPUS THAT RECORDS ITS OWN INVOCATION: open the raw artefact and read the "args" object for the exact values that produced it, including the checkpoint directory this command leaves as a placeholder.
Preconditions
- The checkpoint directory named in the artefact's own args.checkpoint_dir. It is outside this repository.
- The artefact records device "cuda": this ran on a GPU, and the selection set it reports is a SELECTION set, not a test set.
This family records its own invocation in the artefact under
args. Open the raw file for the exact values that produced it, rather than relying on the documented invocation above.Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a checkpoint directory and the OTA capture set.
RF channel quality of the capture set
RF channel quality of the capture set
rf-channel-quality@1 contract v1 2 artefactsHow much signal is actually in each captured channel, relative to a control floor?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | ABSENT from all | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | ABSENT from all | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | 1 / 2 | scripts/analyze-618mhz.py | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | 1 / 2 | benchmarks/datasets/ota/nairobi-region-2026-07-29 | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | ABSENT from all | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | ABSENT from all | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary estimated in-window SNR *.est_in_window_snr_db | dB scalar | Higher is better. | scalar A single stated value, not a summary of a distribution. | one channel, estimated against a scalar control floor | 20 | 1 / 2 |
Pass/fail thresholds — and where they came from
None. This suite characterises the capture set; it does not grade it.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking antenna matched to the captured band fact dataset.antenna_matched_to_band · exactly true
Was the receive antenna specified for the band these captures were taken in?
- Why it matters
- Every capture in this set came through an 860-930 MHz antenna used at 470-700 MHz. That bounds what the RECEIVER could hear: a channel sitting at the receiver's own floor cannot be told apart from one the antenna is simply deaf to, so vacancy and dynamic range are not measurable through it. It is NOT an explanation for MISSED OCCUPIED channels. On the 2026-08-04 UHF survey the two strongest carriers in the band — UHF 32 at +42.20 dB and UHF 33 at +41.39 dB over the session floor — were classified Noise on 20 of 20 records at 0.969 and 0.957. Detection here does not track signal strength, so a matched UHF antenna will not fix the occupied-channel failure.
- Blocks the claim
- any vacant-channel or dynamic-range claim about this site that does not carry the antenna mismatch
- On fail
- The captures were received through an antenna that is not specified for this band. Every vacant-channel and dynamic-range claim inherits that. Occupied-channel detection failures do not: they are reproduced on carriers more than 40 dB above the floor, which no antenna explains.
- On absent
- This artefact does not state whether the antenna matched the band. Absent, not confirmed.
Statistical method
- Central tendency
- median_iqr
Median with the interquartile range (Tukey hinges). The default here because latency and accuracy distributions in this corpus are small-n and not symmetric, and a mean hides that.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 1 / 2The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- none
No proportion interval is defined for this suite.
- What counts as a repetition
- A recapture at the same site and gain with control IQ retained.
- Minimum sample count
- 20
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- treating the estimated SNR as a measured SNR
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- In-window SNR figures are ESTIMATES: no control IQ was retained, only a scalar floor.
- Receive-only captures. Nothing was transmitted.
Caveats carried verbatim from the artefacts (6)
- caveats[0]
in-window SNR figures are ESTIMATES: no control IQ was retained, only the scalar -59.9 dBFS floor. They assume in-capture noise == that floor at the same port and gain.
1 file(s): ch39-618mhz-quality-2026-07-31.json - caveats[1]
captures came through a SenseCAP LoRa 860-930 MHz antenna used at 470-700 MHz. Every spectral claim carries that.
1 file(s): ch39-618mhz-quality-2026-07-31.json - caveats[2]
610 and 626 MHz were never captured, so the doubly-adjacent hypothesis cannot be settled from this dataset at all.
1 file(s): ch39-618mhz-quality-2026-07-31.json - caveats[3]
n=20 records, one site, one antenna, five channels, all occupied.
1 file(s): ch39-618mhz-quality-2026-07-31.json - caveats[4]
receive-only RX2 captures. Nothing was transmitted.
1 file(s): ch39-618mhz-quality-2026-07-31.json - target
uhf39_618MHz — never classified DVB-T at any converged epoch
1 file(s): ch39-618mhz-quality-2026-07-31.json
Reproduction
- Command offered Channel quality capture 1 artefact(s)Producer scripts/analyze-618mhz.py verified
Analyses the 618 MHz capture: quality stages, a synthetic SNR ladder and a bandlimit sweep.
The artefact names it in its own harness field.
python scripts/analyze-618mhz.py --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --records 20 --confidence-threshold 0.7Assembled from the producer's own argparse (lines 511-527). It is the producer's documented invocation, not a transcript of the run.
Preconditions
- The capture directory. The default in the script is /ota, the in-container mount.
- This is ONE channel at ONE site on ONE day. It supports no statement about any other channel, site or day.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It reads an IQ capture that is not in the repository.
- Command offered Detector scores 1 artefact(s)Producer scripts/score-detectors.py verified
Runs the detector scorer and writes the artefact with a full provenance envelope.
The 2026-08-04 artefact names it in its own harness field.
python scripts/score-detectors.py --ota-dir benchmarks/datasets/ota/nairobi-region-2026-07-29 --synth-records 100 --seed0 20260803Assembled from the producer's own argparse defaults, so it is the documented invocation rather than a transcript of any particular run. The 2026-08-04 artefact separately records its own invocation under args, which is where the exact values that produced it are readable.
Preconditions
- The capture set at benchmarks/datasets/ota/nairobi-region-2026-07-29. It exists in this tree.
- CPU only, about 13 seconds, no GPU and no running service. It is safe to run while the demo is up.
- The captures came through an antenna that is not matched to the band, and every record in the set is labelled occupied. Re-running measures the detector, never those two facts.
This family records its own invocation in the artefact under
args. Open the raw file for the exact values that produced it, rather than relying on the documented invocation above.Continuous integration No CI jobNo workflow in .github/workflows runs this producer, and nothing about the producer stops one: it is a CPU-only numpy job that finishes in about 13 seconds. What it needs is the 5 IQ captures under benchmarks/datasets/ota/, roughly 400 MB that this repository does not carry, so a GitHub runner has nothing to score.
cuPHY LDPC decode latency
cuPHY LDPC decode latency
cuphy-ldpc-decode@1 contract v1 2 artefactsHow long does ONE cuPHY LDPC decode take, timed on the decoder’s own CUDA stream?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | 2 / 2 | BG1/Z384/81cb/rate0.5 | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | 2 / 2 | 1000 timed decodes | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | 2 / 2 | 20 · 50 | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | 2 / 2 | pyaerial LdpcDecoder + cudaEvent per call | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | ABSENT from all | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | ABSENT from all | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | ABSENT from all | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary cuPHY decode p50 cuphy_decode.p50 | us p50 | Lower is better. | raw_sample_percentile A percentile computed over the RETAINED RAW SAMPLES. This is the only aggregation from which a p99 or p99.9 is legitimate. | one cuPHY LDPC decode over 81 code blocks | 100 | 2 / 2 |
| Secondary cuPHY decode p99 cuphy_decode.p99 | us p99 | Lower is better. | raw_sample_percentile A percentile computed over the RETAINED RAW SAMPLES. This is the only aggregation from which a p99 or p99.9 is legitimate. | one cuPHY LDPC decode over 81 code blocks | 100 | 2 / 2 |
| Secondary decodes over the slot budget cuphy_decode.deadline_misses | decodes count | Lower is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one cuPHY LDPC decode over 81 code blocks | 100 | 2 / 2 |
Pass/fail thresholds — and where they came from
slot_budget_us in the artefact (500 us), which is the producer’s stated budget, not a 3GPP requirement.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking the measurement covers a full slot fact scope.is_a_full_slot · exactly true
Does this number cover a whole 5G slot, or only the LDPC decode?
- Why it matters
- A slot also carries control channels, DMRS and the rest of the PUSCH/PDSCH chain. A deadline conclusion drawn from the decoder alone is a lower bound, not a slot result.
- Blocks the claim
- that the platform holds a full-slot L1 deadline
- On fail
- LDPC decode only. This is a microbenchmark and a LOWER BOUND on real slot overrun — it cannot support a full-slot claim.
- On absent
- The artefact does not state whether this covers a full slot.
- Blocking the device is recorded in the artefact fact device.stated · exactly true
Does the file say which GPU produced the number?
- Why it matters
- Schema v2 carries no device block, so its number cannot be attributed to a machine and cannot be ranked against one that can.
- Blocks the claim
- attributing this number to a particular GPU
- On fail
- No device block: the GPU behind this number is unknowable from the file.
- On absent
- The artefact does not record whether a device block is present.
Statistical method
- Central tendency
- median_iqr
Median with the interquartile range (Tukey hinges). The default here because latency and accuracy distributions in this corpus are small-n and not symmetric, and a mean hides that.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- raw_samples
The contract permits a percentile, and only from RETAINED RAW SAMPLES. The sample count still gates which percentile is expressible, and a run that retained no samples cannot supply one however the contract is written — see the reality check beside this.
Raw samples retained by 2 / 2All 2 artefact(s) retain a raw sample distribution, so the contract's permission is backed by the evidence.
- Proportion interval
- none
No proportion interval is defined for this suite.
- What counts as a repetition
- A fresh process on the same device, same image, same code configuration and same warmup, at the same repetition count.
- Minimum sample count
- 100
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- reading a throughput claim off a random-LLR run: early termination never fires, so it is the WORST case by construction
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- LDPC decode only — no MAC, no scheduler, no fronthaul, no cell, no OTA.
- Random Gaussian LLRs: correct for deadline analysis, wrong for throughput claims.
Caveats carried verbatim from the artefacts (5)
- llr_mode_note
random: Gaussian LLRs, no valid codeword, early termination never fires -> full iteration count -> WORST case. Correct for deadline analysis, wrong for throughput claims.
2 file(s): cuphy-per-slot-20260729-005456.json, cuphy-per-slot-20260730-080736.json - measures
GPU elapsed time of ONE cuPHY LDPC decode over 81 code blocks, timed on the decoder's own CUDA stream.
2 file(s): cuphy-per-slot-20260729-005456.json, cuphy-per-slot-20260730-080736.json - scope_warning
LDPC decode only. A 5G slot also carries control channels, DMRS and the rest of the PUSCH/PDSCH chain, so a miss counted here is a miss BY THE DECODER ALONE and is a LOWER BOUND on real slot overrun.
2 file(s): cuphy-per-slot-20260729-005456.json, cuphy-per-slot-20260730-080736.json - series_note
cuphy_decode is cuPHY's LDPC decode. pyaerial_decode is the same work wrapped in pyAerial's public decode(), which re-copies the LLR block and the output on every call. Quote cuphy_decode for PHY latency; the difference is Python-binding overhead the real L1 does not pay.
2 file(s): cuphy-per-slot-20260729-005456.json, cuphy-per-slot-20260730-080736.json - device.co_tenancy_visibility
namespace-limited: measured from inside the Aerial container, so host processes are INVISIBLE here. An empty list does NOT mean the GPU was idle — verify from the host with nvidia-smi before claiming a clean run.
1 file(s): cuphy-per-slot-20260730-080736.json
Reproduction
- Command withheld cuPHY per-slot latency 2 artefact(s)Producer scripts/cuphy-per-slot-latency.py verified
Times one pyaerial LdpcDecoder.decode() call per sample with cudaEvent, inside the Aerial container, and writes the per-slot artefact to benchmarks/reports/.
The script writes to "benchmarks/reports/" (line 293) and its own docstring documents the percentile gate this family carries.
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- A GB10 or H200 with the Aerial cuBB container image available locally.
- Exclusive use of the GPU for the duration: the harness re-execs itself into the Aerial container and times decodes.
- The documented invocation is ./scripts/cuphy-per-slot-latency.py --decodes 1000 (from the script header).
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.
cuPHY LDPC under a capped MPS AI tenant
cuPHY LDPC under a capped MPS AI tenant
mps-cotenancy@1 contract v1 1 artefactsWhat happens to cuPHY LDPC invocation-average latency while a capped MPS-managed AI client loads the same GPU?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | 1 / 1 | caps=[100,50,25] sizes=[8192] | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | 1 / 1 | 12 invocations/level x 100 inner decodes, 240s AI load | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | ABSENT from all | none stated | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | ABSENT from all | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | ABSENT from all | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | ABSENT from all | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary p50 of the invocation averages *.p50 | us p50 | Lower is better. | invocation_average_percentile A percentile computed over INVOCATION AVERAGES. Each sample is already a mean, so the tail of the underlying distribution has been averaged away before the percentile was taken. No tail claim survives this. | one cuphy_ex_ldpc invocation, itself already an average across 100 inner decodes | 12 | 1 / 1 |
| Secondary block errors *.block_errors | blocks count | Lower is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one decoded code block | 1 | 1 / 1 |
| Safety uncontrolled tenants on the GPU unmanaged_tenants_at_start | processes count | Lower is better. | count A count of events. Whole numbers; no interpolation is meaningful. | one process holding a CUDA context outside the MPS server | 1 | 1 / 1 |
Pass/fail thresholds — and where they came from
slot_budget_us in the artefact, applied to invocation averages. Not a deadline guarantee.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking every GPU client was MPS-managed fact mps.all_clients_managed · exactly true
Was the GPU free of processes outside the MPS server for the whole run?
- Why it matters
- Uncontrolled default-mode tenants share the same SMs. With any present, the measurement is of the whole machine, not of the capped client.
- Blocks the claim
- that this run demonstrates isolation between the RAN and AI clients
- On fail
- Uncontrolled tenants were on the GPU for this run. No isolation claim is available from it, and the latency it reports is not attributable to the capped client.
- On absent
- The artefact does not census the processes on the GPU.
- Blocking the GPU was verified clean at the end of the run fact mps.gpu_clean_at_end · exactly true
Was the GPU free of uncontrolled tenants when the run finished?
- Why it matters
- When end_of_run_gpu_check is null, an empty at-end tenant list is not-collected. Empty is not clean.
- Blocks the claim
- that the GPU was clean for the duration of this run
- On fail
- The end-of-run check ran and did not come back clean.
- On absent
- The end-of-run co-tenancy check did not run, so the empty at-end tenant list proves nothing.
- Blocking every planned cell was measured fact mps.partial_run · exactly false
Did the sweep complete the cells it planned?
- Why it matters
- Caps that were never measured are ABSENT, not passing. No envelope may be read across a partial sweep.
- Blocks the claim
- a largest-passing-cap envelope
- On fail
- PARTIAL SWEEP. The caps that were never measured are absent, not passing, so no envelope may be read from this document.
- On absent
- The artefact does not state whether the sweep completed.
Statistical method
- Central tendency
- median_iqr
Median with the interquartile range (Tukey hinges). The default here because latency and accuracy distributions in this corpus are small-n and not symmetric, and a mean hides that.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 0 / 1The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- none
No proportion interval is defined for this suite.
- What counts as a repetition
- A fresh run on the same host with the same cap set, load sizes, invocation count and AI load duration — and a clean GPU census at both ends.
- Minimum sample count
- 12
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- any tail percentile: each sample is already an average across 100 inner decodes
- reading an isolation guarantee from a provisioning knob
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- CUDA_MPS_ACTIVE_THREAD_PERCENTAGE PROVISIONS but does not RESERVE: kernels from different clients may still execute on the same SM.
- CUDA_MPS_CLIENT_PRIORITY is a documented HINT, not a guarantee.
- cuPHY LDPC only: no MAC, scheduler, fronthaul, cell or OTA.
Caveats carried verbatim from the artefacts (6)
- cap_semantics
CUDA_MPS_ACTIVE_THREAD_PERCENTAGE PROVISIONS but does not RESERVE: the documentation is explicit that kernels from different clients may still execute on the same SM. cuPHY is 100%; the AI client is capped.
1 file(s): mps-sweep-2026-08-03-livebox.json - partial_run_note
PARTIAL SWEEP: 2 of 4 planned cells completed. The caps that were never measured are absent, not passing, so no envelope may be read from this document and largest_observed_passing_ai_cap_percent is withheld.
1 file(s): mps-sweep-2026-08-03-livebox.json - priority_semantics
CUDA_MPS_CLIENT_PRIORITY is a documented HINT, not a guarantee. A null result for this arm is expected, not anomalous.
1 file(s): mps-sweep-2026-08-03-livebox.json - run_error
AI load client exited before MPS join: container=ulap-cuphy-ai-2936710-50-8192; state=exited; exit_code=1; logs_tail='Traceback (most recent call last):\n File "<string>", line 4, in <module>\ntorch.AcceleratorError: CUDA error: CUDA-capable device(s) is/are busy or unavailable\nSearch for `cudaErrorDevicesUnavailable\' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.\nFor more detailed error information, run with CUDA_LOG_FILE=stderr'
1 file(s): mps-sweep-2026-08-03-livebox.json - scope
cuPHY LDPC invocation-average latency under a capped MPS-managed PyTorch matmul client; no MAC, scheduler, fronthaul, cell, or OTA
1 file(s): mps-sweep-2026-08-03-livebox.json - timing_warning
Each sample is cuphy_ex_ldpc's average across 100 inner decodes. This is not a per-slot latency distribution and cannot support p99.9.
1 file(s): mps-sweep-2026-08-03-livebox.json
Reproduction
- Command withheld MPS co-tenancy sweep 1 artefact(s)Producer scripts/cuphy-mps-sweep.py verified
Sweeps an MPS active-thread cap on an AI client while cuPHY LDPC runs uncapped.
Writes to "benchmarks/reports/" and emits partial_run / reportable_clean_mps_run.
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- A scheduled clean-GPU window. The producer's own docstring says: "Use a scheduled clean-GPU window, or move every competing client into the same MPS server, before removing the gate."
- Both managed clients must join the same private MPS server, which means starting an MPS control daemon.
- The Aerial cuBB image and the AI load image (amini/wg3-sr-worker:latest) present locally.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.
GPU partitioning mechanism probe
GPU partitioning mechanism probe
gpu-partition-mechanism@1 contract v1 2 artefactsDoes the partitioning mechanism itself do what its documentation says?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | ABSENT from all | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | ABSENT from all | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | 1 / 2 | scripts/mps-cap-test.sh | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | ABSENT from all | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | ABSENT from all | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | ABSENT from all | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary fp16 matmul throughput *.tflops | TFLOPS scalar | Higher is better. | scalar A single stated value, not a summary of a distribution. | one 4096-square fp16 matmul run | 1 | 1 / 2 |
Pass/fail thresholds — and where they came from
None. The probe characterises a mechanism; it grades nothing.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
This contract declares no gate. It reports; it does not grade. That is not a pass.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 0 / 2The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- none
No proportion interval is defined for this suite.
- What counts as a repetition
- A re-run of a harness that does not currently exist.
- Minimum sample count
- 1
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- reading a concurrency result from sequential runs
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- Concurrent non-interference was NOT demonstrated: two instances produced no output when all three ran at once.
- The host is not this box.
Caveats carried verbatim from the artefacts (4)
- NOT_measured.cuphy_ldpc_inside_a_mig_instance
not attempted in this window
1 file(s): mig-h200-2026-08-03.json - NOT_measured.three_way_concurrent_isolation
Instances 2 and 3 produced NO output when all three ran at once, though each works sequentially. The concurrent non-interference claim is therefore UNPROVEN and is not made. Cause not diagnosed; the window had to close to restore a user-facing service.
1 file(s): mig-h200-2026-08-03.json - note
Unmanaged default-mode tenants were present and untouched; they are why the cuPHY co-tenancy sweep could not isolate. This test scopes the question to the mechanism by putting both clients inside one private MPS server.
1 file(s): mps-cap-result.json - teardown
compute instances and GPU instances destroyed, MIG mode returned to Disabled, akili trio restarted and verified serving in 56 s with 95772 MiB restored
1 file(s): mig-h200-2026-08-03.json
Reproduction
- Command withheld GPU partitioning probe 1 artefact(s)Producer scripts/cotenancy/mig-h200-probe.py verified
Regenerates the observable half of this family and refuses the destructive half unless an operator explicitly asks and the GPU is free.
Emits benchmarks/reports/mig-h200-observed-<date>.json.
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- The default `--mode observe` is read-only and safe, but it needs ssh to the H200, which this UI cannot offer as a one-click command.
- `--mode measure` requires a compute-free H200 and an explicit operator flag. It is a privileged one-way device reconfiguration, not a run.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.
- Command withheld MPS thread-cap enforcement 1 artefact(s)Producer scripts/mps-cap-test.sh verified
Runs both clients inside a private MPS server to test whether an active-thread cap partitions the GB10.
Line 16 writes benchmarks/reports/mps-cap-result.json by name.
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- It starts its own MPS control daemon (CUDA_MPS_PIPE_DIRECTORY) and quits it on cleanup.
- A python at $HOME/.venvs/wg3-test/bin/python, or PY= pointing at one.
- The script states it "adds only its own processes and never touches an existing service" — the exclusive-control clause is still tripped, because an MPS control daemon is a device-wide object.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.
LDPC latency under a stepped AI load
LDPC latency under a stepped AI load
ldpc-step-load@1 contract v1 2 artefactsHow does LDPC decode latency move as an AI tenant’s load steps up on the same GPU?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | ABSENT from all | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | ABSENT from all | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | 1 / 2 | scripts/cuphy-cotenancy-sweep.py | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | ABSENT from all | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | ABSENT from all | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | ABSENT from all | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary p90 decode latency at load level *.p90 | us p90 | Lower is better. | raw_sample_percentile A percentile computed over the RETAINED RAW SAMPLES. This is the only aggregation from which a p99 or p99.9 is legitimate. | one LDPC decode at one AI load level | 15 | 2 / 2 |
Pass/fail thresholds — and where they came from
None recorded in the artefact.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
This contract declares no gate. It reports; it does not grade. That is not a pass.
Statistical method
- Central tendency
- median_iqr
Median with the interquartile range (Tukey hinges). The default here because latency and accuracy distributions in this corpus are small-n and not symmetric, and a mean hides that.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- raw_samples
The contract permits a percentile, and only from RETAINED RAW SAMPLES. The sample count still gates which percentile is expressible, and a run that retained no samples cannot supply one however the contract is written — see the reality check beside this.
Raw samples retained by 0 / 2The contract permits raw-sample percentiles and 2 of 2 artefact(s) retain NO raw sample distribution. For those runs the percentile in the file was computed by the producer and cannot be recomputed, re-checked or extended here; the evaluation raises a blocking statistical finding for each one.
- Proportion interval
- none
No proportion interval is defined for this suite.
- What counts as a repetition
- A second sweep on a recorded host with a recorded revision. runA and runB differ only by filename and mtime, which is not identity.
- Minimum sample count
- 15
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- a tail percentile beyond p90 at n=15
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- LDPC decode only: no fronthaul, no MAC, no scheduler, no cell.
- runA and runB are distinguishable only by filename and mtime.
Caveats carried verbatim from the artefacts (0)
No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.
Reproduction
- Command withheld LDPC step-load window 2 artefact(s)Producer scripts/cuphy-cotenancy-sweep.py verified
Steps an AI load against cuPHY LDPC and records one object per load level.
Names window-clean-gpu-runA.json in its own comments (lines 277-279).
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- Exclusive GPU access while a stepped AI load runs against cuPHY.
- These two artefacts are a BARE JSON ARRAY with no envelope: the producer must be changed to emit a schema_version, a measurement instant and a device block before a re-run is comparable with anything.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.
Synthetic held-out classification accuracy
Synthetic held-out classification accuracy
heldout-accuracy@1 contract v1 6 artefactsHow accurately does the checkpoint classify a held-out split of the SAME generator?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | 6 / 6 | /repo/models/checkpoints/best.pt · /ck/best.pt · /ma/tvws-sensing/best.pt | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | 6 / 6 | single pass over the held-out split | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | ABSENT from all | none stated | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | ABSENT from all | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | 6 / 6 | 42 | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | ABSENT from all | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary overall accuracy overall_accuracy | fraction fraction | Higher is better. | proportion A proportion of a countable population. A confidence interval is expressible only when the numerator and denominator are whole records. | one synthetic held-out sample | 200 | 6 / 6 |
| Secondary random chance floor random_chance_floor | fraction scalar | No direction. This number is not rankable: "better" is undefined for it. | scalar A single stated value, not a summary of a distribution. | the class prior | 1 | 6 / 6 |
Pass/fail thresholds — and where they came from
None. This suite reports; it does not gate.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking the endpoint is proven to serve this checkpoint fact model.endpoint_weight_hash_proven · exactly true
Is there evidence that the service under test served the weights named here?
- Why it matters
- model_label is a human ASSERTION about what the URL serves; checkpoint_sha256 pins a FILE ON DISK. Both can be wrong together if the operator points --checkpoint at one model and --url at another.
- Blocks the claim
- that the deployed service serves the checkpoint named here
- On fail
- The endpoint is not proven to serve this checkpoint.
- On absent
- The service exposes no weight hash, so nothing here can bind the file to the endpoint. Provenance, not proof.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 0 / 6The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- wilson
Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.
- What counts as a repetition
- A run at an independent seed over a split with the SAME dataset, split and preprocessing hashes. Two files whose names differ is not a repetition.
Independent seeds are required, recorded IN the artefact and not inferred from a file name. - Minimum sample count
- 200
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- reading this as an OTA result
- ranking two arms whose dataset, split and preprocessing hashes are not proven identical
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- synthetic-only; NOT OTA; NOT FCC.
- A held-out split of the SAME generator that produced the training set cannot detect a leak it shares.
Caveats carried verbatim from the artefacts (1)
- scope
synthetic-only; NOT OTA; NOT FCC; NOT paper 94.2%
6 file(s): heldout-committed-8eca022f-2026-07-31.json, heldout-committed-8eca022f-IMAGE-APP-SRC-trap-2026-07-31.json, heldout-committed-8eca022f-prefixgen-17517e57-2026-07-31.json, heldout-v1-828dc46e-gen-9dc4f892-2026-07-31.json, heldout-v1-828dc46e-headgen-c23711ab-2026-07-31.json, heldout-v1-828dc46e-prefixgen-17517e57-2026-07-31.json
Reproduction
- Command offered Held-out split eval 6 artefact(s)Producer external/tvws-sensing/scripts/evaluate_heldout.py verified
Scores a checkpoint against a freshly generated held-out split.
Lines 165-166 emit random_chance_floor and held_out_test_samples.
python external/tvws-sensing/scripts/evaluate_heldout.py --checkpoint models/checkpoints/best.pt --samples-per-class 1800 --seed 42 --snr-min -10 --snr-max 30 --batch-size 32Assembled from the producer's own argparse defaults (lines 49-55). It is the producer's documented invocation, NOT a transcript of the invocation that produced these artefacts — none of them record one.
Preconditions
- A checkpoint at the --checkpoint path. The corpus records checkpoint paths but no artefact binds a checkpoint sha256 to a served endpoint.
- The split is REGENERATED from the seed, not read from disk. Re-running with the same seed does not prove the same data: only a dataset hash would, and this producer does not emit one.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a model checkpoint, which is not in the repository.
Sensing generalization across synthetic conditions
Sensing generalization across synthetic conditions
sensing-generalization@1 contract v1 15 artefactsHow does the served sensing model behave as synthetic conditions move away from the training distribution?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it. It is also served through an endpoint, so the serving seam is part of the measurement.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | ABSENT from all | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | ABSENT from all | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | ABSENT from all | none stated | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | 1 / 15 | 63c971b4d0fc967ec500675afe13d6765dd3178b | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | ABSENT from all | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | 1 / 15 | 9dc4f892f5ff34beccb01ad2d6641abcf59765bc4df0e645fc11597ecda19ede | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | 13 / 15 | 4242 · 9137 | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | 1 / 15 | 8eca022f09b897a579ef74225881343949a1a4a53b880edabf82225aacbb34ee | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | 15 / 15 | http://localhost:8103/api/v1/sense · http://localhost:8002/api/v1/sense | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary per-condition accuracy *.accuracy | fraction fraction | Higher is better. | proportion A proportion of a countable population. A confidence interval is expressible only when the numerator and denominator are whole records. | one synthetic sample at one condition, classified through the service | 200 | 15 / 15 |
Pass/fail thresholds — and where they came from
random_floor in the artefact — the chance level, not a pass mark.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking the endpoint is proven to serve this checkpoint fact model.endpoint_weight_hash_proven · exactly true
Is there evidence that the service under test served the weights named here?
- Why it matters
- model_label is a human ASSERTION about what the URL serves; checkpoint_sha256 pins a FILE ON DISK. Both can be wrong together if the operator points --checkpoint at one model and --url at another.
- Blocks the claim
- that the deployed service serves the checkpoint named here
- On fail
- The endpoint is not proven to serve this checkpoint.
- On absent
- The service exposes no weight hash, so nothing here can bind the file to the endpoint. Provenance, not proof.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 0 / 15The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- wilson
Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.
- What counts as a repetition
- The same arm at an independent seed against the same endpoint and generator hash. 0.125 is seed 4242 only; seed 9137 is 0.050.
Independent seeds are required, recorded IN the artefact and not inferred from a file name. - Minimum sample count
- 200
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- a single number averaged across the conditions
- quoting one seed as though it were the measurement
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- This is NOT over-the-air validation and does not substitute for it. A model can pass every condition here and still collapse on air.
Caveats carried verbatim from the artefacts (4)
- class_excluded
Unknown (5) — not a generatable ground truth
15 file(s): sensing-generalization-new-control-seed4242.json, sensing-generalization-new-fixA-seed4242.json, sensing-generalization-new-fixB-seed4242.json, sensing-generalization-new-legacy-seed4242.json, sensing-generalization-old-control-seed4242.json, sensing-generalization-old-control-seed9137.json, sensing-generalization-old-fixA-seed4242.json, sensing-generalization-old-fixA-seed9137.json, sensing-generalization-old-fixB-DEPLOYED-seed4242.json, sensing-generalization-old-fixB-seed4242.json, sensing-generalization-old-fixB-seed9137.json, sensing-generalization-old-legacy-seed4242.json, sensing-generalization-old-legacy-seed9137.json, sensing-generalization-runA.json, sensing-generalization-runB.json - measured_through
live service POST /api/v1/sense (deployed weights)
15 file(s): sensing-generalization-new-control-seed4242.json, sensing-generalization-new-fixA-seed4242.json, sensing-generalization-new-fixB-seed4242.json, sensing-generalization-new-legacy-seed4242.json, sensing-generalization-old-control-seed4242.json, sensing-generalization-old-control-seed9137.json, sensing-generalization-old-fixA-seed4242.json, sensing-generalization-old-fixA-seed9137.json, sensing-generalization-old-fixB-DEPLOYED-seed4242.json, sensing-generalization-old-fixB-seed4242.json, sensing-generalization-old-fixB-seed9137.json, sensing-generalization-old-legacy-seed4242.json, sensing-generalization-old-legacy-seed9137.json, sensing-generalization-runA.json, sensing-generalization-runB.json - ota_caveat
This is NOT over-the-air validation and does not substitute for it. Real OTA carries front-end nonlinearity, phase noise, ADC quantisation and real interference that no generator reproduces. A model can pass every condition here and still collapse on air. The OTA holdout in this repo is a 4 KB README stub with no IQ data.
15 file(s): sensing-generalization-new-control-seed4242.json, sensing-generalization-new-fixA-seed4242.json, sensing-generalization-new-fixB-seed4242.json, sensing-generalization-new-legacy-seed4242.json, sensing-generalization-old-control-seed4242.json, sensing-generalization-old-control-seed9137.json, sensing-generalization-old-fixA-seed4242.json, sensing-generalization-old-fixA-seed9137.json, sensing-generalization-old-fixB-DEPLOYED-seed4242.json, sensing-generalization-old-fixB-seed4242.json, sensing-generalization-old-fixB-seed9137.json, sensing-generalization-old-legacy-seed4242.json, sensing-generalization-old-legacy-seed9137.json, sensing-generalization-runA.json, sensing-generalization-runB.json - model_identity.note
model_label is a human ASSERTION about what --url serves. checkpoint_sha256 pins a FILE ON DISK — it does NOT prove the endpoint is serving that file; nothing here can, because the service exposes no weight hash. Both can therefore be wrong together if the operator points --checkpoint at one model and --url at another. Treat them as provenance, not proof, and keep candidate models on their own port so the endpoint corroborates the label. generator_sha256 IS exact: this project generates its training data in-process, so that hash pins the training distribution itself. If self_describing is false the arm is knowable only from the filename or an external manifest — see benchmarks/reports/SENSING-ARMS-MANIFEST.md.
1 file(s): sensing-generalization-old-fixB-DEPLOYED-seed4242.json
Reproduction
- Command withheld Sensing generalization arm 15 artefact(s)Producer scripts/sensing-generalization.py verified
Generates per-class samples and scores them through the deployed sensing endpoint.
Argument parser carries --per-class, --seed, --output, --url.
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT drive an endpoint that is serving the live demo.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- The sensing service reachable at --url. On this host that endpoint serves the live demo, so a sweep against it is load on a running service.
- The arm identity is only in the FILE NAME for this family (control / fixA / fixB / legacy / old / new). The producer must be changed to write the arm and the checkpoint sha256 INTO the artefact before two arms can be compared at all.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs the sensing service running and a model checkpoint that is not in the repository.
Training history
Training history
training-history@1 contract v1 3 artefactsWhat did the training curve do, and which epoch was best on the validation split?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | ABSENT from all | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | ABSENT from all | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | ABSENT from all | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | ABSENT from all | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | ABSENT from all | none stated | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | ABSENT from all | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | ABSENT from all | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | ABSENT from all | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | ABSENT from all | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | ABSENT from all | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | ABSENT from all | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | 2 / 3 | 42 · 7 | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | ABSENT from all | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | ABSENT from all | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary best validation accuracy best_val_accuracy | fraction scalar | Higher is better. | proportion A proportion of a countable population. A confidence interval is expressible only when the numerator and denominator are whole records. | one synthetic validation sample at the best epoch | 1 | 3 / 3 |
Pass/fail thresholds — and where they came from
None.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking the reported number is not a selection score fact selection.held_out_from_selection · exactly true
Was the split that produced this number also used to choose the epoch?
- Why it matters
- best_val_accuracy is the maximum over the epochs it selected from. It is optimistically biased by construction.
- Blocks the claim
- that this is an unbiased estimate of the checkpoint’s accuracy
- On fail
- This is a selection score, not a test score.
- On absent
- The artefact does not state that the scoring split was held out of selection. Treat it as a selection score.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 3 / 3The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- wilson
Wilson score interval, and only where the numerator and denominator are whole records. Where the counts are not recoverable the interval is refused in words rather than fabricated.
- What counts as a repetition
- An independent training seed with the same generator hash.
Independent seeds are required, recorded IN the artefact and not inferred from a file name. - Minimum sample count
- 1
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- quoting best_val_accuracy as a test result
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- A validation maximum is a selection score, not a test score.
- The arm identity exists only in the filename.
Caveats carried verbatim from the artefacts (0)
No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.
Reproduction
- Command withheld Training history 3 artefact(s)Producer external/tvws-sensing/scripts/train_model.py verified
Trains the sensing CNN and dumps the per-epoch history.
Line 380 writes "synthetic_history": history.to_dict().
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT take exclusive control of the GPU or the MPS control daemon.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- A full training run on a GPU. These are training curves, not a benchmark: a validation accuracy from a training loop is a SELECTION signal, not a test result.
- The producer must be changed to write a measurement instant and a dataset hash into the history file; it currently writes neither.
Continuous integration No CI jobNo workflow in .github/workflows runs this producer. It needs a GPU, a radio or an RF capture that GitHub-hosted runners do not have, so the evidence for this suite can only be produced by hand on a host that has them.
Bench harness run
Bench harness run
bench-harness-run@1 contract v1 no artefacts in this snapshotWhat was the environment, revision and dataset set that a harness run executed against?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | no artefacts | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | no artefacts | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | no artefacts | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | no artefacts | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | no artefacts | none stated | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | no artefacts | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | no artefacts | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | no artefacts | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | no artefacts | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | no artefacts | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | no artefacts | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | no artefacts | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | no artefacts | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | no artefacts | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary tests recorded in this run tests_recorded | tests count | No direction. This number is not rankable: "better" is undefined for it. | count A count of events. Whole numbers; no interpolation is meaningful. | one graded test | 1 | no artefacts |
Pass/fail thresholds — and where they came from
Per-test thresholds live in the individual tests, not in this metadata record.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
This contract declares no gate. It reports; it does not grade. That is not a pass.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- not_expressible
PERCENTILES ARE NOT EXPRESSIBLE for this suite. No p95, p99 or p99.9 may be inferred from the numbers it produces, whatever their field names say.
Raw samples retained by 0 / 0The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- none
No proportion interval is defined for this suite.
- What counts as a repetition
- A run at the same git sha, image digests and dataset hashes, executing the SAME test set.
- Minimum sample count
- 1
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- an aggregate pass/fail across a heterogeneous test set
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- The two harness runs on disk contain DIFFERENT test sets; no aggregate pass/fail is expressible across them.
- started_at is a run start and is never a measurement instant.
Caveats carried verbatim from the artefacts (0)
No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.
Reproduction
- Command withheld Bench harness run 0 artefact(s)Producer scripts/run-bench.sh verified
Creates the run directory, runs the harness and post-processes the report.
Line 64 creates reports/run-$RUN_ID; line 225 runs benchmarks/reports/make_report.py.
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT drive an endpoint that is serving the live demo.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- The demo stack up. The harness drives the running services, which on this host are serving the live demo.
- CI runs a reduced form of this on GitHub runners with TVWS_FORCE_CPU=1 and SR_USE_CPU=true — see .github/workflows/wg3-ci.yml, job e2e-bench-quick.
Continuous integration .github/workflows/wg3-ci.yml · e2e-bench-quickStep
./scripts/e2e/test-bench-quick.sh— benchmarks/reports/run-*/ uploaded as the bench-reports-${{ github.run_id }} artifact. It runs the harness on GitHub runners with TVWS_FORCE_CPU=1 and SR_USE_CPU=true, so it produces service-tier evidence and cannot produce any GPU, lab-RF or over-the-air evidence.
Bench harness per-test result
Bench harness per-test result
bench-per-test@1 contract v1 no artefacts in this snapshotWhat did one graded test measure, and did it record a status at all?
Workload and dataset identity — as the artefacts state it
Observed, not assumed. A contract may require a field; that does not make one appear. This suite is dataset-sensitive: the dataset, split and preprocessing identity are part of comparability for it.
| Field | Stated in | Distinct values | If absent |
|---|---|---|---|
workload id facts["workload.id"] | no artefacts | none stated | Without a workload id, two runs are not known to have answered the same question. |
workload hash identity.workloadHash | no artefacts | none stated | Without a workload hash, a shared workload is asserted rather than proven. |
measurement protocol facts["protocol.id"] | no artefacts | none stated | A different repetition or timing protocol produces a different quantity, however similar the field names are. |
warmup discarded facts["protocol.warmup_discarded"] | no artefacts | none stated | Discarding a different number of warmup iterations changes the distribution that was measured. |
harness harness | no artefacts | none stated | Without the producing harness named in the artefact, the measurement path is unknown. |
harness revision harnessRevision | no artefacts | none stated | Without a revision, a difference between runs may be a change in the measurement rather than in the thing measured. |
dataset id facts["dataset.id"] | no artefacts | none stated | Without a dataset id the inputs are unnamed. |
dataset hash identity.datasetHash | no artefacts | none stated | The same seed and the same sample count are not the same data. Only a hash proves two runs saw the same inputs. |
split id facts["dataset.split_id"] | no artefacts | none stated | A different split is a different test. Accuracy across two splits is not a comparison. |
preprocessing id facts["preprocessing.id"] | no artefacts | none stated | Preprocessing changes what the model sees, so two numbers computed after different preprocessing are not the same measurement. |
generator sha256 identity.generatorSha256 | no artefacts | none stated | For generated data the generator IS the dataset. Without its hash, a regenerated split is a different split. |
seed identity.seed | no artefacts | none stated | A seed without a generator hash reproduces a procedure, not a dataset. |
checkpoint sha256 identity.checkpointSha256 | no artefacts | none stated | A checkpoint hash pins a FILE. It does not prove that any endpoint served those weights. |
serving endpoint identity.endpoint | no artefacts | none stated | For a served model the endpoint is part of the measurement: a different endpoint may be a different model. |
Metric definitions and direction
| Metric | Unit | Direction | Aggregation | One sample is | Min n | Carried by |
|---|---|---|---|---|---|---|
| Primary the test’s own primary metric primary_metric | none scalar | No direction. This number is not rankable: "better" is undefined for it. | scalar A single stated value, not a summary of a distribution. | defined by the individual test | 1 | no artefacts |
Pass/fail thresholds — and where they came from
The individual test’s own threshold, recorded beside its status.
Two runs judged against different threshold versions carry different verdicts by construction, so the threshold version is part of run identity for comparison.
- Blocking opt-in components are not counted as failures fact test.all_components_engaged · exactly true
Were the components in this test actually asked to run?
- Why it matters
- In sr_model_matrix two of the four entries are opt-in and were never driven this run — ESPCN runs only in mode=fast, and the quality worker needs a POST to reload. Counting them as failures is the same defect as counting a failure as a pass, inverted.
- Blocks the claim
- reading this test’s entry count as a failure count
- On fail
- Components in this test were never engaged; their entries are absent, not failed.
- On absent
- The artefact does not record whether every component was engaged, so entries that did not run cannot be told apart from entries that failed.
Statistical method
- Central tendency
- none
No central-tendency summary is defined for this suite.
- Mean ± SD
- Not permitted
Mean ± SD is not permitted for this suite. Asking for it raises an error rather than returning a plausible number, because a symmetric summary of a skewed distribution reads as more precise than the data is.
- Percentile source
- approved_histogram
Percentiles may be read only off the named, approved histogram.
Raw samples retained by 0 / 0The contract does not source percentiles from raw samples, so retention is not required of these artefacts.
- Proportion interval
- none
No proportion interval is defined for this suite.
- What counts as a repetition
- The same test id in a run at the same revision and dataset hashes.
- Minimum sample count
- 1
Below this the number does not mean what its name says.
Treatments this suite's numbers can never support
- reading a null metric as zero
- reading a percentile with no source field as though it came from raw samples
Known limitations
Scope limits — claims this suite can never support, whatever its numbers say
- A result with no status is ABSENT: neither a pass nor a failure.
- A pass recorded against a NaN metric is a pass-by-absence, not a real pass.
Caveats carried verbatim from the artefacts (0)
No artefact of this suite carries a caveat field. That is an absence of recorded caveats, not an absence of limitations — the scope limits above still hold.
Reproduction
- Command withheld Bench harness per-test result 0 artefact(s)Producer benchmarks/bench_harness.py verified
Writes one result.json per test under reports/run-<id>/per-test/<test>/.
The per-test layout is created by the harness scripts/run-bench.sh invokes.
No command is printed. The safety policy clause it trips: A reproduction command is offered only when running it does NOT drive an endpoint that is serving the live demo.
The producer is named above. The command is not printed here because the safety policy clause below was tripped.
Preconditions
- The demo stack up, as for the parent run.
- A per-test result with no status field reads as ABSENT here, and the harness must be changed to always write a status before absence and failure can be told apart.
Continuous integration .github/workflows/wg3-ci.yml · e2e-bench-quickStep
./scripts/e2e/test-bench-quick.sh— benchmarks/reports/run-*/ uploaded as the bench-reports-${{ github.run_id }} artifact. It runs the harness on GitHub runners with TVWS_FORCE_CPU=1 and SR_USE_CPU=true, so it produces service-tier evidence and cannot produce any GPU, lab-RF or over-the-air evidence.