agentsclimarketplace

Sensys reproducibility

Skill brycewang-stanford/Awesome-Journal-Skills/SenSys-Skills/skills/sensys-reproducibility

Use when making a SenSys result reproducible across a different testbed — capturing energy-measurement method, hardware and firmware provenance, sensor ground-truth protocol, and deployment conditions while the testbed is still live, and deciding early which traces and firmware can legally and safely ship.From its SKILL.md

Install
npx -y skills add brycewang-stanford/Awesome-Journal-Skills --skill sensys-reproducibility

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

4.6 KB, 957 tokens by cl100k_base, as published. Nobody here has run it

SenSys Reproducibility

A SenSys result is reproducible when someone on different hardware can understand — and, from your traces, re-obtain — your numbers. The threat is not messy code; it is provenance that evaporates when the deployment ends. Energy measured without its method, accuracy scored against forgotten ground truth, and firmware whose toolchain is unrecorded cannot be reproduced at any price once the testbed is torn down. Capture it while the nodes are still powered.

Capture it live, not later

ProvenanceCapture while liveWhy it cannot be reconstructed
Energy methodInstrument model, sampling rate, integration boundaries, sleep floorA stored current number loses its method; "42 µA" of what phases?
HardwareMCU/SoC part + clock, sensor config, radio + TX power, board revision, battery/harvester specA later "same board" is rarely bit-identical
FirmwareSources + toolchain version + compiler flags, pinnedA rebuild on a new toolchain shifts timing and energy
Ground truthReference protocol + the reference's own errorLabels without their protocol are unfalsifiable
DeploymentPlacement, environment, duration, node uptime/failuresConditions are gone once the deployment ends
Harvest inputEnergy-source trace (light/RF/vibration), buffer sizingBehavior across brownouts is unreproducible without the input

The reproducibility that matters is cross-hardware

Reproducing your own result on your own testbed proves little. The SenSys bar is that a different lab can interpret a mismatch: when their number differs from yours, the provenance tells them why (different MCU clock, different harvest input) rather than leaving them to guess. Ship recorded traces so the analysis can be re-run even by someone who lacks your hardware — this is what turns a deployment result into a reproducible one and underpins the trace-replay path in sensys-artifact-evaluation.

Minimum reproducible bundle for a SenSys figure:
  traces/fig4_power.csv        # raw power samples + timestamps
  traces/fig4_sensor.csv       # the sensor stream behind the same figure
  analysis/reproduce_fig4.py   # regenerates Fig. 4 from the two traces
  ENERGY.md                    # instrument, rate, integration boundaries
  HARDWARE.md                  # part numbers, clocks, board rev, harvester spec
  firmware/ + toolchain.lock   # pinned build that produced the traces

Decide early what can ship

Some SenSys artifacts carry release constraints that a purely computational paper never faces:

  • Sensor data of people or spaces may need consent, anonymization, or aggregation before it can be released — decide at collection time, because retroactive consent is usually impossible.
  • Firmware and hardware designs may touch third-party IP or a vendor NDA; confirm what is releasable before promising an Available badge.
  • Deployment locations can be sensitive (critical infrastructure, private property); scrub identifying detail from released traces.

Resolving these late forces a choice between a weak artifact and a broken promise. Resolve them in the experiment plan (sensys-experiments).

Reproducibility checklist

[ ] Energy method (instrument, rate, boundaries) recorded, not just the number.
[ ] Hardware provenance: parts, clocks, board revision, battery/harvester spec.
[ ] Firmware sources + pinned toolchain + compiler flags archived.
[ ] Ground-truth protocol and the reference's own error documented.
[ ] Deployment conditions + honest uptime/failure log captured while live.
[ ] Harvest-input traces + buffer sizing stored for batteryless results.
[ ] Recorded traces shipped so figures re-generate without your hardware.
[ ] Release constraints (consent, NDA, location) resolved at collection time.

Output format

[Live-capture] which provenance is captured vs. still at risk while the testbed runs
[Cross-HW]     can a different lab interpret a mismatch from what you shipped? Y/N
[Traces]       do shipped traces regenerate the headline figures without hardware? Y/N
[Release]      consent/NDA/location constraints resolved? open items
[Open]         the provenance whose loss would most damage reproducibility

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,871. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.