Today we are releasing the Cascade Ambulatory Reference Set: roughly 4,000 hours of continuous single-lead ECG from 310 participants, recorded on Axon Sense 2 devices over an eighteen-month period, de-identified and free to download.

It is not the largest ECG dataset in the world and it is not trying to be. What makes it useful is what is attached to it.

Provenance, attached to every recording

Public physiological datasets tend to arrive as signal plus label. That is enough to train something and not enough to explain a disagreement. When two groups get different results from the same data, the interesting question is usually about the instrument, and the instrument is exactly what has been stripped out.

Every recording in this set carries:

  • The device model and serial number.
  • The exact firmware version, by build hash, not by release name.
  • The device's calibration record, including the date of its last calibration and the drift measured at that point.
  • The sensor front-end configuration in force, including gain and filter settings.
  • Time synchronisation quality for the recording window, in milliseconds.
  • Every signal-quality flag the device raised at the time, rather than a cleaned-up version.

That last one matters more than it sounds. Most released datasets are quietly cleaned. Segments the device flagged as poor quality have been dropped, which makes the set easier to work with and makes it a bad proxy for what an algorithm will meet in the field. We have left them in and flagged them, so you can choose.

How consent worked

Participants consented specifically to open release, at enrolment, in plain language, with the option to participate in the study without contributing to the public set. Forty-one people took that option and their data is not here.

The consent language, the participant information sheet, and the de-identification procedure are all published alongside the data. If you are designing a study and want to do the same thing, that material is more reusable than the signals are.

We asked people whether their heart data could be given away, and a surprising number said yes, provided we said plainly who would get it and what they could do with it. Plainly was the operative word.

Dr. Naomi Achterberg, Cascade University

The dataset

Duration
4,012 hours across 310 participants
Signal
Single-lead ECG, 512 Hz, 16-bit
Formats
EDF+, Parquet, and raw binary with a reader
Annotations
Beat-level, two independent annotators, disagreements preserved rather than resolved
License
CC BY 4.0, no citation of Nexaform required
Size
218 GB uncompressed

Disagreements are preserved

Two annotators labeled every recording independently. Where they disagreed, we have kept both labels rather than adjudicating to a single truth. About 2.3 percent of beats have a disagreement attached.

That subset is the most interesting part of the release. It is small, it is hard, and it is where the difference between an algorithm that works and an algorithm that reports that it works tends to live.

What we would like people to do with it

Anything, honestly. The license is CC BY 4.0 and we have deliberately not asked to be cited, because a citation requirement from a device manufacturer creates an incentive we would rather not create.

If you find something wrong with the data, tell us, and we will publish the correction with your name on it if you want it there.