yasa.EpochByEpochAgreement#
- class yasa.EpochByEpochAgreement(ref_hyps, obs_hyps)[source]#
Evaluate agreement between two hypnograms or two collections of hypnograms.
Evaluation includes averaged agreement scores, one-vs-rest agreement scores, agreement scores summarized across all sleep and summarized by sleep stage, and various plotting options to visualize the two hypnograms simultaneously. See examples for more detail.
For a complete walkthrough of the evaluation pipeline of Menghini et al. (2021), see the Evaluating a wearable or staging algorithm against a reference tutorial.
Added in version 0.7.0.
- Parameters:
- ref_hypsiterable of
yasa.Hypnogram A collection of reference hypnograms (i.e., those considered ground-truth).
Each
yasa.Hypnograminref_hypsmust have the samescorer.If a
dict, key values are used to generate unique sleep session IDs. If any other iterable (e.g.,listortuple), then unique sleep session IDs are automatically generated.- obs_hypsiterable of
yasa.Hypnogram A collection of observed hypnograms (i.e., those to be evaluated).
Each
yasa.Hypnograminobs_hypsmust have the samescorer, and this scorer must be different than the scorer of hypnograms inref_hyps.If a
dict, key values must match those ofref_hyps.Important
It is assumed that the order of hypnograms are the same in
ref_hypsandobs_hyps. For example, the third hypnogram inref_hypsandobs_hypsmust come from the same sleep session, and they must only differ in that they have different scorers.See also
For comparing just two hypnograms, use
yasa.Hypnogram.evaluate.
- ref_hypsiterable of
Notes
Epochs scored as artefact (
ART) or unscored (UNS) by either scorer are excluded from all epoch-by-epoch analyses (agreement scores and confusion matrices), since they are not a sleep stage. They are kept in the hypnograms themselves, e.g. inget_sleep_statsandplot_hypnograms.Changed in version 0.8.0:
ARTandUNSepochs were previously counted as regular stages.References
[Menghini2021]Menghini, L., Cellini, N., Goldstone, A., Baker, F. C., & de Zambotti, M. (2021). A standardized framework for testing the performance of sleep-tracking technology: step-by-step guidelines and open-source code. SLEEP, 44(2), zsaa170. https://doi.org/10.1093/sleep/zsaa170
Examples
>>> import yasa >>> ref_hyps = [yasa.simulate_hypnogram(tib=600, scorer="Human", seed=i) for i in range(10)] >>> obs_hyps = [h.simulate_similar(scorer="YASA", seed=i) for i, h in enumerate(ref_hyps)] >>> ebe = yasa.EpochByEpochAgreement(ref_hyps, obs_hyps) >>> agr = ebe.get_agreement() >>> agr.head(5).round(1) accuracy balanced_acc kappa mcc precision f1 sleep_id 1 30.6 26.0 0.1 0.1 30.6 30.6 2 33.3 32.7 0.1 0.1 34.9 33.5 3 35.1 23.9 0.1 0.1 34.8 34.6 4 22.5 21.4 0.0 0.0 20.6 20.5 5 21.4 16.9 -0.1 -0.1 20.3 20.7
>>> ebe.get_agreement_bystage().head(12).round(1) fbeta npv precision recall specificity support stage sleep_id WAKE 1 39.1 ... 37.1 41.3 ... 189.0 2 29.9 ... 27.6 32.6 ... 184.0 ... N1 1 18.5 ... 18.5 18.5 ... 124.0 2 12.1 ... 13.1 11.2 ... 160.0
>>> ebe.get_confusion_matrix(sleep_id=1) YASA WAKE N1 N2 N3 REM Human WAKE 78 24 50 3 34 N1 23 23 43 15 20 N2 60 58 183 43 139 N3 30 10 50 5 32 REM 19 9 121 50 78
>>> import yasa >>> import matplotlib.pyplot as plt >>> ref_hyps = [yasa.simulate_hypnogram(tib=600, scorer="Human", seed=i) for i in range(10)] >>> obs_hyps = [h.simulate_similar(scorer="YASA", seed=i) for i, h in enumerate(ref_hyps)] >>> ebe = yasa.EpochByEpochAgreement(ref_hyps, obs_hyps) >>> fig, ax = plt.subplots(figsize=(6, 3), constrained_layout=True) >>> _ = ebe.plot_hypnograms(sleep_id=10)
>>> import yasa >>> import matplotlib.pyplot as plt >>> ref_hyps = [yasa.simulate_hypnogram(tib=600, scorer="Human", seed=i) for i in range(10)] >>> obs_hyps = [h.simulate_similar(scorer="YASA", seed=i) for i, h in enumerate(ref_hyps)] >>> ebe = yasa.EpochByEpochAgreement(ref_hyps, obs_hyps) >>> fig, ax = plt.subplots(figsize=(6, 3)) >>> _ = ebe.plot_hypnograms( ... sleep_id=8, ax=ax, obs_kwargs={"color": "red", "lw": 2, "ls": "dotted"} ... ) >>> plt.tight_layout()
>>> import yasa >>> import matplotlib.pyplot as plt >>> ref_hyps = [yasa.simulate_hypnogram(tib=600, scorer="Human", seed=i) for i in range(10)] >>> obs_hyps = [h.simulate_similar(scorer="YASA", seed=i) for i, h in enumerate(ref_hyps)] >>> ebe = yasa.EpochByEpochAgreement(ref_hyps, obs_hyps) >>> session = 8 >>> fig, ax = plt.subplots(figsize=(6.5, 2.5), constrained_layout=True) >>> style_a = dict(alpha=1, lw=2.5, ls="solid", color="gainsboro", label="Michel") >>> style_b = dict(alpha=1, lw=2.5, ls="solid", color="cornflowerblue", label="Jouvet") >>> legend_style = dict( ... title="Scorer", frameon=False, ncol=2, loc="lower center", bbox_to_anchor=(0.5, 0.9) ... ) >>> ax = ebe.plot_hypnograms( ... sleep_id=session, ref_kwargs=style_a, obs_kwargs=style_b, legend=legend_style, ax=ax ... ) >>> acc = ebe.get_agreement().at[session, "accuracy"] >>> _ = ax.text( ... 0.01, 1, f"Accuracy = {acc:.0f}%", ha="left", va="bottom", transform=ax.transAxes ... )
When comparing only 2 hypnograms, use the
evaluatemethod:>>> hypno_a = yasa.simulate_hypnogram(tib=90, scorer="RaterA", seed=8) >>> hypno_b = hypno_a.simulate_similar(scorer="RaterB", seed=9) >>> ebe = hypno_a.evaluate(hypno_b) >>> ebe.get_confusion_matrix() RaterB WAKE N1 N2 N3 RaterA WAKE 71 2 20 8 N1 1 0 9 0 N2 12 4 25 0 N3 24 0 1 3
Methods
__init__(ref_hyps, obs_hyps)get_agreement([sample_weight, scorers, pooled])Return a
pandas.DataFrameof weighted (i.e., averaged) agreement scores.get_agreement_bystage([beta, zero_division])Return a
pandas.DataFrameof unweighted (i.e., one-vs-rest) agreement scores.get_confusion_matrix([sleep_id, agg_func])Return a
ref_hyp/obs_hypconfusion matrix from either a single session or all sessions concatenated together.Return the group-level proportional confusion (error) matrix, i.e. the mean, standard deviation, and confidence interval across sessions of the row-normalized per-session confusion matrices, as reported in Menghini et al. (2021).
Return a
pandas.DataFrameof sleep statistics for each hypnogram derived from both reference and observed scorers.multi_scorer(df, scorers)Compute multiple agreement scores from a 2-column dataframe (an optional 3rd column may contain sample weights).
plot_hypnograms([sleep_id, legend, ax, ...])Plot the two hypnograms of one session overlapping on the same axis.
summary([by_stage, ci_method, confidence, ...])Return group-level agreement scores.
Attributes
A
pandas.DataFrameincluding all hypnograms, without the epochs scored asARTorUNSby either scorer.The number of unique sleep sessions.
The name of the observed scorer.
The name of the reference scorer.