yasa.SleepStatsAgreement#

class yasa.SleepStatsAgreement(ref_data, obs_data=None, *, ref_scorer='Reference', obs_scorer='Observed', confidence=0.95, alpha=0.05, effect_size_gates=None, bootstrap_kwargs=None, log_transform=False)[source]#

Evaluate agreement between sleep statistics reported by two different scorers. Evaluation includes bias and limits of agreement (as well as both their confidence intervals), various plotting options, and a calibration function for correcting a constant bias of the observed scorer.

Features include:

  • Get summary calculations of bias, limits of agreement, and their confidence intervals.

  • Test statistical assumptions of bias, limits of agreement, and their confidence intervals, and apply corrective procedures when the assumptions are not met.

  • Get bias and limits of agreement in a string-formatted table.

  • Calibrate new data to correct for a constant bias in observed data.

  • Visualize Bland-Altman plots.

For a complete walkthrough of the evaluation pipeline of Menghini et al. (2021), see the Evaluating a wearable or staging algorithm against a reference tutorial.

Added in version 0.7.0.

Parameters:
ref_datapandas.DataFrame

A pandas.DataFrame with sleep statistics from the reference scorer. Rows are unique observations and columns are unique sleep statistics.

Alternatively, the output of yasa.EpochByEpochAgreement.get_sleep_stats, i.e. a single pandas.DataFrame whose first index level is the scorer. In that case obs_data must be None and the two scorers are taken from the index in order of appearance (reference first).

obs_datapandas.DataFrame or None

A pandas.DataFrame with sleep statistics from the observed scorer. Rows are unique observations and columns are unique sleep statistics. Shape, index, and columns must be identical to ref_data.

ref_scorerstr

Name of the reference scorer. Ignored when obs_data is None.

obs_scorerstr

Name of the observed scorer. Ignored when obs_data is None.

confidencefloat

Confidence level (between 0 and 1) for the confidence intervals applied to bias and limits of agreement. Default is 0.95 (i.e., 95%). The same level is used for both parametric and bootstrapped confidence intervals.

alphafloat

Alpha cutoff used for all assumption tests. Default is 0.05.

effect_size_gatesdict or None

Effect-size thresholds that must be exceeded, in addition to pvalue < alpha, for the normality, constant-bias and homoscedasticity assumptions to be considered not met. Because the power of these tests grows with the number of sessions, the p-value alone would flag small and practically irrelevant deviations in large samples: the p-value assesses the statistical evidence for an association, the effect size its practical magnitude. Pass a dict to override one or more of the defaults, or None (the default) to use all of them. Setting a threshold to None disables that criterion. Keys:

  • 'skew' (default 1.0) and 'kurtosis' (default 2.0) — normality is only considered not met if the differences are markedly asymmetric (|skew| > threshold) or heavy-tailed (excess kurtosis > threshold).

  • 'r2' (default 0.1) — proportional bias is only present if at least 10% of the variability of the differences is explained by the reference value.

  • 'sd_ratio' (default 1.5) — heteroscedasticity is only present if the fitted absolute residual changes by more than 50% between the lower and the upper end of the observed reference range, i.e. if the typical error varies materially across the range. The criterion is applied symmetrically (a ratio above 1.5 or below 1/1.5), because the variability can either grow or shrink with the reference value.

Added in version 0.8.0.

bootstrap_kwargsdict

Optional keyword arguments passed to scipy.stats.bootstrap. Defaults use n_resamples=1000 and method='BCa'. The keys 'confidence_level', 'vectorized', and 'paired' cannot be overridden.

log_transformbool

If True, apply the Euser et al. (2008) log-transformation method to all sleep statistics. Limits of agreement are then expressed as bias ± slope × ref, where slope is derived from the standard deviation of log-ratio differences. This is appropriate when measurement variability is proportional to the measurement magnitude (heteroscedasticity), which is common for duration statistics such as TST, SOL, and WASO. As in the reference pipeline of Menghini et al. (2021), bias is the mean difference on the original scale and the slope is multiplied by the reference value. This is an approximation: Euser et al. (2008) derived the slope for differences relative to the mean of both scorers, and the exact back-transformed limits relative to the reference value are slightly asymmetric. When True, loa_method='auto' in report and plot_blandaltman will automatically select the Euser method for all statistics, bypassing the homoscedasticity assumption test. Statistics with a value of exactly zero in either scorer (e.g. SOL for a subject who fell asleep in the first epoch) can not be log-transformed; they are excluded with a warning and keep the regular LoA methods. Passing loa_method='log' for such statistics, or when log_transform=False, raises a ValueError. Default is False.

Notes

Sleep statistics that are identical between scorers are removed from analysis. Sessions with a missing value in either scorer (e.g. Lat_REM for a night without REM sleep) are removed separately for each sleep statistic, with a warning, so the number of sessions can differ between statistics.

References

[Menghini2021]

Menghini, L., Cellini, N., Goldstone, A., Baker, F. C., & de Zambotti, M. (2021). A standardized framework for testing the performance of sleep-tracking technology: step-by-step guidelines and open-source code. SLEEP, 44(2), zsaa170. https://doi.org/10.1093/sleep/zsaa170

Examples

>>> import pandas as pd
>>> import yasa
>>>
>>> # Generate fake reference and observed datasets with similar sleep statistics
>>> ref_scorer = "Henri"
>>> obs_scorer = "Piéron"
>>> ref_hyps = [yasa.simulate_hypnogram(tib=600, scorer=ref_scorer, seed=i) for i in range(20)]
>>> obs_hyps = [h.simulate_similar(scorer=obs_scorer, seed=i) for i, h in enumerate(ref_hyps)]
>>> # Generate sleep statistics from hypnograms using EpochByEpochAgreement
>>> eea = yasa.EpochByEpochAgreement(ref_hyps, obs_hyps)
>>> sstats = eea.get_sleep_stats()
>>> # Create SleepStatsAgreement instance
>>> ssa = yasa.SleepStatsAgreement(sstats)
>>> ssa.summary(ci_method="param").round(1).head(3)
variable   bias_intercept             bias_mean  ... loa_slope loa_upper
interval           center lower upper    center  ...     upper    center lower upper
sleep_stat                                       ...
%N1                  -5.4 -13.9   3.2       0.3  ...       0.4       6.1   3.7   8.5
%N2                 -27.3 -49.1  -5.6      -0.2  ...       0.2      12.4   7.2  17.6
%N3                  -9.1 -23.8   5.5       1.4  ...       0.6      20.4  12.6  28.3

[3 rows x 24 columns]
>>> ssa.report(ci_method="param").head(3)[["Bias [95% CI]", "LoA [95% CI]"]]
>>> ssa.assumptions["constant_bias"].head(3).round(3)
metric      slope  pvalue     r2  passed method
sleep_stat
%N1         0.370   0.181  0.097    True  param
%N2         0.553   0.017  0.279   False   regr
%N3         0.613   0.131  0.122    True  param
>>> ssa.assumptions.xs("method", level="metric", axis=1).head(3)
assumption normal constant_bias homoscedastic
sleep_stat
%N1         param         param         param
%N2         param          regr         param
%N3         param         param         param
>>> new_hyps = [h.simulate_similar(scorer="Kelly", seed=i) for i, h in enumerate(obs_hyps)]
>>> new_sstats = pd.Series(new_hyps).map(lambda h: h.sleep_statistics()).apply(pd.Series)
>>> new_sstats[["N1", "TST", "WASO"]].round(1).head(5)
     N1    TST   WASO
0  42.5  439.5  147.5
1  84.0  550.0   38.5
2  53.5  489.0  103.0
3  57.0  469.5  120.0
4  71.0  531.0   69.0
>>> new_stats_calibrated = ssa.calibrate(new_sstats[["N1", "TST", "WASO"]], bias_method="auto")
>>> new_stats_calibrated.round(1).head(5)
     N1    TST   WASO
0  42.1  445.2  145.0
1  83.6  555.8   36.0
2  53.1  494.8  100.5
3  56.6  475.2  117.5
4  70.6  536.8   66.5

Methods

__init__(ref_data[, obs_data, ref_scorer, ...])

calibrate(data[, bias_method])

Calibrate a DataFrame of sleep statistics from a new scorer based on observed biases in obs_data/obs_scorer.

plot_blandaltman([sleep_stats, bias_method, ...])

Plot Bland-Altman agreement plots for one or more sleep statistics.

report([bias_method, loa_method, ci_method, ...])

Return a human-readable DataFrame for reporting bias, limits of agreement, and statistical assumption results, following the reporting format proposed by Menghini et al. (2021).

summary([ci_method, sleep_stats])

Return a DataFrame that includes all calculated metrics:

Attributes

assumptions

A pandas.DataFrame with the results of the statistical assumption tests for each sleep statistic.

data

A long-format pandas.DataFrame containing all raw sleep statistics from ref_data and obs_data, with a MultiIndex with levels sleep_stat and session_id (or the original index name from the input data).

effect_size_gates

The effect-size thresholds used together with alpha in the assumption tests.

n_sessions

The number of sessions.

obs_scorer

The name of the observed scorer.

ref_scorer

The name of the reference scorer.

sleep_statistics

Return a list of all sleep statistics included in the agreement analyses.