yasa.SleepStatsAgreement#
- class yasa.SleepStatsAgreement(ref_data, obs_data=None, *, ref_scorer='Reference', obs_scorer='Observed', confidence=0.95, alpha=0.05, effect_size_gates=None, bootstrap_kwargs=None, log_transform=False)[source]#
Evaluate agreement between sleep statistics reported by two different scorers. Evaluation includes bias and limits of agreement (as well as both their confidence intervals), various plotting options, and a calibration function for correcting a constant bias of the observed scorer.
Features include:
Get summary calculations of bias, limits of agreement, and their confidence intervals.
Test statistical assumptions of bias, limits of agreement, and their confidence intervals, and apply corrective procedures when the assumptions are not met.
Get bias and limits of agreement in a string-formatted table.
Calibrate new data to correct for a constant bias in observed data.
Visualize Bland-Altman plots.
For a complete walkthrough of the evaluation pipeline of Menghini et al. (2021), see the Evaluating a wearable or staging algorithm against a reference tutorial.
See also
Added in version 0.7.0.
- Parameters:
- ref_data
pandas.DataFrame A
pandas.DataFramewith sleep statistics from the reference scorer. Rows are unique observations and columns are unique sleep statistics.Alternatively, the output of
yasa.EpochByEpochAgreement.get_sleep_stats, i.e. a singlepandas.DataFramewhose first index level is the scorer. In that caseobs_datamust beNoneand the two scorers are taken from the index in order of appearance (reference first).- obs_data
pandas.DataFrameor None A
pandas.DataFramewith sleep statistics from the observed scorer. Rows are unique observations and columns are unique sleep statistics. Shape, index, and columns must be identical toref_data.- ref_scorerstr
Name of the reference scorer. Ignored when
obs_dataisNone.- obs_scorerstr
Name of the observed scorer. Ignored when
obs_dataisNone.- confidencefloat
Confidence level (between 0 and 1) for the confidence intervals applied to bias and limits of agreement. Default is 0.95 (i.e., 95%). The same level is used for both parametric and bootstrapped confidence intervals.
- alphafloat
Alpha cutoff used for all assumption tests. Default is 0.05.
- effect_size_gatesdict or None
Effect-size thresholds that must be exceeded, in addition to
pvalue < alpha, for the normality, constant-bias and homoscedasticity assumptions to be considered not met. Because the power of these tests grows with the number of sessions, the p-value alone would flag small and practically irrelevant deviations in large samples: the p-value assesses the statistical evidence for an association, the effect size its practical magnitude. Pass a dict to override one or more of the defaults, orNone(the default) to use all of them. Setting a threshold toNonedisables that criterion. Keys:'skew'(default 1.0) and'kurtosis'(default 2.0) — normality is only considered not met if the differences are markedly asymmetric (|skew| >threshold) or heavy-tailed (excesskurtosis >threshold).'r2'(default 0.1) — proportional bias is only present if at least 10% of the variability of the differences is explained by the reference value.'sd_ratio'(default 1.5) — heteroscedasticity is only present if the fitted absolute residual changes by more than 50% between the lower and the upper end of the observed reference range, i.e. if the typical error varies materially across the range. The criterion is applied symmetrically (a ratio above 1.5 or below 1/1.5), because the variability can either grow or shrink with the reference value.
Added in version 0.8.0.
- bootstrap_kwargsdict
Optional keyword arguments passed to
scipy.stats.bootstrap. Defaults usen_resamples=1000andmethod='BCa'. The keys'confidence_level','vectorized', and'paired'cannot be overridden.- log_transformbool
If
True, apply the Euser et al. (2008) log-transformation method to all sleep statistics. Limits of agreement are then expressed asbias ± slope × ref, whereslopeis derived from the standard deviation of log-ratio differences. This is appropriate when measurement variability is proportional to the measurement magnitude (heteroscedasticity), which is common for duration statistics such as TST, SOL, and WASO. As in the reference pipeline of Menghini et al. (2021),biasis the mean difference on the original scale and the slope is multiplied by the reference value. This is an approximation: Euser et al. (2008) derived the slope for differences relative to the mean of both scorers, and the exact back-transformed limits relative to the reference value are slightly asymmetric. WhenTrue,loa_method='auto'inreportandplot_blandaltmanwill automatically select the Euser method for all statistics, bypassing the homoscedasticity assumption test. Statistics with a value of exactly zero in either scorer (e.g. SOL for a subject who fell asleep in the first epoch) can not be log-transformed; they are excluded with a warning and keep the regular LoA methods. Passingloa_method='log'for such statistics, or whenlog_transform=False, raises aValueError. Default isFalse.
- ref_data
Notes
Sleep statistics that are identical between scorers are removed from analysis. Sessions with a missing value in either scorer (e.g.
Lat_REMfor a night without REM sleep) are removed separately for each sleep statistic, with a warning, so the number of sessions can differ between statistics.References
[Menghini2021]Menghini, L., Cellini, N., Goldstone, A., Baker, F. C., & de Zambotti, M. (2021). A standardized framework for testing the performance of sleep-tracking technology: step-by-step guidelines and open-source code. SLEEP, 44(2), zsaa170. https://doi.org/10.1093/sleep/zsaa170
Examples
>>> import pandas as pd >>> import yasa >>> >>> # Generate fake reference and observed datasets with similar sleep statistics >>> ref_scorer = "Henri" >>> obs_scorer = "Piéron" >>> ref_hyps = [yasa.simulate_hypnogram(tib=600, scorer=ref_scorer, seed=i) for i in range(20)] >>> obs_hyps = [h.simulate_similar(scorer=obs_scorer, seed=i) for i, h in enumerate(ref_hyps)] >>> # Generate sleep statistics from hypnograms using EpochByEpochAgreement >>> eea = yasa.EpochByEpochAgreement(ref_hyps, obs_hyps) >>> sstats = eea.get_sleep_stats() >>> # Create SleepStatsAgreement instance >>> ssa = yasa.SleepStatsAgreement(sstats) >>> ssa.summary(ci_method="param").round(1).head(3) variable bias_intercept bias_mean ... loa_slope loa_upper interval center lower upper center ... upper center lower upper sleep_stat ... %N1 -5.4 -13.9 3.2 0.3 ... 0.4 6.1 3.7 8.5 %N2 -27.3 -49.1 -5.6 -0.2 ... 0.2 12.4 7.2 17.6 %N3 -9.1 -23.8 5.5 1.4 ... 0.6 20.4 12.6 28.3 [3 rows x 24 columns]
>>> ssa.report(ci_method="param").head(3)[["Bias [95% CI]", "LoA [95% CI]"]]
>>> ssa.assumptions["constant_bias"].head(3).round(3) metric slope pvalue r2 passed method sleep_stat %N1 0.370 0.181 0.097 True param %N2 0.553 0.017 0.279 False regr %N3 0.613 0.131 0.122 True param
>>> ssa.assumptions.xs("method", level="metric", axis=1).head(3) assumption normal constant_bias homoscedastic sleep_stat %N1 param param param %N2 param regr param %N3 param param param
>>> new_hyps = [h.simulate_similar(scorer="Kelly", seed=i) for i, h in enumerate(obs_hyps)] >>> new_sstats = pd.Series(new_hyps).map(lambda h: h.sleep_statistics()).apply(pd.Series) >>> new_sstats[["N1", "TST", "WASO"]].round(1).head(5) N1 TST WASO 0 42.5 439.5 147.5 1 84.0 550.0 38.5 2 53.5 489.0 103.0 3 57.0 469.5 120.0 4 71.0 531.0 69.0
>>> new_stats_calibrated = ssa.calibrate(new_sstats[["N1", "TST", "WASO"]], bias_method="auto") >>> new_stats_calibrated.round(1).head(5) N1 TST WASO 0 42.1 445.2 145.0 1 83.6 555.8 36.0 2 53.1 494.8 100.5 3 56.6 475.2 117.5 4 70.6 536.8 66.5
Methods
__init__(ref_data[, obs_data, ref_scorer, ...])calibrate(data[, bias_method])Calibrate a
DataFrameof sleep statistics from a new scorer based on observed biases inobs_data/obs_scorer.plot_blandaltman([sleep_stats, bias_method, ...])Plot Bland-Altman agreement plots for one or more sleep statistics.
report([bias_method, loa_method, ci_method, ...])Return a human-readable
DataFramefor reporting bias, limits of agreement, and statistical assumption results, following the reporting format proposed by Menghini et al. (2021).summary([ci_method, sleep_stats])Return a
DataFramethat includes all calculated metrics:Attributes
A
pandas.DataFramewith the results of the statistical assumption tests for each sleep statistic.A long-format
pandas.DataFramecontaining all raw sleep statistics fromref_dataandobs_data, with aMultiIndexwith levelssleep_statandsession_id(or the original index name from the input data).The effect-size thresholds used together with
alphain the assumption tests.The number of sessions.
The name of the observed scorer.
The name of the reference scorer.
Return a list of all sleep statistics included in the agreement analyses.