Runs one of the ten DataFrame hypothesis-test facades over the exact series of an
analysis_data() frame (or a plain numeric vector), or all ten at once with
method = "summary_hypothesis". The four two-sample tests split the record at split_index,
comparing observations with a data index below it against those at or above it – the split is
on the record's INDEX, not on array position, so it agrees with the split a caller would get by
re-running the test on a record whose observations were supplied out of order.
Usage
analysis_data_hypothesis_test(
data,
method,
split_index = NULL,
lag_max = NULL,
use_log10 = FALSE
)Arguments
- data
a numeric vector of observations, or a
corehydro_dataobject fromanalysis_data()(only its exact series is read).- method
one of
"jarque_bera"(normality),"ljung_box"(autocorrelation),"equal_variance_t"/"unequal_variance_t"(difference in means, Student's / Welch's),"f"(difference in variances),"linear_trend"(trend),"wald_wolfowitz"(runs test for independence),"mann_whitney"(homogeneity / jump),"mann_kendall"(homogeneity / trend),"unimodality"(a Gaussian-mixture likelihood-ratio test, where a SMALL p-value is evidence AGAINST unimodality), or"summary_hypothesis"(all ten at once – see Details).- split_index
the record index to split the sample at; required by
"equal_variance_t","unequal_variance_t","f", and"mann_whitney", ignored otherwise.- lag_max
the maximum lag for
"ljung_box";NULL(the default) uses the library's own default rule. Ignored by every other method.- use_log10
logical; test the log10-transformed values instead of the real-space values.
Value
A named numeric vector of length 1, named method: the 2-sided p-value. For
method = "summary_hypothesis", a named numeric vector of length 10 carrying every test's
p-value, named by the library's own descriptive test names and in its own order.
Details
method = "summary_hypothesis" is the library's own ten-test summary, and it behaves
differently from calling the ten tests individually in three ways worth knowing, all inherited
from upstream:
split_indexis OPTIONAL. LeftNULL(or given a value outside the record's index range) it selects the midpoint split rather than erroring.Any single test that fails turns the WHOLE result to
NaN, rather than reporting the nine that worked. A record shorter than the 20 observations"mann_whitney"needs will therefore come back all-NaN."ljung_box"is run at the library's default lag, ignoringlag_max.
See also
analysis_data_statistics() for the summary-statistics facades,
analysis_data_summary() for plotting positions.
Examples
peaks <- c(122, 244, 214, 173, 229, 156, 212, 263, 146, 183, 161, 205)
analysis_data_hypothesis_test(peaks, "mann_kendall")
#> mann_kendall
#> 0.8370115
analysis_data_hypothesis_test(peaks, "equal_variance_t", split_index = 6)
#> equal_variance_t
#> 0.8396897