Caps the weights into an absolute interval [lower, upper] and hands the
removed mass back to the units that were not capped, so the weighted total is
preserved. This is the step to use when you have not calibrated yet (or
will calibrate afterwards) and you want the cutoff chosen from the data rather
than argued for: with upper = NULL it picks one by the Tukey far-out fence or
by Potter's MSE rule.
Arguments
- spec
a weighting_spec.
- lower
numeric. Lower floor (default 1: no weight below 1).
- upper
numeric or NULL. Upper cap. If NULL, the cap is chosen automatically by
method.- method
rule for the automatic cap when
upper = NULL: "tukey" (default, Q3 + 3*IQR far-out fence) or "potter" (Potter's MSE-optimal cutoff, which over a grid of candidate cutoffs minimizes an estimate of bias^2 + variance and so balances the bias of trimming against the variance from extreme weights). Ignored whenupperis supplied. See Details for what the Potter criterion assumes.- redistribute
how the trimmed mass is shared among the untrimmed units: "proportional" (default; in proportion to their weights, preserving relative sizes) or "uniform" (an equal amount to each untrimmed unit, and units already trimmed are not reused, exactly reproducing survey::trimWeights()).
- strict
logical. If TRUE (default), iterate cap+redistribution until no weight is outside
[lower, upper](like survey's strict = TRUE). If FALSE, a single pass (redistribution may push some weights slightly past the cap).- maxit
integer. Maximum iterations when strict = TRUE.
- kappa
numeric,
method = "potter"only: the relative price of bias against variance in the criterion minimized,kappa * bias^2 + variance. The default 1 is Potter's own weighting; above 1 the cutoff moves up (trim less, keep the bias down), below 1 it moves down (trim more).- id
optional string: a stable identifier for this step, shown in the recipe print-out and usable to select it in
collect_step_detail(); defaults to a derived"<class>_<k>".
Value
The input weighting_spec with this step appended to its recipe. The
step is recorded only; it is evaluated when prep() is called.
Details
Two things are worth knowing about method = "potter" before reading
its cutoff as optimal, because both are assumptions rather than results.
The criterion is the mean squared error of a total, so its two terms are deliberately on different orders: the bias of capping is the trimmed mass, which grows like the sample size, and the variance term is the sum of squared remaining weights, which also grows like the sample size, so bias squared grows like its square. The cutoff therefore rises with the sample size for the same weight distribution – replicating a 300-unit sample to 30,000 moved it from 126 to 183 in one test, capping a third as many units. That is the correct behaviour for a total (the relative bias stays put while the relative variance shrinks, so trimming buys less), not a defect, but it does mean the rule is not a property of the weight distribution alone. The criterion is invariant to the scale of the weights.
The bias it charges is the bias of capping without redistribution, while
this step always redistributes: the weighted total is preserved exactly, so the
real bias is only the difference between the study variable's mean among the
capped units and among the units receiving their mass. The criterion therefore
overstates the bias and, other things equal, caps less than the MSE it is
named after would. kappa is the handle: set it below 1 to buy back that
conservatism.
See also
Other weighting steps:
step_assert(),
step_calibrate(),
step_cre(),
step_drop_ineligible(),
step_model_calibration(),
step_nonresponse(),
step_nr_sensitivity(),
step_pseudoweight(),
step_rescale(),
step_round(),
step_select_within(),
step_subsample(),
step_trim(),
step_trim_calibrated(),
step_unknown_eligibility()
Examples
weighting_spec(sample_survey, base_weights = pw) |>
step_nonresponse(respondent = responded, method = "weighting_class", by = "region") |>
step_trim_weights(lower = 1, strict = TRUE) |> prep()
#>
#> == Weighting specification (weightflow) ==
#> Data : 467 cases
#> Base wts: pw
#> Steps :
#> 1. nonresponse (weighting class) [nonresponse_1]
#> 2. auto weight trimming [trim_weights_1]
#> Status : estimated (prep)
#>
#> Stage summary:
#> stage n_active sum_wts cv_wts deff_kish n_eff
#> base 467 4371 0.236 1.056 442
#> stage_1_step_nonresponse 270 4371 0.144 1.021 265
#> stage_2_step_trim_weights 270 4371 0.144 1.021 265
#>
#> deff_kish = 1 + CV^2 (Kish design effect from unequal weighting);
#> n_eff = n_active / deff_kish. Both worsen with each adjustment and
#> improve with trimming.
#>
# Potter MSE-optimal cutoff chosen from the data
weighting_spec(sample_survey, base_weights = pw) |>
step_nonresponse(respondent = responded, method = "weighting_class", by = "region") |>
step_trim_weights(method = "potter") |> prep()
#>
#> == Weighting specification (weightflow) ==
#> Data : 467 cases
#> Base wts: pw
#> Steps :
#> 1. nonresponse (weighting class) [nonresponse_1]
#> 2. auto weight trimming (Potter MSE) [trim_weights_1]
#> Status : estimated (prep)
#>
#> Stage summary:
#> stage n_active sum_wts cv_wts deff_kish n_eff
#> base 467 4371 0.236 1.056 442
#> stage_1_step_nonresponse 270 4371 0.144 1.021 265
#> stage_2_step_trim_weights 270 4371 0.137 1.019 265
#>
#> deff_kish = 1 + CV^2 (Kish design effect from unequal weighting);
#> n_eff = n_active / deff_kish. Both worsen with each adjustment and
#> improve with trimming.
#>
