Skip to contents

Caps the weights into an absolute interval [lower, upper] and hands the removed mass back to the units that were not capped, so the weighted total is preserved. This is the step to use when you have not calibrated yet (or will calibrate afterwards) and you want the cutoff chosen from the data rather than argued for: with upper = NULL it picks one by the Tukey far-out fence or by Potter's MSE rule.

Usage

step_trim_weights(
  spec,
  lower = 1,
  upper = NULL,
  method = c("tukey", "potter"),
  redistribute = c("proportional", "uniform"),
  strict = TRUE,
  maxit = 50L,
  kappa = 1,
  id = NULL
)

Arguments

spec

a weighting_spec.

lower

numeric. Lower floor (default 1: no weight below 1).

upper

numeric or NULL. Upper cap. If NULL, the cap is chosen automatically by method.

method

rule for the automatic cap when upper = NULL: "tukey" (default, Q3 + 3*IQR far-out fence) or "potter" (Potter's MSE-optimal cutoff, which over a grid of candidate cutoffs minimizes an estimate of bias^2 + variance and so balances the bias of trimming against the variance from extreme weights). Ignored when upper is supplied. See Details for what the Potter criterion assumes.

redistribute

how the trimmed mass is shared among the untrimmed units: "proportional" (default; in proportion to their weights, preserving relative sizes) or "uniform" (an equal amount to each untrimmed unit, and units already trimmed are not reused, exactly reproducing survey::trimWeights()).

strict

logical. If TRUE (default), iterate cap+redistribution until no weight is outside [lower, upper] (like survey's strict = TRUE). If FALSE, a single pass (redistribution may push some weights slightly past the cap).

maxit

integer. Maximum iterations when strict = TRUE.

kappa

numeric, method = "potter" only: the relative price of bias against variance in the criterion minimized, kappa * bias^2 + variance. The default 1 is Potter's own weighting; above 1 the cutoff moves up (trim less, keep the bias down), below 1 it moves down (trim more).

id

optional string: a stable identifier for this step, shown in the recipe print-out and usable to select it in collect_step_detail(); defaults to a derived "<class>_<k>".

Value

The input weighting_spec with this step appended to its recipe. The step is recorded only; it is evaluated when prep() is called.

Details

Two things are worth knowing about method = "potter" before reading its cutoff as optimal, because both are assumptions rather than results.

The criterion is the mean squared error of a total, so its two terms are deliberately on different orders: the bias of capping is the trimmed mass, which grows like the sample size, and the variance term is the sum of squared remaining weights, which also grows like the sample size, so bias squared grows like its square. The cutoff therefore rises with the sample size for the same weight distribution – replicating a 300-unit sample to 30,000 moved it from 126 to 183 in one test, capping a third as many units. That is the correct behaviour for a total (the relative bias stays put while the relative variance shrinks, so trimming buys less), not a defect, but it does mean the rule is not a property of the weight distribution alone. The criterion is invariant to the scale of the weights.

The bias it charges is the bias of capping without redistribution, while this step always redistributes: the weighted total is preserved exactly, so the real bias is only the difference between the study variable's mean among the capped units and among the units receiving their mass. The criterion therefore overstates the bias and, other things equal, caps less than the MSE it is named after would. kappa is the handle: set it below 1 to buy back that conservatism.

Examples

weighting_spec(sample_survey, base_weights = pw) |>
  step_nonresponse(respondent = responded, method = "weighting_class", by = "region") |>
  step_trim_weights(lower = 1, strict = TRUE) |> prep()
#> 
#> == Weighting specification (weightflow) ==
#> Data    : 467 cases
#> Base wts: pw
#> Steps   :
#>   1. nonresponse (weighting class)  [nonresponse_1]
#>   2. auto weight trimming  [trim_weights_1]
#> Status  : estimated (prep)
#> 
#> Stage summary:
#>                      stage n_active sum_wts cv_wts deff_kish n_eff
#>                       base      467    4371  0.236     1.056   442
#>   stage_1_step_nonresponse      270    4371  0.144     1.021   265
#>  stage_2_step_trim_weights      270    4371  0.144     1.021   265
#> 
#> deff_kish = 1 + CV^2 (Kish design effect from unequal weighting);
#> n_eff = n_active / deff_kish. Both worsen with each adjustment and
#> improve with trimming.
#> 

# Potter MSE-optimal cutoff chosen from the data
weighting_spec(sample_survey, base_weights = pw) |>
  step_nonresponse(respondent = responded, method = "weighting_class", by = "region") |>
  step_trim_weights(method = "potter") |> prep()
#> 
#> == Weighting specification (weightflow) ==
#> Data    : 467 cases
#> Base wts: pw
#> Steps   :
#>   1. nonresponse (weighting class)  [nonresponse_1]
#>   2. auto weight trimming (Potter MSE)  [trim_weights_1]
#> Status  : estimated (prep)
#> 
#> Stage summary:
#>                      stage n_active sum_wts cv_wts deff_kish n_eff
#>                       base      467    4371  0.236     1.056   442
#>   stage_1_step_nonresponse      270    4371  0.144     1.021   265
#>  stage_2_step_trim_weights      270    4371  0.137     1.019   265
#> 
#> deff_kish = 1 + CV^2 (Kish design effect from unequal weighting);
#> n_eff = n_active / deff_kish. Both worsen with each adjustment and
#> improve with trimming.
#>