Four small, reproducible household-panel datasets that ship with weightflow to test and
illustrate the panel tools. They share one structure and differ only in the rotation
system, so the same code runs on a pure panel and on each rotating design. All are in
long format: one row per person and per wave the person is in sample. Continuing units
keep the same household_id / person_id across waves, so they link the panel; the household
is the natural cluster and c(household_id, person_no) the person-level key.
Format
A data.frame with one row per person-wave and the columns:
- household_id
household id, persistent across waves (the panel link / cluster).
- person_id, person_no
person id and person-number within household;
c(household_id, person_no)is the person key.- stratum, psu
design stratum and primary sampling unit (PSU nested in stratum, at least two PSUs per stratum), for the coordinated bootstrap / jackknife.
- region, sex, age
covariates usable as estimation domains.
- wave
wave (month) index.
- rotation_group, month_in_sample
rotation group and order-in-sample.
- pw
design (base) weight.
- disposition
between-wave disposition:
"R","NR","OS","UNK".- lf_status
labour status within the working-age population, a factor with levels
"emp"/"unemp"/"inact";NAfor non-respondents. This is thestatusargument ofstep_cre()(the previous-wave composite auxiliary).- employed, unemployed
employed / unemployed indicators (labour force;
NAif not"R"or not in the labour force).- income
labour income (
NAif not"R").
Details
The between-wave disposition disposition follows the four-state taxonomy the longitudinal
cascade needs: "R" responded, "NR" eligible nonresponse (reweight), "OS" out of scope
– left the target population between waves, so the household exits permanently and is not
reweighted – and "UNK" unknown eligibility. The variables of interest (employed,
unemployed, income) are observed only when disposition == "R" (and, for the labour-force
items, when the person is in the labour force), otherwise NA; they repeat across waves for
continuing persons, with within-person persistence, so net change and gross flows are
meaningful.
The datasets differ only in the rotation calendar (which waves each rotation group is in sample), giving different overlap structures:
panel_puroPure panel, no rotation: every unit is followed across all 4 waves; the sample shrinks only through attrition. Style of EU-SILC / SLID pure panels.
panel_clChile ENE, 2-2-2 (in-out-in): 3 waves, consecutive overlap ~1/2, and units that return (in sample in waves 1 and 3 but not 2). The real ENE has no public rotation group;
rotation_groupis included for teaching, but the panel also links through the persistent ids alone.panel_ineINE Uruguay ECH / StatCan LFS, 6-month rotation: 6 groups in sample each wave, one sixth rotating out per wave, so consecutive overlap ~5/6.
panel_usUS CPS, 4-8-4: 3 waves, consecutive overlap ~3/4, with a cohort that leaves and returns after the 8-month gap.
Examples
# rotation structure of the 6-month panel
panel_design(panel_ine, unit = c("household_id", "person_no"), wave = "wave",
rotation_group = "rotation_group", pattern = "6")
#> <weightflow panel design>
#> waves : 3 (1, 2, 3)
#> unit : household_id + person_no
#> rotation : rotation_group pattern: 6
#> units : 2783 (linked in >=2 waves: 2063, 74%)
#> overlap (row wave retained in column wave):
#> 1 2 3
#> 1 1.00 0.83 0.64
#> 2 0.83 1.00 0.81
#> 3 0.65 0.82 1.00
#> Pr(panel selection), adjacent : 0.833, 0.833 (full combination: 0.667)
#> overlap implied by pattern : 0.83 0.67 (lag 1 2)
#> pattern : 6 group(s) in sample, cycle 6, useful lags 1, 2, 3, 4, 5
# coordinated change of the unemployment rate between two waves
t1 <- subset(panel_ine, wave == 1 & disposition == "R")
t2 <- subset(panel_ine, wave == 2 & disposition == "R")
wb <- wave_bootstrap(
list(T1 = weighting_spec(t1, base_weights = pw),
T2 = weighting_spec(t2, base_weights = pw)),
replicates = 100, strata = "stratum", psu = "psu", seed = 1, progress = FALSE)
change_mean(wb, "unemployed")
#> <weightflow net change>
#> T1 -> T2
#> change : -0.0149512 SE 0.010056
#> 95% CI : [-0.0346606, 0.00475807]
#> V1 0.0001126 | V2 0.0001035 | Cov 5.75e-05 | rho 0.533 (levels)
#> V = 0.0001011 vs V1+V2 = 0.0002161 (deff_change 0.468: overlap saved 53%)
