Skip to contents

Re-estimates a propensity score model using only the observations retained after trimming. This is the recommended intermediate step between ps_trim() and weight calculation (e.g. wt_ate()):

ps_trim() -> ps_refit() -> wt_*()

Trimming changes the target population by removing observations with extreme propensity scores. Refitting the model on the retained subset produces propensity scores that better reflect this population, improving both model fit and downstream weight estimation. Weight functions warn if a trimmed propensity score has not been refit.

Usage

ps_refit(trimmed_ps, model, .data = NULL, ...)

Arguments

trimmed_ps

A ps_trim object returned by ps_trim(). Refitting reads the retained positions out of the trimming record, so an object whose record was dropped or no longer covers it raises an error of class propensity_missing_meta_error; see ps_trim(). A single column of a data frame made from a trimmed score matrix raises an error of class propensity_method_error; refit the matrix before converting it. Truncated scores from ps_trunc() keep every unit, so refitting them would return the original model's scores, and they raise an error of class propensity_method_error.

model

The original fitted model used to estimate the propensity scores (e.g. a glm or multinom object). The model is refit via update() on the retained subset. Trimmed propensity scores are refit only by a model of the probability of the exposure; a model of a conditional mean, such as a stats::lm() fit or a gaussian stats::glm(), never produced them and raises an error of class propensity_model_family_error. A dose model trimmed with method = "density" or method = "resid" is refit only by a model of the dose's conditional mean, such as a gaussian stats::glm(), stats::lm(), mgcv::gam(), or MASS::rlm() fit, and any other model, such as one of a probability, raises the same error.

.data

A data frame with one row per observation in trimmed_ps, in the same order. If NULL (the default), the data are recovered from model: its model.frame() when that already holds every variable the refit reads, and otherwise the data the model names, restricted by row name to the rows the model analyzed. A model fit without a data argument names none, and its variables are read out of the formula's environment instead. A formula that transforms a term, such as z ~ log(x) or a spline basis, stores that term already computed, so only the underlying variables let the transformation be recomputed from the retained rows. Pass .data when the data the model was fit on can no longer be reached.

...

Additional arguments passed to update(), such as subset or weights. Each is evaluated against the retained rows, the data the refit is handed: a name is read from their columns first and otherwise from the environment ps_refit() was called from, so subset = x1 > 0 refits on the retained rows where x1 is positive. A logical or index vector supplied instead indexes the retained rows, not the full data. The .data and .env pronouns are available to tell a column from a variable of the same name. Without .data, the retained rows hold only the variables model reads, so an argument naming any other column needs the data frame passed to .data.

Value

A ps_trim object with re-estimated propensity scores for retained observations and NA for trimmed observations. Use is_refit() to confirm refitting was applied. Refitting replaces the values a calibrated score held with predictions from the refit model, so a score trimmed after ps_calibrate() is no longer calibrated and is_ps_calibrated() answers FALSE for the result.

For a trimmed dose model, the retained values are the refit model's conditional means, and the spread in the record is re-estimated from the retained residuals under the recorded density family, unless the trim was made at a .sigma the caller supplied, which is kept. The recorded threshold describes the cut that was made and is left as it was, so the refit model can place a retained unit's conditional density below it; nothing is trimmed again.

Details

Composing with a subset

A subset in the original call has already chosen the sample the propensity scores are about, and the trimming record indexes that sample rather than every row the data carry. Refitting narrows that sample further, to the retained rows, so the original subset is dropped from the call rather than put to work a second time on rows it was never about. A subset passed through ... is an instruction of its own and is honored.

Which level a refit predicts

For a vector of scores, the refit predicts the probability of the level the scores describe: when ps_trim() was given a model with its first level named as focal, the refit reports one minus the model's prediction, and for scores supplied directly it reports the level the model predicts by default.

Arguments read from outside the formula

weights, offset, and na.action in the original call are re-evaluated against the retained rows. A weights or offset naming a column of the data the model was fit on is read from that column and follows the retained rows, whether the data are recovered from model or passed to .data. A vector held outside the data cannot follow them: it keeps the length it had and raises an error about differing variable lengths.

Scores predicted from a fit with na.action = na.exclude are padded back to the full length of the data, so they describe more observations than the fit read and the trimming record indexes a sample the model never analyzed. ps_refit() refuses such scores. Trim scores from a fit whose na.action drops those rows instead.

See also

ps_trim() for the trimming step, is_refit() to check refit status, wt_ate() and other weight functions for the next step in the pipeline.

Examples

set.seed(2)
n <- 200
x <- rnorm(n)
z <- rbinom(n, 1, plogis(0.4 * x))

# fit a propensity score model
ps_model <- glm(z ~ x, family = binomial)
ps <- predict(ps_model, type = "response")

# trim -> refit -> weight pipeline
trimmed <- ps_trim(ps, lower = 0.1, upper = 0.9)
refit <- ps_refit(trimmed, ps_model)
wts <- wt_ate(refit, .exposure = z)
#> ℹ Treating `.exposure` as binary

is_refit(refit)
#> [1] TRUE