Re-estimates a propensity score model using only the observations retained
after trimming. This is the recommended intermediate step between
ps_trim() and weight calculation (e.g. wt_ate()):
ps_trim() -> ps_refit() -> wt_*()
Trimming changes the target population by removing observations with extreme propensity scores. Refitting the model on the retained subset produces propensity scores that better reflect this population, improving both model fit and downstream weight estimation. Weight functions warn if a trimmed propensity score has not been refit.
Arguments
- trimmed_ps
A
ps_trimobject returned byps_trim(). Refitting reads the retained positions out of the trimming record, so an object whose record was dropped or no longer covers it raises an error of classpropensity_missing_meta_error; seeps_trim(). A single column of a data frame made from a trimmed score matrix raises an error of classpropensity_method_error; refit the matrix before converting it. Truncated scores fromps_trunc()keep every unit, so refitting them would return the original model's scores, and they raise an error of classpropensity_method_error.- model
The original fitted model used to estimate the propensity scores (e.g. a glm or multinom object). The model is refit via update() on the retained subset. Trimmed propensity scores are refit only by a model of the probability of the exposure; a model of a conditional mean, such as a
stats::lm()fit or a gaussianstats::glm(), never produced them and raises an error of classpropensity_model_family_error. A dose model trimmed withmethod = "density"ormethod = "resid"is refit only by a model of the dose's conditional mean, such as a gaussianstats::glm(),stats::lm(),mgcv::gam(), orMASS::rlm()fit, and any other model, such as one of a probability, raises the same error.- .data
A data frame with one row per observation in
trimmed_ps, in the same order. IfNULL(the default), the data are recovered frommodel: its model.frame() when that already holds every variable the refit reads, and otherwise the data the model names, restricted by row name to the rows the model analyzed. A model fit without a data argument names none, and its variables are read out of the formula's environment instead. A formula that transforms a term, such asz ~ log(x)or a spline basis, stores that term already computed, so only the underlying variables let the transformation be recomputed from the retained rows. Pass.datawhen the data the model was fit on can no longer be reached.- ...
Additional arguments passed to update(), such as
subsetorweights. Each is evaluated against the retained rows, the data the refit is handed: a name is read from their columns first and otherwise from the environmentps_refit()was called from, sosubset = x1 > 0refits on the retained rows wherex1is positive. A logical or index vector supplied instead indexes the retained rows, not the full data. The.dataand.envpronouns are available to tell a column from a variable of the same name. Without.data, the retained rows hold only the variablesmodelreads, so an argument naming any other column needs the data frame passed to.data.
Value
A ps_trim object with re-estimated propensity scores for retained
observations and NA for trimmed observations. Use is_refit() to
confirm refitting was applied. Refitting replaces the values a calibrated
score held with predictions from the refit model, so a score trimmed after
ps_calibrate() is no longer calibrated and is_ps_calibrated() answers
FALSE for the result.
For a trimmed dose model, the retained values are the refit model's
conditional means, and the spread in the record is re-estimated from the
retained residuals under the recorded density family, unless the trim was
made at a .sigma the caller supplied, which is kept. The recorded
threshold describes the cut that was made and is left as it was, so the
refit model can place a retained unit's conditional density below it;
nothing is trimmed again.
Details
Composing with a subset
A subset in the original call has already chosen the sample the propensity
scores are about, and the trimming record indexes that sample rather than
every row the data carry. Refitting narrows that sample further, to the
retained rows, so the original subset is dropped from the call rather than
put to work a second time on rows it was never about. A subset passed
through ... is an instruction of its own and is honored.
Which level a refit predicts
For a vector of scores, the refit predicts the probability of the level the
scores describe: when ps_trim() was given a model with its first level
named as focal, the refit reports one minus the model's prediction, and for
scores supplied directly it reports the level the model predicts by default.
Arguments read from outside the formula
weights, offset, and na.action in the original call are re-evaluated
against the retained rows. A weights or offset naming a column of the
data the model was fit on is read from that column and follows the retained
rows, whether the data are recovered from model or passed to .data. A
vector held outside the data cannot follow them: it keeps the length it had
and raises an error about differing variable lengths.
Scores predicted from a fit with na.action = na.exclude are padded back to
the full length of the data, so they describe more observations than the fit
read and the trimming record indexes a sample the model never analyzed.
ps_refit() refuses such scores. Trim scores from a fit whose na.action
drops those rows instead.
See also
ps_trim() for the trimming step, is_refit() to check refit
status, wt_ate() and other weight functions for the next step in the
pipeline.
Examples
set.seed(2)
n <- 200
x <- rnorm(n)
z <- rbinom(n, 1, plogis(0.4 * x))
# fit a propensity score model
ps_model <- glm(z ~ x, family = binomial)
ps <- predict(ps_model, type = "response")
# trim -> refit -> weight pipeline
trimmed <- ps_trim(ps, lower = 0.1, upper = 0.9)
refit <- ps_refit(trimmed, ps_model)
wts <- wt_ate(refit, .exposure = z)
#> ℹ Treating `.exposure` as binary
is_refit(refit)
#> [1] TRUE
