Skip to contents

bw_entropy() specifies entropy balancing for balance(). The weights minimize the Kullback-Leibler divergence from a set of base weights subject to the covariate constraints, so among all reweightings that achieve balance the solution stays as close as possible to the base weights. Entropy balancing supports binary, categorical, and continuous exposures.

Usage

bw_entropy(
  ...,
  base_weights = NULL,
  distribution_moments = NULL,
  convergence_tolerance = 1e-10,
  max_iterations = NULL
)

Arguments

...

Reserved for future extensions; must be empty. Tuning parameters must be passed by name.

base_weights

A numeric vector of base weights, one per observation, or NULL for uniform base weights. The estimated weights minimize sum(w * log(w / base_weights)).

distribution_moments

For continuous exposures, the number of exposure and covariate marginal moments held equal to the sample under the base measure, or NULL for the constraint moments. Raised automatically when smaller than the constraint moments. The base measure is the product of the sampling weights and base_weights, so without either the marginals are held equal to the unweighted sample.

convergence_tolerance

The solver convergence tolerance. What it measures depends on which problem is solved. The exact problem, chosen when every tolerance in balance_terms() is zero, measures the gradient sup norm and responds to this value across its range. A positive tolerance selects the inexact problem, solved by FISTA against the relative change in the loss; that criterion is the weaker of the two, so the value is tightened to at most 1e-14 to hold the achieved balance inside the requested box, and anything above 1e-14 is inert there. 1e-10 is both this argument's default and the value the solver resolves for NULL.

max_iterations

The maximum solver iterations, or NULL for the resolved default of 1000. When the L-BFGS then Newton hybrid runs, either as the automatic retry of a Newton solve that came back short or because options(balancing.entropy_solver = "lbfgs_then_newton") asked for it, the cap applies to each phase separately and the reported iteration count is the sum of the two, so such a fit can report more iterations than the cap.

Value

An bw_entropy specification, a balance_method.

Details

For a binary exposure the average treatment effect reweights each exposure group to the pooled covariate means, and the average treatment effect on the treated reweights the control group to the treated covariate means while the treated group is left unreweighted. Its reported weights are its base weights carried to the group's sampling-weighted total rather than the base weights themselves, so they are proportional to the base weights and constant base weights come back as ones whatever level they were set at. When every requested tolerance is zero the constraints hold exactly and the weights solve smooth estimating equations, which balance() records for the M-estimation variance in ipw(). A positive tolerance in balance_terms() selects the inexact problem, which balances each constraint to within the tolerance and does not produce estimating equations.

The exact problem is solved by Newton's method, which starts from the base measure and is the only solver that drives the estimating equations to machine precision. A flat or badly scaled constraint set can leave that cold start short of its tolerance, so a failed Newton solve is retried once with the L-BFGS-then-Newton hybrid, which reaches a neighborhood with L-BFGS before polishing it with Newton and so ends at the same precision. The retry announces itself, and @solver_status records the solver the returned fit came from. Pinning the balancing.entropy_solver option, described in balancing_options, selects one solver and disables the retry.

References

Hainmueller, J. (2012). Entropy balancing for causal effects: A multivariate reweighting method to produce balanced samples in observational studies. Political Analysis, 20(1), 25-46.

Examples

n <- 200
x1 <- rnorm(n)
x2 <- rnorm(n)
df <- data.frame(
  exposure = rbinom(n, 1, plogis(0.5 * x1 - 0.5 * x2)),
  x1 = x1,
  x2 = x2
)
fit <- balance(df, exposure, c(x1, x2), method = bw_entropy())
#> ℹ Treating `.exposure` as binary
fit
#> 
#> ── Entropy balancing ───────────────────────────────────────────────────────────
#> Exposure: "exposure" (binary)
#> Estimand: "ate"
#> Observations: 200
#> Solver: converged in 4 iterations
#> Constraints: 2 terms (tolerance 0)
#> Largest imbalance: 4.38e-11 (standardized mean difference)