UpEvidence · Independent Published Methods Example

Benchmarking MA Utilization by Individual Service

An independent published methods example comparing service counts across MA payment arrangements rather than relying on encounter spending.

Home / UpEvidence / Benchmarking MA Utilization by Individual Service

Source published October 2, 2026 · Resource reviewed October 8, 2026

The research question

Which billed services differ between full-risk and nonrisk MA arrangements?

Data files and cohort construction

Humana enrollment, claims, encounters, and contract information from 2015–2019 formed beneficiary-year cohorts. Monthly primary-care attribution identified full-risk HMO and nonrisk PPO arrangements. Group changes, hospice, missing information, and delegated claims processing were excluded.

Measures and analytic design

Professional and outpatient records were deduplicated into service counts for 1,271 HCPCS billing codes, excluding physician-administered drugs. Augmented inverse-probability weighting combined contract-group probability and expected-use models; shrinkage stabilized sparse services. Adjustment included demographics, enrollment-based disadvantage, pharmacy-based health measures, geography, and year. A fee-schedule simulation assumed supply responsiveness and unchanged total spending.

Robustness checks and interpretation

Checks used partial-risk and nonrisk HMO comparisons, ordinary regression, alternative deduplication, geography, engagement, and separate years. Alternative elasticities tested the simulation.

Practical application: count care before interpreting dollars

The following are FastHSR implementation considerations, not additional procedures claimed for the study. A payer could use service-level benchmarking to identify categories for clinical review. Keep contract exposure separate from the claims table; a zero-dollar encounter can still represent a delivered service. Display raw and adjusted service counts per eligible person-year together, with uncertainty and a minimum reporting threshold.

Decisions to settle before reuse

Confirm that contract rosters identify actual financial risk, not just product labels. Prespecify attribution, observation time, code versions, and whether facility and professional bills describe one service or two. Reconcile delegated processing and encounter submission before comparing groups. A contract with missing records should not look artificially efficient.

Limitations and local application

Selection, HMO/PPO differences, and encounter completeness can bias comparisons. Lower utilization does not establish better care; simulated payment responses are assumptions, not observed effects. For a local application, start with a small, clinically interpretable service set and review differences with practitioners. Pair utilization with safety or quality indicators before proposing payment changes. If modeling a fee schedule, show the budget constraint and vary behavioral assumptions; do not present a single simulation as a forecast.

Source article and supplement

Schwartz AL, Bulat T, Ben-Michael E, et al. Service-Level Utilization in Risk-Based Contracts and a Benchmarking Approach for Fee-Schedule Reforms. JAMA Health Forum. 2026;7(10):e263533. doi:10.1001/jamahealthforum.2026.3533.

The full article and supplement listings were reviewed on October 8, 2026. Separate supplementary downloads could not be retrieved. The summary uses methods described in the article; verify deduplication rules and model details in the supplements before replication.

Frequently asked question

Does lower service use prove a risk contract improves care?

No. Lower use can reflect efficiency, unmet need, patient differences, or incomplete records. Clinical and data-quality review are needed to interpret a utilization benchmark.

Explore more UpEvidence