Statistical twin of patient claims data

The patterns are real. The patients aren't.

ClaimsTwin generates statistical twins of high-value patient populations: every record is synthetic, while every statistic is measured from 200 million real patients across ten years of claims.

Built by FastHSR, the claims-analytics team behind peer-reviewed research in JAMA Network Open, Health Affairs, Health Services Research, Medical Care, Journal of Clinical Oncology, and AJMC.

Claims data is either locked or broken.

Real claims sit behind HIPAA, data-use agreements, and operational controls that can take months or years to clear. Even then, patient-level files often cannot be redistributed to engineers, vendors, or models.

Common substitutes lose the parts that matter. Deidentified files may remove dates of service. Date-shifted files may preserve sequence but are often old, narrow, or difficult to share. Static extracts miss new NDCs and J-codes, new CPT/HCPCS codes for emerging care models, newly enrolled providers, and post-policy-change utilization patterns.

Public alternatives do not solve the problem. Perturbed files can distort the statistics analysts need. Rule-based synthetic files can create plausible-looking patients without matching the empirical structure of real claims.

Until now, realistic and shareable were opposites.

A statistical twin, not a stand-in.

ClaimsTwin takes a third path. Instead of scrambling real records or inventing patients from rules, it measures the statistical structure of real claims: frequencies, relationships, sequences, costs, and utilization patterns. It then generates entirely new records that preserve those structures.

The records are generated. The statistics are measured. Nothing is copied, and nothing is guessed.

Complete claims architecture.

Enrollment spine plus inpatient, outpatient, carrier, SNF, home health, hospice, and pharmacy files, with full referential integrity.

Longitudinal patient timelines.

Multi-year histories with realistic care sequences, enrollment transitions, and end-of-life patterns.

Real vocabularies, real prices.

Current ICD-10, CPT/HCPCS, and NDC code sets, with dollars and payment rules.

Cohorts at any scale.

Define the population, geography, and period, then generate the cohort size your project needs.

The scorecard is the product.

Every ClaimsTwin release ships with evidence, not promises:

  • Structural validity: eligibility logic, date logic, code-set vintages, edit compliance, referential integrity, and pricing reconciliation.
  • Statistical fidelity: code frequencies, comorbidity co-occurrence, cost distributions, utilization rates, and HCC RAF distributions compared with real benchmarks under explicit tolerances.
  • Analytic utility: train-on-synthetic/test-on-real model performance and replication of completed real-data studies, with confidence-interval overlap reported.
  • Distinguishability: machine-discriminator audits.

Things only a twin can do.

  • Cohorts on demand. Generate the population you actually need, such as dual-eligible male CHF patients aged 70-74 in Henrico County, Virginia, instead of settling for the sample you can access.
  • Policy-regime scenarios. Generate the same population under different rule eras, such as PBM benefit redesign or risk-model transitions, and study the difference.
  • The recipe attached. Real claims are a mixture of medicine, economics, and administration. ClaimsTwin generates that mixture in layers, so the same underlying patient population can be rendered under different coding and payment conditions.

Back to Claims Data Analytics