Claims data vs EHR data: which one fits your research question?

Claims data shows what was billed. EHR data shows what happened in the exam room. Here is how to pick the right real-world data source for your study.


A research team plans a retrospective study on treatment patterns in type 2 diabetes. They buy a claims dataset, pull two million patients, and then find out they cannot tell which patients had an HbA1c above 9. The codes say the patients have diabetes. Nothing says how sick they were. 

That is the claims vs EHR tradeoff in one example. Both are real-world data. They were built for completely different purposes, and that purpose shows up in every analysis you run.

What claims data actually is

Claims data is billing data. When a provider treats a patient and sends a bill to a payer, that transaction creates a record: diagnosis codes (ICD-10), procedure codes (CPT and HCPCS), pharmacy fills (National Drug Codes), dates of service, place of service, and paid amounts.

Nobody created that record for research. It exists so someone gets paid.

What EHR data actually is 

EHR data is created during care. Vitals, lab results, medication orders, problem lists, imaging reports, procedure notes, and the clinician's own narrative in the chart. Some of it sits in structured fields. A lot of it sits in free text and PDFs that need work before anyone can analyze them.

For research use, this data is de-identified and analyzed at the cohort level.

Side by side 

Claims data

EHR data

Built for

Payment

Care delivery

Clinical depth

Codes, unless linked to a clinical or lab feed

Labs, vitals, notes, results

Follows the patient across providers

Yes, within the plan

Only within the health system or network

Drug information

What was dispensed

What was prescribed

Disease severity

Usually not available directly

Usually available

Cost and utilization

Strong

Weak or absent

Time lag

Weeks to months after service

Days, depending on refresh cadence

Continuity

Breaks when coverage changes

Breaks when the patient changes providers

Analysis effort

Lower, already structured

Higher, much of it unstructured

Where claims data wins

Claims follow the patient wherever they use their coverage. If a patient sees a cardiologist in one system, fills a prescription at a retail pharmacy, and ends up in an urgent care across town, the payer sees all three. No single EHR does.

That makes claims the better source when your question is about utilization, total cost of care, medication adherence, or care that happens outside the practice you have data from.

Denominators are cleaner too. If you know who was enrolled and for how long, you know who was at risk during the study window.

Where claims data falls down

Codes are not clinical facts. A rule-out code for chest pain looks identical to a confirmed diagnosis. Coding also responds to reimbursement rules, so what gets recorded is shaped partly by what gets paid. You also lose the patient when coverage changes, since someone who switches employers in March disappears from the dataset in April and reappears somewhere else, or not at all.

And unless the claims are linked to a lab or clinical feed, there are no labs, no vitals, no stage, no BMI, no smoking status. For a study that needs any of those as an inclusion criterion or an endpoint, standard administrative claims will not get you there.

Where EHR data wins

EHR data tells you what the clinician saw. Lab values, dosing changes, blood pressure trending over eight visits, the note explaining why a therapy was stopped. If you are phenotyping a cohort, validating an endpoint, or screening for trial eligibility, that detail is the whole point.

It is also fast. A lab resulted today shows up today, not after a payer adjudicates a claim six weeks later.

Ambulatory EHR data specifically captures the part of the patient journey that happens between hospital visits, which is where most chronic disease management happens.

Where EHR data falls down

Fragmentation. Your dataset covers the practices in it and nothing else. Care delivered elsewhere is invisible, and the patient looks healthier than they are.

A prescription written is not a prescription filled. If your question involves adherence or persistence, EHR data on its own will overstate both.

Documentation habits also vary between clinicians and specialties. One doctor records smoking status at every visit, another records it once in 2019.

What happened when researchers compared EHR data with Medicare claims

A team from the University of Florida, Purdue, Regenstrief, and Penn checked one academic health system's EHR against Medicare claims for the same patients, roughly 12,900 adults aged 65 and over with type 2 diabetes per year. The results were published in JAMIA Open in June 2026.

The clinical measurements were generally reliable within the EHR. Among patients with an HbA1c at or above 6.5%, 98.6% also had a type 2 diabetes diagnosis recorded the same year. The Medicare data did not contain HbA1c or BMI (body mass index) values at all, so for those variables the EHR was the only source.

Coverage was the weak spot. The EHR captured 38% of the encounters and 68% of the deaths that appeared in claims for the same patients. Spotting when something started was harder still, with accuracy of 7.9% for incident hypertension and 20.2% for incident chronic kidney disease. A patient who began a drug elsewhere looks like a new start in your data. Linking the two sources improved completeness on the billing-derived elements while keeping the clinical variables claims did not have.

The practical lesson attaches to a dataset, not to a data type. Coverage varies by network and population, so the thing to establish early is how much of your cohort's care a dataset actually sees, and whether the variables you need are populated in it.

How to choose between claims and EHR data

Let the endpoint decide, and be specific about what the endpoint is made of.

When cohort definitions or outcomes depend on labs, vitals, medication detail, or what the clinician wrote in the note, EHR data is the stronger starting point. Clinical detail is the harder thing to recreate later. If you have labs and notes, you can define a cohort and measure an outcome. If you only have codes, you can count events but you cannot describe the patient.

When the question turns on utilization, filled prescriptions, cost, or care delivered outside the network you have data from, claims is the stronger starting point. For studies where healthcare cost or paid utilization is a primary outcome, claims data is typically essential.

The order matters more than the choice. Adding claims to an EHR study fills gaps in the patient journey. Adding clinical depth to a claims-only study after the fact is much harder, and sometimes not possible at all.

When to link both

Linked data is worth the effort when your question spans care and cost at the same time. A few cases where it pays off:

  • Safety studies where an adverse event may have been treated outside the network

  • Adherence research comparing what was prescribed against what was actually filled

  • Health economics work that needs a clinical severity measure alongside spend

  • Long-term outcomes where patients move between providers over a decade

The cost of linking is complexity. Tokenization, overlap rates, and privacy review all take time, and the overlap between two datasets is usually smaller than either one alone. Budget for a feasibility count before you design the study around it.

A quick way to decide

Ask what your primary endpoint is made of.

If it is built from codes and dollars, use claims. If it is built from lab values, vitals, notes, or patient-reported outcomes, use EHR. If it is built from both, plan for linked data and a longer timeline.

Then ask where the patient's care happens. Concentrated in one network points to EHR. Spread across systems and pharmacies points to claims.

Where Sidus Insights fits

Sidus Insights works with de-identified ambulatory EHR data across specialties including cardiology, oncology, OB/GYN, behavioral health, primary care, gastroenterology, and pediatrics. That includes structured records plus clinical notes, lab results, prescribing data, and de-identified patient reports, with longitudinal history going back many years for a large share of patients.

If your eligibility criteria, endpoints, or subgroup definitions depend on clinical variables that billing codes don't carry, EHR is the place to start. The question to settle early is whether those variables are actually populated in the data you plan to use. The team works with research groups on exactly that, helping them assess which data sources and clinical variables suit a given study. A feasibility count answers it before the study design is locked.

Discuss your research question

Frequently Asked Questions

What is the main difference between claims data and EHR data?

Claims data records what was billed to a payer. EHR data records what happened during care. Claims give you broad coverage with limited clinical detail. EHR gives you clinical detail with narrower coverage. 

Which is better for clinical research, claims or EHR data?

It depends on the endpoint. Studies built on lab values, vitals, or disease severity need EHR data. Studies built on cost, utilization, or medication adherence need claims. Neither is better in general.

Can claims data show lab results?

Administrative claims usually show that a lab service happened and was billed, but not the resulting value. Some commercial datasets link claims to a lab feed and do carry results. If your study needs the value, check whether the specific dataset includes it rather than assuming either way.

What is linked claims and EHR data?

It is a dataset where records from both sources are matched to the same de-identified individual using privacy-preserving tokens. It gives you clinical detail plus the full picture of care across providers, at the cost of a smaller usable population.

How far back does longitudinal data usually go?

It varies by source. Claims history is generally limited by enrollment periods. EHR history depends on how long the practice has been on the system, and some ambulatory datasets like Sidus Insights carry more than 20 years of records for many patients.

Which data source is better for clinical trial recruitment?

EHR data, because eligibility criteria usually depend on lab values, medications, and recent visit activity that claims do not carry. Claims can help confirm whether a patient is still actively receiving care.

Similar posts