Structured vs Unstructured Healthcare Data: A Researcher's Guide

Structured EHR fields and clinical notes hold different information. Here's what each gives a research team and how to choose.


Structured vs. Unstructured Healthcare Data: What Researchers Need to Know

A research team scoping a treatment patterns study pulls a dataset. Diagnosis codes, medication orders, lab values, all sitting neatly in columns.

Then they look for why patients stopped the therapy. The answer isn't in the dataset. It's in the notes, because nobody codes "patient said the side effects were not worth it."

Structured and unstructured healthcare data are not two formats of the same information. They hold different things.

The short answer

Structured data is stored in defined fields with a fixed format. Diagnosis and procedure codes, medication orders, lab results, vitals, demographics. Because every field has a fixed meaning, a computer can query it directly.

Unstructured data is recorded as free text or as an image. Progress notes, consultation and referral letters, discharge summaries, radiology and pathology reports. A computer cannot query any of it until something reads it first.

Structured data gives you speed and consistency. Unstructured data gives you clinical detail and reasoning. Many studies need both, and a common and expensive mistake is assuming one can stand in for the other.

Side by side

Structured data

Unstructured data

What it looks like

Coded fields, numeric values, dropdowns

Free text, scanned documents, images

Why it exists

Billing, reporting, order entry

The clinician writing down what happened

Strength for research

Fast to query, consistent, easy to validate

Nuance, reasoning, symptoms, negation

Main weakness

Only captures what someone coded

Varies by author, needs processing first

Effort to use

Low

High. Needs natural language processing (NLP) plus human review

What happens when structured and unstructured EHR data are used together

A September 2026 study in Nature Medicine used de-identified ambulatory EHR data from Sidus Insights to test this. The team extracted clinical data from note text using large language models, then checked accuracy through physician review against a blinded expert reference standard. The team extracted clinical data from note text using large language models, then checked accuracy through physician review against a blinded expert reference standard.

The extracted information was merged with the structured records and mapped to medical ontologies, so variables from both sources could be queried together rather than analyzed separately.

They applied it to patients starting "glucagon-like peptide-1 (GLP-1) receptor agonists. The approach reconstructed individual treatment trajectories, surfaced outcomes that appeared only in the clinical notes, and supported time-to-event analysis instead of measuring at fixed points.

That last part matters because dated events pulled from clinical notes show when outcomes occurred, not just whether they had happened by a specific visit.

For researchers, the bigger question is where the variables a study depends on actually live.

What structured data gives you

Structured data is where to start for anything that involves counting:

  • Cohort sizing and prevalence

  • Treatment patterns and medication switching

  • Utilization and resource use

  • Anything that has to be reproducible by somebody else, since a code list can be audited and a text search is much harder to defend

The limitation is that coded data reflects why the record was created, and diagnosis codes exist so a claim gets paid. Severity, duration, functional status, and patient-reported symptoms usually have no field to live in. They end up in the note, or nowhere.

What unstructured data gives you

The notes are where the clinical reasoning sits:

  • Why a therapy was stopped

  • What the patient reported between visits

  • What the clinician considered and ruled out

  • Symptoms that never turned into a diagnosis

Negation belongs on that list too, and it is the one people underestimate. A note saying the patient denies chest pain is clinically meaningful, and to a keyword search it looks like nothing at all.

The cost is processing. Pulling usable variables out of notes takes NLP plus a validation step against human-reviewed records, and then a real conversation about error rates. A model that is 90% accurate at finding ejection fraction in a cardiology note is genuinely useful, and also wrong one time in ten.

How to decide what your study needs

Start from the variable list, not the data type. Write out what your protocol depends on, then ask of each item where it would realistically get recorded during routine care:

  1. Inclusion and exclusion criteria

  2. Exposure

  3. Primary endpoint

  4. Key covariates

A prescribed medication has a field. A reason for discontinuation usually does not.

If most of your list lives in coded fields, structured data will carry the study and you can move quickly. If the variables that define your population or your endpoint exist only in narrative, you need one of three things: notes access, an NLP-derived variable with documented validation, or a different endpoint. Finding this out during feasibility is cheap. Finding it out after site selection is not.

The FDA's July 2024 guidance on EHR and claims data asks much the same question. Can you show the data source captures your population, exposures, outcomes, and covariates? Have you validated those definitions, or are you trusting the coded label?

One more question worth putting to any vendor is what de-identification removes from the notes. Identifiers sit inside sentences rather than in a known column, so some contextual detail goes with them. Ask before you build an endpoint that depends on it.

Where Sidus Insights fits

Sidus Insights provides both structured and unstructured de-identified ambulatory EHR data. The structured side covers diagnoses, medications, labs, and procedures. The unstructured side is clinical notes and medical reports, which carry the context that coded fields leave out.

Unstructured files are de-identified through expert determination on a project-specific basis. That means the scope gets set against your research question rather than decided in advance, and you can ask upfront what the process keeps and what it removes.

For a research team, the useful question is which of those sources holds the variables a specific protocol needs, at the population level, in enough records to support the analysis. Better to answer that during feasibility than three months in.

Talk to us about your research question.

Frequently Asked Questions

What is structured healthcare data?

Information stored in predefined fields with a consistent format, such as diagnosis codes, procedure codes, medication orders, lab results, and demographics. Each field has a fixed meaning, so it can be queried and analyzed directly. 

Which is better for research?

Neither in general. Structured data is better for counting, cohort sizing, and reproducibility. Unstructured data is better for clinical nuance and anything with no billing code attached. It depends on which variables your protocol needs. 

Can unstructured data be converted into structured data?

Yes, using natural language processing to extract clinical concepts and map them to standard terminologies. The output is imperfect, so any NLP-derived variable should come with a validated accuracy figure measured against human review. 

What is NLP in healthcare research?

Computational methods for reading clinical text and pulling out structured information such as conditions, medications, or symptoms. It is mostly used to recover variables that exist only in notes. 

Does the FDA accept EHR data in regulatory submissions?

Yes, under conditions set out in its July 2024 final guidance on assessing EHR and medical claims data. Sponsors have to justify the data source, show it captures the relevant population and outcomes, and validate their study design definitions. 

How do I know if a dataset has the variables my study needs?

Ask for field-level completeness for your specific population, not the dataset as a whole. A dataset can be 95% complete on a variable overall and much thinner in your subgroup. 

 

Similar posts