Structured vs. Unstructured Healthcare Data: What Researchers Need to Know
A research team scoping a treatment patterns study pulls a dataset. Diagnosis codes, medication orders, lab values, all sitting neatly in columns.
Then they look for why patients stopped the therapy. The answer isn't in the dataset. It's in the notes, because nobody codes "patient said the side effects were not worth it."
Structured and unstructured healthcare data are not two formats of the same information. They hold different things.
The short answer
Structured data is stored in defined fields with a fixed format. Diagnosis and procedure codes, medication orders, lab results, vitals, demographics. Because every field has a fixed meaning, a computer can query it directly.
Unstructured data is recorded as free text or as an image. Progress notes, consultation and referral letters, discharge summaries, radiology and pathology reports. A computer cannot query any of it until something reads it first.
Structured data gives you speed and consistency. Unstructured data gives you clinical detail and reasoning. Many studies need both, and a common and expensive mistake is assuming one can stand in for the other.
Side by side
|
|
Structured data
|
Unstructured data
|
|
What it looks like
|
Coded fields, numeric values, dropdowns
|
Free text, scanned documents, images
|
|
Why it exists
|
Billing, reporting, order entry
|
The clinician writing down what happened
|
|
Strength for research
|
Fast to query, consistent, easy to validate
|
Nuance, reasoning, symptoms, negation
|
|
Main weakness
|
Only captures what someone coded
|
Varies by author, needs processing first
|
|
Effort to use
|
Low
|
High. Needs natural language processing (NLP) plus human review
|
What happens when structured and unstructured EHR data are used together
A September 2026 study in Nature Medicine used de-identified ambulatory EHR data from Sidus Insights to test this. The team extracted clinical data from note text using large language models, then checked accuracy through physician review against a blinded expert reference standard. The team extracted clinical data from note text using large language models, then checked accuracy through physician review against a blinded expert reference standard.
The extracted information was merged with the structured records and mapped to medical ontologies, so variables from both sources could be queried together rather than analyzed separately.
They applied it to patients starting "glucagon-like peptide-1 (GLP-1) receptor agonists. The approach reconstructed individual treatment trajectories, surfaced outcomes that appeared only in the clinical notes, and supported time-to-event analysis instead of measuring at fixed points.
That last part matters because dated events pulled from clinical notes show when outcomes occurred, not just whether they had happened by a specific visit.
For researchers, the bigger question is where the variables a study depends on actually live.
What structured data gives you
Structured data is where to start for anything that involves counting:
-
Cohort sizing and prevalence
-
Treatment patterns and medication switching
-
Utilization and resource use
-
Anything that has to be reproducible by somebody else, since a code list can be audited and a text search is much harder to defend
The limitation is that coded data reflects why the record was created, and diagnosis codes exist so a claim gets paid. Severity, duration, functional status, and patient-reported symptoms usually have no field to live in. They end up in the note, or nowhere.
What unstructured data gives you
The notes are where the clinical reasoning sits:
-
Why a therapy was stopped
-
What the patient reported between visits
-
What the clinician considered and ruled out
-
Symptoms that never turned into a diagnosis
Negation belongs on that list too, and it is the one people underestimate. A note saying the patient denies chest pain is clinically meaningful, and to a keyword search it looks like nothing at all.
The cost is processing. Pulling usable variables out of notes takes NLP plus a validation step against human-reviewed records, and then a real conversation about error rates. A model that is 90% accurate at finding ejection fraction in a cardiology note is genuinely useful, and also wrong one time in ten.
How to decide what your study needs
Start from the variable list, not the data type. Write out what your protocol depends on, then ask of each item where it would realistically get recorded during routine care:
-
Inclusion and exclusion criteria
-
Exposure
-
Primary endpoint
-
Key covariates
A prescribed medication has a field. A reason for discontinuation usually does not.
If most of your list lives in coded fields, structured data will carry the study and you can move quickly. If the variables that define your population or your endpoint exist only in narrative, you need one of three things: notes access, an NLP-derived variable with documented validation, or a different endpoint. Finding this out during feasibility is cheap. Finding it out after site selection is not.
The FDA's July 2024 guidance on EHR and claims data asks much the same question. Can you show the data source captures your population, exposures, outcomes, and covariates? Have you validated those definitions, or are you trusting the coded label?
One more question worth putting to any vendor is what de-identification removes from the notes. Identifiers sit inside sentences rather than in a known column, so some contextual detail goes with them. Ask before you build an endpoint that depends on it.
Where Sidus Insights fits
Sidus Insights provides both structured and unstructured de-identified ambulatory EHR data. The structured side covers diagnoses, medications, labs, and procedures. The unstructured side is clinical notes and medical reports, which carry the context that coded fields leave out.
Unstructured files are de-identified through expert determination on a project-specific basis. That means the scope gets set against your research question rather than decided in advance, and you can ask upfront what the process keeps and what it removes.
For a research team, the useful question is which of those sources holds the variables a specific protocol needs, at the population level, in enough records to support the analysis. Better to answer that during feasibility than three months in.
Talk to us about your research question.
Frequently Asked Questions