Care gap analyzer: find unmet care needs in the full clinical record
The Care Gap Analyzer identifies patients who are not receiving evidence-based care: a preventive screening never done, chronic disease monitoring that has lapsed, a guideline-recommended medication never started. It evaluates the full OMOP record, structured fields and clinical notes together, against published clinical practice guidelines, and presents every finding with its supporting evidence for clinical review before anyone is contacted.
Care Gap Analysis for Your Patient Cohort
Quality measurement built on billing codes counts what was coded, which is a narrower thing than what was done. A mammogram read at an imaging center and faxed back as a PDF, a statin started during an inpatient stay and described in the discharge summary, a colonoscopy documented in a scanned consultation letter: each of these closes a gap, and none of them reliably appears in a structured query. The Care Gap Analyzer reads them, credits them against the measure, and keeps the patient off the outreach list.
That accuracy runs in both directions. A worklist that names patients who already had the test loses the trust of the nurse working it, and once the third entry is wrong, the fourth stops being called.
Why care gaps persist even where the guideline is unambiguous
Every figure below comes from a peer-reviewed study published since 2023. The pattern holds across specialties, payer types, and care settings, including conditions where the recommended action is neither expensive nor clinically contested.
Half of US adults with hypertension do not have it controlled
Across NHANES cycles covering 25,128 adults, 51.1% of US adults with hypertension had controlled blood pressure in 2021 to 2023, defined as systolic below 140 and diastolic below 90 mm Hg. Among those already taking an antihypertensive, 68.3% were controlled, so roughly a third of treated patients need a change that has not happened yet.
Hardy et al., Am J Hypertens 2025
8.2% of adults with diabetes received none of the six recommended services
Among 25,616 US adults with diagnosed diabetes, receipt of any individual ADA-recommended preventive service ranged from 32.6% to 89.9%, and in 2020, 8.2% received none of the six: dental exam, dilated eye exam, foot exam, two or more A1C tests, cholesterol test, and influenza vaccination. Apart from vaccination, the trend was flat for more than a decade.
Wittman et al., Diabetes Care 2023
38.6% of adults are not up to date on colorectal cancer screening
Only 61.4% of US adults aged 45 to 75 were up to date in 2022, below the Healthy People 2030 target of 72.8%. The shortfall is mostly absence rather than lapse: 32.3% had never been screened and 6.3% were overdue. Among adults aged 45 to 49, newly eligible since 2021, 65.7% had never been screened.
King et al., Prev Chronic Dis 2025
More than 4 in 5 eligible adults are not up to date on lung cancer screening
USPSTF gives annual low-dose CT a Grade B recommendation for adults aged 50 to 80 with a 20 pack-year history who currently smoke or quit within 15 years. National survey analysis of that eligible population found fewer than 1 in 5 up to date, leaving the majority of high-risk adults unscreened for a cancer whose survival depends heavily on stage at detection.
Ezenwankwo et al., CHEST 2026 · USPSTF recommendation
28.3% of eligible atrial fibrillation patients receive no anticoagulant
Among 9,513 atrial fibrillation patients with a CHADS₂ score of 2 or higher in an integrated delivery system, 71.7% received guideline-recommended oral anticoagulation, leaving 28.3% on none as of 2021. An integrated system holds one record for the whole patient, so this figure is closer to a floor than to a typical rate.
Malik et al., BMC Cardiovasc Disord 2023
Antihypertensive non-adherence costs $1,441 per patient per year
Among 379,503 commercially insured adults with hypertension, 45.6% were non-adherent to antihypertensive therapy. Non-adherent patients incurred $1,441 more in annual medical costs than adherent patients (95% CI, $1,172 to $1,709), $19,210 against $17,770.
Lee et al., J Am Heart Assoc 2024
Where these gaps sit is as informative as how large they are. Across 126,306 primary care patients at 18 clinics in one health system, 79.8% had a documented mental health screening over a 14-month period, but the clinic-level rate ranged from 51.3% to 98.6% (Müller et al., PLOS ONE 2024). Same guideline, same organization, a 47-point spread.
Gaps also close once an organization can see them. Across 20 Ohio primary care practices, 17 of them federally qualified health centers, depression screening at visits by women aged 18 to 44 rose from 78.4% to 91.1% after a systematic protocol was put in place, an absolute gain of 12.7 percentage points (Bose-Brill et al., J Gen Intern Med 2026). Nothing changed about the recommendation. What changed was whether the practice could tell which patients were missing.
What the analyzer compares: the guideline against the record
Care gap detection is a comparison between what the evidence recommends for a patient with this profile and what the patient's record shows actually happened.
The recommendation
Guideline retrieval runs through the Clinical Guidelines module, under Agents and Tools, which indexes published clinical practice guidelines and returns the passage carrying the recommendation along with its citation. Coverage includes:
- ACC/AHA for cardiovascular risk, heart failure, hypertension, and lipid management
- ADA Standards of Care for diabetes screening, monitoring, and treatment targets
- GOLD for COPD diagnosis, staging, and pharmacological management
- NCCN for oncology screening and surveillance
- USPSTF for preventive services in adults and children
- CMS and NCQA HEDIS measure specifications for payer-reported quality
The record
Detection evaluates every relevant OMOP CDM table: condition_occurrence, drug_exposure, procedure_occurrence, measurement, observation, and note_nlp. Structured codes and NLP-extracted note evidence are evaluated in the same SQL pass, so a fact carries the same weight whether it was coded at billing time or written in a narrative.
Concept hierarchy expansion through concept_ancestor handles the distance between how a guideline is written and how care is recorded. A rule specifying moderate-intensity statin therapy matches a patient on atorvastatin 20 mg without anyone maintaining a code list by hand.
How much this matters depends on the measure. For exclusion and eligibility criteria that live in narrative text, the difference between reading notes and reading codes is not marginal: language models applied to clinical notes identified 93.8% of patients with an adverse social determinant of health in an annotated cohort, against 2.0% for ICD-10 Z-codes (Guevara et al., npj Digital Medicine 2024).
Why unstructured note coverage changes the answer
Across 119,492 clinical notes from 15,000 encounters at a US academic medical center, 80% to 95% of the clinical concepts found in the notes were absent from that encounter's ICD-10-CM codes (Smith et al., JAMIA Open 2026). A separate study of 1.8 million patients found that only 13% of concepts extracted from free text had a similar structured counterpart (Seinen et al., J Med Internet Res 2025).
The error runs in both directions. Across roughly 17 million patient-years at 56 US healthcare organizations, about 45% of person-diagnosis data appeared only in claims and was missing from the EHR (Xiong et al., JAMIA Open 2026). No single source holds the whole patient, which is why gap detection runs against the harmonized OMOP record rather than any one feed.
For care gaps the consequence is asymmetric: missing evidence produces a false open gap, and a false open gap produces a phone call to a patient who already had the test. See the clinical data accuracy gap for the full body of research.
How every finding is classified: four statuses, HEDIS-aligned
Each patient-gap pair receives one of four statuses. The taxonomy follows HEDIS measurement methodology, so results feed quality reporting without a second translation step.
Closed
Qualifying evidence was found within the lookback window. The patient received the care. The source document and extraction details are recorded so the decision can be re-examined later.
Open - never done
No evidence of the intervention anywhere in the patient record, in any document type or time window. This is the status most likely to reflect missing source data rather than missing care, which is why review comes before outreach.
Open - overdue
Evidence was found, but outside the required lookback window. The care was delivered previously and has lapsed; compliance requires a new qualifying event.
Excluded
The record documents an exclusion: a contraindication, a prior adverse event, a competing diagnosis, or a recorded patient preference. The gap does not count against the quality metric, and the exclusion carries its source reference.
The population gap rate follows the same convention:
Gap rate = open gaps / (eligible patients − excluded patients)
Every exclusion in that denominator adjustment is documented with a source reference and included in the audit trail, which is what allows the rate to be defended in a quality audit rather than merely reported.
The five-step detection pipeline
Each finding is produced by a deterministic sequence with inspectable intermediate output. The language model reads guidelines and writes reports; SQL decides which patients have a gap.
Retrieve the guideline passage
For each care gap definition, the pipeline retrieves the relevant sections of the indexed clinical practice guideline corpus through retrieval-augmented generation. The retrieved passages, with their citations, define what counts as a compliant intervention.
Structure the recommended intervention
A Medical LLM reads the retrieved passages and extracts the intervention type, the recommended frequency, the eligible population characteristics, and the acceptable forms of evidence into a JSON schema you can read and correct before anything runs against patient data.
Resolve concepts through the Terminology Server
Each extracted intervention is mapped to OMOP concept IDs through the JSL Terminology Server, which resolves synonyms, brand and generic equivalences, and hierarchical relationships. A rule written as 'ACE inhibitor or ARB' resolves to the full concept set covering both drug classes.
Detect gaps with SQL against the OMOP CDM
Detection runs as SQL over the patient's OMOP records with concept_ancestor joins covering the full concept hierarchy. Structured records and note_nlp evidence are evaluated together in the same pass, against the cohort and lookback window the measure specifies.
Generate the patient and population reports
A Medical LLM at low temperature turns the SQL results into patient-level and population-level reports: gap status, the evidence found or absent, source references, and a plain-language summary for the reviewer. Because the clinical decision was made in SQL, the model is describing a result rather than producing one.
Audit lineage on every finding
Each gap carries its full provenance: the guideline source and version, the retrieved passage, the extracted intervention JSON, the resolved OMOP concept IDs, the SQL evaluation result, the source documents reviewed, the NLP extraction timestamps, and the reviewer's decision. That chain is what makes a gap count defensible in a quality audit rather than a number someone has to take on trust. See Data Governance for how lineage is recorded and retained across the platform.
Running a care gap analysis
Step 1: Scope the patient population
Define the population the measure applies to using any of the OMOP cohort tools in Patient Journey Intelligence: the visual filter builder, the conversational Assistant, or a saved template. Care gap cohorts are typically scoped by:
- Condition, for patients carrying a specific chronic diagnosis
- Age and demographics, for age-banded preventive measures
- Attribution, for a practice, panel, or payer population
- Last contact date, to separate patients who are unreachable from patients who are simply overdue
Cancer Registry - step 1: define the patient cohort describes all three cohort definition methods in full; the same tools apply here. When the measure comes from a published guideline rather than an established measure set, turn the guideline into cohort criteria first, then return to this step.
Scope the denominator before you run anything
A gap rate is only interpretable against a stated denominator. Exclude patients who should not be in it - wrong attribution, deceased, out of age band, enrolled in hospice - at cohort definition time, and record the exclusion reasoning in the cohort description. A gap count without a written definition is a number nobody can reproduce next quarter.
Step 2: Select or configure the care gap definition
A care gap definition states the clinical criteria a patient must meet to be considered compliant, and the evidence the analyzer accepts as proof that the care was delivered.
Start from a template
Pre-built templates encode common quality measure criteria aligned with HEDIS, ACC/AHA, ADA, and USPSTF specifications.
Preventive screening
Breast cancer screening (mammography), colorectal cancer screening (colonoscopy, FIT, stool DNA), cervical cancer screening (cytology, HPV testing), lung cancer screening (low-dose CT), diabetic retinal exam, and diabetic foot exam.
Chronic disease monitoring
HbA1c testing intervals in diabetes, blood pressure monitoring in hypertension, LDL testing for patients at cardiovascular risk, and kidney function monitoring (eGFR and urine albumin-to-creatinine ratio) in chronic kidney disease.
Medication initiation and adherence
Statin therapy after a cardiovascular event, ACE inhibitor or ARB in heart failure with reduced ejection fraction, anticoagulation in non-valvular atrial fibrillation, inhaled corticosteroid in persistent asthma, and SGLT2 inhibitor in HFrEF.
Immunization and counseling
Annual influenza vaccination, pneumococcal and herpes zoster vaccination for eligible patients, tobacco cessation counseling for current smokers, and depression screening with PHQ-9.
Define a custom rule
For institution-defined or program-specific measures, configure a rule by specifying:
- Name and plain-language definition of the unmet need, written so a clinician who was not in the room can tell what is being measured
- Eligible population criteria, the patient characteristics that make the gap applicable, for example age 50 or older with a type 2 diabetes diagnosis
- Compliance evidence, the structured codes (ICD-10, CPT/HCPCS, LOINC, RxNorm) and the documentation patterns in narrative text that count as proof
- Lookback window, the interval within which evidence must fall, for example HbA1c within 12 months or mammography within 24 months
- Evidence threshold, whether one qualifying event closes the gap or several are required, as HEDIS specifies for some measures
- Exclusion criteria, the contraindications, competing diagnoses, and documented patient preferences that remove a patient from the denominator
Reusable calculations that a gap rule depends on, such as eGFR or CHA₂DS₂-VASc, should be defined once in Clinical Measures rather than restated inside the rule. One definition, one threshold, every team querying the same logic.
Step 3: Run the analysis
Go to Agents and Tools > Care Gap Analyzer and start a new analysis.
- Select the saved cohort.
- Select the target condition, care topic, or measure to evaluate.
- Review the suggested criteria or care gap definition, and adjust them where the configuration allows.
- Start the analysis.
- Wait until the job status reaches Completed, then open the results.
The five-step pipeline runs for every patient in the cohort against every active definition, and each patient receives a status for each applicable gap: Closed, Open - never done, Open - overdue, or Excluded. Analysis runs as a background job, so you can leave the page and return to it.
Step 4: Review the population and patient results
Population gap summary
The summary shows how gaps are distributed across the analyzed cohort:
- Gap type breakdown, which definitions have the highest prevalence in this population
- Affected patient counts, the number and share of patients with each gap open, by status
- Severity distribution, gaps stratified by clinical urgency and by time elapsed since the care was due
- Trend comparison, gap rates measured against a previous run on the same cohort, which is what tells you whether an intervention is working
Patient-level gap detail
For each patient, the analyzer shows the open gaps with whether each is a first-time gap or lapsed compliance, the closed gaps with the evidence found and a link to the source document, and a priority score combining gap severity, patient risk factors, and time since the care was due.
Selecting any open gap opens the full evidence search result: what was looked for, what was found, which documents were reviewed, and which OMOP concept IDs were queried. A care manager gets the context needed to act without opening the EHR.
Step 5: Confirm gaps against the patient timeline
Before a gap becomes a phone call, open the Patient Journey timeline behind it. Check whether the timeline holds evidence that supports or resolves the gap, open the source documents where more detail is needed, and record the review decision. Missing events in a timeline can reflect missing source data rather than absent care, and this is the step that tells the two apart.
Human-in-the-loop is not optional here
Patient Journey Intelligence identifies and evidences the gap. Every finding is presented for clinician or quality reviewer confirmation before it counts against a quality metric or drives outreach, and override decisions are written to the audit trail. Clinical judgment about what to do for the patient stays with the care team.
Step 6: Export for care management and quality reporting
Click Export CSV to generate a structured file of all gap findings for the analyzed cohort. Individual patient records can also be exported from the patient detail view.
Exported records carry the provenance metadata the downstream workflow needs: guideline source, OMOP concept IDs queried, source document references, extraction timestamps, gap status, due date, explanation, and reviewer notes. That metadata is what supports HEDIS and CMS quality program submission, where the finding has to be traceable back to the document that justified it.
How this compares to the alternatives
| Capability | Manual chart review | Claims and code-based rules | General-purpose LLM | Care Gap Analyzer |
|---|---|---|---|---|
| Unstructured note coverage | Full, at human speed | None | Partial, unvalidated | Full, with extraction provenance |
| Clinical entity extraction accuracy | Varies with reviewer training | Not applicable | Below healthcare-specific models on clinical benchmarks | 96%+ medical information extraction; 98% entity coding |
| Concept hierarchy expansion | Reviewer-dependent | Hand-maintained code lists | None | concept_ancestor joins over OMOP vocabularies |
| Standards alignment | Manual mapping | Varies by vendor | None | Native OMOP CDM v5.4, FHIR R4 export |
| Audit lineage for quality submission | Manual documentation | Partial | None | Guideline version through source document |
| Scaling to a full attributed population | Limited by reviewer hours | High | Cost scales per token | Batch job on your own infrastructure |
Extraction and coding accuracy figures are published on the Patient Journey Intelligence product page. The underlying models are the same medical language models benchmarked on johnsnowlabs.com/healthcare-llm, which reports first place across 15 medical benchmarks against the current frontier models from OpenAI, Anthropic, and Google.
Who runs care gap analysis, and why
Providers and health systems
Find patients in an attributed panel who are overdue for screening, monitoring, or guideline-recommended therapy. Prioritize outreach by clinical urgency and time since last contact, and push confirmed gaps into care management workflows or EHR task queues.
Payers and health plans
Run HEDIS and CMS Star Ratings gap closure across attributed populations without a manual chart chase. Exports carry the provenance a quality submission has to survive, so the evidence package is produced by the same run that found the gap.
Quality and population health teams
Measure where delivered care diverges from the evidence base, track closure rates across successive runs on the same cohort, and compare gap rates across care settings, demographics, and provider panels to surface disparities that a single population-level number hides.
Life sciences and real-world evidence
Quantify guideline concordance in real patient populations for comparative effectiveness research and post-market surveillance, identify trial-eligible patients from documented treatment gaps, and use gap rate as an outcome variable in RWE studies.
In production
SelectData applies John Snow Labs medical language models to home health documentation at scale, increasing overall production by 10% while maintaining 95% accuracy with existing staff. See the public case studies: SelectData uses AI to better understand home health patients and SelectData interprets millions of patient stories with deep-learned OCR and NLP.
Security, compliance, and deployment
Data never leaves your environment
The analyzer runs inside your deployment. No PHI is sent to an external API, and no external model dependency sits in the detection path. Data handling meets the HIPAA Privacy and Security Rules and GDPR Article 25 data protection by design.
Immutable audit lineage
Every finding carries guideline version, retrieved passage, extracted intervention JSON, resolved OMOP concept IDs, SQL result, source document references, and reviewer decisions, retained for quality program audit and regulatory submission.
Role-based access control
Field-level RBAC scopes what care managers, quality analysts, and administrators can see. Access to sensitive gap categories such as behavioral health, HIV, and substance use is configured independently of general access.
Human review before action
No gap counts against a quality metric or triggers outreach without a reviewer confirming it. Overrides, along with their reasoning, are written to the audit trail alongside the original finding.
Deployment options
Runs on-premises, on AWS, Azure, Databricks, or Snowflake, and in air-gapped environments. Compatible with existing EHR integrations and with an OMOP CDM you already maintain.
Open standards throughout
OMOP CDM v5.4 for patient data, FHIR R4 for interoperability, and SNOMED CT, LOINC, RxNorm, ICD-10, and CPT for concept standardization. No proprietary format anywhere in the pipeline.
Where to go next
- Run care gap analysis on a patient cohort - the end-to-end walkthrough, from saved cohort to validated outreach worklist
- Turn a clinical guideline into measurable cohort criteria - for measures that come from a published recommendation rather than an established measure set
- Cohort Builder - define the population the measure applies to
- Clinical Measures - define the calculations a gap rule depends on, once
- Patient Journey - the timeline reviewers open before confirming a gap
The Care Gap Analyzer is a Patient Journey Intelligence agent that identifies patients with unmet care needs by comparing published clinical practice guidelines against the full OMOP patient record, including facts extracted from clinical notes and scanned documents. It runs a five-step pipeline - guideline retrieval, Medical LLM intervention extraction, OMOP concept resolution, SQL gap detection, and report generation - and assigns each patient-gap pair one of four HEDIS-aligned statuses: Closed, Open - never done, Open - overdue, or Excluded.
Claims-based and code-based detection only sees care that was coded. Care delivered at an outside facility, documented in a scanned report, or described in a discharge summary is invisible to it, which produces open gaps for patients who already received the care. Across 119,492 notes from 15,000 encounters at a US academic medical center, 80% to 95% of the clinical concepts documented in the notes were absent from that encounter's ICD-10-CM codes, and a study of 1.8 million patients found only 13% of concepts extracted from free text had a structured counterpart. The Care Gap Analyzer evaluates structured records and NLP-extracted note evidence in the same SQL pass, so documented care closes the gap regardless of where it was written down.
Guideline retrieval covers ACC/AHA for cardiovascular care, ADA Standards of Care for diabetes, GOLD for COPD, NCCN for oncology screening and surveillance, USPSTF for preventive services, and CMS and NCQA HEDIS measure specifications. You can also define custom rules for institution-specific or program-specific measures by specifying eligible population criteria, accepted compliance evidence, lookback window, evidence threshold, and exclusions.
Three mechanisms. It reads unstructured documents as well as structured fields, so care recorded only in narrative text is credited. It expands concepts through OMOP concept_ancestor joins, so a rule written at the class level matches the specific drug or procedure that was recorded. And it separates Open - never done from Open - overdue, which distinguishes a patient with no history of the intervention from one whose compliance has lapsed. Every finding is then reviewed against the Patient Journey timeline and its source documents before any outreach.
Gap rate is open gaps divided by eligible patients minus excluded patients. A patient is excluded when the record documents a contraindication, a prior adverse event, a competing diagnosis, or a recorded patient preference. Every exclusion carries a source reference and appears in the audit trail, so the denominator adjustment can be inspected rather than taken on trust.
Guideline source and version, the retrieved guideline passage, the extracted intervention JSON, the resolved OMOP concept IDs, the SQL evaluation result, the source documents reviewed with the supporting passage highlighted, NLP extraction timestamps, and the reviewer's decision and any override reasoning. This lineage is what makes a gap count defensible in a CMS or HEDIS quality submission audit.
No. A Medical LLM reads the guideline and writes the report; the gap decision itself is made by SQL evaluated against the OMOP CDM. The intermediate outputs of each pipeline step - the retrieved passage, the extracted intervention JSON, the resolved concept IDs - are inspectable, so a disagreement about a result becomes a disagreement about a stated criterion rather than an argument about what a model did.
Yes, and review is required rather than optional. Every identified gap is presented for clinician or quality reviewer confirmation before it counts against a quality metric or generates outreach. Overrides and their reasoning are written to the audit trail next to the original finding, so the record shows both what the pipeline found and what a human decided.
Results export to CSV for the whole analyzed cohort or for individual patients. Each record carries patient identifier, gap status, due date, explanation, reviewer notes, guideline source, the OMOP concept IDs queried, source document references, and extraction timestamps, which is the metadata a care management worklist and a quality program submission both need.
No. The Care Gap Analyzer runs entirely inside your deployment - on-premises, on AWS, Azure, Databricks, or Snowflake, including air-gapped environments. There is no external API call in the detection path and no PHI transmitted to a third-party service.