Automated ICD-10, CPT, LOINC, and SNOMED mapping breaks on negation, timing, and family context. See the technical framework and production evidence behind it.

Entity recognition alone can’t tell a negated symptom from an active one, or a family member’s diagnosis from the patient’s own – assertion status and relation detection close that gap.
Automated clinical code mapping turns diagnoses, procedures, lab results, and clinical findings documented in free text into ICD-10, CPT, LOINC, and SNOMED CT codes that billing, quality-measurement, and research systems can consume. The hard part isn’t spotting the words. A note stating a patient “denies chest pain” or documenting a “family history of colorectal cancer” contains the right clinical term but the wrong code if a model can’t separate presence from absence, or the patient from a relative. That distinction, repeated across negation, timing, and site-to-site documentation habits, is what separates a mapping pipeline that works in production from one that only works in a demo.
Why is mapping clinical text to codes harder than it looks?
Finding “pneumonia” in a clinical note is entity recognition. Deciding whether that mention should ever become a code is a different problem, and it’s the one that actually determines coding accuracy. Three context signals decide it – assertion status, experiencer, and timing – and general-purpose LLMs and standard extraction pipelines handle all three inconsistently, because each requires reasoning about context rather than recognizing a span of text.
Assertion status: whether a finding is present, absent, hypothetical, or conditional. Peer-reviewed results put the ceiling below production tolerance: an LLM-based assertion detector using instruction prompting and LoRA fine-tuning reached 0.74 F1 across the three assertion axes – certainty, temporality, and experiencer – a 0.31 F1 improvement over the prior best method and still well short of reliable at the level of a single coded finding (Ji et al., 2024). Reported accuracy clusters high on straightforward present-versus-absent cases, but dropped to 0.599 on conditional statements like “ruled out for tuberculosis pending culture results.” Model scale does not close the gap: a 2025 evaluation of LLMs on longitudinal clinical records found that longer context windows improved input integration without consistently improving clinical reasoning, leaving hedged and conditional statements mis-assigned even as models grew (Kruse et al., 2025). Conditional and hedged language, not the easy negation cases, is where coding accuracy actually fails – and it fails silently, because the entity itself was recognized correctly.
Experiencer attribution: whose condition is actually being described. “Family history of colorectal cancer,” “personal history of colorectal cancer,” and “active colorectal cancer” can appear in adjacent sentences of one note, and each maps to a different ICD-10 code with different reimbursement and risk-adjustment consequences. Getting this right requires tracking subject references across sentences, not just recognizing the diagnosis term, which is why experiencer – patient versus family member – is scored as its own axis in the assertion-detection literature, and why systems evaluated across that axis top out near 0.74 F1 rather than the high-0.9s routinely reported for plain entity recognition (Ji et al., 2024).
Temporal anchoring. A discharge summary can describe a diagnosis from three years ago, a procedure performed last week, and a follow-up scheduled next month, all in one paragraph, and correct coding depends on anchoring each event to when it actually happened rather than to the date the note was written. Recent work applying graph transformers to clinical temporal-relation extraction improved the standard TempEval F1 score by 5.5% over the prior best result, and by 8.9% specifically on long-range relations linking events across distant parts of a document (Chaturvedi et al., 2025). A 2025 JAMIA Open evaluation of LLMs on temporal relation extraction from clinical reports is blunter about the ceiling: the best configuration reached 0.70 F1 on simple relations, while complex relations fell to between 0.03 and 0.40 F1 (Andrew et al., 2025). Temporal anchoring is not a detail a modern language model handles by default — it is the axis where reported performance is weakest, and the one that decides whether a historical diagnosis is coded as active.
Why the same model performs differently at every hospital

A model validated at one hospital can lose more than 20 points of accuracy at another – documentation style, EHR structure, and workflow differences all break generalization.
Even a model that handles negation, family context, and timing correctly on validation data can fail once it moves to a new institution’s notes, because clinical documentation style is not standardized across hospitals, EHR systems, or even departments within the same health system. A review of NLP methods applied to cancer notes found extraction targets ranging from TNM staging and biomarker status to medication administration and social determinants of health, with reported performance depending heavily on “variations in clinical documentation practices across institutions” (Bilal et al., 2024).
The clearest evidence of how large this gap gets comes from CPT coding itself. A study training anesthesiology procedure-coding models on 1.6 million procedures across 44 U.S. hospitals found that a model built and validated at a single institution reached 92.5% accuracy on its own data, then lost roughly 22.4 percentage points when applied to a different hospital’s documentation. A model trained across multiple institutions traded a small amount of internal accuracy for a 17-percentage-point gain in external performance (Pandian et al., 2025). A vendor benchmark validated at one hospital says very little about how that same model will perform once it hits a health system it has never seen, which is the scenario every real deployment faces on day one.
Entity resolution, the step that maps a recognized concept to a specific SNOMED CT, ICD-10, or LOINC code, shows the same fragility with general-purpose models. A 2025 evaluation of large language models mapping medical terms to SNOMED CT concepts found GPT-4o reaching 93.75% precision, while an open-source alternative, Llama 3.3 70B, managed only 19.19% on the identical task, with both models weakest on complex, multi-parameter clinical measurements, the kind of composite lab findings that LOINC exists to standardize (Mavridis et al., 2025). A 74-point precision gap between two general models on the same mapping task is not a rounding error; it’s evidence that entity resolution needs purpose-built terminology infrastructure rather than a general model’s best guess.
What a production mapping pipeline actually requires
Closing these gaps takes a pipeline built around the specific failure modes above, not a single accurate model. Healthcare NLP extracts clinical entities, then applies assertion detection to classify each one as present, absent, historical, conditional, or family-associated before any code gets assigned. Kocaman et al. (2025) define what that layer has to look like: assertion status treated as a dedicated, domain-adapted classification task rather than something a general model is trusted to infer. John Snow Labs’ fine-tuned clinical model reaches 0.962 combined accuracy against GPT-4o’s 0.901, with the widest margins on precisely the categories that break coding – hypothetical findings and conditions belonging to someone other than the patient – and this deep-learning variant holds comparable accuracy at a fraction of LLM inference cost, which is what makes assertion classification affordable to run on every entity in every note rather than on a sample. Temporal and relation models anchor each finding to when it happened and connect related concepts, such as a diagnosis and its complication, so that a code representing a causal relationship (like diabetes with diabetic chronic kidney disease) isn’t assembled by guessing at co-occurring entities.
Entity resolution then maps the correctly qualified concept to ICD-10, CPT, LOINC, or SNOMED CT through John Snow Labs’ Terminology Server, which handles synonyms, abbreviations, and cross-vocabulary translation between code systems, and which John Snow Labs has benchmarked directly against general-purpose LLMs on terminology mapping. Because ICD-10-CM, CPT, SNOMED CT, and LOINC are each updated on their own schedule, a production pipeline needs mapping tables and models that get revalidated against each release, not a static mapping trained once and left alone. Generative AI Lab supplies the human-in-the-loop layer this requires in practice: uncertain extractions route to a reviewer, corrections get captured, and every coded output keeps an audit trail back to the source note, the model version, and the reviewer decision, which is what makes a coding result defensible when a payer or an auditor asks how it was produced.
Evidence from production deployments
Oncology data curation shows what this pipeline looks like at scale. In a collaboration with MiBA, John Snow Labs built an AI pipeline for oncology EHR data that reached an average F1 score of 0.9 on entity relationships, the kind of causal and temporal links between diagnosis, staging, and treatment that a simple entity list cannot represent. The same pipeline substantially increased how much of each key oncology field was actually documented, relative to the pre-pipeline baseline: histology capture rose by 67.5%, metastasis information by 39.9%, and BRAF mutation status by 81.5%. A hybrid NLP/LLM model for clinical trial matching reached an F1 score of 0.81, and an adverse-event detection model reached 0.93 F1 with 0.95 recall and 0.91 precision (Bonis, 2025). None of those numbers come from entity extraction alone; they depend on correctly resolving extracted findings to structured, comparable values across a million-plus note corpus.
Martlet.ai, John Snow Labs’ HCC coding sub-brand, applies the same assertion-and-resolution architecture to Medicare Advantage risk adjustment, where a code is only defensible if it can be traced to specific, date-bounded documentation from an appropriate provider. Martlet.ai’s RADV audit readiness platform links each HCC code to the documentation that supports it and flags conditions where that support is missing, which is a direct answer to the audit-trail requirement described above: a code and its provenance have to travel together.
For teams evaluating an automated mapping pipeline, the questions worth asking a vendor follow directly from this evidence: what is the model’s accuracy on conditional and family-context statements specifically, not just on straightforward present-absent cases; how much accuracy does the model lose when it moves to your institution’s documentation; and what audit trail exists between a code and the note it came from. Explore Healthcare NLP and the Terminology Server directly, or request a technical demonstration scoped to your own note volume and code systems to get answers specific to your environment rather than a benchmark average.
Frequently asked questions
What is automated clinical code mapping? Automated clinical code mapping is the process of converting clinical concepts documented in free text, such as diagnoses, procedures, lab results, and findings, into standard codes such as ICD-10 for diagnoses, CPT for procedures, LOINC for laboratory observations, and SNOMED CT for clinical concepts generally. It combines entity extraction, assertion and temporal classification, and terminology resolution.
Why isn’t entity recognition enough for accurate code mapping? Entity recognition identifies that a clinical term appears in text but doesn’t determine whether it should be coded. A note can mention “pneumonia” while denying it, describing it in a family member, or listing it as ruled out, and each case requires a different, or no, code. A 2025 study found assertion accuracy dropping from 0.976 on straightforward cases to 0.599 on conditional statements, showing this distinction is where mapping accuracy is actually won or lost (Kocaman et al., 2025).
How does negation affect ICD-10 and CPT coding accuracy? Coding a negated finding as present produces an incorrect diagnostic or procedure code with direct billing and quality-measurement consequences. Negation ranges from direct patterns like “denies chest pain” to indirect ones like “not consistent with malignancy,” and assertion-detection models purpose-built for clinical text handle this range more reliably than general-purpose language models evaluated on the same task (Kocaman et al., 2025).
Why do coding models trained at one hospital perform worse at another? Models learn the documentation habits of the institution they were trained on, including abbreviation conventions and EHR template structure, rather than clinical language generally. A study of CPT coding for anesthesiology found single-institution models lost about 22 percentage points of accuracy when applied to a different hospital’s notes, while models trained across multiple institutions generalized substantially better (Pandian et al., 2025).
What’s the difference between entity extraction and terminology mapping? Entity extraction identifies a clinical concept in text, such as a diagnosis mention. Terminology mapping, also called entity resolution or normalization, assigns that concept a specific code in ICD-10, CPT, LOINC, or SNOMED CT. The two steps have different accuracy profiles: one evaluation found a 74-percentage-point precision gap between two general-purpose models on the same SNOMED CT mapping task, showing that extraction accuracy alone does not predict mapping accuracy (Mavridis et al., 2025).
Can general-purpose LLMs replace purpose-built clinical coding models? General-purpose models vary widely on clinical coding subtasks and are not consistent substitutes. On assertion classification, a clinical-specific model reached 0.962 combined accuracy against GPT-4o’s 0.901 (Kocaman et al., 2025). On SNOMED CT terminology mapping, GPT-4o reached 93.75% precision while an open-source general model reached 19.19% on the same task (Mavridis et al., 2025), a gap wide enough that model choice needs task-specific validation, not a general assumption.
What governance does regulatory-grade code mapping require beyond accuracy? An accurate code still needs to be defensible: the pipeline should log what text supported the code, which model version produced it, and its confidence level, so the result can be traced back to a specific note during an audit. Martlet.ai’s RADV audit readiness platform is built around this requirement, linking each HCC code to the date-bounded documentation that supports it before it reaches a CMS-ready export.
Does automated mapping still need human review? Yes, for the cases where confidence is low or the coding decision carries significant financial, clinical, or regulatory weight. A well-designed pipeline uses confidence scoring to route only those cases to a reviewer, rather than reviewing everything uniformly, which is what makes automation a meaningful gain over manual coding rather than manual coding with an extra step attached.
What is automated clinical code mapping?
Automated clinical code mapping is the process of converting clinical concepts documented in free text, such as diagnoses, procedures, lab results, and findings, into standard codes such as ICD-10 for diagnoses, CPT for procedures, LOINC for laboratory observations, and SNOMED CT for clinical concepts generally. It combines entity extraction, assertion and temporal classification, and terminology resolution.
Why isn’t entity recognition enough for accurate code mapping?
Entity recognition identifies that a clinical term appears in text but doesn’t determine whether it should be coded. A note can mention “pneumonia” while denying it, describing it in a family member, or listing it as ruled out, and each case requires a different, or no, code. A 2025 study found assertion accuracy dropping from 0.976 on straightforward cases to 0.599 on conditional statements, showing this distinction is where mapping accuracy is actually won or lost (Kocaman et al., 2025).
How does negation affect ICD-10 and CPT coding accuracy?
Coding a negated finding as present produces an incorrect diagnostic or procedure code with direct billing and quality-measurement consequences. Negation ranges from direct patterns like “denies chest pain” to indirect ones like “not consistent with malignancy,” and assertion-detection models purpose-built for clinical text handle this range more reliably than general-purpose language models evaluated on the same task (Kocaman et al., 2025).
Why do coding models trained at one hospital perform worse at another?
Models learn the documentation habits of the institution they were trained on, including abbreviation conventions and EHR template structure, rather than clinical language generally. A study of CPT coding for anesthesiology found single-institution models lost about 22 percentage points of accuracy when applied to a different hospital’s notes, while models trained across multiple institutions generalized substantially better (Pandian et al., 2025).
What’s the difference between entity extraction and terminology mapping?
Entity extraction identifies a clinical concept in text, such as a diagnosis mention. Terminology mapping, also called entity resolution or normalization, assigns that concept a specific code in ICD-10, CPT, LOINC, or SNOMED CT. The two steps have different accuracy profiles: one evaluation found a 74-percentage-point precision gap between two general-purpose models on the same SNOMED CT mapping task, showing that extraction accuracy alone does not predict mapping accuracy (Mavridis et al., 2025).
Can general-purpose LLMs replace purpose-built clinical coding models?
General-purpose models vary widely on clinical coding subtasks and are not consistent substitutes. On assertion classification, a clinical-specific model reached 0.962 combined accuracy against GPT-4o’s 0.901 (Kocaman et al., 2025). On SNOMED CT terminology mapping, GPT-4o reached 93.75% precision while an open-source general model reached 19.19% on the same task (Mavridis et al., 2025), a gap wide enough that model choice needs task-specific validation, not a general assumption.
What governance does regulatory-grade code mapping require beyond accuracy?
An accurate code still needs to be defensible: the pipeline should log what text supported the code, which model version produced it, and its confidence level, so the result can be traced back to a specific note during an audit. Martlet.ai’s RADV audit readiness platform is built around this requirement, linking each HCC code to the date-bounded documentation that supports it before it reaches a CMS-ready export.
Does automated mapping still need human review?
Yes, for the cases where confidence is low or the coding decision carries significant financial, clinical, or regulatory weight. A well-designed pipeline uses confidence scoring to route only those cases to a reviewer, rather than reviewing everything uniformly, which is what makes automation a meaningful gain over manual coding rather than manual coding with an extra step attached.





























