Prepare a de-identified dataset for third-party sharing
Build the dataset the agreement covers, de-identify it, review original and de-identified text side by side, and export in the form the recipient is entitled to.
De-Identify Your Clinical Data for Third-Party Sharing
Who this is for
Anyone preparing clinical data to leave the organization under an agreement: data sharing offices and research collaboration managers, privacy officers approving the release, study teams delivering a dataset to a pharma or academic partner, teams contributing to a research network or multi-site study, and groups supplying a sample for a vendor evaluation.
Why it matters
A health system signs a research collaboration with a life sciences partner. The agreement obligates delivery of a de-identified note corpus and, separately, documented evidence that identifiers were reviewed before release. The second obligation is the one that stalls projects. A processing job that removed identifiers is not the same as a record showing that a named person examined what was removed and approved it.
This scenario differs from De-Identify Clinical Notes with Human Review in what happens at the end. That one produces reviewed de-identified text for internal use. This one produces a package prepared for handoff outside the organization.
What you gain
Side-by-side comparison of original and de-identified text during review, and export options covering original, masked, or both. That comparison is what your privacy officer signs against, and the export choice lets you match the partner's contractual entitlement exactly rather than sending more than the agreement allows.
Before you start
- An executed data use agreement or equivalent authorization naming what may be released
- A cohort or ingestion results defining the population in scope
- Access to Data Enrichment > De-Identification, and a named reviewer with authority to approve the release
- Agreement on which export format the partner receives: original, masked, or both
Step 1: Create a dataset from a cohort or ingestion filters
- In the left navigation, go to Data Curation Studio > Dataset Builder.
- Select the dataset mode based on the sharing scope, such as Cohort & Clinical Filter or an ingestion-filter-based dataset mode.
- Select the data type required for the external review package, such as clinical documents or images.
- Click Next to continue to the next dataset configuration step.
- Select the cohort, source, document filters, image filters, date range, or note types that should be included.
- Keep the dataset scope limited to the files approved for external review.
- Click Next to continue to Dataset Details.
- Enter the dataset name and description.
- Use a name that clearly identifies the dataset as an external sharing or de-identification package.
- Click Next to continue to Review Dataset.
- Review the selected patients, documents, images, filters, and expected dataset contents.
- Click Create Dataset to create the dataset.
Step 2: Include the documents or images intended for external review
- Open the created dataset from Dataset Explorer.
- Review the included patients, documents, images, note types, dates, and source files.
- Confirm that the dataset includes only the documents or images intended for external review.
- Remove or exclude files that should not be part of the external sharing package where the workflow supports adjustment.
- Confirm that the final dataset scope matches the approved sharing request.
Step 3: Run De-Identification on the dataset
- In the left navigation, go to Data Enrichment > De-Identification.
- Start a new de-identification job or session.
- Select the dataset created for external sharing.
- Select the appropriate de-identification profile and pipeline for the package.
- Review the job configuration before starting.
- Click the start action to begin de-identification.
- Monitor the job until processing is complete.
Step 4: Review detected identifiers and de-identified outputs
- Open the completed de-identification job details page.
- Review detected identifiers at the document or image level.
- Inspect the original and de-identified versions side by side where comparison is available.
- Confirm that patient names, identifiers, dates, contact details, locations, provider details, and other protected or sensitive elements are handled according to the external sharing requirements.
- Approve, correct, mask, unmask, or flag detections according to the available review workflow.
- Pay special attention to scanned documents, embedded tables, headers, footers, and image overlays.
- Save review decisions where supported.
Step 5: Export or save the de-identified package
- Confirm that the de-identification review is complete.
- From the de-identification details page, choose the export option that matches the delivery need: Original images only, Masked images only, or Both original and masked images.
- Export the reviewed de-identified outputs.
- If patient-level or file-level download is available, open the relevant item and download the reviewed de-identified version.
- Store the exported package in the approved delivery location.
- Track any files that were excluded, flagged, or left unresolved during review.
Step 6: Prepare the reviewed output package for approved sharing
- Review the exported package before delivery.
- Confirm that the package contains the expected de-identified documents or images.
- Confirm that excluded or flagged files are not included unless explicitly approved.
- Add any required delivery notes, review notes, or package metadata.
- Share the reviewed package according to the approved external sharing workflow.
- Keep the dataset, de-identification job, and export details available for traceability.
Recipe reference
Each stage of this scenario is also a reusable building block.
These steps have their own pages:
Recipe: De-identify a dataset before sharing or review
When to use on its own: Before sharing data externally, assembling a training or evaluation corpus, or reviewing images and documents that may contain patient identifiers.
Features involved: Dataset Builder, De-Identification, patient/document/image review, side-by-side original and de-identified review where available.
Edge cases / limitations: De-Identification requires a dataset; it isn't cohort-only. Export/save behavior should be confirmed before promising a complete handoff. Validate outputs before external sharing.
Value: Produces a documented privacy review a privacy officer can sign against before anything is released.