Open source · Seiba Risk Scanner

Know how risky your health data is
before it moves.

Seiba Risk Scanner is an open-source toolkit that turns sensitive data detection into a full privacy risk assessment, for tables and clinical text, before data goes into AI, analytics or a partner’s hands.

Sample privacy report · Seiba Risk Scanner
Sensitive Data Risk Report from Seiba Risk Scanner: exposure index 46.9 out of 100, and how it was calculated.
Not a list of findings. How much risk the dataset carries, where it sits, and what to do about it.
Beyond detection

Detection is only the beginning.

Good tools already find names, dates and record numbers in health data. Before that data moves into an AI system, a research dataset or a partner’s environment, your team still has harder questions to answer.

Where it fits. After sensitive data has been found, and before the dataset is used or shared.

  1. 01How serious is the privacy risk?
  2. 02Which datasets, records or fields need review first?
  3. 03Where is the exposure concentrated?
  4. 04How likely is it that someone could be re-identified?
  5. 05Which de-identification approach fits how the data will be used?
  6. 06How much useful information is lost once it is anonymized?
  7. 07Can every decision be explained and audited later?
How it works

Every assessment follows the same six steps.

The plan is written and explained before any data changes. Your team reviews it, then applies it.

01

Identify

Find names, dates, record numbers and other sensitive details in tables and clinical text.

CSV · tabular exports · clinical notes · letters
02

Quantify risk

Measure exposure across the whole dataset, not one finding at a time.

Exposure Index · severity · review queues · re-identification indicators
03

Explain

Show what was found, why, and the reasoning behind each recommended action.

Detector evidence · confidence · provenance
04

Optimize

Choose a de-identification strategy that meets your privacy goals and keeps the data useful.

Configurable privacy objectives
05

Apply

Run the plan across tables, documents and notes, with each repeated identifier replaced the same way.

Policy-driven de-identification
06

Measure

Check the privacy risk that remains, and how much useful information you kept.

Residual exposure · analytical utility
In the toolkit

Built to be explained, configured and run on real data.

It works the same way on structured exports and on free-text clinical documents, so one assessment covers the whole dataset.

Risk you can see across the dataset

See where privacy exposure is concentrated and which records need review first, instead of a flat list of findings.

Exposure Index · severity distribution · review queues · privacy heatmaps

A reason behind every decision

Each recommended action records what was found, why it was classified that way, which policy decided it and what replaces it.

Evidence · confidence · provenance · replacement strategy

A plan before any change

The de-identification plan comes first. Once approved, it runs across tables and documents automatically.

Policy-driven de-identification

Privacy balanced with usefulness

Not every sensitive value needs the same treatment. Optimization profiles let you weigh privacy against what the data still has to do.

Utility-aware optimization

Policies you control

Keep specific clinical concepts, keep a year instead of removing a date, or change the default for any entity type, without touching code.

Configurable policies · editable ontology

Whole folders, one report

Point it at a folder of mixed tables and documents and get one consolidated report instead of reviewing files one by one.

Patient datasets · lab exports · clinical notes · consultation reports · referral letters
Pillar 02 · Privacy and governance

Seiba Risk Scanner is one part of one pillar, released as open source. Seiba Pulse watches all seven pillars, continuously, inside your environment.

Explore Seiba Pulse
Seiba Risk Scanner

Start with the code and the examples.

Full documentation is being written. Until then, the repository and its examples are the best place to begin.

GitHub repositorygithub.com/seiba-ai/seiba-risk-scanner Examplesseiba-risk-scanner/examples
Walkthrough · Structured datasetsSoon
Walkthrough · Clinical notesSoon