Seiba Risk Scanner is an open-source toolkit that turns sensitive data detection into a full privacy risk assessment, for tables and clinical text, before data goes into AI, analytics or a partner’s hands.

Good tools already find names, dates and record numbers in health data. Before that data moves into an AI system, a research dataset or a partner’s environment, your team still has harder questions to answer.
Where it fits. After sensitive data has been found, and before the dataset is used or shared.
The plan is written and explained before any data changes. Your team reviews it, then applies it.
Find names, dates, record numbers and other sensitive details in tables and clinical text.
CSV · tabular exports · clinical notes · lettersMeasure exposure across the whole dataset, not one finding at a time.
Exposure Index · severity · review queues · re-identification indicatorsShow what was found, why, and the reasoning behind each recommended action.
Detector evidence · confidence · provenanceChoose a de-identification strategy that meets your privacy goals and keeps the data useful.
Configurable privacy objectivesRun the plan across tables, documents and notes, with each repeated identifier replaced the same way.
Policy-driven de-identificationCheck the privacy risk that remains, and how much useful information you kept.
Residual exposure · analytical utilityIt works the same way on structured exports and on free-text clinical documents, so one assessment covers the whole dataset.
See where privacy exposure is concentrated and which records need review first, instead of a flat list of findings.
Exposure Index · severity distribution · review queues · privacy heatmapsEach recommended action records what was found, why it was classified that way, which policy decided it and what replaces it.
Evidence · confidence · provenance · replacement strategyThe de-identification plan comes first. Once approved, it runs across tables and documents automatically.
Policy-driven de-identificationNot every sensitive value needs the same treatment. Optimization profiles let you weigh privacy against what the data still has to do.
Utility-aware optimizationKeep specific clinical concepts, keep a year instead of removing a date, or change the default for any entity type, without touching code.
Configurable policies · editable ontologyPoint it at a folder of mixed tables and documents and get one consolidated report instead of reviewing files one by one.
Patient datasets · lab exports · clinical notes · consultation reports · referral lettersSeiba Risk Scanner is one part of one pillar, released as open source. Seiba Pulse watches all seven pillars, continuously, inside your environment.
Full documentation is being written. Until then, the repository and its examples are the best place to begin.