avishake/phi-deidentification-pipeline
Clinical De-identification Pipeline.
Hospital archives → MedGemma training corpora, on-prem
model card · description
An on-premise pipeline that turns unstructured hospital archives into MedGemma training data without a single identifier leaving the building. Provider-agnostic multimodal page reading (Anthropic API, any OpenAI-compatible endpoint, or self-hosted vLLM), clinical/non-clinical classification, PHI removal from transcribed text and scanned images, automated leak verification, and AES-256 per-location packaging with SHA-256 manifests.
- Image PHI removed with fractional rotated zonal masking; zero surviving identifiers on a 48-page clinical validation sample.
- Data-driven redaction entities generated from extractor-reported PHI, eliminating O(patients × pages) matching at multi-terabyte scale.
- In-place ZIP reading with header-only content sniffing surveyed an 11,631-file archive ~20× faster than full decompression.
- Batched, checkpointed stages that resume after interruption.
- Hardware and cost analysis (memory bandwidth, dense vs. MoE, per-page token economics) behind the on-prem decision.
evaluation
- 0
- Leaked identifiers
- ~20× faster
- Archive survey
- 11,631
- Files surveyed